Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Deborah A. Agarwal

dblp:a/DeborahAAgarwal · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
3since 2021 · last 2023
0000-0001-5045-2396ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7Software engineering, systems software and programming languages · 6Computer networks · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
High-performance computing · 52% Distributed systems · 30% Storage systems · 16%
Computer networks
2 papers
Network measurement and analytics · 58% Internet of things and sensor networks · 29% Internet architecture and protocols · 13%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.312018
Dac-Man: data change management for scientific datasets on HPC systems · SC 2018
Storage systems
storage reliability
0.112018
Dac-Man: data change management for scientific datasets on HPC systems · SC 2018
Distributed systems
fault tolerance
0.031998
The Totem Multiple-Ring Ordering and Topology Maintenance Protocol · ACM Trans. Comput. Syst. 1998
The Totem Single-Ring Ordering and Membership Protocol · ACM Trans. Comput. Syst. 1995
A reliable ordered delivery protocol for interconnected local area networks · ICNP 1995
Internet of things and sensor networks › wireless sensor network › network diagnosis
bottleneck detection
0.012001
Network Characterization Service (NCS) · HPDC 2001
Network measurement and analytics
network characterization
0.012001
Network Characterization Service (NCS) · HPDC 2001
Distributed systems
group communication
0.011998
The Totem Multiple-Ring Ordering and Topology Maintenance Protocol · ACM Trans. Comput. Syst. 1998
Distributed systems › group communication
total order delivery
0.011998
The Totem Multiple-Ring Ordering and Topology Maintenance Protocol · ACM Trans. Comput. Syst. 1998
Internet architecture and protocols › multicast
multicast protocols
0.011995
A reliable ordered delivery protocol for interconnected local area networks · ICNP 1995
Distributed systems › group communication
atomic broadcast
0.011995
The Totem Single-Ring Ordering and Membership Protocol · ACM Trans. Comput. Syst. 1995
Distributed systems
distributed coordination
0.011995
The Totem Single-Ring Ordering and Membership Protocol · ACM Trans. Comput. Syst. 1995
Distributed systems › group communication
group membership
0.011995
The Totem Single-Ring Ordering and Membership Protocol · ACM Trans. Comput. Syst. 1995
Distributed systems › replication
replica consistency
0.011995
A reliable ordered delivery protocol for interconnected local area networks · ICNP 1995
Distributed systems
replication
0.011995
The Totem Single-Ring Ordering and Membership Protocol · ACM Trans. Comput. Syst. 1995
Performance modeling and evaluation › simulation
discrete-event simulation
0.011994
The Totem protocol development environment · ICNP 1994
Distributed systems › fault tolerance
topology change
0.011998
The Totem Multiple-Ring Ordering and Topology Maintenance Protocol · ACM Trans. Comput. Syst. 1998
Distributed systems › group communication
virtual synchrony
0.011995
The Totem Single-Ring Ordering and Membership Protocol · ACM Trans. Comput. Syst. 1995
Software testing
protocol testing
0.011994
The Totem protocol development environment · ICNP 1994

Methods — techniques the papers use, named apart from their topics

provenance tracking · 0.3change detection · 0.3cooperative information gathering · 0.1totem protocol · 0.0discrete-event simulation · 0.0protocol design · 0.0performance evaluation · 0.0token ring · 0.0flow control · 0.0
YearPublicationVenuePosition
2023 Long-term missing value imputation for time series data using deep neural networks
abstract
Abstract We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods and . The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.
Jangho Park, Juliane Mueller 0002, Bhavna Arora, Boris Faybishenko, Gilberto Zonta Pastorello, Charuleka Varadharajan, Reetik Sahu, Deborah A. Agarwal
Neural Comput. Appl.8
2021 Assessing data change in scientific datasets
abstract
Summary Scientific datasets are growing rapidly and becoming critical to next‐generation scientific discoveries. The validity of scientific results relies on the quality of data used and data are often subject to change, for example, due to observation additions, quality assessments, or processing software updates. The effects of data change are not well understood and difficult to predict. Datasets are often repeatedly updated and recomputing derived data products quickly becomes time consuming and resource intensive and may in some cases not even be necessary, thus delaying scientific advance. Despite its importance, there is a lack of systematic approaches for best comparing data versions to quantify the changes, and ad‐hoc or manual processes are commonly used. In this article, we propose a novel hierarchical approach for analyzing data changes, including real‐time (online) and offline analyses. We employ a variety of fast‐to‐compute numerical analyses, graphical data change representations, and more resource‐intensive recomputations of a subset of the data product. We illustrate the application of our approach using three scientific diverse use cases, namely, satellite, cosmological, and x‐ray data. The results show that a variety of data change metrics should be employed to enable a comprehensive representation and qualitative evaluation of data changes.
Juliane Mueller 0002, Boris Faybishenko, Deborah A. Agarwal, Chongya Jiang, Youngryel Ryu, Craig Tull, Lavanya Ramakrishnan
Concurr. Comput. Pract. Exp.3
2021 Surrogate optimization of deep neural networks for groundwater predictions
Juliane Mueller 0002, Jangho Park, Reetik Sahu, Charuleka Varadharajan, Bhavna Arora, Boris Faybishenko, Deborah A. Agarwal
J. Glob. Optim.7
2019 Understanding Data Similarity in Large-Scale Scientific Datasets
abstract
Today, scientific experiments and simulations produce massive amounts of heterogeneous data that need to be stored and analyzed. Given that these large datasets are stored in many files, formats and locations, how can scientists find relevant data, duplicates or similarities? In this context, we concentrate on developing algorithms to compare similarity of time series for the purpose of search, classification and clustering. For example, generating accurate patterns from climate related time series is important not only for building models for weather forecasting and climate prediction, but also for modeling and predicting the cycle of carbon, water, and energy. We developed the methodology and ran an exploratory analysis of climatic and ecosystem variables from the FLUXNET2015 dataset. The proposed combination of similarity metrics, nonlinear dimension reduction, clustering methods and validity measures for time series data has never been applied to unlabeled datasets before, and provides a process that can be easily extended to other scientific time series data. The dimensionality reduction step provides a good way to identify the optimum number of clusters, detect outliers and assign initial labels to the time series data. We evaluated multiple similarity metrics, in terms of the internal cluster validity for driver as well as response variables. While the best metric often depends on a number of factor, the Euclidean distance seems to perform well for most variables and also in terms of computational expense.
Payton Linton, William Melodia, Alina Lazar, Deborah A. Agarwal, Ludovico Bianchi, Devarshi Ghoshal, Gilberto Zonta Pastorello, Lavanya Ramakrishnan, Kesheng Wu
IEEE BigData4
2018 Dac-Man: data change management for scientific datasets on HPC systems
Devarshi Ghoshal, Lavanya Ramakrishnan, Deborah A. Agarwal
SC3
2017 Hunting Data Rogues at Scale: Data Quality Control for Observational Data in Research Infrastructures
abstract
Data quality control is one of the most time consuming activities within Research Infrastructures (RIs), especially when involving observational data and multiple data providers. In this work we report on our ongoing development of data rogues, a scalable approach to manage data quality issues for observational data within RIs. The motivation for this work started with the creation of the FLUXNET2015 dataset, which includes carbon, water, and energy fluxes plus micrometeorological and ancillary data measured in over 200 sites around the world. To create an uniform dataset, including derived data products, extensive work on data quality control was needed. The unpredictable nature of observational data quality issues makes the automation of data quality control inherently difficult. Developed based on this experience, the data rogues methodology allows for increased automation of quality control activities by systematically identifying, cataloging, and documenting implementations of solutions to data issues. We believe this methodology can be extended and applied to others domains and types of data, making the automation of data quality control a more tractable problem.
Gilberto Zonta Pastorello, Dan Gunter, Housen Chu, Danielle Christianson, Carlo Trotta, Eleonora Canfora, Boris Faybishenko, You-Wei Cheah, Norm Beekwilder, Stephen Chan, Sigrid Dengel, Trevor Keenan, Fianna O'Brien, Abdelrahman Elbashandy, Cristina Poindexter, Marty Humphrey, Dario Papale, Deborah A. Agarwal
eScience18
2017 How Much Energy Can Green HPC Cloud Users Save?
abstract
Cloud computing has become an attractive and easy-to-use solution for users who want to externalize the run of their applications. However, data centers hosting cloud systems consume enormous amounts of energy. Reducing this consumption becomes an urgent challenge with the rapid growth of cloud utilization. In this paper, we explore a way for energy-aware HPC cloud users to reduce their footprint on cloud infrastructures by reducing the size of the virtual resources they are asking for. We study the influence of green users on the system energy consumption and compare it with the consumption of more aggressive users in terms of resource utilization. We found that larger resources are more energy demanding even if they are faster in executing the applications. But, reducing too much the resources' size is also not beneficial for the energy consumption. A tradeoff lies in between these two options.
David Guyon, Anne-Cécile Orgerie, Christine Morin, Deborah A. Agarwal
PDP4
2016 Data management and simulation support accelerating carbon capture through computing
abstract
The Carbon Capture Simulation Initiative (CCSI) project has developed and deployed scientific infrastructure called the CCSI Toolset. The CCSI Toolset provides state-of-the-art computational modeling and simulation tools to accelerate the commercialization of carbon capture technologies from discovery to development, demonstration, and ultimately the widespread deployment to hundreds of power plants. Carbon capture technologies have the potential to dramatically reduce the carbon emissions from power plants. The CCSI Toolset provides end users in industry with a comprehensive, integrated suite of leading-edge, scientifically validated models with simulation, uncertainty quantification, optimization, risk analysis and decision making support. The CCSI Toolset has at its core an integrated framework that enables execution of simulations and workflows including optimization and uncertainty parameter sweeps using a wide variety of computing platforms including desktops, clusters, Clouds, and HPC systems. The integration framework enables the running of a variety of commercial process simulation packages as well as custom simulators. Moreover, the framework enables scientists to run and manage thousands of concurrent simulations to perform optimizations and uncertainty quantification. Components of the CCSI Toolset are connected through the use of a data management system that stores data to a repository and enables the tracking of provenance for each simulation as well as its associated components. The data management system tracks all the configurations, models, simulations, and results created during the design of a carbon capture system and supports the design life-cycle as well as decision making. The primary contribution of this paper is thus the design and implementation of the integration framework within the CCSI Toolset, which provides both data management and simulation support for CCSI. This integration framework has been deployed and is in use by several groups of researchers and commercial entities.
You-Wei Cheah, Joshua Boverhof, Abdelrahman Elbashandy, Deborah A. Agarwal, James Leek, Thomas Epperly, John C. Eslick, David C. Miller
eScience4
2016 A science data gateway for environmental management
abstract
Summary Science data gateways are effective in providing complex science data collections to the world‐wide user communities. In this paper we describe a gateway for the Advanced Simulation Capability for Environmental Management (ASCEM) framework. Built on top of established web service technologies, the ASCEM data gateway is specifically designed for environmental modeling applications. Its key distinguishing features include (1) handling of complex spatiotemporal data, (2) offering a variety of selective data access mechanisms, (3) providing state‐of‐the‐art plotting and visualization of spatiotemporal data records, and (4) integrating seamlessly with a distributed workflow system using a RESTful interface. ASCEM project scientists have been using this data gateway since 2011. Copyright © 2015 John Wiley & Sons, Ltd.
Deborah A. Agarwal, Boris Faybishenko, Vicky L. Freedman, Harinarayan Krishnan, Gary Kushner, Carina Lansing, Ellen Porter, Alexandru Romosan, Arie Shoshani, Haruko M. Wainwright, Arthur Weidmer, Kesheng Wu
Concurr. Comput. Pract. Exp.1
2014 Observational Data Patterns for Time Series Data Quality Assessment
abstract
Observational data are fundamental for scientific research in almost any domain. Recent advances in sensor and data management technologies are enabling unprecedented amounts of observational data to be collected and analyzed. However, an essential part of using observational data is not currently as scalable as data collection and analysis methods: data quality assurance and control. While specialized tools for very narrow domains do exist, general methods are harder to create. This paper explores the identification of data issues that lead to the creation of data tests and tools to perform data quality control activities. Developing this identification step in a systematic manner allows for better and more general quality control tools. As our case study, we use carbon, water, and energy fluxes as well as micro-meteorological data collected at field sites that are part of FLUXNET, a network of over 400 ecosystem-level monitoring stations. In an effort toward the release of a new global data set of fluxes, we are doing data quality control for these data. The experience from this work led to the creation of a catalog of issues identified in the data. This paper presents this catalog and its generalization into a set of patterns of data quality issues that can be detected in observational data.
Gilberto Zonta Pastorello, Deborah A. Agarwal, Dario Papale, Taghrid Samak, Carlo Trotta, Alessio Ribeca, Cristina Poindexter, Boris Faybishenko, Dan Gunter, Rachel Hollowgrass, Eleonora Canfora
eScience2
2014 Experiences with User-Centered Design for the Tigres Workflow API
abstract
Scientific data volumes have been growing exponentially. This has resulted in the need for new tools that enable users to operate on and analyze data. Cyber infrastructure tools, including workflow tools, that have been developed in the last few years has often fallen short if user needs and suffered from lack of wider adoption. User-centered Design (UCD) process has been used as an effective approach to develop usable software with high adoption rates. However, UCD has largely been applied for user-interfaces and there has been limited work in applying UCD to application program interfaces and cyber infrastructure tools. We use an adapted version of UCD that we refer to as Scientist-Centered Design (SCD) to engage with users in the design and development of Tigres, a workflow application programming interface. Tigres provides a simple set of programming templates (e.g., sequence, parallel, split, merge) that can be can used to compose and execute computational and data transformation pipelines. In this paper, we describe Tigres and discuss our experiences with the use of UCD for the initial development of Tigres. Our experience-to-date is that the UCD process not only resulted in better requirements gathering but also heavily influenced the architecture design and implementation details. User engagement during the development of tools such as Tigres is critical to ensure usability and increase adoption.
Lavanya Ramakrishnan, Sarah S. Poon, Valerie C. Hendrix, Dan Gunter, Gilberto Zonta Pastorello, Deborah A. Agarwal
eScience6
2014 CAMP: Community Access MODIS Pipeline
Valerie C. Hendrix, Lavanya Ramakrishnan, Youngryel Ryu, Catharine van Ingen, Keith R. Jackson, Deborah A. Agarwal
Future Gener. Comput. Syst.6
2010 On-demand Overlay Networks for Large Scientific Data Transfers
abstract
Large scale scientific data transfers are central to scientific processes. Data from large experimental facilities have to be moved to local institutions for analysis or often data needs to be moved between local clusters and large supercomputing centers. In this paper, we propose and evaluate a network overlay architecture to enable high-throughput, on-demand, coordinated data transfers over wide-area networks. Our work leverages Phoebus and On-demand Secure Circuits and Advance Reservation System (OSCARS) to provide high performance wide-area network connections. OSCARS enables dynamic provisioning of network paths with guaranteed bandwidth and Phoebus enables the coordination and effective utilization of the OSCARS network paths. Our evaluation shows that this approach leads to improved end-to-end data transfer throughput with minimal overheads. The achieved throughput using our overlay was limited only by the ability of the end hosts to sink the data.
Lavanya Ramakrishnan, Chin Guok, Keith R. Jackson, Ezra Kissel, D. Martin Swany, Deborah A. Agarwal
CCGRID6
2010 Fault Tolerance and Scaling in e-Science Cloud Applications: Observations from the Continuing Development of MODISAzure
abstract
It can be natural to believe that many of the traditional issues of scale have been eliminated or at least greatly reduced via cloud computing. That is, if one can create a seemingly well functioning cloud application that operates correctly on small or moderate-sized problems, then the very nature of cloud programming abstractions means that the same application will run as well on potentially significantly larger problems. In this paper, we present our experiences taking MODISAzure, our satellite data processing system built on the Windows Azure cloud computing platform, from the proof-of-concept stage to a point of being able to run on significantly larger problem sizes (e.g., from national-scale data sizes to global-scale data sizes). To our knowledge, this is the longest-running eScience application on the nascent Windows Azure platform. We found that while many infrastructure-level issues were thankfully masked from us by the cloud infrastructure, it was valuable to design additional redundancy and fault-tolerance capabilities such as transparent idempotent task retry and logging to support debugging of user code encountering unanticipated data issues. Further, we found that using a commercial cloud means anticipating inconsistent performance and black-box behavior of virtualized compute instances, as well as leveraging changing platform capabilities over time. We believe that the experiences presented in this paper can help future eScience cloud application developers on Windows Azure and other commercial cloud providers.
Marty Humphrey, You-Wei Cheah, Youngryel Ryu, Deborah A. Agarwal, Keith R. Jackson, Catharine van Ingen
eScience5
2010 eScience in the cloud: A MODIS satellite data reprojection and reduction pipeline in the Windows Azure platform
abstract
The combination of low-cost sensors, low-cost commodity computing, and the Internet is enabling a new era of data-intensive science. The dramatic increase in this data availability has created a new challenge for scientists: how to process the data. Scientists today are envisioning scientific computations on large scale data but are having difficulty designing software architectures to accommodate the large volume of the often heterogeneous and inconsistent data. In this paper, we introduce a particular instance of this challenge, and present our design and implementation of a MODIS satellite data reprojection and reduction pipeline in the Windows Azure cloud computing platform. This cloud-based pipeline is designed with a goal of hiding data complexities and subsequent data processing and transformation from end users. This pipeline is highly flexible and extensible to accommodate different science data processing tasks, and can be dynamically scaled to fulfill scientists' various computational requirements in a cost-efficient way. Experiments show that by running a practical large-scale science data processing job in the pipeline using 150 moderately-sized Azure virtual machine instances, we were able to produce analytical results in nearly 90× less time than was possible with a high-end desktop machine. To our knowledge, this is one of the first eScience applications to use the Windows Azure platform.
Marty Humphrey, Deborah A. Agarwal, Keith R. Jackson, Catharine van Ingen, Youngryel Ryu
IPDPS3
2010 Relationships and data sanitization: a study in scarlet
abstract
Research in data sanitization (including anonymization) emphasizes ways to prevent an adversary from desanitizing data. Most work focuses on using mathematical mappings to sanitize data. A few papers examine incorporation of privacy requirements, either in the guise of templates or prioritization. Essentially these approaches reduce the information that can be gleaned from a data set. In contrast, this paper considers both the need to ``desanitize'' and the need to support privacy. We consider conflicts between privacy requirements and the needs of analysts examining the redacted data. Our goal is to enable an informed decision about the effects of redacting, and failing to redact data. We begin with relationships among the data being examined, including relationships with a known data set and other, additional, external data. By capturing these relationships, desanitization techniques that exploit them can be identified, and the information that must be concealed in order to thwart them can be determined. Knowing that, a realistic assessment of whether the information and relationships are already widely known or available will enable the sanitizers to assess whether irreversible sanitization is possible, and if so, what to conceal to prevent desanitization.
Matt Bishop, Justin Cummins, Sean Peisert, Anhad Singh, Bhume Bhumiratana, Deborah A. Agarwal, Deborah A. Frincke, Michael A. Hogarth
NSPW6
2010 A data-centered collaboration portal to support global carbon-flux analysis
abstract
Abstract Carbon‐climate, like other environmental sciences, has been changing. Large‐scale synthesis studies are becoming more common. These synthesis studies are often conducted by science teams that are geographically distributed and on data sets that are global in scale. A broad array of collaboration and data analytics tools are now available that could support these science teams. However, building tools that scientists actually use is difficult. Also, moving scientists from an informal collaboration structure to one mediated by technology often exposes inconsistencies in the understanding of the rules of engagement between collaborators. We have developed a scientific collaboration portal, calledfluxdata.org, which serves the community of scientists providing and analyzing the global FLUXNET carbon‐flux synthesis data set. The key things we learned or re‐learned during our portal development include: minimize the barrier to entry, provide features on a just‐in‐time basis, development of requirements is an on‐going process, provide incentives to change leaders and leverage the opportunity they represent, automate as much as possible, and you can only learn how to make it better if people depend on it enough to give you feedback. In addition, we also learned that splitting the portal roles between scientists and computer scientists improved user adoption and trust. The fluxdata.org portal has now been in operation for ∼2 years and has become central to the FLUXNET synthesis efforts. Published in 2010 by John Wiley & Sons, Ltd.
Deborah A. Agarwal, Marty Humphrey, Norm Beekwilder, Keith R. Jackson, Monte M. Goode, Catharine van Ingen
Concurr. Comput. Pract. Exp.1
2009 Fluxdata.org: Publication and Curation of Shared Scientific Climate and Earth Sciences Data
abstract
Many of today's large-scale scientific projects attempt to collect data from a diverse set of sources. The traditional campaign-style approach to ¿synthesis¿ efforts gathers data through a single concentrated effort, and the data contributors know in advance exactly who will use their data and why. At even moderate scales, the cost and time required to find, gather, collate, normalize, and customize data in order to build a synthesis dataset can quickly outweigh the value of the resulting dataset. By explicitly identifying and addressing the different requirements for each data role (author, publisher, curator, and consumer), our data management architecture for large-scale shared scientific data enables the creation of such synthesis datasets that continue to grow and evolve with new data, data annotations, participants, and use rules. We show the effectiveness of our approach in the context of the FLUXNET Synthesis Dataset, one of the largest ongoing biogeophysical experiments.
Marty Humphrey, Deborah A. Agarwal, Catharine van Ingen
eScience2
2002 A practical approach to the InterGroup protocols
K. Berket, Deborah A. Agarwal, Olivier Chevassut
Future Gener. Comput. Syst.2
2001 Network Characterization Service (NCS)
abstract
Distributed applications require information to effectively utilize the network. Some of the information they require is the current and maximum bandwidth, current and minimum latency, bottlenecks, burst frequency and congestion extent. This type of information allows applications to determine parameters like the optimal TCP buffer size. In this paper, we present a cooperative information-gathering tool called the Network Characterization Service (NCS). NCS runs in the user space and is used to acquire network information. Its protocol is designed for scalable and distributed deployment, similar to DNS. Its algorithms provide efficient, speedy and accurate detection of bottlenecks, especially dynamic bottlenecks. On current and future networks, dynamic bottlenecks do and will affect network performance dramatically.
Guojun Jin, George Yang, Brian R. Crowley, Deborah A. Agarwal
HPDC4
2001 An Integrated Solution for Secure Group Communication in Wide-Area Networks
abstract
Many distributed applications require a secure reliable group communication system to provide coordination among the application components. This paper describes a secure group layer (SGL) which bundles a reliable group communication system, a group authorization and access control mechanism, and a group key agreement protocol to provide a comprehensive and practical secure group communication platform. The SGL also encapsulates the standard message security services (i.e., confidentiality, authenticity and integrity). A number of challenging issues encountered in the design of SGL are brought to light and experimental results obtained with a prototype implementation are discussed.
Deborah A. Agarwal, Olivier Chevassut, Mary R. Thompson, Gene Tsudik
ISCC1
1998 The Totem Multiple-Ring Ordering and Topology Maintenance Protocol
abstract
The Totem multiple-ring protocol provides reliable totally ordered delivery of messages across multiple local-area networks interconnected by gateways. This consistent message order is maintained in the presence of network partitioning and remerging, and of processor failure and recovery. The protocol provides accurate topology change information as part of the global total order of messages. It addresses the issue of scalability and achieves a latency that increases logarithmically with system size by exploiting process group locality and selective forwarding of messages through the gateways. Pseudocode for the protocol and an evaluation of its performance are given. —Authors' Abstract
Deborah A. Agarwal, Louise E. Moser, P. M. Melliar-Smith, Ravi K. Budhia
ACM Trans. Comput. Syst.1
1995 A reliable ordered delivery protocol for interconnected local area networks
abstract
We present-the Totem multiple-ring protocol, a novel reliable ordered multicast protocol for multiple interconnected local-area networks. The protocol exhibits excellent performance and maintains a consistent network-wide total order of messages despite network partitioning and remerging, or processor failure and recovery with stable storage intact. The Totem protocol is designed for fault-tolerant distributed systems, which replicate data to guard against failures and must ensure that replicated data remain consistent despite failures. The network-wide total order of messages provided by Totem simplifies the maintenance of consistency of replicated data, and, thus, eases the development of fault-tolerant distributed systems.
Deborah A. Agarwal, Louise E. Moser, P. M. Melliar-Smith, Ravi K. Budhia
ICNP1
1995 The Totem Single-Ring Ordering and Membership Protocol
abstract
Fault-tolerant distributed systems are becoming more important, but in existing systems, maintaining the consistency of replicated data is quite expensive. The Totem single-ring protocol supports consistent concurrent operations by placing a total order on broadcast messages. This total order is derived from the sequence number in a token that circulates around a logical ring imposed on a set of processors in a broadcast domain. The protocol handles reconfiguration of the system when processors fail and restart or when the network partitions and remerges. Extended virtual synchrony ensures that processors deliver messages and configuration changes to the application in a consistent, systemwide total order. An effective flow control mechanism enables the Totem single-ring protocol to achieve message-ordering rates significantly higher than the best prior total-ordering protocols.
Yair Amir, Louise E. Moser, P. M. Melliar-Smith, Deborah A. Agarwal, P. Ciarfella
ACM Trans. Comput. Syst.4
1994 Extended Virtual Synchrony
abstract
We formulate a model of extended virtual synchrony that defines a group communication transport service for multicast and broadcast communication in a distributed system. The model extends the virtual synchrony model of the Isis system to support continued operation in all components of a partitioned network. The significance of extended virtual synchrony is that, during network partitioning and remerging and during process failure and recovery, it maintains a consistent relationship between the delivery of messages and the delivery of configuration changes across all processes in the system and provides well-defined self-delivery and failure atomicity properties. We describe an algorithm that implements extended virtual synchrony and construct a filter that reduces extended virtual synchrony to virtual synchrony.>
Louise E. Moser, Yair Amir, P. M. Melliar-Smith, Deborah A. Agarwal
ICDCS4
1994 The Totem protocol development environment
abstract
Introduction The development of communication protocols is significantly more difficult than the development of sequential programs. Communication protocols are harder to develop because of the many possible executions and ordering of events, the difficulty of recording the history of an execution, the inability to stop the system so that its state can be examined, the difficulty of injecting a stimulus or fault at exactly the right moment, and the lack of reproducibility. Fault-tolerance presents additional problems, particularly the need to recover correctly from rare failure modes and the appearance of correct operation even though the system contains serious defects. The protocol development environment presented here was designed to aid in the development of the Totem reliable ordered broadcast protocol, but its concepts are also applicable to other communication protocols. The development environment is a distributed discrete-event simulator designed to provide a testbed
P. Ciarfella, Louise E. Moser, P. M. Melliar-Smith, Deborah A. Agarwal
ICNP4
1993 Fast Message Ordering and Membership Using a Logical Token-Passing Ring
abstract
The Totem protocol supports consistent concurrent operations by placing a total order on broadcast messages. This total order is achieved by including a sequence number in a token circulated around a logical ring that is imposed on a set of processors in a broadcast domain. A membership algorithm handles reconfiguration, including restarting of a failed processor and remerging of a partitioned network. Effective flow-control allows the protocol to achieve message ordering rates two to three times higher than the best prior protocols. The single-ring total ordering protocol of Totem provides fault-tolerant agreed and safe delivery of messages within a broadcast domain.>
Yair Amir, Louise E. Moser, P. M. Melliar-Smith, Deborah A. Agarwal, P. Ciarfella
ICDCS4