VLDB 2026 Research / reviewers in the wild / expert
Flaviu Cristian
dblp:43/190
· DBLP profile ↗
42ranked-venue papers
18as first author
0since 2021 · last 2003
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 8 first-authorSoftware engineering, systems software and programming languages · 8 · 3 first-authorSecurity and privacy · 7 · 3 first-authorDatabases, data management, data science and information retrieval · 3Theory of computation · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
14 papers |
Distributed systems · 68% Storage systems · 28% Hardware reliability and fault tolerance · 4% | |
| Software engineering, system software, and programming languages
6 papers |
Program verification · 55% Programming languages and type systems · 21% Software testing · 15% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
fault tolerance |
0.1 | 11 | 2002 | The Timewheel Group Communication System · IEEE Trans. Computers 2002 A Highly Available Local Leader Election Service · IEEE Trans. Software Eng. 1999 Fault-Tolerance in Air Traffic Control Systems · ACM Trans. Comput. Syst. 1996 |
Storage systems
storage reliability |
0.0 | 2 | 1998 | Declustered Disk Array Architectures with Optimal and Near-Optimal Parallelism · ISCA 1998 Tolerating Multiple Failures in RAID Architectures with Optimal Storage and Uniform Declustering · ISCA 1997 |
Distributed systems
group communication |
0.0 | 1 | 2002 | The Timewheel Group Communication System · IEEE Trans. Computers 2002 |
Distributed systems › group communication
reliable group communication |
0.0 | 1 | 2002 | The Timewheel Group Communication System · IEEE Trans. Computers 2002 |
Distributed systems › distributed coordination
leader election |
0.0 | 2 | 1999 | A Highly Available Local Leader Election Service · IEEE Trans. Software Eng. 1999 The Timed Asynchronous Distributed System Model · IEEE Trans. Parallel Distributed Syst. 1999 |
Distributed systems › group communication
atomic broadcast |
0.0 | 3 | 1995 | Atomic Broadcast: From Simple Message Diffusion to Byzantine Agreement · Inf. Comput. 1995 New Latency Bounds for Atomic Broadcast · RTSS 1990 Early-Delivery Atomic Broadcast · PODC 1990 |
Storage systems
declustering |
0.0 | 2 | 1998 | Declustered Disk Array Architectures with Optimal and Near-Optimal Parallelism · ISCA 1998 Tolerating Multiple Failures in RAID Architectures with Optimal Storage and Uniform Declustering · ISCA 1997 |
Storage systems
disk array |
0.0 | 2 | 1998 | Declustered Disk Array Architectures with Optimal and Near-Optimal Parallelism · ISCA 1998 Tolerating Multiple Failures in RAID Architectures with Optimal Storage and Uniform Declustering · ISCA 1997 |
Distributed systems
distributed system modeling |
0.0 | 1 | 1999 | The Timed Asynchronous Distributed System Model · IEEE Trans. Parallel Distributed Syst. 1999 |
Distributed systems
consensus |
0.0 | 2 | 1999 | Atomic Broadcast: From Simple Message Diffusion to Byzantine Agreement · Inf. Comput. 1995 The Timed Asynchronous Distributed System Model · IEEE Trans. Parallel Distributed Syst. 1999 |
Hardware reliability and fault tolerance › network fault tolerance
multinode failure tolerance |
0.0 | 1 | 1998 | Declustered Disk Array Architectures with Optimal and Near-Optimal Parallelism · ISCA 1998 |
Storage systems › disk array
parity placement |
0.0 | 1 | 1998 | Declustered Disk Array Architectures with Optimal and Near-Optimal Parallelism · ISCA 1998 |
Storage systems › storage reliability
erasure coding |
0.0 | 1 | 1997 | Tolerating Multiple Failures in RAID Architectures with Optimal Storage and Uniform Declustering · ISCA 1997 |
Storage systems › storage reliability
RAID |
0.0 | 1 | 1997 | Tolerating Multiple Failures in RAID Architectures with Optimal Storage and Uniform Declustering · ISCA 1997 |
Distributed systems › fault tolerance
high availability |
0.0 | 1 | 1996 | Fault-Tolerance in Air Traffic Control Systems · ACM Trans. Comput. Syst. 1996 |
Distributed systems › consensus
byzantine agreement |
0.0 | 1 | 1995 | Atomic Broadcast: From Simple Message Diffusion to Byzantine Agreement · Inf. Comput. 1995 |
Distributed computing theory › distributed synchronization
clock synchronization |
0.0 | 1 | 1995 | Lower Bounds for Convergence Function Based Clock Synchronization · PODC 1995 |
Distributed systems
clock synchronization |
0.0 | 1 | 1990 | Continuous Clock Amortization Need Not Affect the Precision of a Clock Synchronization Algorithm · PODC 1990 |
Distributed systems
replication |
0.0 | 1 | 1990 | Early-Delivery Atomic Broadcast · PODC 1990 |
Storage systems › repair
data reconstruction |
0.0 | 1 | 1998 | Declustered Disk Array Architectures with Optimal and Near-Optimal Parallelism · ISCA 1998 |
Programming languages and type systems › control structures
exception handling |
0.0 | 2 | 1984 | Correct and Robust Programs · IEEE Trans. Software Eng. 1984 Exception Handling and Software Fault Tolerance · IEEE Trans. Computers 1982 |
Embedded and real-time systems
cyber-physical system platforms |
0.0 | 1 | 1996 | Fault-Tolerance in Air Traffic Control Systems · ACM Trans. Comput. Syst. 1996 |
Distributed and cloud data management
data replication |
0.0 | 1 | 1985 | An Efficient, Fault-Tolerant Protocol for Replicated Data Management · PODS 1985 |
Transaction processing and concurrency control › serializability
one-copy serializability |
0.0 | 1 | 1985 | An Efficient, Fault-Tolerant Protocol for Replicated Data Management · PODS 1985 |
Distributed systems › replication
replicated data management |
0.0 | 1 | 1985 | An Efficient, Fault-Tolerant Protocol for Replicated Data Management · PODS 1985 |
Storage systems › storage reliability › fault-tolerant storage
stable storage |
0.0 | 1 | 1985 | A Rigorous Approach to Fault-Tolerant Programming · IEEE Trans. Software Eng. 1985 |
Program verification
deductive verification |
0.0 | 1 | 1984 | Correct and Robust Programs · IEEE Trans. Software Eng. 1984 |
Program verification › correctness proof
total correctness |
0.0 | 1 | 1984 | Correct and Robust Programs · IEEE Trans. Software Eng. 1984 |
Operating systems
fault tolerance |
0.0 | 2 | 1987 | Masking System Crashes in Database Application Programs · VLDB 1987 Correct and Robust Programs · IEEE Trans. Software Eng. 1984 |
Software testing
software fault tolerance |
0.0 | 1 | 1982 | Exception Handling and Software Fault Tolerance · IEEE Trans. Computers 1982 |
Methods — techniques the papers use, named apart from their topics
timed asynchronous model · 0.0protocol design · 0.0performance measurement · 0.0message delay measurement · 0.0clock drift measurement · 0.0simulation · 0.0information dispersal algorithm · 0.0fault-tolerance design · 0.0axiomatic reasoning · 0.0amortization · 0.0formal validation · 0.0fault modeling · 0.0deductive system · 0.0backward error recovery · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2003 | Fail-Awareness: An Approach to Construct Fail-Safe Systems
Christof Fetzer, Flaviu Cristian |
Real Time Syst. | 2 |
| 2002 | The Timewheel Group Communication SystemabstractDescribes the timewheel group communication system, which has been designed for a timed asynchronous distributed system model. All protocols in the timewheel group communication system have been designed to be fail-aware in the sense that a process can detect, at any point in time, whether any of its properties is violated. Although these protocols have been designed to operate in an asynchronous distributed computing environment, they provide timeliness properties. The timewheel group communication system provides nine group communication semantics that a user can dynamically choose from while broadcasting an update. This system provides high throughput, fast delivery and stability times, uses a small number of messages per update broadcast, and evenly distributes the processing load among group members. Shivakant Mishra, Christof Fetzer, Flaviu Cristian |
IEEE Trans. Computers | 3 |
| 2000 | Simulation-based Testing of Communication Protocols for Dependable Embedded Systems
Guillermo A. Alvarez, Flaviu Cristian |
J. Supercomput. | 2 |
| 1999 | The Timed Asynchronous Distributed System ModelabstractWe propose a formal definition for the timed asynchronous distributed system model. We present extensive measurements of actual message and process scheduling delays and hardware clock drifts. These measurements confirm that this model adequately describes current distributed systems such as a network of workstations. We also give an explanation of why practically needed services, such as consensus or leader election, which are not implementable in the time-free model, are implementable in the timed asynchronous system model. Flaviu Cristian, Christof Fetzer |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1999 | A Highly Available Local Leader Election ServiceabstractWe define the highly available local leader election problem (G. LeLann, 1977), a generalization of the leader election problem for partitionable systems. We propose a protocol that solves the problem efficiently and give some performance measurements of our implementation. The local leader election service has been proven useful in the design and implementation of several fail-aware services for partitionable systems. Christof Fetzer, Flaviu Cristian |
IEEE Trans. Software Eng. | 2 |
| 1998 | Evaluating the Performance of Group Membership ProtocolsabstractGroup membership protocols are designed to achieve agreement among a set of replicated servers on a global service state, despite random message transmission delays, component failures, and server joins. This paper proposes a set of performance metrics for group membership protocols and investigates the performance characteristics of a suite of five asynchronous group membership protocols with regard to those metrics. After a brief description of the protocol suite, we propose a systematic procedure for determining the timeout delays used by the protocols, so as to achieve the best compromise between protocol stability and the speed with which failures are detected and joins are processed. The paper then discusses the measured performance of the membership protocols in an environment consisting of UNIX workstations interconnected by a 10 Megabit/sec Ethernet local area network. Oliver Suciu, Flaviu Cristian |
ICECCS | 2 |
| 1998 | Declustered Disk Array Architectures with Optimal and Near-Optimal ParallelismabstractThis paper investigates the placement of data and parity on redundant disk arrays. Declustered organizations have been traditionally used to achieve fast reconstruction of a failed disk's contents. In previous work, Holland and Gibson identified six desirable properties for ideal layouts; however no declustered layout satisfying all properties has been published in the literature. We present a complete, constructive characterization of the collection of ideal declustered layouts possessing all six properties. Given that ideal layouts exist only for a limited set of configurations, we also present two novel layout families. PRIME and RELPR can tolerate multiple failures in a wide variety of configurations with slight deviations from the ideal. Our simulation studies show that the new layouts provide excellent parallel access performance and reduced incremental loads during degraded operation, when compared with previously published layouts. For large accesses and under high loads, response times for the new layouts are typically smaller than those of previously published declustered layouts by a factor of 2.5. Guillermo A. Alvarez, Walter A. Burkhard, Larry J. Stockmeyer, Flaviu Cristian |
ISCA | 4 |
| 1997 | Centralized Failure Injection for Distributed, Fault-Tolerant Protocol TestingabstractWe describe a centralized approach to testing that distributed fault-tolerant protocols satisfy their safety and timeliness specifications in the presence of the very failures they are designed to tolerate. CESIUM is a testing environment based on the centralized simulation of distributed executions and failures. Processes are run in a single address space while providing the appearance of a truly distributed execution. The human tester can force the occurrence of arbitrary failures and security attacks. The implementations under test are not instrumented for testing purposes, and their source codes need not be available. We prove that CESIUM can execute exactly the set of runs feasible in the real distributed system being simulated. We also show that there are safety and timeliness properties in the specifications of many existing distributed protocols that cannot be tested in practical distributed systems. All of these properties can, however, be accurately tested by CESIUM without introducing any perturbation in test experiments. Guillermo A. Alvarez, Flaviu Cristian |
ICDCS | 2 |
| 1997 | Tolerating Multiple Failures in RAID Architectures with Optimal Storage and Uniform DeclusteringabstractWe present DATUM, a novel method for tolerating multiple disk failures in disk arrays. DATUM is the first known method that can mask any given number of failures, requires an optimal amount of redundant storage space, and spreads reconstruction accesses uniformly over disks in the presence of failures without needing large layout tables in controller memory. Our approach is based on information dispersal, a coding technique that admits an efficient hardware implementation. As the method does not restrict the configuration parameters of the disk array, many existing RAID organizations are particular cases of DATUM. A detailed performance comparison with two other approaches shows that DATUM'S response times are similar to those of the best competitor when two or less disks fail, and that the performance degrades gracefully when more than two disks fail. Guillermo A. Alvarez, Walter A. Burkhard, Flaviu Cristian |
ISCA | 3 |
| 1997 | Applying Simulation to the Design and Performance Evaluation of Fault-tolerant SystemsabstractThe paper illustrates how the CESIUM simulation tool can be used for design and performance evaluation of fault tolerant and real time systems, in addition to testing the correctness of protocol implementations. We calibrate three increasingly accurate simulation models of a network of workstations using independently obtained data. For a sample group membership protocol, the predictions of the simulator are very close to the actual performance measured in the real system. We also apply CESIUM to the evaluation of two potential improvements for the protocol, performing experiments that would have been difficult to implement in the real system. The results of the simulations give us valuable insight on how to tune configuration parameters, as well as on the performance gains of the improved versions. Our experience shows that CESIUM can be used to develop best effort services which adapt their quality of service according to the failures that occur during operation. Guillermo A. Alvarez, Flaviu Cristian |
SRDS | 2 |
| 1997 | A Fail-Awar Membership ServiceabstractWe propose a new protocol that can be used to implement a partitionable membership service for timed asynchronous systems. The protocol is fail-aware in the sense that a process p knows at all times if its approximation of the set of processes in its partition is up-to-date or out-of-date. The protocol minimizes wrong suspicions of processes by giving processes a second chance to stay in the membership before they are removed. Our measurements show that the exclusion of live processes is rare and the crash detection times are good. The protocol guarantees that the memberships of two partitions never overlap. Christof Fetzer, Flaviu Cristian |
SRDS | 2 |
| 1997 | Integrating External and Internal Clock Synchronization
Christof Fetzer, Flaviu Cristian |
Real Time Syst. | 2 |
| 1996 | Fail-Awareness in Timed Asynchronous SystemsabstractWe address the problem of the impossibdity of implementing synchronous fault-tolerant service specifications in asynchronous distributed systems.We introduce a method for weakening a synchronous service specification so that it becomes implementable in "timed" asynchronous systems, that q This research was partially sponsored by a grant from the Air Force Office of Scientific Research Fe fmkioo to meted@d/bcrd copies of cll or pert of W:s rnetericl for perm-mel or clcssroarn use is grcnted withcwt &c provided Urct the copies not IIU& or dktdwtrd fw protit or commerc"ml q dvmtege, the.c~yrigbt rdce, the title of the publkxtion q nd ite dcte q ppecr, q nd aottce u given tbct copyright is by pcrmisrkM of tbe ACM, inc.To copy othcnviee, to republicb, to poA 00 aetvers or to rdetribute to Iietcj requires epccitic pcrmhioo cndhr f-.PODC'%, Pttiladelphis PA, USA O l% ACM &SgT$)l.~~%/OS..$3.50 asynchronous distributed systems in which processes have access to local hardware clocks.Hardware clocks and the notion of "performance failures" are essential for our approach.This work is therefore based on the timed asynchronous system model[11] Christof Fetzer, Flaviu Cristian |
PODC | 2 |
| 1996 | Implementation and Performance of a Stable-Storage Service in UnixabstractThis paper describes the design, implementation, and performance of a stable-storage service that has been implemented on top of the Unix operating system. This service allows servers to create, access, and delete persistent memory that survives server crashes. We describe its functionality and exported operations, discuss the experiences and performance of its implementation, and offer concrete examples of its use in implementing some real fault-tolerant distributed protocols. Flaviu Cristian, Shivakant Mishra, Young S. Hyun |
SRDS | 1 |
| 1996 | Fail-Aware Failure DetectorsabstractIn existing asynchronous distributed systems it is impossible to implement failure detectors which are perfect, i.e. they only suspect crashed processes and eventually suspect all crashed processes. Some recent research has however proposed that any "reasonable" failure detector for solving the election problem must be perfect. We address this problem by introducing two new classes of fail-aware failure detectors that are (1) implementable in existing asynchronous distributed systems, (2) not necessarily perfect, and (3) can be used to solve the election problem. In particular we show that there exists a fail-aware failure detector that allows to solve the election problem and which is strictly weaker than a perfect failure detector. Christof Fetzer, Flaviu Cristian |
SRDS | 2 |
| 1996 | Fault-Tolerance in Air Traffic Control SystemsabstractThe distributed real-time system services developed by Lockheed Martin's Air Traffic Management group serve the infrastructure for a number of air traffic control systems. Either completed development or under development are the US Federal Aviation Administration's Display System Replacement (DSR) system, the UK Civil Aviation Authority's New Enroute Center (NERC) system, and the Republic of China's Air Traffic Control Automated System (ATCAS). These systems are intended to replace present en route systems over the next decade. High availability of air traffic control services is an essential requirement of these systems. This article discusses the general approach to fault-tolerance adopted in this infrastructure, by reviewing some of the questions which were asked during the system design, various alternative solutions considered, and the reasons for the design choices made. The aspects of this infrastructure chosen for the individual ATC systems mentioned above, along with the status of those systems, are presented in the Section 11 of the article. Flaviu Cristian, Bob Dancey, Jon Dehn 0001 |
ACM Trans. Comput. Syst. | 1 |
| 1995 | Fault-Tolerant External Clock SynchronizationabstractWe address the problem of how to integrate fault-tolerant internal and external clock synchronization. We propose a new algorithm which provides both external and internal clock synchronization for as long as no more than F reference time servers out of a total of 2F+1 are faulty. When the number of faulty reference time servers exceeds F, the algorithm degrades to a fault-tolerant internal clock synchronization algorithm. We prove that at least 2F+1 reference time servers are necessary for achieving external clock synchronization when up to F reference time servers can suffer arbitrary failures, thus our algorithm provides maximum fault-tolerance. The algorithm is also optimal in another sense: we show that the maximum deviation between reference time and the clocks of nonreference time servers is minimal. Flaviu Cristian, Christof Fetzer |
ICDCS | 1 |
| 1995 | The pinwheel asynchronous atomic broadcast protocolsabstractWe discuss two asynchronous atomic broadcast protocols that provide fast delivery and stability times, use a small number of messages to accomplish a broadcast, distribute evenly the processing load, use efficient flow control techniques, and provide gracefully degraded performance in the presence of communication failures. In a prototype implementation on top of UDP and Ethernet, for a group of three broadcast servers, these protocols achieve a throughput of up to a thousand independent broadcasts per second and measure average delivery and stability times of 2.9 and 4.7 msec.> Flaviu Cristian, Shivakant Mishra |
ISADS | 1 |
| 1995 | Lower Bounds for Convergence Function Based Clock SynchronizationabstractArticle Free Access Share on Lower bounds for convergence function based clock synchronization Authors: Christof Fetzer Department of Computer Science & Engineering, University of California, San Diego, La Jolla, CA Department of Computer Science & Engineering, University of California, San Diego, La Jolla, CAView Profile , Flaviu Cristian Department of Computer Science & Engineering, University of California, San Diego, La Jolla, CA Department of Computer Science & Engineering, University of California, San Diego, La Jolla, CAView Profile Authors Info & Claims PODC '95: Proceedings of the fourteenth annual ACM symposium on Principles of distributed computingAugust 1995 Pages 137–143https://doi.org/10.1145/224964.224980Online:20 August 1995Publication History 10citation288DownloadsMetricsTotal Citations10Total Downloads288Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Christof Fetzer, Flaviu Cristian |
PODC | 2 |
| 1995 | Atomic Broadcast: From Simple Message Diffusion to Byzantine Agreement
Flaviu Cristian, Houtan Aghili, Ray Strong, Danny Dolev |
Inf. Comput. | 1 |
| 1994 | Probabilistic Internal Clock SynchronizationabstractWe propose an improved probabilistic method for reading remote clocks in systems subject to unbounded communication delays and use this method to design a fault-tolerant probabilistic internal clock synchronization protocol. This protocol masks clock reading failures and arbitrary failures of processes. Because of probabilistic reading, our protocol achieves better synchronization precisions than those achievable by previously known deterministic algorithms. Another advantage of the proposed protocol is that it uses a linear, instead of quadratic, number of messages, and that message exchanges are staggered in time instead of all happening in narrow synchronization intervals. The drift rate of the synchronized clocks is optimal.> Flaviu Cristian, Christof Fetzer |
SRDS | 1 |
| 1993 | Automatic service availability managementabstractA new kind of distributed system service called availability management service is introduced. It is responsible for ensuring that the critical services of a distributed system remain continuously available to users despite arbitrary numbers of concurrent node removals and node restarts caused by failures, maintenance, and growth. The description of many details involved in a realistic design is sacrificed to make the underlying concepts easily understandable. To this end, the availability management service is designed on top of an easy-to-understand synchronous communication environment, and only one kind of service availability policy is considered. It is indicated how the initial specification and design can be extended to deal with asynchronous systems subject to partitioning as well as with other kinds of service availability policies.> Flaviu Cristian |
ISADS | 1 |
| 1993 | Coordinator Log Transaction Execution Protocol
James W. Stamos, Flaviu Cristian |
Distributed Parallel Databases | 2 |
| 1991 | A Timestamp-Based Checkpointing Protocol for Long-Lived Distributed ComputationsabstractThe authors present a timestamp-based protocol for checkpointing the global state of a long-lived distributed computation in an environment in which processor clocks are approximately synchronized. The protocol is based on periodic checkpointing of local process states and logging of incoming messages during a short bounded interval. It tolerates process crash and performance failures as well as network omission and performance failures. The proposed approach has the advantage of optimistic logging protocols in that it does not require synchronous logging of each message on stable storage. The approach also has the advantage of pessimistic logging protocols in that it avoids the domino effect by recovering to the most recent successful local checkpoint.> Flaviu Cristian, Farnam Jahanian |
SRDS | 1 |
| 1991 | Reaching Agreement on Processor-Group Memebership in Synchronous Distributed Systems
Flaviu Cristian |
Distributed Comput. | 1 |
| 1990 | Early-Delivery Atomic BroadcastabstractArticle Free Access Share on Early-delivery atomic broadcast Authors: Ajei Gopal Cornell University Cornell UniversityView Profile , Ray Strong IBM ARC IBM ARCView Profile , Sam Toueg Cornell University Cornell UniversityView Profile , Flaviu Cristian IBM ARC IBM ARCView Profile Authors Info & Claims PODC '90: Proceedings of the ninth annual ACM symposium on Principles of distributed computingAugust 1990 Pages 297–309https://doi.org/10.1145/93385.93430Published:01 August 1990Publication History 16citation324DownloadsMetricsTotal Citations16Total Downloads324Last 12 Months9Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Ajei S. Gopal, Ray Strong, Sam Toueg, Flaviu Cristian |
PODC | 4 |
| 1990 | Continuous Clock Amortization Need Not Affect the Precision of a Clock Synchronization AlgorithmabstractIntroduction Clock s~nc:hrc~rti~nt.ion is ncedcd in many distribt~I,ecl s~~s~,crns (.o ~rreasurc~ the dvration of c~istril~n~ccl a. Frank B. Schmuck, Flaviu Cristian |
PODC | 2 |
| 1990 | New Latency Bounds for Atomic BroadcastabstractTighter bounds are provided on the time required to reach agreement in a distributed system as a function of the failure model. After describing the model of a distributed system that is a context for this work the authors define several failure classes. They define a partial order on classes of failures that involves whether there is a latency penalty in converting from tolerance of one failure class to another. In this setting they distinguish clock and timing failures, showing that there can be a penalty in converting from timing failure tolerance to clock failure tolerance. The authors leave open the exact expression for the optimal latency for timing failure tolerant atomic broadcast, though it is conjectured that there is some penalty in converting from omission failure tolerance to timing failure tolerance.> Ray Strong, Danny Dolev, Flaviu Cristian |
RTSS | 3 |
| 1990 | A Low-Cost Atomic Commit ProtocolabstractThe proposed coordinator log transaction execution protocol centralizes logging on a per-transaction basis and exploits piggybacking to provide the semantics of a distributed atomic commit without the associated costs. This protocol eliminates two rounds of messages (one phase) from the presumed commit protocol and dramatically reduces the number of log forces needed for distributed atomic commit. The authors compare the coordinator log transaction execution protocol with existing protocols, describe when it is desirable, and discuss how it affects the write-ahead log protocol and the database crash recovery algorithm.> James W. Stamos, Flaviu Cristian |
SRDS | 2 |
| 1990 | Synchronous Atomic Broadcast for Redundant Broadcast Channels
Flaviu Cristian |
Real Time Syst. | 1 |
| 1989 | A probabilistic approach to distributed clock synchronizationabstractA probabilistic method is proposed for reading remote clocks in distributed systems subject to unbounded random communication delays. The method can achieve clock synchronization precisions superior to those attainable by previously published clock synchronization algorithms. The method can be used to improve the precision of both internal and external synchronization algorithms. The approach is probabilistic because it does not guarantee that a processor can always read a remote clock with an a priori specified precision; however, by retrying a sufficient number of times, a process can read the clock of another process with a given precision with a probability as close to one as desired. An important characteristic of the method is that, when a process succeeds in reading a remote clock, it knows the actual reading precision achieved. The use of the remote clock reading methods is illustrated by presenting a time service which maintains externally (and, hence, internally) synchronized clocks in the presence of process, communication, and clock failures.> Flaviu Cristian |
ICDCS | 1 |
| 1989 | Probabilistic Clock Synchronization
Flaviu Cristian |
Distributed Comput. | 1 |
| 1987 | Handshake Protocols
Ray Strong, Dale Skeen, Flaviu Cristian, Houtan Aghili |
ICDCS | 3 |
| 1987 | Masking System Crashes in Database Application Programs
Johann-Christoph Freytag, Flaviu Cristian, Bo Kähler |
VLDB | 2 |
| 1985 | An Efficient, Fault-Tolerant Protocol for Replicated Data ManagementabstractA data management protocol for executing transactions on a replicated database is presented. The protocol ensures one-copy serializability. i.e., the concurrent execution of transactions on a replicated database is equivalent to some serial execution of the same transactions on a non-replicated database. The protocol tolerates a large class of failures, including: processor and communication link crashes, partitioning of the communication network, lost messages, and slow responses of processors and communication links. Processor and link recoveries are also handled. The protocol implements the reading of a replicated object efficiently by reading the nearest available copy of the object. When reads outnumber writes, the protocol performs better than other known protocols. Amr El Abbadi, Dale Skeen, Flaviu Cristian |
PODS | 3 |
| 1985 | Comments on "Self-Stabilizing Programs: The Fault-Tolerant Capability of Self-Checking Programs"abstractIn the above correspondence,1 A. Mili aims "to introduce the theoretical basis for the design and validation of self-checking programs." A theoretically sound basis is indeed needed for designing and validating robust fault-tolerant programs, and we follow progress made in this area with great interest. To our disappointment, we found that the formalism presented in the above correspondence1has not been properly worked out. Eike Best, Flaviu Cristian |
IEEE Trans. Computers | 2 |
| 1985 | A Rigorous Approach to Fault-Tolerant ProgrammingabstractThe design of programs that are tolerant of hardware fault occurrences and processor crashes is investigated. Using a stable storage management system as a running example, a new approach is suggested for specifying, understanding, and verifying the correctness of fault-tolerant software. The approach extends previously developed axiomatic reasoning methods to the design of fault-tolerant systems by modeling faults as being operations that are performed at random time intervals on any computing system by the system's adverse environment. Flaviu Cristian |
IEEE Trans. Software Eng. | 1 |
| 1984 | Correct and Robust ProgramsabstractThe design of programs which are both correct and robust is investigated. It is argued that the notion of an exception is a valuable tool for structuring the specification, design, verification, and modification of such programs. The syntax and semantics of a language with procedures and exception handling are presented. A deductive system is proposed for proving total correctness and robustness properties of programs written in this language. The system is both sound and complete. It supports proof modularization, in that it allows one to reason separately about fault-free and fault-tolerant system properties. Since the programming languages considered closely resembles CLU or Ada, the presented deductive system is easily adaptable for verifying total correctness and robustness properties of programs written in these, or similar, languages. Flaviu Cristian |
IEEE Trans. Software Eng. | 1 |
| 1982 | Robust Data Types
Flaviu Cristian |
Acta Informatica | 1 |
| 1982 | Exception Handling and Software Fault ToleranceabstractSome basic concepts underlying the issue of fault-tolerant software design are investigated. Relying on these concepts, a unified point of view on programmed exception handling and default exception handling based on automatic backward recovery is constructed. The cause–effect relationship between software design faults and failure occurrences is explored and a class of faults for which default exception handling can provide effective fault tolerance is characterized. It is also shown that there exists a second class of design faults which cannot be tolerated by using default exception handling. The role that software verification methods can play in avoiding the production of such faults is discussed. Flaviu Cristian |
IEEE Trans. Computers | 1 |
| 1981 | Systematic Detection of Exception Occurrences
Eike Best, Flaviu Cristian |
Sci. Comput. Program. | 2 |
| 1979 | A Recovery Mechanism for Modular Software
Flaviu Cristian |
ICSE | 1 |