VLDB 2026 Research / reviewers in the wild / expert
Brian Randell
dblp:r/BrianRandell
· DBLP profile ↗
39ranked-venue papers
16as first author
1since 2021 · last 2025
0000-0002-5863-0107ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 5 first-author · 1 since 2021Security and privacy · 8 · 3 first-authorSystems, architecture and hardware · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 5 first-authorTheory of computation · 5 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
8 papers |
Operating systems · 50% Software testing · 50% Concurrent programming · 0% | |
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Distributed systems · 90% Processor architecture and microarchitecture · 6% Embedded and real-time systems · 2% |
Topics — the 24 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems › fault tolerance
recovery blocks |
0.9 | 2 | 2025 | Looking Back on Recovery Blocks and Conversations · IEEE Trans. Software Eng. 2025 System Structure for Software Fault Tolerance · IEEE Trans. Software Eng. 1975 |
Software testing
software fault tolerance |
0.9 | 2 | 2025 | Looking Back on Recovery Blocks and Conversations · IEEE Trans. Software Eng. 2025 System Structure for Software Fault Tolerance · IEEE Trans. Software Eng. 1975 |
Distributed systems
fault tolerance |
0.3 | 5 | 2025 | Looking Back on Recovery Blocks and Conversations · IEEE Trans. Software Eng. 2025 Basic Concepts and Taxonomy of Dependable and Secure Computing · IEEE Trans. Dependable Secur. Comput. 2004 Concurrent Exception Handling and Resolution in Distributed Object Systems · IEEE Trans. Parallel Distributed Syst. 2000 |
Distributed systems
distributed object systems |
0.0 | 1 | 2000 | Concurrent Exception Handling and Resolution in Distributed Object Systems · IEEE Trans. Parallel Distributed Syst. 2000 |
Processor architecture and microarchitecture
exception handling |
0.0 | 1 | 2000 | Concurrent Exception Handling and Resolution in Distributed Object Systems · IEEE Trans. Parallel Distributed Syst. 2000 |
Privacy and data protection
data confidentiality |
0.0 | 1 | 2004 | Basic Concepts and Taxonomy of Dependable and Secure Computing · IEEE Trans. Dependable Secur. Comput. 2004 |
Embedded and real-time systems › critical systems
safety-critical systems |
0.0 | 1 | 2002 | Rigorous Development of an Embedded Fault-Tolerant System Based on Coordinated Atomic Actions · IEEE Trans. Computers 2002 |
Distributed systems › distributed system architecture
distributed operating systems |
0.0 | 1 | 1987 | The architecture of UNIX united · Proc. IEEE 1987 |
Concurrent programming › atomicity
atomic operations |
0.0 | 1 | 1986 | Error Recovery in Asynchronous Systems · IEEE Trans. Software Eng. 1986 |
Hardware reliability and fault tolerance
error recovery |
0.0 | 1 | 1986 | Error Recovery in Asynchronous Systems · IEEE Trans. Software Eng. 1986 |
Internet architecture and protocols › naming and addressing
network addressing |
0.0 | 1 | 1985 | Interfacing UNIX to Data Communications Networks · IEEE Trans. Software Eng. 1985 |
Operating systems › system security › operating system security › secure operating system
multilevel secure operating system |
0.0 | 1 | 1983 | A Distributed Secure System · S&P 1983 |
Operating systems › operating system family
UNIX |
0.0 | 1 | 1987 | The architecture of UNIX united · Proc. IEEE 1987 |
Systems and software security › trusted computing
trusted computing base |
0.0 | 1 | 1986 | Building Reliable Secure Computing Systems out of Unreliable Insecure Components · S&P 1986 |
Distributed systems
asynchronous systems |
0.0 | 1 | 1986 | Error Recovery in Asynchronous Systems · IEEE Trans. Software Eng. 1986 |
Distributed systems › distributed system dependability
distributed system reliability |
0.0 | 1 | 1986 | Building Reliable Secure Computing Systems out of Unreliable Insecure Components · S&P 1986 |
Distributed systems › distributed programming
communication abstraction |
0.0 | 1 | 1985 | Interfacing UNIX to Data Communications Networks · IEEE Trans. Software Eng. 1985 |
Distributed systems › distributed programming
distributed application design |
0.0 | 1 | 1985 | Interfacing UNIX to Data Communications Networks · IEEE Trans. Software Eng. 1985 |
Operating systems › resource management › memory management
dynamic memory allocation |
0.0 | 1 | 1967 | Dynamic storage allocation systems · SOSP 1967 |
Operating systems › resource management
memory management |
0.0 | 1 | 1967 | Dynamic storage allocation systems · SOSP 1967 |
Storage systems › storage management
storage allocation |
0.0 | 1 | 1967 | Dynamic storage allocation systems · SOSP 1967 |
Compilers and program optimization
code generation |
0.0 | 1 | 1964 | Single-Scan Techniques for the Translation of Arithmetic Expressions into ALGOL 60 · J. ACM 1964 |
Programming languages and type systems › language design
ALGOL 60 |
0.0 | 1 | 1964 | Single-Scan Techniques for the Translation of Arithmetic Expressions into ALGOL 60 · J. ACM 1964 |
Compilers and program optimization › compiler optimization › local optimization
constant folding |
0.0 | 1 | 1964 | Single-Scan Techniques for the Translation of Arithmetic Expressions into ALGOL 60 · J. ACM 1964 |
Methods — techniques the papers use, named apart from their topics
retrospective analysis · 1.7taxonomy · 0.1definitions · 0.1message complexity analysis · 0.1distributed algorithm · 0.1model checking · 0.0coordinated atomic actions · 0.0exception tree resolution · 0.0exception handling · 0.0trusted security mechanism · 0.0newcastle connection · 0.0fault-tolerance techniques · 0.0fault tolerance techniques · 0.0datagram primitives · 0.0system structuring · 0.0reverse polish notation · 0.0pushdown store · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Looking Back on Recovery Blocks and ConversationsabstractOur 1975 paper “System Structure for Software Fault Tolerance” introduced “Recovery Blocks” (a backward error recovery strategy for use in isolated processes), the “Domino Effect” (the problem that a single error could cause multiple interacting processes with uncoordinated error recovery strategies to lose all their recovery capability) and “Conversations” (an error recovery strategy for interacting processes motivated by the danger of the domino effect). This retrospective account describes how these ideas were developed by the Newcastle group and its collaborators, and what further research ensued. A tentative assessment is then provided of the impact of this research. Brian Randell, Jie Xu 0007 |
IEEE Trans. Software Eng. | 1 |
| 2012 | A Turing Enigma
Brian Randell |
CONCUR | 1 |
| 2011 | Occurrence Nets Then and Now: The Path to Structured Occurrence Nets
Brian Randell |
Petri Nets | 1 |
| 2009 | Structured Occurrence Nets: A Formalism for Aiding System Failure Prevention and Analysis TechniquesabstractThis paper introduces the concept of a 'structured occurrence net', which as its name indicates is based on that of an 'occurrence net', a well-established formalism for an abstract record that represents causality and concurrency information concerning a single execution of a system. Structured occurrence nets consist of multiple occurrence nets, associated together by means of various types of relationship, and are intended for recording or predicting, either the actual behaviour of complex systems as they communicate and evolve, or evidence that is being gathered and analysed concerning their alleged past behaviour. We provide a formal basis for the new formalism and show how it can be used to gain better understanding of complex fault-error-failure chains (i) among co-existing communicating systems, (ii) between systems and their sub-systems, and (iii) involving systems that are controlling, creating ormodifying other systems. We then go on to discuss how, with appropriate tools support, perhaps using extended versions of existing tools, structured occurrence nets could form a basis for improved techniques of system failure prevention and analysis. Maciej Koutny, Brian Randell |
Fundam. Informaticae | 2 |
| 2008 | Position Statement: How Far Have We Come?abstractIn the late 1960s computers were mainly very big (room size) and expensive, and were mainly used by wealthy organisations for large-scale commercial data processing and scientific calculations, though the first minicomputers were starting to appear in well-funded laboratories. Since then not just one, but rather many, types of software industry have come into existence, in particular those that design or tailor "bespoke" software for particular clients and environments, and those that produce "of-the-peg" software packages that are sold to thousands or even millions of customers. Among the many looming technical developments that were discussed enthusiastically at the NATO conferences, two that now stand out as still of great interest and constituting a considerable challenge are software components and software development environments. A third technical development that was already under way in the late 1960s, though less prominent at the conferences, was that of multiprocessor design - a technical challenge that has been revitalised by the arrival of multi-core processor chips. Brian Randell |
COMPSAC | 1 |
| 2007 | Distributed Secure Systems: Then and NowabstractThe early 1980s saw the development of some rather sophisticated distributed systems. These were not merely networked file systems: rather, using remote procedure calls, hierarchical naming, and what would now be called middleware, they allowed a collection of systems to operate as a coherent whole. One such system in particular was developed at Newcastle that allowed pre-existing applications and (Unix) systems to be used, completely unchanged, as components of an apparently standard large (multiprocessor) Unix system. The distributed secure system (DSS) described in our 1983 paper proposed a new way to construct secure systems by exploiting the design freedom created by this form of distributed computing. The DSS separated the security concerns of policy enforcement from those due to resource sharing and used a variety of mechanisms (dedicated components, cryptography, periods processing, separation kernels) to manage resource sharing in ways that were simpler than before. In this retrospective, we provide the full original text of our DSS paper, prefaced by an introductory discussion of the DSS in the context of its time, and followed by an account of the subsequent implementation and deployment of an industrial prototype of DSS, and a description of its modern interpretation in the form of the MILS architecture. We conclude by outlining current opportunities and challenges presented by this approach to security. Brian Randell, John M. Rushby |
ACSAC | 1 |
| 2007 | The National Programme for Information Technology in the UK Health Service: Dependability Challenges and StrategiesabstractThe National Health Service (NHS) provides the majority of health-care in the UK. Its main section, that for England, serves a population of over 50 million, employs 40,000 general practitioners (family physicians), 80,000 other doctors, and 350,000 nurses, and includes over 300 hospitals. Its National Programme for Information Technology (NPfIT) is the largest civil IT project in the world. (Estimates of its total cost have ranged from £6.2 billion up to £20 billion.) This project, which was launched in 2002, aims to implement electronic care records for all patients and to provide a reliable and secure information service, for medical records, radiography, patient administration, etc., for all the hospitals, and all general practitioners' premises, to which all the NHS health professionals in England will have strictly-controlled access. This Special Plenary Session will provide an overview of NPfIT, and its dependability challenges and strategies. Speakers will, it is hoped, include representatives of Connecting for Health (the NHS Agency responsible for NPfIT), the medical profession, and the dependability research community. Brian Randell |
DSN | 1 |
| 2007 | Failures: Their Definition, Modelling and Analysis
Brian Randell, Maciej Koutny |
ICTAC | 1 |
| 2004 | Dependable Pervasive SystemsabstractSummary form only given. Present trends indicate that huge networked computer systems are likely to become pervasive, as information technology is embedded into virtually everything, and to be required to function essentially continuously. I believe that even today's (underused) "best practice" regarding the achievement of high dependability - reliability, availability, security, safety, etc. - from large networked computer systems will not suffice for future pervasive systems. I will give my perspective on the current state of research into the four basic dependability technologies: (i) fault prevention (to avoid the occurrence or introduction of faults), (ii) fault removal (through validation and verification), (iii) fault tolerance (so that failures do not necessarily occur even if faults remain), and (iv) fault forecasting (the means of assessing progress towards achieving adequate dependability). I will then argue that much further research is required on all four dependability technologies in order to cope with pervasive systems, identify some priorities, and discuss how this research could best be aimed at making system dependability into a "commodity" that industry can value and from which it can profit. Brian Randell |
SRDS | 1 |
| 2004 | Basic Concepts and Taxonomy of Dependable and Secure ComputingabstractThis paper gives the main definitions relating to dependability, a generic concept including a special case of such attributes as reliability, availability, safety, integrity, maintainability, etc. Security brings in concerns for confidentiality, in addition to availability and integrity. Basic definitions are given first. They are then commented upon, and supplemented by additional definitions, which address the threats to dependability and security (faults, errors, failures), their attributes, and the means for their achievement (fault prevention, fault tolerance, fault removal, fault forecasting). The aim is to explicate a set of general concepts, of relevance across a wide range of situations and, therefore, helping communication and cooperation among a number of scientific and technical communities, including ones that are concentrating on particular types of system, of system failures, or of causes of system failures. Algirdas Avizienis, Jean-Claude Laprie, Brian Randell, Carl E. Landwehr |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2002 | Rigorous Development of an Embedded Fault-Tolerant System Based on Coordinated Atomic ActionsabstractDescribes our experience using coordinated atomic (CA) actions as a system structuring tool to design and validate a sophisticated and embedded control system for a complex industrial application that has high reliability and safety requirements. Our study is based on an extended production cell model, the specification and simulator for which were defined and developed by FZI (Forschungszentrum Informatik, Germany). This "fault-tolerant production cell" represents a manufacturing process involving redundant mechanical devices (provided in order to enable continued production in the presence of machine faults). The challenge posed by the model specification is to design a control system that maintains specified safety and liveness properties even in the presence of a large number and variety of device and sensor failures. Based on an analysis of such failures, we provide details of: (1) a design for a control program that uses CA actions to deal with both safety-related and fault tolerance concerns and (2) the formal verification of this design based on the use of model checking. We found that CA action structuring facilitated both the design and verification tasks by enabling the various safety problems (involving possible clashes of moving machinery) to be treated independently. Even complex situations involving the concurrent occurrence of any pairs of the many possible mechanical and sensor failures can be handled simply yet appropriately. The formal verification activity was performed in parallel with the design activity, and the interaction between them resulted in a combined exercise in "design for validation"; formal verification was very valuable in identifying some very subtle residual bugs in early versions of our design which would have been difficult to detect otherwise. Jie Xu 0007, Brian Randell, Alexander B. Romanovsky, Robert J. Stroud, Avelino Francisco Zorzo, Ercument Canver, Friedrich W. von Henke |
IEEE Trans. Computers | 2 |
| 2001 | Building Reliable Secure Computing Systems out of Unreliable Insecure ComponentsabstractParallels are drawn between the problems and techniques associated with achieving high reliability, and those associated with the provision of security, in distributed computing systems. Some limitations of the concept of a Trusted Computing Base are discussed, and an alternative approach to the design of highly secure computing systems is put forward, based on fault tolerance concepts and techniques. John E. Dobson, Brian Randell |
ACSAC | 2 |
| 2001 | On Applying Coordinated Atomic Actions and Dependable Software Architectures for Developing Complex SystemsabstractModern concurrent and distributed applications are becoming increasingly complex; so, in order to provide fault tolerance, special structuring mechanisms are required to help reduce this complexity. Unfortunately, such structuring techniques are mostly introduced as design and implementation features, which complicates their employment. The approach we propose relies on introducing the appropriate software structuring together with associated fault tolerance measures at the earlier phases of software development and on supporting it with special software architectures and design patterns. Delano M. Beder, Cecília M. F. Rubira, Brian Randell, Alexander B. Romanovsky |
ISORC | 3 |
| 2000 | Turing Memorial Lecture Facing Up to FaultsabstractAbstract As individuals, organizations and indeed the world at large have become more dependent on computer-based systems, so there has been an ever-growing amount of research into means for improving the dependability of these systems. In particular, there has been much work on trying to gain an increased understanding of the many and varied types of faults that need to be prevented or tolerated in order to reduce the probability and severity of system failures. In this talk I discuss the assumptions that are often made by computing system designers regarding faults, survey a number of continuing issues related to fault tolerance, and identify some of the latest challenges facing researchers in this arena. Brian Randell |
Comput. J. | 1 |
| 2000 | Concurrent Exception Handling and Resolution in Distributed Object SystemsabstractWe address the problem of how to handle exceptions in distributed object systems. In a distributed computing environment, exceptions may be raised simultaneously in different processing nodes and thus need to be treated in a coordinated manner. Mishandling concurrent exceptions can lead to catastrophic consequences. We take two kinds of concurrency into account: 1) Several objects are designed collectively and invoked concurrently to achieve a global goal and 2) multiple objects (or object groups) that are designed independently compete for the same system resources. We propose a new distributed algorithm for resolving concurrent exceptions and show that the algorithm works correctly even in complex nested situations, and is an improvement over previous proposals in that it requires only O(n/sub max/N/sup 2/) messages, thereby permitting quicker response to exceptions. Jie Xu 0007, Alexander B. Romanovsky, Brian Randell |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1999 | Fault Tolerance in Decentralized SystemsabstractIn a decentralised system the problems of fault tolerance, and in particular error recovery, vary greatly depending on the design assumptions. For example, in a distributed database system, if one disregards the possibility of undetected invalid inputs or outputs, the errors that have to be recovered from will just affect the database, and backward error recovery will be feasible and should suffice. Such a system is typically supporting a set of activities that are competing for access to a shared database, but which are otherwise essentially independent of each other in such circumstances conventional database transaction processing and distributed protocols enable backward recovery to be provided very effectively. But in more general systems the multiple activities will often not simply be competing against each other, but rather will at times be attempting to co-operate with each other, in pursuit of some common goal. Moreover, the activities in decentralised systems typically involve not just computers, but also external entities that are not capable of backward error recovery. Such additional complications make the task of error recovery more challenging, and indeed more interesting. This paper provides a brief analysis of the consequences of various such complications, and outlines some recent work on advanced error recovery techniques that they have motivated. Brian Randell |
ISADS | 1 |
| 1999 | Using Coordinated Atomic Actions to Design Safety-Critical Systems: a Production Cell Case StudyabstractCoordinated Atomic actions (CA actions) are a unified approach to structuring complex concurrent activities and supporting error recovery between multiple interacting objects in object-oriented systems. This paper explains how we have used the CA action concept to design and implement a safety-critical application. We have used the Production Cell model that was developed in the Forschungszentrum Informatik (FZI), Karlsruhe, Germany, to present a realistic industry-oriented problem, where safety requirements play a significant role. Our design consists of two levels: the first level deals with the scheduling of CA actions, and the second level deals with the interactions between devices. Both the scheduling mechanism and the device interactions are enclosed by CA actions. Exception handling and error recovery are incorporated into CA actions in order to satisfy high safety and fault tolerance requirements. A controlling program based on our design was developed in the Java language and used to drive a graphical simulator provided by the FZI. Copyright © 1999 John Wiley & Sons, Ltd. Avelino Francisco Zorzo, Alexander B. Romanovsky, Jie Xu 0007, Brian Randell, Robert J. Stroud, Ian Welch |
Softw. Pract. Exp. | 4 |
| 1998 | Coordinated Exception Handling in Distributed Object Systems: From Model to System ImplementationabstractException handling in concurrent and distributed programs is a difficult task though it is often necessary. In many cases traditional exception mechanisms for sequential programs are no longer appropriate. One major difficulty is that the process of handling an exception may need to involve multiple concurrent components that are cooperating in pursuit of some global goal. Another complication is that several exceptions may be raised concurrently in different nodes of a distributed environment. Existing proposals and actual concurrent languages either ignore these difficulties or only cope with a limited form of them. The paper attempts a general solution, developed especially for distributed object systems, starting from a conceptual model, together with algorithms for coordinating concurrent components and resolving multiple exceptions, through to an actual system implementation. An industrial production cell is chosen as a case study to demonstrate the usefulness of the proposed model and algorithms. A system that supports coordinated atomic actions and exception resolution is implemented in distributed Ada 95 and examined through several performance-related experiments. Jie Xu 0007, Alexander B. Romanovsky, Brian Randell |
ICDCS | 3 |
| 1998 | Exception Handling in Object-Oriented Real-Time Distributed SystemsabstractException handling in a complex concurrent and distributed system (e.g. one involving cooperating rather than just competing activities) is often a necessary, but a very difficult, task. No widely accepted models or approaches exist in this area. The object-oriented paradigm, for all its structuring benefits, and real-time requirements each add further difficulties to the design and implementation of exception handling in such systems. In this paper, we develop a general structuring framework based on the coordinated atomic (CA) action concept for handling exceptions in an object-oriented distributed system, in which exceptions in both the value and the time domain are taken into account. In particular, we attempt to attack several difficult problems related to real-time system design and error recovery, including action-level timing constraints, time-triggered CA actions, and time-dependent exception handling. The proposed framework is then demonstrated and assessed using an industrial real-time application-the Production Cell III case study. Alexander B. Romanovsky, Jie Xu 0007, Brian Randell |
ISORC | 3 |
| 1998 | Protecting IT Systems from Cyber CrimeabstractLarge-scale commercial, industrial and financial operations are becoming ever more interdependent, and ever more dependent on IT. At the same time, the rapidly growing interconnectivity of IT systems, and the convergence of their technology towards industry-standard hardware and software components and sub-systems, renders these IT systems increasingly vulnerable to malicious attack. This paper is aimed particularly at readers concerned with major systems employed in medium to large commercial or industrial enterprises. It examines the nature and significance of the various potential attacks, and surveys the defence options available. It concludes that IT owners need to think of the threat in more global terms, and to give a new focus and priority to their defence. Prompt action can ensure a major improvement in IT resilience at a modest marginal cost, both in terms of finance and in terms of normal IT operation. R. Benjamin, B. Gladman, Brian Randell |
Comput. J. | 3 |
| 1997 | Implementation of blocking coordinated atomic actions based on forward error recovery
Alexander B. Romanovsky, Brian Randell, Robert J. Stroud, Jie Xu 0007, Avelino Francisco Zorzo |
J. Syst. Archit. | 2 |
| 1996 | Exception Handling and Resolution in Distributed Object-oriented SystemsabstractWe address the problem of how to handle exceptions in distributed object-oriented systems. In a distributed computing environment exceptions may be raised simultaneously and thus need to be treated in a coordinated manner. We take two kinds of concurrency into account: 1) several objects are designed collectively and invoked concurrently to achieve a global goal, and 2) concurrent objects or object groups that are designed independently compete for the same system resources. We propose a new distributed algorithm for resolving concurrent exceptions and show that the algorithm works correctly even in complex nested situations, and is an improvement over previous proposals in that it requires only O(N/sup 2/) messages, and is fully object-oriented. Jie Xu 0007, Alexander B. Romanovsky, Brian Randell |
ICDCS | 3 |
| 1996 | Roll-forward error recovery in embedded real-time systemsabstractRoll-forward checkpointing schemes are developed in order to avoid rollback in the presence of independent faults and to increase the possibility that a task completes within a tight deadline. However, despite of the adoption of roll-forward recovery, these schemes are not necessarily appropriate for time-critical applications because interactions with the external environment and communications between processes must be deferred during checkpoint validation steps (typically, two checkpoint intervals) until the fault-free processors are identified. The deadlines on providing services may thus be violated. In this paper we present and discuss two alternative roll-forward recovery schemes, especially for time-critical and interaction-intensive applications, that deliver correct, timely results even when checkpoint validation is required. Jie Xu 0007, Brian Randell |
ICPADS | 2 |
| 1993 | The Duality of Fault-tolerant System StructuresabstractAbstract An examination of the structure of fault‐tolerant systems incorporating backward error recovery indicates a partitioning into two broad classes. Two canonical models, each representing a particular class of systems, have been constructed. The first model incorporates objects and actions as the entities for program construction whereas the second model employs communicating processes and conversations. Applications in areas such as office information and banking systems are typically described and built in terms of the first model whereas applications in the area of process control are usually described and built in terms of the second model. The paper claims that the two models are duals of each other and presents arguments and examples to substantiate this claim. It will be shown that the techniques that have been developed within the context of one model turn out to have interesting and hitherto unexplored duals in the other model. Santosh K. Shrivastava, Luigi V. Mancini, Brian Randell |
Softw. Pract. Exp. | 3 |
| 1992 | An Object-Oriented View of Fragmented Data Processing for Fault and Intrusion Tolerance in Distributed Systems
Jean-Charles Fabre, Brian Randell |
ESORICS | 2 |
| 1987 | The architecture of UNIX unitedabstractUNIX United is an architecture for a distributed system based on UNIX. As it is compatible with UNIX at the system call level, any program written for a normal UNIX system can be transparently extended to exploit the richer environment of UNIX United. As it relies on having a UNIX system beneath it, the implementation of UNIX United, called the Newcastle Connection, provides an interesting example of the construction of a very powerful distributed system with only a modicum of effort. A description of the basic semantics of UNIX United is followed by that of the architecture implied by the protocol between components in a UNIX United system, and of a software structure appropriate to the architecture and the protocol. James P. Black, Lindsay F. Marshall, Brian Randell |
Proc. IEEE | 3 |
| 1986 | Building Reliable Secure Computing Systems out of Unreliable Insecure ComponentsabstractParallels are drawn between the problems and techniques associated with achieving high reliability, and those associated with the provision of security, in distributed computing systems. Some limitations of the concept of a Trusted Computing Base are discussed, end an alternative approach to thedesign of highly secure computing systems is put forward,based on fault tolerance concepts and techniques. John E. Dobson, Brian Randell |
S&P | 2 |
| 1986 | System Design and StructuringabstractThe task of implementing a large and sophisticated computing system is often unduly costly and time-consuming, with the resulting system exhibiting inadequate performance and reliability, because of excessive system complexity. Such complexity can be reduced significantly by ensuring that the system is constructed out of a well-chosen set of largely independent components, which interact in well-understood ways. However, the task of structuring a system, i.e. of choosing and defining appropriate components, can be very difficult. This paper describes a technique of system structuring which involved distinguishing the functionality which a system is intended to have from other desirable attributes, such as reliability and security, and then using separate components to provide each of these attributes. Various UNIX-based systems which have been implemented at Newcastle are used to illustrate this structuring technique. Brian Randell |
Comput. J. | 1 |
| 1986 | Error Recovery in Asynchronous SystemsabstractA framework for the provision of fault tolerance in asynchronous systems is introduced. The proposal generalizes the form of simple recovery facilities supported by nested atomic actions in which the exception mechanisms only permit backward error recovery. It allows the construction of systems using both forward and backward error recovery and thus allows the exploitation of the complementary benefits of the two schemes. Backward recovery, forward recovery, and normal processing activities can occur concurrently within the organization proposed. Exception handling is generalized to provide a uniform basis for fault tolerance schemes with the atomic action structure. The generalization includes a resolution scheme for concurrently raised exceptions based on an exception tree and an abortion scheme that permits the termination of the internal atomic actions. An automatic resolution mechanism is outlined for exceptions in atomic actions which allows users to separate their recovery schemes from the details of the underlying algorithms. Roy H. Campbell, Brian Randell |
IEEE Trans. Software Eng. | 2 |
| 1985 | Interfacing UNIX to Data Communications NetworksabstractWe propose an interface for use from within UNIX1 user programs for communicating over multiple and varied local and wide area networks. This interface aids the design of a distributed application program by hiding the actual communications protocols used over each network, and providing instead simple primitives for sending and receiving (possibly large) datagrams, using a simple standardized network addressing scheme based on apair. Fabio Panzieri, Brian Randell |
IEEE Trans. Software Eng. | 2 |
| 1983 | A Distributed Secure SystemabstractWe describe the design of a distributed general-purpose computing system that enforces a multilevel security policy. The system is composed of standard UNIX systems and small trustworthy security mechanisms linked together in such a way as to provide a total system which, is not only demonstrably secure, but also highly efficient and cost effective. Despite the heterogeneity of its components, the system as a whole appears to be a single multilevel secure UNIX system, since the fact that it is actually a distributed system is completely hidden from its users and their programs.This is achieved through the use of the "Newcastle Connection", a software subsystem that links together multiple UNIX or UNIX-look-alike systems, without requiring any changes to the source code of either the operating system or any user programs. Construction of a prototype implementation is in progress. John M. Rushby, Brian Randell |
S&P | 2 |
| 1982 | The Newcastle Connection or UNIXes of the World Unite!abstractAbstract In this paper we describe a software subsystem that can be added to each of a set of physically interconnected UNIX or UNIX look‐alike systems, so as to construct a distributed system which is functionally indistinguishable at both the user and the program level from a conventional single‐processor UNIX system. The techniques used are applicable to a variety and multiplicity of both local and wide area networks, and enable all issues of inter‐processor communication, network protocols, etc., to be hidden. A brief account is given of experience with such a distributed system, which is currently operational on a set of PDPlls connected by a Cambridge Ring. The final sections compare our scheme to various precursor schemes and discuss its potential relevance to other operating systems. David R. Brownbridge, Lindsay F. Marshall, Brian Randell |
Softw. Pract. Exp. | 3 |
| 1981 | A Formal Model of Atomicity in Asynchronous Systems
Eike Best, Brian Randell |
Acta Informatica | 2 |
| 1979 | Software Engineering: As it was in 1968
Brian Randell |
ICSE | 1 |
| 1975 | System Structure for Software Fault ToleranceabstractPresents and discusses the rationale behind a method for structuring complex computing systems by the use of what is termed `recovery blocks,' `conversations,' and `fault-tolerant interfaces.' The aim is to facilitate the provision of dependable error detection and recovery facilities which can cope with errors caused by residual design inadequacies, particularly in the system software, rather than merely the occasional malfunctioning of hardware components. Brian Randell |
IEEE Trans. Software Eng. | 1 |
| 1971 | Performance Predictions for Extended Paged Memories
Edward G. Coffman Jr., Brian Randell |
Acta Informatica | 2 |
| 1971 | Ludgate's Analytical Machine of 1909abstractThis paper discusses the little known analytical machine, or program-controlled mechanical calculator, designed by Percy E. Ludgate in Ireland during the years 1903 to 1909, and documents the results of a search for information about his life and work. Brian Randell |
Comput. J. | 1 |
| 1967 | Dynamic storage allocation systemsabstractIn many recent computer system designs, hardware facilities have been provided for easing the problems of storage allocation. This paper presents a method of characterizing dynamic storage allocation systems, according to the functional capabilities provided, and the underlying techniques used. The basic purpose of the paper is to provide a useful perspective from which the utility of various hardware facilities may be assessed. The paper includes as an appendix, a brief survey of storage allocation facilities in several representative computer systems. Brian Randell, C. J. Kuehner |
SOSP | 1 |
| 1964 | Single-Scan Techniques for the Translation of Arithmetic Expressions into ALGOL 60abstractThe first section of the paper contains a brief description of the well-known technique of using a stack, or pushdown store, to re-order the operators of an arithmetic expression, as defined in ALGOL 60, in order to transform the expression into Reverse Polish parenthesis-free form. It is shown that improvements to this Reverse Polish form can be made quite simply, by extending the use of the stack to include information about the operands of the expression. Firstly, information gained from the declarations of the operands can be used to control the generation of real-integer conversion instructions. Secondly, operators whose operands are numerical constants can be computed during translation, using the partially generated Reverse Polish object program as a second stack. Brian Randell, L. J. Russell |
J. ACM | 1 |