Brian Randell

dblp:r/BrianRandell · DBLP profile ↗
← Back
39ranked-venue papers
16as first author
1since 2021 · last 2025
0000-0002-5863-0107ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 5 first-author · 1 since 2021Security and privacy · 8 · 3 first-authorSystems, architecture and hardware · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 5 first-authorTheory of computation · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
8 papers
Operating systems · 50% Software testing · 50% Concurrent programming · 0%
Computer architecture, parallel and distributed computing, and storage systems
10 papers
Distributed systems · 90% Processor architecture and microarchitecture · 6% Embedded and real-time systems · 2%

Topics — the 24 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Operating systems › fault tolerance
recovery blocks
0.922025
Looking Back on Recovery Blocks and Conversations · IEEE Trans. Software Eng. 2025
System Structure for Software Fault Tolerance · IEEE Trans. Software Eng. 1975
Software testing
software fault tolerance
0.922025
Looking Back on Recovery Blocks and Conversations · IEEE Trans. Software Eng. 2025
System Structure for Software Fault Tolerance · IEEE Trans. Software Eng. 1975
Distributed systems
fault tolerance
0.352025
Looking Back on Recovery Blocks and Conversations · IEEE Trans. Software Eng. 2025
Basic Concepts and Taxonomy of Dependable and Secure Computing · IEEE Trans. Dependable Secur. Comput. 2004
Concurrent Exception Handling and Resolution in Distributed Object Systems · IEEE Trans. Parallel Distributed Syst. 2000
Distributed systems
distributed object systems
0.012000
Concurrent Exception Handling and Resolution in Distributed Object Systems · IEEE Trans. Parallel Distributed Syst. 2000
Processor architecture and microarchitecture
exception handling
0.012000
Concurrent Exception Handling and Resolution in Distributed Object Systems · IEEE Trans. Parallel Distributed Syst. 2000
Privacy and data protection
data confidentiality
0.012004
Basic Concepts and Taxonomy of Dependable and Secure Computing · IEEE Trans. Dependable Secur. Comput. 2004
Embedded and real-time systems › critical systems
safety-critical systems
0.012002
Rigorous Development of an Embedded Fault-Tolerant System Based on Coordinated Atomic Actions · IEEE Trans. Computers 2002
Distributed systems › distributed system architecture
distributed operating systems
0.011987
The architecture of UNIX united · Proc. IEEE 1987
Concurrent programming › atomicity
atomic operations
0.011986
Error Recovery in Asynchronous Systems · IEEE Trans. Software Eng. 1986
Hardware reliability and fault tolerance
error recovery
0.011986
Error Recovery in Asynchronous Systems · IEEE Trans. Software Eng. 1986
Internet architecture and protocols › naming and addressing
network addressing
0.011985
Interfacing UNIX to Data Communications Networks · IEEE Trans. Software Eng. 1985
Operating systems › system security › operating system security › secure operating system
multilevel secure operating system
0.011983
A Distributed Secure System · S&P 1983
Operating systems › operating system family
UNIX
0.011987
The architecture of UNIX united · Proc. IEEE 1987
Systems and software security › trusted computing
trusted computing base
0.011986
Building Reliable Secure Computing Systems out of Unreliable Insecure Components · S&P 1986
Distributed systems
asynchronous systems
0.011986
Error Recovery in Asynchronous Systems · IEEE Trans. Software Eng. 1986
Distributed systems › distributed system dependability
distributed system reliability
0.011986
Building Reliable Secure Computing Systems out of Unreliable Insecure Components · S&P 1986
Distributed systems › distributed programming
communication abstraction
0.011985
Interfacing UNIX to Data Communications Networks · IEEE Trans. Software Eng. 1985
Distributed systems › distributed programming
distributed application design
0.011985
Interfacing UNIX to Data Communications Networks · IEEE Trans. Software Eng. 1985
Operating systems › resource management › memory management
dynamic memory allocation
0.011967
Dynamic storage allocation systems · SOSP 1967
Operating systems › resource management
memory management
0.011967
Dynamic storage allocation systems · SOSP 1967
Storage systems › storage management
storage allocation
0.011967
Dynamic storage allocation systems · SOSP 1967
Compilers and program optimization
code generation
0.011964
Single-Scan Techniques for the Translation of Arithmetic Expressions into ALGOL 60 · J. ACM 1964
Programming languages and type systems › language design
ALGOL 60
0.011964
Single-Scan Techniques for the Translation of Arithmetic Expressions into ALGOL 60 · J. ACM 1964
Compilers and program optimization › compiler optimization › local optimization
constant folding
0.011964
Single-Scan Techniques for the Translation of Arithmetic Expressions into ALGOL 60 · J. ACM 1964

Methods — techniques the papers use, named apart from their topics

retrospective analysis · 1.7taxonomy · 0.1definitions · 0.1message complexity analysis · 0.1distributed algorithm · 0.1model checking · 0.0coordinated atomic actions · 0.0exception tree resolution · 0.0exception handling · 0.0trusted security mechanism · 0.0newcastle connection · 0.0fault-tolerance techniques · 0.0fault tolerance techniques · 0.0datagram primitives · 0.0system structuring · 0.0reverse polish notation · 0.0pushdown store · 0.0
YearPublicationVenuePosition
2025 Looking Back on Recovery Blocks and Conversations
abstract
Our 1975 paper “System Structure for Software Fault Tolerance” introduced “Recovery Blocks” (a backward error recovery strategy for use in isolated processes), the “Domino Effect” (the problem that a single error could cause multiple interacting processes with uncoordinated error recovery strategies to lose all their recovery capability) and “Conversations” (an error recovery strategy for interacting processes motivated by the danger of the domino effect). This retrospective account describes how these ideas were developed by the Newcastle group and its collaborators, and what further research ensued. A tentative assessment is then provided of the impact of this research.
Brian Randell, Jie Xu 0007
IEEE Trans. Software Eng.1
2012 A Turing Enigma
Brian Randell
CONCUR1
2011 Occurrence Nets Then and Now: The Path to Structured Occurrence Nets
Brian Randell
Petri Nets1
2009 Structured Occurrence Nets: A Formalism for Aiding System Failure Prevention and Analysis Techniques
abstract
This paper introduces the concept of a 'structured occurrence net', which as its name indicates is based on that of an 'occurrence net', a well-established formalism for an abstract record that represents causality and concurrency information concerning a single execution of a system. Structured occurrence nets consist of multiple occurrence nets, associated together by means of various types of relationship, and are intended for recording or predicting, either the actual behaviour of complex systems as they communicate and evolve, or evidence that is being gathered and analysed concerning their alleged past behaviour. We provide a formal basis for the new formalism and show how it can be used to gain better understanding of complex fault-error-failure chains (i) among co-existing communicating systems, (ii) between systems and their sub-systems, and (iii) involving systems that are controlling, creating ormodifying other systems. We then go on to discuss how, with appropriate tools support, perhaps using extended versions of existing tools, structured occurrence nets could form a basis for improved techniques of system failure prevention and analysis.
Maciej Koutny, Brian Randell
Fundam. Informaticae2
2008 Position Statement: How Far Have We Come?
abstract
In the late 1960s computers were mainly very big (room size) and expensive, and were mainly used by wealthy organisations for large-scale commercial data processing and scientific calculations, though the first minicomputers were starting to appear in well-funded laboratories. Since then not just one, but rather many, types of software industry have come into existence, in particular those that design or tailor "bespoke" software for particular clients and environments, and those that produce "of-the-peg" software packages that are sold to thousands or even millions of customers. Among the many looming technical developments that were discussed enthusiastically at the NATO conferences, two that now stand out as still of great interest and constituting a considerable challenge are software components and software development environments. A third technical development that was already under way in the late 1960s, though less prominent at the conferences, was that of multiprocessor design - a technical challenge that has been revitalised by the arrival of multi-core processor chips.
Brian Randell
COMPSAC1
2007 Distributed Secure Systems: Then and Now
abstract
The early 1980s saw the development of some rather sophisticated distributed systems. These were not merely networked file systems: rather, using remote procedure calls, hierarchical naming, and what would now be called middleware, they allowed a collection of systems to operate as a coherent whole. One such system in particular was developed at Newcastle that allowed pre-existing applications and (Unix) systems to be used, completely unchanged, as components of an apparently standard large (multiprocessor) Unix system. The distributed secure system (DSS) described in our 1983 paper proposed a new way to construct secure systems by exploiting the design freedom created by this form of distributed computing. The DSS separated the security concerns of policy enforcement from those due to resource sharing and used a variety of mechanisms (dedicated components, cryptography, periods processing, separation kernels) to manage resource sharing in ways that were simpler than before. In this retrospective, we provide the full original text of our DSS paper, prefaced by an introductory discussion of the DSS in the context of its time, and followed by an account of the subsequent implementation and deployment of an industrial prototype of DSS, and a description of its modern interpretation in the form of the MILS architecture. We conclude by outlining current opportunities and challenges presented by this approach to security.
Brian Randell, John M. Rushby
ACSAC1
2007 The National Programme for Information Technology in the UK Health Service: Dependability Challenges and Strategies
abstract
The National Health Service (NHS) provides the majority of health-care in the UK. Its main section, that for England, serves a population of over 50 million, employs 40,000 general practitioners (family physicians), 80,000 other doctors, and 350,000 nurses, and includes over 300 hospitals. Its National Programme for Information Technology (NPfIT) is the largest civil IT project in the world. (Estimates of its total cost have ranged from £6.2 billion up to £20 billion.) This project, which was launched in 2002, aims to implement electronic care records for all patients and to provide a reliable and secure information service, for medical records, radiography, patient administration, etc., for all the hospitals, and all general practitioners' premises, to which all the NHS health professionals in England will have strictly-controlled access. This Special Plenary Session will provide an overview of NPfIT, and its dependability challenges and strategies. Speakers will, it is hoped, include representatives of Connecting for Health (the NHS Agency responsible for NPfIT), the medical profession, and the dependability research community.
Brian Randell
DSN1
2007 Failures: Their Definition, Modelling and Analysis
Brian Randell, Maciej Koutny
ICTAC1
2004 Dependable Pervasive Systems
abstract
Summary form only given. Present trends indicate that huge networked computer systems are likely to become pervasive, as information technology is embedded into virtually everything, and to be required to function essentially continuously. I believe that even today's (underused) "best practice" regarding the achievement of high dependability - reliability, availability, security, safety, etc. - from large networked computer systems will not suffice for future pervasive systems. I will give my perspective on the current state of research into the four basic dependability technologies: (i) fault prevention (to avoid the occurrence or introduction of faults), (ii) fault removal (through validation and verification), (iii) fault tolerance (so that failures do not necessarily occur even if faults remain), and (iv) fault forecasting (the means of assessing progress towards achieving adequate dependability). I will then argue that much further research is required on all four dependability technologies in order to cope with pervasive systems, identify some priorities, and discuss how this research could best be aimed at making system dependability into a "commodity" that industry can value and from which it can profit.
Brian Randell
SRDS1
2004 Basic Concepts and Taxonomy of Dependable and Secure Computing
abstract
This paper gives the main definitions relating to dependability, a generic concept including a special case of such attributes as reliability, availability, safety, integrity, maintainability, etc. Security brings in concerns for confidentiality, in addition to availability and integrity. Basic definitions are given first. They are then commented upon, and supplemented by additional definitions, which address the threats to dependability and security (faults, errors, failures), their attributes, and the means for their achievement (fault prevention, fault tolerance, fault removal, fault forecasting). The aim is to explicate a set of general concepts, of relevance across a wide range of situations and, therefore, helping communication and cooperation among a number of scientific and technical communities, including ones that are concentrating on particular types of system, of system failures, or of causes of system failures.
Algirdas Avizienis, Jean-Claude Laprie, Brian Randell, Carl E. Landwehr
IEEE Trans. Dependable Secur. Comput.3
2002 Rigorous Development of an Embedded Fault-Tolerant System Based on Coordinated Atomic Actions
abstract
Describes our experience using coordinated atomic (CA) actions as a system structuring tool to design and validate a sophisticated and embedded control system for a complex industrial application that has high reliability and safety requirements. Our study is based on an extended production cell model, the specification and simulator for which were defined and developed by FZI (Forschungszentrum Informatik, Germany). This "fault-tolerant production cell" represents a manufacturing process involving redundant mechanical devices (provided in order to enable continued production in the presence of machine faults). The challenge posed by the model specification is to design a control system that maintains specified safety and liveness properties even in the presence of a large number and variety of device and sensor failures. Based on an analysis of such failures, we provide details of: (1) a design for a control program that uses CA actions to deal with both safety-related and fault tolerance concerns and (2) the formal verification of this design based on the use of model checking. We found that CA action structuring facilitated both the design and verification tasks by enabling the various safety problems (involving possible clashes of moving machinery) to be treated independently. Even complex situations involving the concurrent occurrence of any pairs of the many possible mechanical and sensor failures can be handled simply yet appropriately. The formal verification activity was performed in parallel with the design activity, and the interaction between them resulted in a combined exercise in "design for validation"; formal verification was very valuable in identifying some very subtle residual bugs in early versions of our design which would have been difficult to detect otherwise.
Jie Xu 0007, Brian Randell, Alexander B. Romanovsky, Robert J. Stroud, Avelino Francisco Zorzo, Ercument Canver, Friedrich W. von Henke
IEEE Trans. Computers2
2001 Building Reliable Secure Computing Systems out of Unreliable Insecure Components
abstract
Parallels are drawn between the problems and techniques associated with achieving high reliability, and those associated with the provision of security, in distributed computing systems. Some limitations of the concept of a Trusted Computing Base are discussed, and an alternative approach to the design of highly secure computing systems is put forward, based on fault tolerance concepts and techniques.
John E. Dobson, Brian Randell
ACSAC2
2001 On Applying Coordinated Atomic Actions and Dependable Software Architectures for Developing Complex Systems
abstract
Modern concurrent and distributed applications are becoming increasingly complex; so, in order to provide fault tolerance, special structuring mechanisms are required to help reduce this complexity. Unfortunately, such structuring techniques are mostly introduced as design and implementation features, which complicates their employment. The approach we propose relies on introducing the appropriate software structuring together with associated fault tolerance measures at the earlier phases of software development and on supporting it with special software architectures and design patterns.
Delano M. Beder, Cecília M. F. Rubira, Brian Randell, Alexander B. Romanovsky
ISORC3
2000 Turing Memorial Lecture Facing Up to Faults
abstract
Abstract As individuals, organizations and indeed the world at large have become more dependent on computer-based systems, so there has been an ever-growing amount of research into means for improving the dependability of these systems. In particular, there has been much work on trying to gain an increased understanding of the many and varied types of faults that need to be prevented or tolerated in order to reduce the probability and severity of system failures. In this talk I discuss the assumptions that are often made by computing system designers regarding faults, survey a number of continuing issues related to fault tolerance, and identify some of the latest challenges facing researchers in this arena.
Brian Randell
Comput. J.1
2000 Concurrent Exception Handling and Resolution in Distributed Object Systems
abstract
We address the problem of how to handle exceptions in distributed object systems. In a distributed computing environment, exceptions may be raised simultaneously in different processing nodes and thus need to be treated in a coordinated manner. Mishandling concurrent exceptions can lead to catastrophic consequences. We take two kinds of concurrency into account: 1) Several objects are designed collectively and invoked concurrently to achieve a global goal and 2) multiple objects (or object groups) that are designed independently compete for the same system resources. We propose a new distributed algorithm for resolving concurrent exceptions and show that the algorithm works correctly even in complex nested situations, and is an improvement over previous proposals in that it requires only O(n/sub max/N/sup 2/) messages, thereby permitting quicker response to exceptions.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
IEEE Trans. Parallel Distributed Syst.3
1999 Fault Tolerance in Decentralized Systems
abstract
In a decentralised system the problems of fault tolerance, and in particular error recovery, vary greatly depending on the design assumptions. For example, in a distributed database system, if one disregards the possibility of undetected invalid inputs or outputs, the errors that have to be recovered from will just affect the database, and backward error recovery will be feasible and should suffice. Such a system is typically supporting a set of activities that are competing for access to a shared database, but which are otherwise essentially independent of each other in such circumstances conventional database transaction processing and distributed protocols enable backward recovery to be provided very effectively. But in more general systems the multiple activities will often not simply be competing against each other, but rather will at times be attempting to co-operate with each other, in pursuit of some common goal. Moreover, the activities in decentralised systems typically involve not just computers, but also external entities that are not capable of backward error recovery. Such additional complications make the task of error recovery more challenging, and indeed more interesting. This paper provides a brief analysis of the consequences of various such complications, and outlines some recent work on advanced error recovery techniques that they have motivated.
Brian Randell
ISADS1
1999 Using Coordinated Atomic Actions to Design Safety-Critical Systems: a Production Cell Case Study
abstract
Coordinated Atomic actions (CA actions) are a unified approach to structuring complex concurrent activities and supporting error recovery between multiple interacting objects in object-oriented systems. This paper explains how we have used the CA action concept to design and implement a safety-critical application. We have used the Production Cell model that was developed in the Forschungszentrum Informatik (FZI), Karlsruhe, Germany, to present a realistic industry-oriented problem, where safety requirements play a significant role. Our design consists of two levels: the first level deals with the scheduling of CA actions, and the second level deals with the interactions between devices. Both the scheduling mechanism and the device interactions are enclosed by CA actions. Exception handling and error recovery are incorporated into CA actions in order to satisfy high safety and fault tolerance requirements. A controlling program based on our design was developed in the Java language and used to drive a graphical simulator provided by the FZI. Copyright © 1999 John Wiley & Sons, Ltd.
Avelino Francisco Zorzo, Alexander B. Romanovsky, Jie Xu 0007, Brian Randell, Robert J. Stroud, Ian Welch
Softw. Pract. Exp.4
1998 Coordinated Exception Handling in Distributed Object Systems: From Model to System Implementation
abstract
Exception handling in concurrent and distributed programs is a difficult task though it is often necessary. In many cases traditional exception mechanisms for sequential programs are no longer appropriate. One major difficulty is that the process of handling an exception may need to involve multiple concurrent components that are cooperating in pursuit of some global goal. Another complication is that several exceptions may be raised concurrently in different nodes of a distributed environment. Existing proposals and actual concurrent languages either ignore these difficulties or only cope with a limited form of them. The paper attempts a general solution, developed especially for distributed object systems, starting from a conceptual model, together with algorithms for coordinating concurrent components and resolving multiple exceptions, through to an actual system implementation. An industrial production cell is chosen as a case study to demonstrate the usefulness of the proposed model and algorithms. A system that supports coordinated atomic actions and exception resolution is implemented in distributed Ada 95 and examined through several performance-related experiments.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
ICDCS3
1998 Exception Handling in Object-Oriented Real-Time Distributed Systems
abstract
Exception handling in a complex concurrent and distributed system (e.g. one involving cooperating rather than just competing activities) is often a necessary, but a very difficult, task. No widely accepted models or approaches exist in this area. The object-oriented paradigm, for all its structuring benefits, and real-time requirements each add further difficulties to the design and implementation of exception handling in such systems. In this paper, we develop a general structuring framework based on the coordinated atomic (CA) action concept for handling exceptions in an object-oriented distributed system, in which exceptions in both the value and the time domain are taken into account. In particular, we attempt to attack several difficult problems related to real-time system design and error recovery, including action-level timing constraints, time-triggered CA actions, and time-dependent exception handling. The proposed framework is then demonstrated and assessed using an industrial real-time application-the Production Cell III case study.
Alexander B. Romanovsky, Jie Xu 0007, Brian Randell
ISORC3
1998 Protecting IT Systems from Cyber Crime
abstract
Large-scale commercial, industrial and financial operations are becoming ever more interdependent, and ever more dependent on IT. At the same time, the rapidly growing interconnectivity of IT systems, and the convergence of their technology towards industry-standard hardware and software components and sub-systems, renders these IT systems increasingly vulnerable to malicious attack. This paper is aimed particularly at readers concerned with major systems employed in medium to large commercial or industrial enterprises. It examines the nature and significance of the various potential attacks, and surveys the defence options available. It concludes that IT owners need to think of the threat in more global terms, and to give a new focus and priority to their defence. Prompt action can ensure a major improvement in IT resilience at a modest marginal cost, both in terms of finance and in terms of normal IT operation.
R. Benjamin, B. Gladman, Brian Randell
Comput. J.3
1997 Implementation of blocking coordinated atomic actions based on forward error recovery
Alexander B. Romanovsky, Brian Randell, Robert J. Stroud, Jie Xu 0007, Avelino Francisco Zorzo
J. Syst. Archit.2
1996 Exception Handling and Resolution in Distributed Object-oriented Systems
abstract
We address the problem of how to handle exceptions in distributed object-oriented systems. In a distributed computing environment exceptions may be raised simultaneously and thus need to be treated in a coordinated manner. We take two kinds of concurrency into account: 1) several objects are designed collectively and invoked concurrently to achieve a global goal, and 2) concurrent objects or object groups that are designed independently compete for the same system resources. We propose a new distributed algorithm for resolving concurrent exceptions and show that the algorithm works correctly even in complex nested situations, and is an improvement over previous proposals in that it requires only O(N/sup 2/) messages, and is fully object-oriented.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
ICDCS3
1996 Roll-forward error recovery in embedded real-time systems
abstract
Roll-forward checkpointing schemes are developed in order to avoid rollback in the presence of independent faults and to increase the possibility that a task completes within a tight deadline. However, despite of the adoption of roll-forward recovery, these schemes are not necessarily appropriate for time-critical applications because interactions with the external environment and communications between processes must be deferred during checkpoint validation steps (typically, two checkpoint intervals) until the fault-free processors are identified. The deadlines on providing services may thus be violated. In this paper we present and discuss two alternative roll-forward recovery schemes, especially for time-critical and interaction-intensive applications, that deliver correct, timely results even when checkpoint validation is required.
Jie Xu 0007, Brian Randell
ICPADS2
1993 The Duality of Fault-tolerant System Structures
abstract
Abstract An examination of the structure of fault‐tolerant systems incorporating backward error recovery indicates a partitioning into two broad classes. Two canonical models, each representing a particular class of systems, have been constructed. The first model incorporates objects and actions as the entities for program construction whereas the second model employs communicating processes and conversations. Applications in areas such as office information and banking systems are typically described and built in terms of the first model whereas applications in the area of process control are usually described and built in terms of the second model. The paper claims that the two models are duals of each other and presents arguments and examples to substantiate this claim. It will be shown that the techniques that have been developed within the context of one model turn out to have interesting and hitherto unexplored duals in the other model.
Santosh K. Shrivastava, Luigi V. Mancini, Brian Randell
Softw. Pract. Exp.3
1992 An Object-Oriented View of Fragmented Data Processing for Fault and Intrusion Tolerance in Distributed Systems
Jean-Charles Fabre, Brian Randell
ESORICS2
1987 The architecture of UNIX united
abstract
UNIX United is an architecture for a distributed system based on UNIX. As it is compatible with UNIX at the system call level, any program written for a normal UNIX system can be transparently extended to exploit the richer environment of UNIX United. As it relies on having a UNIX system beneath it, the implementation of UNIX United, called the Newcastle Connection, provides an interesting example of the construction of a very powerful distributed system with only a modicum of effort. A description of the basic semantics of UNIX United is followed by that of the architecture implied by the protocol between components in a UNIX United system, and of a software structure appropriate to the architecture and the protocol.
James P. Black, Lindsay F. Marshall, Brian Randell
Proc. IEEE3
1986 Building Reliable Secure Computing Systems out of Unreliable Insecure Components
abstract
Parallels are drawn between the problems and techniques associated with achieving high reliability, and those associated with the provision of security, in distributed computing systems. Some limitations of the concept of a Trusted Computing Base are discussed, end an alternative approach to thedesign of highly secure computing systems is put forward,based on fault tolerance concepts and techniques.
John E. Dobson, Brian Randell
S&P2
1986 System Design and Structuring
abstract
The task of implementing a large and sophisticated computing system is often unduly costly and time-consuming, with the resulting system exhibiting inadequate performance and reliability, because of excessive system complexity. Such complexity can be reduced significantly by ensuring that the system is constructed out of a well-chosen set of largely independent components, which interact in well-understood ways. However, the task of structuring a system, i.e. of choosing and defining appropriate components, can be very difficult. This paper describes a technique of system structuring which involved distinguishing the functionality which a system is intended to have from other desirable attributes, such as reliability and security, and then using separate components to provide each of these attributes. Various UNIX-based systems which have been implemented at Newcastle are used to illustrate this structuring technique.
Brian Randell
Comput. J.1
1986 Error Recovery in Asynchronous Systems
abstract
A framework for the provision of fault tolerance in asynchronous systems is introduced. The proposal generalizes the form of simple recovery facilities supported by nested atomic actions in which the exception mechanisms only permit backward error recovery. It allows the construction of systems using both forward and backward error recovery and thus allows the exploitation of the complementary benefits of the two schemes. Backward recovery, forward recovery, and normal processing activities can occur concurrently within the organization proposed. Exception handling is generalized to provide a uniform basis for fault tolerance schemes with the atomic action structure. The generalization includes a resolution scheme for concurrently raised exceptions based on an exception tree and an abortion scheme that permits the termination of the internal atomic actions. An automatic resolution mechanism is outlined for exceptions in atomic actions which allows users to separate their recovery schemes from the details of the underlying algorithms.
Roy H. Campbell, Brian Randell
IEEE Trans. Software Eng.2
1985 Interfacing UNIX to Data Communications Networks
abstract
We propose an interface for use from within UNIX1 user programs for communicating over multiple and varied local and wide area networks. This interface aids the design of a distributed application program by hiding the actual communications protocols used over each network, and providing instead simple primitives for sending and receiving (possibly large) datagrams, using a simple standardized network addressing scheme based on apair.
Fabio Panzieri, Brian Randell
IEEE Trans. Software Eng.2
1983 A Distributed Secure System
abstract
We describe the design of a distributed general-purpose computing system that enforces a multilevel security policy. The system is composed of standard UNIX systems and small trustworthy security mechanisms linked together in such a way as to provide a total system which, is not only demonstrably secure, but also highly efficient and cost effective. Despite the heterogeneity of its components, the system as a whole appears to be a single multilevel secure UNIX system, since the fact that it is actually a distributed system is completely hidden from its users and their programs.This is achieved through the use of the "Newcastle Connection", a software subsystem that links together multiple UNIX or UNIX-look-alike systems, without requiring any changes to the source code of either the operating system or any user programs. Construction of a prototype implementation is in progress.
John M. Rushby, Brian Randell
S&P2
1982 The Newcastle Connection or UNIXes of the World Unite!
abstract
Abstract In this paper we describe a software subsystem that can be added to each of a set of physically interconnected UNIX or UNIX look‐alike systems, so as to construct a distributed system which is functionally indistinguishable at both the user and the program level from a conventional single‐processor UNIX system. The techniques used are applicable to a variety and multiplicity of both local and wide area networks, and enable all issues of inter‐processor communication, network protocols, etc., to be hidden. A brief account is given of experience with such a distributed system, which is currently operational on a set of PDPlls connected by a Cambridge Ring. The final sections compare our scheme to various precursor schemes and discuss its potential relevance to other operating systems.
David R. Brownbridge, Lindsay F. Marshall, Brian Randell
Softw. Pract. Exp.3
1981 A Formal Model of Atomicity in Asynchronous Systems
Eike Best, Brian Randell
Acta Informatica2
1979 Software Engineering: As it was in 1968
Brian Randell
ICSE1
1975 System Structure for Software Fault Tolerance
abstract
Presents and discusses the rationale behind a method for structuring complex computing systems by the use of what is termed `recovery blocks,' `conversations,' and `fault-tolerant interfaces.' The aim is to facilitate the provision of dependable error detection and recovery facilities which can cope with errors caused by residual design inadequacies, particularly in the system software, rather than merely the occasional malfunctioning of hardware components.
Brian Randell
IEEE Trans. Software Eng.1
1971 Performance Predictions for Extended Paged Memories
Edward G. Coffman Jr., Brian Randell
Acta Informatica2
1971 Ludgate's Analytical Machine of 1909
abstract
This paper discusses the little known analytical machine, or program-controlled mechanical calculator, designed by Percy E. Ludgate in Ireland during the years 1903 to 1909, and documents the results of a search for information about his life and work.
Brian Randell
Comput. J.1
1967 Dynamic storage allocation systems
abstract
In many recent computer system designs, hardware facilities have been provided for easing the problems of storage allocation. This paper presents a method of characterizing dynamic storage allocation systems, according to the functional capabilities provided, and the underlying techniques used. The basic purpose of the paper is to provide a useful perspective from which the utility of various hardware facilities may be assessed. The paper includes as an appendix, a brief survey of storage allocation facilities in several representative computer systems.
Brian Randell, C. J. Kuehner
SOSP1
1964 Single-Scan Techniques for the Translation of Arithmetic Expressions into ALGOL 60
abstract
The first section of the paper contains a brief description of the well-known technique of using a stack, or pushdown store, to re-order the operators of an arithmetic expression, as defined in ALGOL 60, in order to transform the expression into Reverse Polish parenthesis-free form. It is shown that improvements to this Reverse Polish form can be made quite simply, by extending the use of the stack to include information about the operands of the expression. Firstly, information gained from the declarations of the operands can be used to control the generation of real-integer conversion instructions. Secondly, operators whose operands are numerical constants can be computed during translation, using the partially generated Reverse Polish object program as a second stack.
Brian Randell, L. J. Russell
J. ACM1