VLDB 2026 Research / reviewers in the wild / expert
Algirdas Avizienis
dblp:74/5913
· DBLP profile ↗
36ranked-venue papers
20as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 10 first-authorSecurity and privacy · 8 · 5 first-authorSoftware engineering, systems software and programming languages · 7 · 3 first-authorTheory of computation · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
21 papers |
Distributed systems · 42% Hardware reliability and fault tolerance · 35% Interconnection networks and networks-on-chip · 14% | |
| Network and information security
2 papers |
Privacy and data protection · 58% Malware analysis · 21% Systems and software security · 21% | |
| Databases, data mining, and information retrieval
1 paper |
Database system architecture and tuning · 50% Transaction processing and concurrency control · 50% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
fault tolerance |
0.1 | 4 | 2004 | Basic Concepts and Taxonomy of Dependable and Secure Computing · IEEE Trans. Dependable Secur. Comput. 2004 A fault tolerance approach to computer viruses · S&P 1988 Dependable computing: From concepts to design diversity · Proc. IEEE 1986 |
Privacy and data protection
data confidentiality |
0.0 | 1 | 2004 | Basic Concepts and Taxonomy of Dependable and Secure Computing · IEEE Trans. Dependable Secur. Comput. 2004 |
Hardware reliability and fault tolerance › software fault tolerance
n-version programming |
0.0 | 2 | 1988 | A fault tolerance approach to computer viruses · S&P 1988 The N-Version Approach to Fault-Tolerant Software · IEEE Trans. Software Eng. 1985 |
Interconnection networks and networks-on-chip
network topology |
0.0 | 2 | 1985 | Reliable Loop Topologies for Large Local Computer Networks · IEEE Trans. Computers 1985 Fault Tolerance in Binary Tree Architectures · IEEE Trans. Computers 1984 |
Malware analysis › malware detection
computer virus detection |
0.0 | 1 | 1988 | A fault tolerance approach to computer viruses · S&P 1988 |
Systems and software security › runtime security
control flow monitoring |
0.0 | 1 | 1988 | A fault tolerance approach to computer viruses · S&P 1988 |
Hardware reliability and fault tolerance
error detection and correction |
0.0 | 3 | 1984 | On the Effectiveness of Fault-Tolerance Techniques in Parallel Associative Database Processors · ICDE 1984 Arithmetic Algorithms for Error-Coded Operands · IEEE Trans. Computers 1973 Arithmetic Error Codes: Cost and Effectiveness Studies for Application in Digital System Design · IEEE Trans. Computers 1971 |
Hardware reliability and fault tolerance › software fault tolerance
design diversity |
0.0 | 1 | 1986 | Dependable computing: From concepts to design diversity · Proc. IEEE 1986 |
Routing and switching › routing
distributed routing |
0.0 | 1 | 1985 | Reliable Loop Topologies for Large Local Computer Networks · IEEE Trans. Computers 1985 |
Interconnection networks and networks-on-chip › network topology › loop networks
double-loop network |
0.0 | 1 | 1985 | Reliable Loop Topologies for Large Local Computer Networks · IEEE Trans. Computers 1985 |
Interconnection networks and networks-on-chip › network topology
loop networks |
0.0 | 1 | 1985 | Reliable Loop Topologies for Large Local Computer Networks · IEEE Trans. Computers 1985 |
Hardware reliability and fault tolerance
software fault tolerance |
0.0 | 1 | 1985 | The N-Version Approach to Fault-Tolerant Software · IEEE Trans. Software Eng. 1985 |
Interconnection networks and networks-on-chip › network topology › tree networks
binary tree architecture |
0.0 | 1 | 1984 | Fault Tolerance in Binary Tree Architectures · IEEE Trans. Computers 1984 |
Hardware reliability and fault tolerance
fault tolerance techniques |
0.0 | 1 | 1984 | On the Effectiveness of Fault-Tolerance Techniques in Parallel Associative Database Processors · ICDE 1984 |
Hardware reliability and fault tolerance
fault-tolerant architecture |
0.0 | 1 | 1984 | Fault Tolerance in Binary Tree Architectures · IEEE Trans. Computers 1984 |
Processor architecture and microarchitecture › SIMD
parallel associative processor |
0.0 | 1 | 1984 | On the Effectiveness of Fault-Tolerance Techniques in Parallel Associative Database Processors · ICDE 1984 |
Interconnection networks and networks-on-chip › network topology
tree networks |
0.0 | 1 | 1984 | Fault Tolerance in Binary Tree Architectures · IEEE Trans. Computers 1984 |
Hardware reliability and fault tolerance
reliability modeling |
0.0 | 2 | 1980 | A Unified Reliability Model for Ault-Tolerant Computers · IEEE Trans. Computers 1980 Fault-Tolerant Systems · IEEE Trans. Computers 1976 |
Database system architecture and tuning
database machine |
0.0 | 1 | 1983 | Performance of Recovery Architectures in Parallel Associative Database Processors · ACM Trans. Database Syst. 1983 |
Transaction processing and concurrency control
recovery |
0.0 | 1 | 1983 | Performance of Recovery Architectures in Parallel Associative Database Processors · ACM Trans. Database Syst. 1983 |
Hardware reliability and fault tolerance
reliability analysis |
0.0 | 2 | 1984 | A Unified Reliability Model for Ault-Tolerant Computers · IEEE Trans. Computers 1980 Fault Tolerance in Binary Tree Architectures · IEEE Trans. Computers 1984 |
Hardware reliability and fault tolerance
arithmetic codes |
0.0 | 3 | 1978 | Detection of Storage Errors in Mass Memories Using Low-Cost Arithmetic Error Codes · IEEE Trans. Computers 1978 Arithmetic Algorithms for Error-Coded Operands · IEEE Trans. Computers 1973 Arithmetic Error Codes: Cost and Effectiveness Studies for Application in Digital System Design · IEEE Trans. Computers 1971 |
Hardware reliability and fault tolerance › fault-tolerant architecture
fault-tolerant computer |
0.0 | 1 | 1980 | A Unified Reliability Model for Ault-Tolerant Computers · IEEE Trans. Computers 1980 |
Hardware reliability and fault tolerance › soft errors
soft error detection |
0.0 | 1 | 1988 | A fault tolerance approach to computer viruses · S&P 1988 |
Hardware reliability and fault tolerance
error detection |
0.0 | 1 | 1978 | Detection of Storage Errors in Mass Memories Using Low-Cost Arithmetic Error Codes · IEEE Trans. Computers 1978 |
Processor architecture and microarchitecture
computer arithmetic |
0.0 | 3 | 1973 | Arithmetic Algorithms for Error-Coded Operands · IEEE Trans. Computers 1973 A Universal Arithmetic Building Element (ABE) and Design Methods for Arithmetic Processors · IEEE Trans. Computers 1970 Signed-Digit Numbe Representations for Fast Parallel Arithmetic · IRE Trans. Electron. Comput. 1961 |
Distributed systems › fault tolerance › resilience
graceful degradation |
0.0 | 1 | 1976 | A Design Study of a Shared Resource Computing System · ISCA 1976 |
Hardware reliability and fault tolerance
network fault tolerance |
0.0 | 1 | 1985 | Reliable Loop Topologies for Large Local Computer Networks · IEEE Trans. Computers 1985 |
Parallel and multicore computing
parallel architecture |
0.0 | 1 | 1976 | A Design Study of a Shared Resource Computing System · ISCA 1976 |
Integrated circuit design
digital circuit design |
0.0 | 2 | 1971 | Arithmetic Error Codes: Cost and Effectiveness Studies for Application in Digital System Design · IEEE Trans. Computers 1971 A Universal Arithmetic Building Element (ABE) and Design Methods for Arithmetic Processors · IEEE Trans. Computers 1970 |
Methods — techniques the papers use, named apart from their topics
taxonomy · 0.1definitions · 0.1reliability analysis · 0.0program flow monitoring · 0.0n-version programming · 0.0throughput analysis · 0.0fault removal · 0.0fault forecasting · 0.0fault avoidance · 0.0workload modeling · 0.0queueing analysis · 0.0distributed supervisor · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | The architecture of a resilience infrastructure for computing and communication systemsabstractThe resilience infrastructure is a physically and functionally separate add-on to a “Client” computing and/or communication system that provides resilience to the Client system. This short paper summarizes the main features of the architecture of a resilience infrastructure. Algirdas Avizienis |
DSN | 1 |
| 2008 | On representing knowledge in the dependability domain: a panel discussionabstractThe objective of the panel is to discuss the urgent need, the means, the progress, the obstacles, and the challenges in creating a structured representation of the contents of the large and very rapidly increasing set of documents that represent knowledge in the technical domain of dependability. Algirdas Avizienis |
DSN | 1 |
| 2004 | Panel Summary StatementsabstractNext generation space-based systems will necessitate onboard high performance computing, which is the key to enabling spacecraft autonomy, onboard sensor/science data processing, and multi-spacecraft interactive-cooperative robotics. For example, while conceptually straightforward, “Internet in the sky” communications will require multiple GOPS (giga-operation per second) to perform high-speed routing and protocoltranslation at real-time rates. The viability of high performance space computing stems from the advent of new enabling technologies. Such enabling technologies encompass wireless network communications, non-semiconductor-based (e.g., magnetic, carbon nanotube, ferroelectric and MEMS) components, and deep submicron semiconductors. In addition, it is anticipated that high performance space computing will leads to 1) more extensive use of COTS (commercially-of-the-shelf) products and standards, and 2) the emergence of larger, more complex embedded software running on multi-threaded, file-oriented operating systems. Accordingly, high performance space computing will bring space system dependability concepts and challenges into new, more sophisticated settings. Moreover, as space-based systems are rapidly becoming pervasive and crucial to our worldwide infrastructure, high performance space computing will further increase our reliance on space system operations. Hence, we are reaching the point where failures of space-based systems will have far-reaching and potentially catastrophic consequences in such areas as hazardous weather prediction, communications, finance, aircraft control, military operations, homeland security, and disaster relief and recovery. With the above motivation, this panel brings distinguished researchers and practitioners with wide ranging expertise in space systems and applications to a discussion. The thrust is to foster debating, exchanging, and integrating opinions and solutions for dependable high performance space computing. We particularly solicit different views from various perspectives on the following issues: • What are the critical dependability issues in high performance space computing? • What types of fault tolerance strategies and capabilities that are currently being developed for terrestrial high performance computing systems will be applicable to high performance computing in space? • Among the various means (e.g., V&V, fault tolerance) for dependable high performance space computing, which should we give higher priority? • How do we define fault models and conduct benchmarking to predict COTS performance degradation in the presence of faults? Raphael R. Some, Algirdas Avizienis, Jiri Gaisler, Hirokazu Ihara, Shubu Mukherjee, Neeraj Suri |
PRDC | 2 |
| 2004 | Basic Concepts and Taxonomy of Dependable and Secure ComputingabstractThis paper gives the main definitions relating to dependability, a generic concept including a special case of such attributes as reliability, availability, safety, integrity, maintainability, etc. Security brings in concerns for confidentiality, in addition to availability and integrity. Basic definitions are given first. They are then commented upon, and supplemented by additional definitions, which address the threats to dependability and security (faults, errors, failures), their attributes, and the means for their achievement (fault prevention, fault tolerance, fault removal, fault forecasting). The aim is to explicate a set of general concepts, of relevance across a wide range of situations and, therefore, helping communication and cooperation among a number of scientific and technical communities, including ones that are concentrating on particular types of system, of system failures, or of causes of system failures. Algirdas Avizienis, Jean-Claude Laprie, Brian Randell, Carl E. Landwehr |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2000 | A Fault Tolerance Infrastructure for Dependable Computing with High-Performance COTS ComponentsabstractThe failure rates of current COTS processors have dropped to 100 FITs (failures per 10/sup 9/ hours), indicating a potential MTTF of over 1100 years. However our recent study of Intel P6 family processors has shown that they have very limited error detection and recovery capabilities and contain numerous design faults ("errata"). Other limitations are susceptibility to transient faults and uncertainty about "wearout" that could increase the failure rate in time. Because of these limitations, an external fault tolerance infrastructure is needed to assure the dependability of a system with such COTS components. The paper describes a fault-tolerant "infrastructure" system of fault tolerance functions that makes possible the use of low-coverage COTS processors in a fault-tolerant, self-repairing system. The custom hardware supports transient recovery design fault tolerance, and self-repair by scaring and replacement. Fault tolerance functions are implemented by four types of hardware are processors of low complexity that are fault-tolerant. High error detection coverage, including design faults, is attained by diversity and replication. Algirdas Avizienis |
DSN | 1 |
| 2000 | Assessment of the Applicability of COTS Microprocessors in High-Confidence Computing Systems: A Case StudyabstractCommercial-off-the-shelf (COTS) components are increasingly used in building high-confidence systems to assure their dependability in an affordable way. The effectiveness of such a COTS-based design critically depends on the design of the COTS components. Their dependability attributes need a thorough understanding and rigorous assessment yet has received relatively little attention. The research presented in this paper investigates the error detection and recovery features of contemporary high-performance COTS microprocessors. A method of assessment is proposed and two state-of-the-art microprocessors are studied and compared. Yutao He, Algirdas Avizienis |
DSN | 2 |
| 1996 | Systematic Design of Fault-Tolerant Computers
Algirdas Avizienis |
SAFECOMP | 1 |
| 1995 | Dependable computing depends on structured fault toleranceabstractFault tolerance is a fundamental technique for the attainment of dependable computing. This paper discusses a general paradigm for the design of fault-tolerant systems and illustrates it by a design paradigm for fault-tolerant software. Algirdas Avizienis |
ISSRE | 1 |
| 1992 | Software diversity metrics and measurementsabstractThe authors define and formalize the concept of software diversity which characterizes N-Version software (NVS) from four different points of view that are designated as structural diversity, fault diversity, tough-spot diversity, and failure diversity. The goals are to find a way to quantify software diversity and to investigate the measurements which can be applied during the life cycle of NVS to gain confidence that operation will be dependable when NVS is actually used. The versions from a six-language N-Version programming project for fault-tolerant flight control software were used in the software diversity measurement.> Michael R. Lyu, Jia-Hong Chen, Algirdas Avizienis |
COMPSAC | 3 |
| 1988 | A fault tolerance approach to computer virusesabstractExtensions of program flow monitors and n-version programming can be combined to provide a solution to the detection and containment of computer viruses. The consequence is that a computer can tolerate both deliberate faults and random physical faults by one common mechanism. Specifically, the technique detects control flow errors due to physical faults as well as the presence of viruses.> Mark K. Joseph, Algirdas Avizienis |
S&P | 2 |
| 1986 | Dependable computing: From concepts to design diversityabstractThis paper is composed of two sections. The first provides a conceptual framework for expressing the attributes of what constitutes dependable and reliable computing: a) the impairments to dependability (faults, errors, and failures), b) the means for dependability (fault avoidance, tolerance, removal, and forecasting), and c) the measures of dependability (reliability, availability, safety). The second section focuses on one of the most challenging problems for dependable computing: coping with design faults. Algirdas Avizienis, Jean-Claude Laprie |
Proc. IEEE | 1 |
| 1985 | Arithmetic algorithms for operands encoded in two-dimensional low-cost arithmetic error codesabstractA generalization of low-cost residue codes into two-dimensional encodings was presented and error detecting and error correcting properties of two dimensional inverse residue codes were discussed previously. This paper presents byte-serial checking, additive inverse (complementation), and addition algorithms for operands encoded in two-dimensional residue and inverse residue codes. Algirdas Avizienis |
IEEE Symposium on Computer Arithmetic | 1 |
| 1985 | Reliable Loop Topologies for Large Local Computer NetworksabstractSingle-loop networks tend to become unreliable when the number of nodes in the network becomes large. Reliability can be improved using double loops. In this paper a highly reliable and efficient double-loop network architecture is proposed and analyzed. This network is based on forward loop backward hop topology, with a loop in the forward direction connecting all the neighboring nodes, and a backward loop connecting nodes that are separated by a distance ⌊√N⌋where N is the number of nodes in the network. It is shown that this topology is optimal, among this class of double-loop networks, in terms of diameter, average hop distance, processing overhead, delay, throughput, and reliability. The paper includes derivation of closed form expressions for diameter and average hop distance, throughput, and number of distinct routes between two farthest nodes. For fault-tolerance study, the effect of node and link failures on the performance of the network is analyzed. A simple distributed routing algorithm for reliable loop network operation is also presented. Cauligi S. Raghavendra, Mario Gerla, Algirdas Avizienis |
IEEE Trans. Computers | 3 |
| 1985 | The N-Version Approach to Fault-Tolerant SoftwareabstractEvolution of the N-version software approach to the tolerance of design faults is reviewed. Principal requirements for the implementation of N-version software are summarized and the DEDIX distributed supervisor and testbed for the execution of N-version software is described. Goals of current research are presented and some potential benefits of the N-version approach are identified. Algirdas Avizienis |
IEEE Trans. Software Eng. | 1 |
| 1984 | An Event-Synchronized System Architecture for Integrated Hardware and Software Fault-Tolerance
Srinivas V. Makam, Algirdas Avizienis |
ICDCS | 2 |
| 1984 | On the Effectiveness of Fault-Tolerance Techniques in Parallel Associative Database ProcessorsabstractFault-tolerance is systematically applied to the organization of a parallel associative (content addressable) database processor. Storage areas are protected by duplication and by error detecting and/or correcting codes. A duplex Checking Unit is introduced to check the reading and writing operations. Individual Area Processors are employed to search the storage areas. They are protected by duplexing or triplexing with or without sparing. Effectiveness of these fault-tolerance techniques is assessed by means of reliability predictions for a 100-storage area system. The UCLA ARIES reliability modeling program is employed to obtain the reliability predictions. Major gains in system reliability are demonstrated, and relative effectiveness of alternate approaches is analyzed. Algirdas Avizienis, Alfonso F. Cardenas, Farid Alavian |
ICDE | 1 |
| 1984 | Fault Tolerance in Binary Tree ArchitecturesabstractBinary tree network architectures are applicable in the design of hierarchical computing systems and in specialized high-performance computers. In this correspondence, the reliability and fault tolerance issues in binary tree architecture with spares are considered. Two different fault-tolerance mechanisms are described and studied, namely: 1) scheme with spares; and 2) scheme with performance degradation. Reliability analysis and estimation of the fault-tolerant binary tree structures are performed using the interactive ARIES 82 program. The discussion is restricted to the topological level, and certain extensions of the schemes are also discussed. Cauligi S. Raghavendra, Algirdas Avizienis, Milos D. Ercegovac |
IEEE Trans. Computers | 2 |
| 1983 | Applications for arithmetic error codes in large, high-performance computersabstractLarge, high-performance computers are too costly to allow full replication for fault detection and error correction in the communication and processing of numerical information. For this reason more cost-effective arithmetic error code applications offer an attractive alternative. Algirdas Avizienis, Cauligi S. Raghavendra |
IEEE Symposium on Computer Arithmetic | 1 |
| 1983 | Frameworks for a Taxonomy of Fault-Tolerance Attributes in Computer SystemsabstractA conceptual framework is presented that relates various aspects of fault-tolerance in the context of system structure and architecture. Such a framework is an essential first step for the construction of a taxonomy of fault-tolerance. Algirdas Avizienis |
ISCA | 1 |
| 1983 | Performance of Recovery Architectures in Parallel Associative Database ProcessorsabstractThe need for robust recovery facilities in modern database management systems is quite well known. Various authors have addressed recovery facilities and specific techniques, but none have delved into the problem of recovery in database machines. In this paper, the types of undesirable events that occur in a database environment are classified and the necessary recovery information, with subsequent actions to recover the correct state of the database, is summarized. A model of the “processor-per-track” class of parallel associative database processor is presented. Three different types of recovery mechanisms that may be considered for parallel associative database processors are identified. For each architecture, both the workload imposed by the recovery mechanisms on the execution of database operations (i.e., retrieve, modify, delete, and insert) and the workload involved in the recovery actions (i.e., rollback, restart, restore, and reconstruct) are analyzed. The performance of the three architectures is quantitatively compared. This comparison is made in terms of the number of extra revolutions of the database area required to process a transaction versus the number of records affected by a transaction. A variety of different design parameters of the database processor, of the database, and of a mix of transaction types (modify, insert, and delete) are considered. A large number of combinations is selected and the effects of the parameters on the extra processing time are identified. Alfonso F. Cardenas, Farid Alavian, Algirdas Avizienis |
ACM Trans. Database Syst. | 3 |
| 1982 | Reliability optimization in the design of distributed systems
Cauligi S. Raghavendra, Mario Gerla, Algirdas Avizienis |
ICDCS | 3 |
| 1982 | Fault-Tolerant Design for VLSI: Effect of Interconnect Requirements on Yield Improvement of VLSI DesignsabstractIn order to take full advantage of VLSI, new design methods are necessary to improve the yield and testability. Designs which incorporate redundancy to improve the yields of high density memory chips are well known. The goal of this paper is to motivate the extension of this technique to other types of VLSI logic circuits. The benefits and the limitations of on-chip modularization and the use of spare elements are presented, and significant yield improvements are shown to be possible. Tülin Erdim Mangir, Algirdas Avizienis |
IEEE Trans. Computers | 2 |
| 1980 | A Unified Reliability Model for Ault-Tolerant ComputersabstractThe diversified nature of fault-tolerant computers led to the development of a multiplicity of reliability models which are seemingly unrelated to each other. As a result, it becomes difficult to develop automated tools for reliability analysis which are both general and efficient. Thus, the potential of reliability modeling as a practical and useful tool in the design process of fault-tolerant computers has not been fully realized. This paper summarizes the results of an extended effort to develop a unified approach to reliability modeling of fault-tolerant computers which strikes a good compromise between generality and practicality. The unified model developed encompasses repairable and nonrepairable systems and models, transient as well as permanent faults, and their recovery. Based on the unified model, a powerful and efficient reliability estimation program ARIES has been developed. Ying W. Ng, Algirdas Avizienis |
IEEE Trans. Computers | 2 |
| 1978 | A modified bi-imaginary number systemabstractIn this paper the properties of p-imaginary number systems are reviewed and a modified bi-imaginary number system is introduced as a special case with p = 2. Major properties, including conversion of integer and floating point operands represented in a radix +p system, range, sign and zero tests, and shifting are discussed. The ability to represent the operands as vectors of radix −2 digits suggests advantages in implementing machine-usable arithmetic algorithms. Arunas G. Slekys, Algirdas Avizienis |
IEEE Symposium on Computer Arithmetic | 2 |
| 1978 | Detection of Storage Errors in Mass Memories Using Low-Cost Arithmetic Error CodesabstractArithmetic error codes constitute a class of error codes that are preserved during most arithmetic operations. Effectiveness studies for arithmetic error codes have shown their value for concurrent detection of faults in arithmetic processors, data transmission subsystems, and main storage units in fault-tolerant computers. In this paper, it is shown that the same class of codes is also quite effective for detecting storage errors in both shift-register and magnetic-recording mass memories. Some of the results are more general and deal with properties of arithmetic error codes in detecting unidirectional failures. For example, it is shown that a low-cost arithmetic error code with check modulus A = 2N - 1 can detect any unidirectional failure which affects fewer than N bits. The use of arithmetic error codes for checking of mass memories is further justified since it eliminates the need for hard-core or self-checking code translators and reduces the number of different types of code checkers required. Behrooz Parhami, Algirdas Avizienis |
IEEE Trans. Computers | 2 |
| 1976 | A Design Study of a Shared Resource Computing SystemabstractThe motivations for the design study of a modular, shared resource computing system are given by discussing fault-tolerance and resource utilization issues in parallel processing architectures. A design is presented which employs an array of pipelined arithmetic processors to perform array operations. The design provides for fault-tolerance (“graceful degradation”) capability and is efficient in using main memory bandwidth. Various architectural tradeoffs of the design are discussed. Some results of simulations used for the verification of design decisions are also reported. Alexander Thomasian, Algirdas Avizienis |
ISCA | 2 |
| 1976 | Fault-Tolerant SystemsabstractBasic concepts, motivation, and techniques of fault tolerance are discussed in this paper. The topics include fault classification, redundancy techniques, reliability modeling and prediction, examples of fault-tolerant computers, and some approaches to the problem of tolerating design faults. Algirdas Avizienis |
IEEE Trans. Computers | 1 |
| 1976 | Comments on "Fault Folding for Irredundant and Redundant Combinational Circuits"
Ying W. Ng, Algirdas Avizienis |
IEEE Trans. Computers | 2 |
| 1975 | Redundancy in number representations as an aspect of computational complexity of arithmetic functionsabstractIntroduction Recent research has led to the derivation of bounds for the time required to perform arithmetic operations by means of logical elements with a limited number of inputs [1]–[4]. The model of a (d, r) logical circuit C employed in these studies consists of a set of (d, r) logical elements and a rule of interconnection with designated sets of input and output lines. The (d, r) logical element has r input lines and one output line; these lines can assume one of d distinct states. The (d, r) logical element has a unit time delay; that is, the state of the output line at the time t+1 is a function of the states of the input lines at time t. Algirdas Avizienis |
IEEE Symposium on Computer Arithmetic | 1 |
| 1973 | Design of Fault-Tolerant Associative ProcessorsabstractRecent advances in computer technology have made the design of large and very flexible associative processors possible. Such systems are extremely complex and must be adequately protected against failures if they are to be used in critical application areas such as air traffic control or for performing control functions in fault-tolerant computers. This paper summarizes the results of a study which has indicated the techniques that are applicable in the design of fault tolerant associative processors. Associative processors are divided into four classes of fully parallel, bit-serial, word-serial, and block-oriented systems. A technique for modularizing the design of an associative processor is given. The detection of errors within modules is discussed for the four classes mentioned above. Several schemes for reconfiguration are discussed which allow us to establish an appropriate inter-communication pattern after replacing the faulty module by a spare. The design of a fault-tolerant associative processor, which uses some of the techniques discussed previously, is presented. Behrooz Parhami, Algirdas Avizienis |
ISCA | 2 |
| 1973 | Arithmetic Algorithms for Error-Coded OperandsabstractA set of arithmetic algorithms is described for operands that are encoded in the ``AN'' error-detecting code with the low-cost check modulus A = 2a- 1. The set includes addition additive inverse (complementation), multiplication, division, roundoff, and two auxiliary algorithms: ``multiply by 2a- 1,'' and ``divide by 2a- 1.'' The design of a serial radix-16 processor is presented in which these algorithms are implemented for the low-cost AN code with A = 15. This processor has been constructed for the Jet Propulsion Laboratory STAR computer. The adaptation of ``two's complement'' arithmetic for an inverse-residue code is also described. Algirdas Avizienis |
IEEE Trans. Computers | 1 |
| 1971 | Arithmetic Error Codes: Cost and Effectiveness Studies for Application in Digital System DesignabstractThe application of error-detecting or error-correcting codes in digital computer design requires studies of cost and effectiveness trade-offs to supplement the knowledge of their theoretical properties. General criteria for cost and effectiveness studies of error codes are developed, and results are presented for arithmetic error codes with the low-cost check modulus 2a-1. Both separate (residue) and nonseparate (AN) codes are considered. The class of multiple arithmetic error codes is developed as an extension of low-cost single codes. Algirdas Avizienis |
IEEE Trans. Computers | 1 |
| 1971 | The STAR (Self-Testing And Repairing) Computer: An Investigation of the Theory and Practice of Fault-Tolerant Computer DesignabstractThis paper presents the results obtained in a continuing investigation of fault-tolerant computing which is being conducted at the Jet Propulsion Laboratory. Initial studies led to the decision to design and construct an experimental computer with dynamic (standby) redundancy, including replaceable subsystems and a program rollback provision to eliminate transient errors. This system, called the STAR computer, began operation in 1969. The following aspects of the STAR system are described: architecture, reliability analysis, software, automatic maintenance of peripheral systems, and adaptation to serve as the central computer of an outerplanet exploration spacecraft. Algirdas Avizienis, George C. Gilley, Francis P. Mathur, David A. Rennels, John A. Rohr, David K. Rubin |
IEEE Trans. Computers | 1 |
| 1970 | A Universal Arithmetic Building Element (ABE) and Design Methods for Arithmetic ProcessorsabstractThe advent of large-scale integration of logic circuits requires the definition of digital computer structure in terms of large functional arrays of logic of very few types. This paper describes a single-package arithmetic processor called the arithmetic building element (ABE). The ABE accepts operands in either conventional or signed-digit radix-r representation and produces signed-digit results, which the ABE can reconvert to conventional form. Radix 16 is chosen for illustrations. Arrays of ABE's may be arranged to implement unit- time parallel addition, all-combinational multiplication, and more complex functions which are presently computed by subroutines. To facilitate such arithmetic design, a graph model is developed which permits a translation of the given arithmetical algorithm into an interconnection diagram of ABE's. The design procedure is illustrated by an array for polynomial evaluation. Speed, cost, and roundoff error of the array are considered. A computer program has been written for the automatic translation of the algorithm graph to an interconnection graph, and for the evaluation of the cost and speed for a given polynomial degree and a given precision requirement. Algirdas Avizienis, Chin Tung |
IEEE Trans. Computers | 1 |
| 1969 | On the Problem of Computational Time and Complexity of Arithmetic FunctionsabstractThe time and incremental complexity required to perform two-operand addition using logical circuitry are compared for nonredundant and minimally redundant encodings of the operands. The comparison is extended to multi-operand addition and two-operand multiplication. Algirdas Avizienis |
STOC | 1 |
| 1961 | Signed-Digit Numbe Representations for Fast Parallel ArithmeticabstractThis paper describes a class of number representations which are called signed-digit representations. Signed-digit representations limit carry-propagation to one position to the left during the operations of addition and subtraction in digital computers. Carry-propagation chains are eliminated by the use of redundant representations for the operands. Redundancy in the number representation allows a method of fast addition and subtraction in which each sum (or difference) digit is the function only of the digits in two adjacent digital positions of the operands. The addition time for signed-digit numbers of any length is equal to the addition time for two digits. The paper discusses the properties of signed-digit representations and arithmetic operations with signed-digit numbers: addition, subtraction, multiplication, division and roundoff. A brief discussion of logical design problems for a signed-digit adder concludes the presentation. Algirdas Avizienis |
IRE Trans. Electron. Comput. | 1 |