Lorenzo Strigini

dblp:19/6375 · DBLP profile ↗
← Back
46ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-4246-2866ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 23 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 18 · 1 first-author · 2 since 2021Systems, architecture and hardware · 7 · 1 first-author · 1 since 2021Computer networks · 3
YearPublicationVenuePosition
2024 Impact of Prior Beliefs on Dependability Prediction for a Changed System Using Pre-change Operational Evidence
abstract
Bayesian inference allows integration of heterogeneous evidence for dependability assessment, e.g., evidence of successful operation together with process quality evidence that determines “prior” probability distributions for parameters of interest. How to specify these “priors” is often non-obvious. There is a concern that results - claims about a system’s dependability - could be determined by somewhat arbitrary aspects of the priors. “Robust” Bayesian methods attempt to mitigate this risk. We report observations from applying one such method, “Conservative Bayesian Inference” (CBI), to a practical problem: how to account, in assessing a system after a change of its design or its use, for evidence of successful operation before that change. CBI, by avoiding other potential difficulties, puts in sharper focus the crucial need for clarity in translating prior evidence into mathematically stated “priors”. Using both previously published and new mathematical results, we give examples of how results may change with changes in priors, like different assumptions on parameter values or even different mathematical translations of the same natural language statement. In one example, the dependability that can be claimed is substantially improved by making the mathematical statement of priors model more closely the evidence behind it. In others, assumptions that appear prudent are shown to counterintuitively cause over-optimistic claims. Besides supporting the case for rigour in scrutinising priors, our examples may serve others by flagging types of potential pitfalls in selecting priors.
Robab Aghazadeh Chakherlou, Lorenzo Strigini
QRS2
2023 Editorial: Software Reliability and Dependability Engineering
abstract
As software plays an increasingly important role in our lives, it is essential to maintain its reliability, and generally dependability. Software bugs can cause huge financial losses and dangerous accidents; the safety risks from software are underscored these days to even the non-technical public by the emergence of autonomous software-based systems. Thus, it is important to explore principled approaches to reduce the harm from defects in software, preferably by removing them as early as possible, but also by fault tolerance and by predicting their effects so as to inform mitigation actions.
Zheng Zheng 0001, Lorenzo Strigini, Nuno Antunes, Kishor S. Trivedi
IEEE Trans. Dependable Secur. Comput.2
2022 Authentication for Operators of Critical Medical Devices: A Contribution to Analysis of Design Trade-offs
abstract
Increasingly evident safety risks due to attacks on safety-critical devices are causing new requirements for authentication of these devices’ human operators. These requirements have now extended to medical devices. However, authentication may also introduce new safety risks, reduce usability, cause delays, and/or encourage user behaviors that compromise the very security it should protect. Thus, design of authentication mechanisms needs to take on a holistic approach that considers such interrelationships, and the effects not just of the general method chosen (say, passwords vs. fingerprints), but also of its implementation details. We illustrate this problem on a medical case study. We report early steps in a trade-off analysis that captures interactions between safety, security, usability and performance issues, to assist designers in choosing and tuning viable solutions. A qualitative analysis to narrow down the field of possible solutions is followed by a probabilistic analysis. The analyses highlight non-obvious links between system attributes, especially links due to the complex way humans interact with, and adapt to, such devices. The probabilistic analysis systematically describes risk as a function of the authentication method and its design parameters. We show example results quantifying how some key design parameters produce opposite effects on risk due to accidental and malicious causes, requiring a trade-off: the quantitative model allows the designer to manage this trade-off to achieve an acceptable level of overall risk, taking into account environmental factors like the expected prevalence of certain attack types. Both the qualitative and quantitative approaches aim to help device designers make rational decisions about authentication options and the tuning of their design parameters.
Marwa Gadala, Lorenzo Strigini, Radek Fujdiak
ARES2
2022 Bootstrapping confidence in future safety from past safe operation
abstract
We propose an approach to managing the rolling out of a new system type so as to contain the risk of mishaps in its operation to an acceptable level, while using the evidence of safe operation to support confidence for extending the scale of operation. This cautious approach of “bootstrapping” confidence in the safety of a system is now widely applied to autonomous vehicle (AVs), our example. AVs are subject to extreme safety requirements; a major concern is the inability to give meaningful quantitative assurance of safety of an AV type, to the extent required by society, before it is used extensively. We exploit a previously published approach to achieving more moderate, but useful, assurance, e.g. about low enough probability of causing accidents in a limited period of operation; and demonstrate how this approach supports choice of production/deployment strategies, so as to manage the growth of a fleet of AVs for a given accepted level of risk. Via a formal mathematical description of “confidence bootstrapping”, we show: (1) that it is a substantially sound approach in the right circumstances, and useful for deciding about the early deployment phase for a new system; (2) how much confidence can be rightly derived from such a “cautious deployment” approach, avoiding over-optimism; (3) under which conditions our sound formulas for future confidence are applicable; (4) thus, which analyses of the concrete situations, and/or constraints on practice, are needed in order to enjoy the advantages of provably correct confidence in adequate future safety over a definite time (“confidence horizon”).
Peter Bishop 0001, Andrey Povyakalo, Lorenzo Strigini
ISSRE3
2022 Impact of Machine Learning on Safety Monitors
Francesco Terrosi, Lorenzo Strigini, Andrea Bondavalli
SAFECOMP2
2021 Conservative Confidence Bounds in Safety, from Generalised Claims of Improvement & Statistical Evidence
abstract
“Proven-in-use”, “globally-at-least-equivalent”, “stress-tested”, are concepts that come up in diverse contexts in acceptance, certification or licensing of critical systems. Their common feature is that dependability claims for a system in a certain operational environment are supported, in part, by evidence – viz of successful operation – concerning different, though related, system[s] and/or environment[s], together with an auxiliary argument that the target system/environment offers the same, or improved, safety. We propose a formal probabilistic (Bayesian) organisation for these arguments. Through specific examples of evidence for the “improvement” argument above, we demonstrate scenarios in which formalising such arguments substantially increases confidence in the target system, and show why this is not always the case. Example scenarios concern vehicles and nuclear plants. Besides supporting stronger claims, the mathematical formalisation imposes precise statements of the bases for “improvement” claims: seemingly similar forms of prior beliefs are sometimes revealed to imply substantial differences in the claims they can support.
Kizito Salako, Lorenzo Strigini, Xingyu Zhao 0001
DSN2
2020 Assessing safety-critical systems from operational testing: A study on autonomous vehicles
abstract
Demonstrating high reliability and safety for safety-critical systems (SCSs) remains a hard problem. Diverse evidence needs to be combined in a rigorous way: in particular, results of operational testing with other evidence from design and verification. Growing use of machine learning in SCSs, by precluding most established methods for gaining assurance, makes evidence from operational testing even more important for supporting safety and reliability claims. We revisit the problem of using operational testing to demonstrate high reliability. We use Autonomous Vehicles (AVs) as a current example. AVs are making their debut on public roads: methods for assessing whether an AV is safe enough are urgently needed. We demonstrate how to answer 5 questions that would arise in assessing an AV type, starting with those proposed by a highly-cited study. We apply new theorems extending our Conservative Bayesian Inference (CBI) approach, which exploit the rigour of Bayesian methods while reducing the risk of involuntary misuse associated (we argue) with now-common applications of Bayesian inference; we define additional conditions needed for applying these methods to AVs. Prior knowledge can bring substantial advantages if the AV design allows strong expectations of safety before road testing. We also show how naive attempts at conservative assessment may lead to over-optimism instead; why extrapolating the trend of disengagements (take-overs by human drivers) is not suitable for safety claims; use of knowledge that an AV has moved to a “less stressful” environment. While some reliability targets will remain too high to be practically verifiable, our CBI approach removes a major source of doubt: it allows use of prior knowledge without inducing dangerously optimistic biases. For certain ranges of required reliability and prior beliefs, CBI thus supports feasible, sound arguments. Useful conservative claims can be derived from limited prior knowledge.
Xingyu Zhao 0001, Kizito Salako, Lorenzo Strigini, Valentin Robu, David Flynn
Inf. Softw. Technol.3
2019 Assessing the Safety and Reliability of Autonomous Vehicles from Road Testing
abstract
There is an urgent societal need to assess whether autonomous vehicles (AVs) are safe enough. From published quantitative safety and reliability assessments of AVs, we know that, given the goal of predicting very low rates of accidents, road testing alone requires infeasible numbers of miles to be driven. However, previous analyses do not consider any knowledge prior to road testing - knowledge which could bring substantial advantages if the AV design allows strong expectations of safety before road testing. We present the advantages of a new variant of Conservative Bayesian Inference (CBI), which uses prior knowledge while avoiding optimistic biases. We then study the trend of disengagements (take-overs by human drivers) by applying Software Reliability Growth Models (SRGMs) to data from Waymo's public road testing over 51 months, in view of the practice of software updates during this testing. Our approach is to not trust any specific SRGM, but to assess forecast accuracy and then improve forecasts. We show that, coupled with accuracy assessment and recalibration techniques, SRGMs could be a valuable test planning aid.
Xingyu Zhao 0001, Valentin Robu, David Flynn, Kizito Salako, Lorenzo Strigini
ISSRE5
2014 Estimating Worst Case Failure Dependency with Partial Knowledge of the Difficulty Function
Peter Bishop 0001, Lorenzo Strigini
SAFECOMP2
2014 When Does "Diversity" in Development Reduce Common Failures? Insights from Probabilistic Modeling
abstract
Fault tolerance via diverse redundancy, with multiple "versions" of a system in a redundant configuration, is an attractive defence against design faults. To reduce the probability of common failures, development and procurement practices pursue "diversity" between the ways the different versions are developed. But difficult questions remain open about which practices are more effective to this aim. About these questions, probabilistic models have helped by exposing fallacies in "common sense" judgements. However, most make very restrictive assumptions. They model well scenarios in which diverse versions are developed in rigorous isolation from each other: A condition that many think desirable, but is unlikely in practice. We extend these models to cover nonindependent development processes for diverse versions. This gives us a rigorous way of framing claims and open questions about how best to pursue diversity, and about the effects - negative and positive - of commonalities between developments, from specification corrections to the choice of test cases. We obtain three theorems that, under specific scenarios, identify preferences between alternative ways of seeking diversity. We also discuss nonintuitive issues, including how expected system reliability may be improved by creating intentional "negative" dependences between the developments of different versions.
Kizito Salako, Lorenzo Strigini
IEEE Trans. Dependable Secur. Comput.2
2013 Software Fault-Freeness and Reliability Predictions
Lorenzo Strigini, Andrey Povyakalo
SAFECOMP1
2012 An Empirical Study of the Effectiveness of "Forcing" Diversity Based on a Large Population of Diverse Programs
abstract
Use of diverse software components is a viable defence against common-mode failures in redundant software-based systems. Various forms of "Diversity-Seeking Decisions" ("DSDs") can be applied to the process of developing, or procuring, redundant components, to improve the chances of the resulting components not failing on the same demands. An open question is how effective these decisions, and their combinations, are for achieving large enough reliability gains. Using a large population of software programs, we studied experimentally the effectiveness of specific "DSDs" (and their combinations) mandating differences between redundant components. Some of these combinations produced much better improvements in system probability of failure per demand (PFD) than "uncontrolled" diversity did. Yet, our findings suggest that the gains from such "DSDs" vary significantly between them and between the application problems studied. The relationship between DSDs and system PFD is complex and does not allow for simple universal rules (e.g. "the more diversity the better") to apply.
Peter T. Popov, Vladimir Stankovic 0002, Lorenzo Strigini
ISSRE3
2010 Assessing Asymmetric Fault-Tolerant Software
abstract
The most popular forms of fault tolerance against design faults use "asymmetric" architectures in which a "primary" part performs the computation and a "secondary" part is in charge of detecting errors and performing some kind of error processing and recovery. In contrast, the most studied forms of software fault tolerance are "symmetric" ones, e.g. N-version programming. The latter are often controversial, the former are not. We discuss how to assess the dependability gains achieved by these methods. Substantial difficulties have been shown to exist for symmetric schemes, but we show that the same difficulties affect asymmetric schemes. Indeed, the latter present somewhat subtler problems. In both cases, to predict the dependability of the fault-tolerant system it is not enough to know the dependability of the individual components. We extend to asymmetric architectures the style of probabilistic modeling that has been useful for describing the dependability of "symmetric" architectures, to highlight factors that complicate the assessment. In the light of these models, we finally discuss fault injection approaches to estimating coverage factors. We highlight the limits of what can be predicted and some useful research directions towards clarifying and extending the range of situations in which estimates of coverage of fault tolerance mechanisms can be trusted.
Peter T. Popov, Lorenzo Strigini
ISSRE2
2009 Why Are People's Decisions Sometimes Worse with Computer Support?
Eugenio Alberdi, Lorenzo Strigini, Andrey Povyakalo, Peter Ayton
SAFECOMP2
2007 Fault Tolerance via Diversity for Off-the-Shelf Products: A Study with SQL Database Servers
abstract
If an off-the-shelf software product exhibits poor dependability due to design faults, then software fault tolerance is often the only way available to users and system integrators to alleviate the problem. Thanks to low acquisition costs, even using multiple versions of software in a parallel architecture, which is a scheme formerly reserved for few and highly critical applications, may become viable for many applications. We have studied the potential dependability gains from these solutions for off-the-shelf database servers. We based the study on the bug reports available for four off-the-shelf SQL servers plus later releases of two of them. We found that many of these faults cause systematic noncrash failures, which is a category ignored by most studies and standard implementations of fault tolerance for databases. Our observations suggest that diverse redundancy would be effective for tolerating design faults in this category of products. Only in very few cases would demands that triggered a bug in one server cause failures in another one, and there were no coincident failures in more than two of the servers. Use of different releases of the same product would also tolerate a significant fraction of the faults. We report our results and discuss their implications, the architectural options available for exploiting them, and the difficulties that they may present.
Ilir Gashi, Peter T. Popov, Lorenzo Strigini
IEEE Trans. Dependable Secur. Comput.3
2006 E-voting: Dependability Requirements and Design for Dependability
abstract
Elections are increasingly dependent on computers and telecommunication systems. Such "e-voting" schemes create socio-technical systems (combinations of technology and human organisations) that are complex and critical, as the future of nations depends on their proper operation. Thus heated debate surrounds their adoption and the possible methods for making them demonstrably dependable. We discuss the dependability requirements for such systems, and the design issues in ensuring their satisfaction, with reference to a recent proposal that uses cryptography for fault tolerance, in order to avoid some of the perceived dangers of electronic voting. Our treatment highlights the need for considering the whole socio-technical system, and for integrating security and fault tolerance viewpoints.
Jeremy W. Bryans, Bev Littlewood, Peter Y. A. Ryan, Lorenzo Strigini
ARES4
2005 On the Effectiveness of Run-Time Checks
Meine van der Meulen, Lorenzo Strigini, Miguel A. Revilla
SAFECOMP2
2004 Fault Diversity among Off-The-Shelf SQL Database Servers
abstract
Fault tolerance is often the only viable way of obtaining the required system dependability from systems built out of "off-the-shelf" (OTS) products. We have studied a sample of bug reports from four off-the-shelf SQL servers so as to estimate the possible advantages of software fault tolerance - in the form of modular redundancy with diversity - in complex off-the-shelf software. We checked whether these bugs would cause coincident failures in more than one of the servers. We found that very few bugs affected two of the four servers, and none caused failures in more than two. We also found that only four of these bugs would cause identical, undetectable failures in two servers. Therefore, a fault-tolerant server, built with diverse off-the-shelf servers, seems to have a good chance of delivering improvements in availability and failure rates compared with the individual off-the-shelf servers or their replicated, nondiverse configurations.
Ilir Gashi, Peter T. Popov, Lorenzo Strigini
DSN3
2004 Workshop on Interdisciplinary Approaches to Achieving and Analysing System Dependability
Michael D. Harrison, Lorenzo Strigini
DSN2
2004 Redundancy and Diversity in Security
Bev Littlewood, Lorenzo Strigini
ESORICS2
2003 Human-Machine Diversity in the Use of Computerised Advisory Systems: A Case Study
abstract
Computer-based advisory systems form with their users composite, human-machine systems. Redundancy and diversity between the human and the machine are often important for the dependability of such systems. We describe a case study on assessing failure probabilities for the analysis of X-ray films for detecting cancer, performed by a person assisted by a computer-based tool. Differently from most approaches to human reliability assessment, we focus on the effects of failure diversity – or correlation – between humans and machines. We illustrate some of the modelling and prediction problems, especially those caused by the presence of the human component. We show two alternative models, with their pros and cons, and illustrate, via numerical examples and analytically, some interesting and non-intuitive answers to questions about reliability assessment and design choices for human-computer systems. Copyright of the authors, 2002 This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. 11
Lorenzo Strigini, Andrey Povyakalo, Eugenio Alberdi
DSN1
2003 Estimating Bounds on the Reliability of Diverse Systems
abstract
We address the difficult problem of estimating the reliability of multiple-version software. The central issue is the degree of statistical dependence between failures of diverse versions. Previously published models of failure dependence described what behavior could be expected "on average" from a pair of "independently generated" versions. We focus instead on predictions using specific information about a given pair of versions. The concept of "variation of difficulty" between situations to which software may be subject is central to the previous models cited, and it turns out to be central for our question as well. We provide new understanding of various alternative imprecise estimates of system reliability and some results of practical use, especially with diverse systems assembled from pre-existing (e.g., "off-the-shelf") subsystems. System designers, users, and regulators need useful bounds on the probability of system failure. We discuss how to use reliability data about the individual diverse versions to obtain upper bounds and other useful information for decision making. These bounds are greatly affected by how the versions' probabilities of failure vary between subdomains of the demand space or between operating regimes-it is even possible in some cases to demonstrate, before operation, upper bounds that are very close to the true probability of failure of the system-and by the level of detail with which these variations are documented in the data.
Peter T. Popov, Lorenzo Strigini, John H. R. May, Silke Kuball
IEEE Trans. Software Eng.2
2001 The Reliability of Diverse Systems: A Contribution Using Modelling of the Fault Creation Process
abstract
Design diversity is a defence against design faults causing common-mode failure in redundant systems, but we badly lack knowledge about how much reliability it will buy in practice, and thus about its cost-effectiveness, the situations in which it is an appropriate solution and how it should be taken into account by assessors and safety regulators. Both current practice and the scientific debate about design diversity depend largely on intuition. More formal probabilistic reasoning would facilitate critical discussion and empirical validation of any predictions: to this aim, we propose a model of the generation of faults and failures in two separately-developed program versions. We show results on: (i) what degree of reliability improvement an assessor can reliably expect from diversity; and (ii) how this reliability improvement may change with higher-quality development processes. We discuss the practical relevance of these results and the degree to which they can be trusted.
Peter T. Popov, Lorenzo Strigini
DSN2
2000 Software reliability (tutorial session): basic concepts and assessment methods
abstract
No abstract available.
Bev Littlewood, Lorenzo Strigini
ICSE2
2000 Fault tolerance via diversity against design faults (tutorial session): design principles and reliability assessment
abstract
Research results indicate that (as usual in software engineering) these question can only be answered with reference to each specific application context and that diversity is no “silver bullet”. But diversity is an attractive option, made more interesting by current trends like the preference for COTS items, and it is important for practitioners to go beyond the summary opinions and misunderstanding that surround it.
Bev Littlewood, Lorenzo Strigini
ICSE2
2000 Assessment of the Reliability of Fault-Tolerant Software: A Bayesian Approach
Bev Littlewood, Peter T. Popov, Lorenzo Strigini
SAFECOMP3
2000 The meaning and role of value in scheduling flexible real-time systems
Alan Burns 0001, Divya Prasad, Andrea Bondavalli, Felicita Di Giandomenico, Krithi Ramamritham, John A. Stankovic, Lorenzo Strigini
J. Syst. Archit.7
2000 Modeling the Effects of Combining Diverse Software Fault Detection Techniques
abstract
Considers what happens when several different fault-finding techniques are used together. The effectiveness of such multi-technique approaches depends upon a quite subtle interplay between their individual efficacies. The modeling tool we use to study this problem is closely related to earlier work on software design diversity which showed that it would be unreasonable even to expect software versions that were developed truly independently to fail independently of one another. The key idea was a "difficulty function" over the input space. Later work extended these ideas to introduce a notion of "forced" diversity. In this paper, we show that many of these results for design diversity have counterparts in diverse fault detection in a single software version. We define measures of fault-finding effectiveness and diversity, and show how these might be used to give guidance for the optimal application of different fault-finding procedures to a particular program. The effects on reliability of repeated applications of a particular fault-finding procedure are not statistically independent; such an incorrect assumption of independence will always give results that are too optimistic. For diverse fault-finding procedures, it is possible for effectiveness to be even greater than it would be under an assumption of statistical independence. Diversity of fault-finding procedures is a good thing and should be applied as widely as possible. The model is illustrated using some data from an experimental investigation into diverse fault-finding on a railway signalling application.
Bev Littlewood, Peter T. Popov, Lorenzo Strigini, Nick Shryane
IEEE Trans. Software Eng.3
1999 Choosing Effective Methods for Design Diversity - How to Progress from Intuition to Science
Peter T. Popov, Lorenzo Strigini, Alexander B. Romanovsky
SAFECOMP2
1999 A Contribution to the Evaluation of the Reliability of Iterative-Execution Software
abstract
This paper deals with the reliability of software executed iteratively, as for example in process control applications. The probability of mission survival is evaluated taking account of two characteristics of iterative software: (a) system failure, defined in terms of the behaviour of the software over successive iterations, because the controlled system can usually tolerate short bursts of errors; (b) the probabilistic correlation between successive executions of the software, which is to be expected for various reasons. The paper presents models accounting for these characteristics and evaluates their effects. The interesting case of fault-tolerant software is considered as well. Using the example of a ‘pair-and-spare’ type fault-tolerant scheme, the relationships between different aspects of failure behaviour that are covered by the models developed here, and those used elsewhere for fault-tolerant software, are shown. Copyright © 1999 John Wiley & Sons, Ltd.
Andrea Bondavalli, Silvano Chiaradonna, Felicita Di Giandomenico, Lorenzo Strigini
Softw. Test. Verification Reliab.4
1998 Comparing the effectiveness of testing methods in improving programs: the effect of variations in program quality
abstract
We compare the efficacy of different testing methods for improving the reliability of software. Specifically, we use modelling to compare "operational" testing, in which test cases are chosen according to their probability of occurring in actual use of the software, against "debug" testing methods, in which the testers look for test cases which they consider likely to cause failure, or that satisfy some coverage criterion. We base our comparisons on the reliability reached by the program at the end of testing. Differently from previous studies, we consider the probability distribution of the achieved reliability, and thus the probability of satisfying specific requirements, rather than just the average reliability achieved. We take account of two sources of variation. The variation between the actual test histories that are possible for a given program and a given test method: and the fact that different programs start testing with different faults and initial reliability levels. By necessity, we use very simplified models of reality. Yet, we can show some interesting conclusions with important practical consequences. In general, there are stronger arguments in favor of operational testing than previous studies have shown.
Michele Pizza, Lorenzo Strigini
ISSRE2
1998 A Holistic View on the Dependability of Software-Intensive Systems
Gerald Sonneck, Erwin Schoitsch, Lorenzo Strigini
SAFECOMP3
1998 Assessing the Risk due to Software Faults: Estimates of Failure Rate versus Evidence of Perfection
abstract
In the debate over the assessment of software reliability (or safety), as applied to critical software, two extreme positions can be discerned: the ‘statistical’ position, which requires that the claims of reliability be supported by statistical inference from realistic testing or operation, and the ‘perfectionist’ position, which requires convincing indications that the software is free from defects. These two positions naturally lead to requiring different kinds of supporting evidence, and actually to stating the dependability requirements in different ways, not allowing any direct comparison. There is often confusion about the relationship between statements about software failure rates and about software correctness, and about which evidence can support either kind of statement. This note clarifies the meaning of the two kinds of statement and how they relate to the probability of failure-free operation, and discusses their practical merits, especially for high required reliability or safety. © 1998 John Wiley & Sons, Ltd.
Antonia Bertolino, Lorenzo Strigini
Softw. Test. Verification Reliab.2
1998 Evaluating Testing Methods by Delivered Reliability
abstract
There are two main goals in testing software: (1) to achieve adequate quality (debug testing), where the objective is to probe the software for defects so that these can be removed, and (2) to assess existing quality (operational testing), where the objective is to gain confidence that the software is reliable. Debug methods tend to ignore random selection of test data from an operational profile, while for operational methods this selection is all-important. Debug methods are thought to be good at uncovering defects so that these can be repaired, but having done so they do not provide a technically defensible assessment of the reliability that results. On the other hand, operational methods provide accurate assessment, but may not be as useful for achieving reliability. This paper examines the relationship between the two testing goals, using a probabilistic analysis. We define simple models of programs and their testing, and try to answer the question of how to attain program reliability: is it better to test by probing for defects as in debug testing, or to assess reliability directly as in operational testing? Testing methods are compared in a model where program failures are detected and the software changed to eliminate them. The "better" method delivers higher reliability after all test failures have been eliminated. Special cases are exhibited in which each kind of testing is superior. An analysis of the distribution of the delivered reliability indicates that even simple models have unusual statistical properties, suggesting caution in interpreting theoretical comparisons.
Phyllis G. Frankl, Richard G. Hamlet, Bev Littlewood, Lorenzo Strigini
IEEE Trans. Software Eng.4
1997 Choosing a Testing Method to Deliver Reliability
abstract
Testing methods are compared in a model where program failures are detected and the software changed to eliminate them.The question considered is whether it is better to use tests that seek out failures ("debug testing") or to simulate usage and find failures along the way ("operational testing'')."Better" is measured by the delivered reliability obtained after all test failures have been eliminated.This comparison extends previous work, where the measure was the probability of detecting a failure.The theoretical treatment of the paper is probabilistic and analytical.Revealing special cases are exhibited in which each kind of testing is superior.
Phyllis G. Frankl, Richard G. Hamlet, Bev Littlewood, Lorenzo Strigini
ICSE4
1996 Acceptance Criteria for Critical Software Based on Testability Estimates and Test Results
Antonia Bertolino, Lorenzo Strigini
SAFECOMP2
1996 On Testing Process Control Software for Reliability Assessment: the Effects of Correlation between Successive Failures
abstract
Statistical testing is the main available means for evaluating the reliability of software products. For ‘batch’ software, one can usually model both testing and operation as sequences of statistically independent trials (Bernoulli trials). Statistical inference from test results to reliability predictions is then straightforward. Things change when non-zero correlation is to be expected between the outcomes (failure or success) of successive executions of the software. This is the case, in particular, for most process control software: the inputs to each execution represent measurements on the controlled plant, and therefore follow quasi-continuous trajectories in the input space. For such software, this paper will: (i) show that the Bernoulli-trial model is inappropriate; (ii) argue that the ‘failure rate’ (probability of failure per execution) is no longer an appropriate indicator of software dependability; (iii) discuss the relationships among the statistical parameters describing the failure behaviour of the software; (iv) argue for direct measurement of the parameters of interest.
Lorenzo Strigini
Softw. Test. Verification Reliab.1
1996 On the Use of Testability Measures for Dependability Assessment
abstract
Program "testability" is informally, the probability that a program will fail under test if it contains at least one fault. When a dependability assessment has to be derived from the observation of a series of failure free test executions (a common need for software subject to "ultra high reliability" requirements), measures of testability can-in theory-be used to draw inferences on program correctness. We rigorously investigate the concept of testability and its use in dependability assessment, criticizing, and improving on, previously published results. We give a general descriptive model of program execution and testing, on which the different measures of interest can be defined. We propose a more precise definition of program testability than that given by other authors, and discuss how to increase testing effectiveness without impairing program reliability in operation. We then study the mathematics of using testability to estimate, from test results: the probability of program correctness and the probability of failures. To derive the probability of program correctness, we use a Bayesian inference procedure and argue that this is more useful than deriving a classical "confidence level". We also show that a high testability is not an unconditionally desirable property for a program. In particular, for programs complex enough that they are unlikely to be completely fault free, increasing testability may produce a program which will be less trustworthy, even after successful testing.
Antonia Bertolino, Lorenzo Strigini
IEEE Trans. Software Eng.2
1995 Using Testability Measures for Dependability Assessment
abstract
Program "testability"is the probability that a fault in a program, if present, will cause the program to fail.Measures of testability can be used to draw inferences on program correctness from the observation of a series of failure-free test executions, a common need for software with "ultra-high reliability" requirements.For a program that has passed a certain number of tests without failing, a high value of testability implies a high probability that the program is correct.We give a general descriptive model of program execution and testing, and propose a more precise definition of program testability than that given by other authors.We then study the use of testability in: i) providing, through testing, confidence in the absence of faults and ii) bounding the probability of failures, from the results of operational testing.We derive the probability of absence of faults through a Bayesian inference procedure, criticise previously proposed derivations of this probability, and study the relationship between the testability of a program and its failure probability in operation.We derive the conditions under which a high testability improves one's expectations about program reliability.Last, we discuss the potential of these methods in practical applications.
Antonia Bertolino, Lorenzo Strigini
ICSE2
1992 Dataflow-Like Languages for Real-Time Systems: Issues of Computational Models and Notations
abstract
The use of dataflow-like models for the in-the-large design of real-time applications is discussed. In these models, modules can only communicate by (asynchronously) receiving messages when activated and transmitting result messages when terminating. This rather restrictive computational model allows the description of typical, cyclic control programs, with predictable, well-verifiable behavior. In particular, important timing properties can be dealt with in the in-the-large design. The case for the use of dataflow-like models is outlined, and the choice of appropriate notations, which implies a tradeoff between predictability of behavior and expressive power, and the potential for an advanced design support environment are discussed.>
Andrea Bondavalli, Lorenzo Strigini, Luca Simoncini
SRDS2
1992 Destination Stripping Dual Ring: A New Protocol for MANs
Andrea Bondavalli, Lorenzo Strigini, Matteo Sereno
Comput. Networks ISDN Syst.2
1991 DSDR: A Fair and Efficient Access Protocol for Ring-Topology MANs
abstract
A new media access control (MAC) protocol for ring-topology metropolitan network (MANs) is defined. The destination-stripping dual ring (DSDR) is defined for slotted-medium dual-ring network. The main feature is the destination-stripping capability, which yields a very high throughput through immediate reuse of the slots. It is shown that the protocol can be tailored for different compromises between guaranteed throughput and total throughput. On each of the two rings, guaranteed throughput, with bounded access delay, can be close to twice the medium capacity, and total throughput, under realistic hypotheses, can be close to four times the medium capacity. The maximum interval between accesses by the same station can be as low as N/2 slot times, with N stations on the network. Multicast can be supported with the same simplicity and the same cost as in other LANs and MANs.>
Andrea Bondavalli, Lorenzo Strigini
INFOCOM2
1991 Flexible Schemes for Application-Level Fault Tolerance
abstract
It is pointed out that the design of fault-tolerance provisions in the application level is normally necessary, but difficult and error-prone due to its ad-hoc nature. Structuring schemes have been proposed to reduce the difficulty of this task, but they appear too restrictive for the building of large, heterogeneous applications. The redundant structures that can be used in the individual components of a system depend on their requirements or inherent characteristics; it would be useful to combine components using different basic schemes. As an example, the authors propose a solution for interfacing components using conversations for backward recovery with components using atomic transactions. Constraints for the designers of the components to be interfaced and requirements on the virtual machine supporting their execution are defined. Ways a classification of components could be organized to allow the formulation of more general solutions are discussed.>
Lorenzo Strigini, Felicita Di Giandomenico
SRDS1
1990 Adjudicators for Diverse-Redundant Components
abstract
The authors define the adjudication problem, summarize the existing literature on the topic, and investigate the use of probabilistic knowledge about error/faults in the subcomponents of a fault-tolerant component to obtain good adjudication functions. They prove the existence of an optimal adjudication function, which is useful both as an upper bound on the probability of correctly adjudged obtainable output and as a guide for design decisions.>
Felicita Di Giandomenico, Lorenzo Strigini
SRDS2
1989 MAC Protocols for High-Speed MANs: Performance Comparisons for a Family of Fasnet-Based Protocols
Andrea Bondavalli, Marco Conti, Enrico Gregori, Luciano Lenzini, Lorenzo Strigini
Comput. Networks ISDN Syst.5
1982 A Distributed algorithm for post-failure load redistribution
G. Barigazzi, Augusto Ciuffoletti, Lorenzo Strigini
ICDCS3