Martin Neil

dblp:95/3981 · DBLP profile ↗
← Back
41ranked-venue papers
9as first author
2since 2021 · last 2025
0000-0002-4922-0843ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 14 · 6 first-authorDatabases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Probabilistic and Bayesian machine learning · 67% Knowledge representation and reasoning · 28% Deep learning architectures and training · 5%
Software engineering, system software, and programming languages
6 papers
Empirical software engineering · 75% Requirements engineering and software design · 10% Software maintenance and evolution · 8%

Topics — the 23 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network
0.952019
Addressing the Practical Limitations of Noisy-OR Using Conditional Inter-Causal Anti-Correlation with Ranked Nodes · IEEE Trans. Knowl. Data Eng. 2019
An Extension to the Noisy-OR Function to Resolve the 'Explaining Away' Deficiency for Practical Bayesian Network Problems · IEEE Trans. Knowl. Data Eng. 2019
Using Ranked Nodes to Model Qualitative Judgments in Bayesian Networks · IEEE Trans. Knowl. Data Eng. 2007
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.522020
The Book of Why: The New Science of Cause and Effect, Judea Pearl, Dana Mackenzie. Basic Books (2018) · Artif. Intell. 2020
Addressing the Practical Limitations of Noisy-OR Using Conditional Inter-Causal Anti-Correlation with Ranked Nodes · IEEE Trans. Knowl. Data Eng. 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network
conditional probability table elicitation
0.412019
Addressing the Practical Limitations of Noisy-OR Using Conditional Inter-Causal Anti-Correlation with Ranked Nodes · IEEE Trans. Knowl. Data Eng. 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning › graphical model learning
bayesian network learning
0.212013
Incorporating Expert Judgement into Bayesian Network Machine Learning · IJCAI 2013
Machine learning › Deep learning architectures and training
binary decomposition
0.112012
Optimizing the Calculation of Conditional Probability Tables in Hybrid Bayesian Networks Using Binary Factorization · IEEE Trans. Knowl. Data Eng. 2012
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
graphical model inference
0.112012
Optimizing the Calculation of Conditional Probability Tables in Hybrid Bayesian Networks Using Binary Factorization · IEEE Trans. Knowl. Data Eng. 2012
Empirical software engineering
software effort estimation
0.122009
Predicting Project Velocity in XP Using a Learning Dynamic Bayesian Network Model · IEEE Trans. Software Eng. 2009
Making Resource Decisions for Software Projects · ICSE 2004
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
explaining away
0.112019
Addressing the Practical Limitations of Noisy-OR Using Conditional Inter-Causal Anti-Correlation with Ranked Nodes · IEEE Trans. Knowl. Data Eng. 2019
Empirical software engineering
agile software development
0.112009
Predicting Project Velocity in XP Using a Learning Dynamic Bayesian Network Model · IEEE Trans. Software Eng. 2009
Empirical software engineering › agile software development
extreme programming
0.112009
Predicting Project Velocity in XP Using a Learning Dynamic Bayesian Network Model · IEEE Trans. Software Eng. 2009
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning
ranked nodes
0.112007
Using Ranked Nodes to Model Qualitative Judgments in Bayesian Networks · IEEE Trans. Knowl. Data Eng. 2007
Empirical software engineering › software project management
software risk analysis
0.112005
Automated population of causal models for improved software risk assessment · ASE 2005
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.012013
Incorporating Expert Judgement into Bayesian Network Machine Learning · IJCAI 2013
Empirical software engineering
software project management
0.012004
Making Resource Decisions for Software Projects · ICSE 2004
Computational complexity › reduction
computational complexity reduction
0.012012
Optimizing the Calculation of Conditional Probability Tables in Hybrid Bayesian Networks Using Binary Factorization · IEEE Trans. Knowl. Data Eng. 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network
dynamic bayesian network
0.012009
Predicting Project Velocity in XP Using a Learning Dynamic Bayesian Network Model · IEEE Trans. Software Eng. 2009
Empirical software engineering › mining software repositories › defect prediction
defect prediction models
0.011999
A Critique of Software Defect Prediction Models · IEEE Trans. Software Eng. 1999
Empirical software engineering
software defect prediction
0.011999
A Critique of Software Defect Prediction Models · IEEE Trans. Software Eng. 1999
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
expert knowledge elicitation
0.012007
Using Ranked Nodes to Model Qualitative Judgments in Bayesian Networks · IEEE Trans. Knowl. Data Eng. 2007
Requirements engineering and software design
formal specification
0.011998
Lessons from Using Z to Specify a Software Tool · IEEE Trans. Software Eng. 1998
Software testing › software reliability
reliability assessment
0.011998
Lessons from Using Z to Specify a Software Tool · IEEE Trans. Software Eng. 1998
Software testing
statistical testing
0.011998
Lessons from Using Z to Specify a Software Tool · IEEE Trans. Software Eng. 1998
Requirements engineering and software design › formal specification
z specification
0.011998
Lessons from Using Z to Specify a Software Tool · IEEE Trans. Software Eng. 1998

Methods — techniques the papers use, named apart from their topics

ranked nodes · 0.4leaky noisy-OR extension · 0.4explaining-away · 0.4conditional anti-correlation · 0.4Noisy-OR · 0.4junction tree inference · 0.3dynamic discretization · 0.3binary factorization · 0.3dynamic bayesian network · 0.2expert judgement elicitation · 0.2bayesian network · 0.1learning from project data · 0.1expert judgment · 0.0statistical modeling · 0.0bayesian belief networks · 0.0statistical testing · 0.0formal methods · 0.0criteria-based evaluation · 0.0
YearPublicationVenuePosition
2025 Stacking Factorizing Partitioned Expressions in Hybrid Bayesian Network Models
abstract
Hybrid Bayesian networks (HBN) contain complex conditional probability distributions (CPD) specified as partitioned expressions over discrete and continuous variables. The size of these CPDs grows exponentially with the number of parent nodes, and when using discrete inference methods, it results in significant execution time and space inefficiency. To reduce the CPD size, a binary factorization (BF) algorithm can be used to decompose the statistical or arithmetic functions in the CPD by factorizing the number of connected parent nodes into sets of size two. However, the BF algorithm was not designed to handle partitioned expressions. Therefore, we propose a new stacking factorization (SF) algorithm to decompose partitioned expressions. The SF algorithm creates intermediate nodes to incrementally reconstruct the conditional densities in the original partitioned expression, ensuring that no more than two continuous parent nodes are connected to each child node in the resulting HBN. It generally applies to both discrete and continuous child nodes with complex partitioned expressions. When we combine SF with a dynamic discretization (DD) inference algorithm, we achieve a significant improvement in inference efficiency. Experimental results demonstrate that the combination of SF and DD can effectively manage HBNs with complex CPDs that may challenge other algorithms, which also outperform competing inference algorithms in accuracy.
Martin Neil, Norman E. Fenton
ACM Trans. Knowl. Discov. Data2
2022 Region-based estimation of the partition functions for hybrid Bayesian network models
abstract
The partition function Ƶ ${\boldsymbol{Ƶ}}$ is a normalization constant for normalizing all the distributions in probabilistic inference. Ƶ ${\boldsymbol{Ƶ}}$ is closely related to the log probability of evidence ( log p ( e ) $\mathrm{log}\unicode{x0200A}p(e)$ ) for Bayesian networks (BNs), which plays an important role in many applications, such as parameter learning, classification, and clustering. However, evaluating Ƶ ${\boldsymbol{Ƶ}}$ and log p ( e ) $\mathrm{log}\unicode{x0200A}p(e)$ for BNs containing both discrete and continuous variables (known as hybrid Bayesian networks [HBNs]) is generally difficult for analytical and simulation-based solutions. This paper describes the work addressing the problem of estimating the partition function for HBNs when exact methods become inefficient or numerically unstable. We use dynamic discretization to provide a discretized version of the model and then perform region-based free energy optimization to obtain the log p ( e ) $\mathrm{log}\unicode{x0200A}p(e)$ for the HBN efficiently, which is called the DDJT-region partition function (RPR) algorithm. It is novel since the region-based methods have previously only been applied to discrete BNs. We show that the log p ( e ) $\mathrm{log}\unicode{x0200A}p(e)$ obtained by DDJT-RPR is exact for discretized models and with space complexity equal to that in marginal inference, thus generally applicable to a wide range of applications. To demonstrate these properties, we combine the DDJT-RPR algorithm with an improved expectation-maximization algorithm to learn Gaussian mixture models (GMMs), for applications requiring data only and with prior knowledge, which performed more robustly than conventional GMM learning algorithms.
Martin Neil, Norman E. Fenton, Eugene Dementiev
Int. J. Intell. Syst.2
2020 The Book of Why: The New Science of Cause and Effect, Judea Pearl, Dana Mackenzie. Basic Books (2018)
Norman E. Fenton, Martin Neil, Anthony C. Constantinou
Artif. Intell.2
2020 A Bayesian network approach for cybersecurity risk assessment implementing and extending the FAIR model
Martin Neil, Norman E. Fenton
Comput. Secur.2
2020 Improved High Dimensional Discrete Bayesian Network Inference using Triplet Region Construction
abstract
Performing efficient inference on high dimensional discrete Bayesian Networks (BNs) is challenging. When using exact inference methods the space complexity can grow exponentially with the tree-width, thus making computation intractable. This paper presents a general purpose approximate inference algorithm, based on a new region belief approximation method, called Triplet Region Construction (TRC). TRC reduces the cluster space complexity for factorized models from worst-case exponential to polynomial by performing graph factorization and producing clusters of limited size. Unlike previous generations of region-based algorithms, TRC is guaranteed to converge and effectively addresses the region choice problem that bedevils other region-based algorithms used for BN inference. Our experiments demonstrate that it also achieves significantly more accurate results than competing algorithms.
Martin Neil, Norman E. Fenton
J. Artif. Intell. Res.2
2020 Medical idioms for clinical Bayesian network development
Evangelia Kyrimi, Mariana R. Neves, Scott McLachlan, Martin Neil, William Marsh 0001, Norman E. Fenton
J. Biomed. Informatics4
2019 An Extension to the Noisy-OR Function to Resolve the 'Explaining Away' Deficiency for Practical Bayesian Network Problems
abstract
The “leaky noisy-OR” function is a common and popular method used to simplify the elicitation of complex conditional probability tables in Bayesian networks involving Boolean variables. It has proven to be useful for approximating the required relationship in many real-world situations where there is a set of two or more variables that are potential causes of a single effect variable. However, one of the properties of leaky noisy-OR is Conditional Inter-causal Independence (CII). This property means that the `explaining away' behavior-one of the most powerful benefits of BN inference-is not present when the effect variable is observed as false. Yet, for many real-world problems where the leaky noisy-OR has been considered, this behavior would be expected, meaning that leaky noisy-OR is deficient as an approximation of the required relationship in such cases. There have been previous attempts to adapt the noisy-OR to resolve this problem. However, they require too many additional parameters to be elicited. We describe a simple but powerful extension to leaky noisy-OR that requires only a single additional parameter. While it does not solve the CII problem in all cases, it does resolve most of the explaining away deficiencies that occur in practice. The problem and solution is illustrated using an example from intelligence analysis.
Norman E. Fenton, Takao Noguchi, Martin Neil
IEEE Trans. Knowl. Data Eng.3
2019 Addressing the Practical Limitations of Noisy-OR Using Conditional Inter-Causal Anti-Correlation with Ranked Nodes
abstract
Numerous methods have been proposed to simplify the problem of eliciting complex conditional probability tables in Bayesian networks. One of the most popular methods -“Noisy-OR”- approximates the required relationship in many real-world situations between a set of variables that are potential causes of an effect variable. However, the Noisy-OR function has the conditional inter-causal independence (CII) property which means that `explaining away' behavior-one of the most powerful benefits of BN inference-is not present when the effect variable is observed as false. Hence, for many real-world problems where the Noisy-OR has been used or proposed, it may be deficient as an approximation of the required relationship. However, there is a very simple alternative solution, namely to define the variables as ranked nodes and to use the ranked node weighted average function. This does not have the CII property-instead, we prove it has the conditional anti-correlation property required to ensure that explaining away works in all cases. Moreover, ranked node variables are not restricted to binary states, and hence provide a more comprehensive and general solution to Noisy-OR in all cases.
Takao Noguchi, Norman E. Fenton, Martin Neil
IEEE Trans. Knowl. Data Eng.3
2018 An improved method for solving Hybrid Influence Diagrams
Barbaros Yet, Martin Neil, Norman E. Fenton, Anthony C. Constantinou, Eugene Dementiev
Int. J. Approx. Reason.2
2017 The opportunity prior: a simple and practical solution to the prior probability problem for legal cases
abstract
One of the greatest impediments to the use of probabilistic reasoning in legal arguments is the difficulty in agreeing on an appropriate prior probability for the ultimate hypothesis, (in criminal cases this is normally "Defendant is guilty of the crime for which he/she is accused"). Even strong supporters of a Bayesian approach prefer to ignore priors and focus instead on considering only the likelihood ratio (LR) of the evidence. But the LR still requires the decision maker (be it a judge or juror during trial, or anybody helping to determine beforehand whether a case should proceed to trial) to consider their own prior; without it the LR has limited value. We show that, in a large class of cases, it is possible to arrive at a realistic prior that is also as consistent as possible with the legal notion of 'innocent until proven guilty'. The approach can be considered as a formalisation of the 'island problem' whereby if it is known the crime took place on an island when n people were present, then each of the people on the island has an equal prior probability 1/n of having carried out the crime. Our prior is based on simple location and time parameters that determine both a) the crime scene/time (within which it is certain the crime took place) and b) the extended crime scene/time which is the 'smallest' within which it is certain the suspect was known to have been 'closest' in location/time to the crime scene. The method applies to cases where we assume a crime has taken place and that it was committed by one person against one other person (e.g. murder, assault, robbery). The paper considers both the practical and legal implications of the approach. We demonstrate how the opportunity prior probability is naturally incorporated into a generic Bayesian network model that allows us to integrate other evidence about the case.
Norman E. Fenton, David A. Lagnado, Christian Dahlman, Martin Neil
ICAIL4
2016 Value of information analysis for interventional and counterfactual Bayesian networks in forensic medical sciences
Anthony C. Constantinou, Barbaros Yet, Norman E. Fenton, Martin Neil, William Marsh 0001
Artif. Intell. Medicine4
2016 Integrating expert knowledge with data in Bayesian networks: Preserving data-driven expectations when the expert variables remain unobserved
Anthony C. Constantinou, Norman E. Fenton, Martin Neil
Expert Syst. Appl.3
2016 A Bayesian network framework for project cost, benefit and risk analysis with an agricultural development case study
Barbaros Yet, Anthony C. Constantinou, Norman E. Fenton, Martin Neil, Eike Luedeling, Keith D. Shepherd
Expert Syst. Appl.4
2016 How to model mutually exclusive events based on independent causal pathways in Bayesian network models
abstract
We show that existing Bayesian network (BN) modelling techniques cannot capture the correct intuitive reasoning in the important case when a set of mutually exclusive events need to be modelled as separate nodes instead of states of a single node. A previously proposed ‘solution’, which introduces a simple constraint node that enforces mutual exclusivity, fails to preserve the prior probabilities of the events, while other proposed solutions involve major changes to the original model. We provide a novel and simple solution to this problem that works in all cases where the mutually exclusive nodes have no common ancestors. Our solution uses a special type of constraint and auxiliary node together with formulas for assigning their necessary conditional probability table values. The solution enforces mutual exclusivity between events and preserves their prior probabilities while leaving all original BN nodes unchanged.
Norman E. Fenton, Martin Neil, David A. Lagnado, William Marsh 0001, Barbaros Yet, Anthony C. Constantinou
Knowl. Based Syst.2
2015 Probabilistic Graphical Models Parameter Learning with Transferred Prior and Constraints
Yun Zhou 0001, Norman E. Fenton, Timothy M. Hospedales, Martin Neil
UAI4
2014 Bayesian network approach to multinomial parameter learning using data and expert judgments
Yun Zhou 0001, Norman E. Fenton, Martin Neil
Int. J. Approx. Reason.3
2013 Incorporating Expert Judgement into Bayesian Network Machine Learning
Yun Zhou 0001, Norman E. Fenton, Martin Neil, Cheng Zhu 0002
IJCAI3
2013 Profiting from an inefficient association football gambling market: Prediction, risk and uncertainty using Bayesian networks
abstract
We present a Bayesian network (BN) model for forecasting Association Football match outcomes. Both objective and subjective information are considered for prediction, and we demonstrate how probabilities transform at each level of model component, whereby predictive distributions follow hierarchical levels of Bayesian inference. The model was used to generate forecasts for each match of the 2011/2012 English Premier League (EPL) season, and forecasts were published online prior to the start of each match. Profitability, risk and uncertainty are evaluated by considering various unit-based betting procedures against published market odds. Compared to a previously published successful BN model, the model presented in this paper is less complex and is able to generate even more profitable returns.
Anthony C. Constantinou, Norman E. Fenton, Martin Neil
Knowl. Based Syst.3
2012 Availability modelling of repairable systems using Bayesian networks
Martin Neil, David Marquez
Eng. Appl. Artif. Intell.1
2012 pi-football: A Bayesian network model for forecasting Association Football match outcomes
Anthony C. Constantinou, Norman E. Fenton, Martin Neil
Knowl. Based Syst.3
2012 Optimizing the Calculation of Conditional Probability Tables in Hybrid Bayesian Networks Using Binary Factorization
abstract
Reducing the computational complexity of inference in Bayesian Networks (BNs) is a key challenge. Current algorithms for inference convert a BN to a junction tree structure made up of clusters of the BN nodes and the resulting complexity is time exponential in the size of a cluster. The need to reduce the complexity is especially acute where the BN contains continuous nodes. We propose a new method for optimizing the calculation of Conditional Probability Tables (CPTs) involving continuous nodes, approximated in Hybrid Bayesian Networks (HBNs), using an approximation algorithm called dynamic discretization. We present an optimized solution to this problem involving binary factorization of the arithmetical expressions declared to generate the CPTs for continuous nodes for deterministic functions and statistical distributions. The proposed algorithm is implemented and tested in a commercial Hybrid Bayesian Network software package and the results of the empirical evaluation show significant performance improvement over unfactorized models.
Martin Neil, Norman E. Fenton
IEEE Trans. Knowl. Data Eng.1
2010 Comparing risks of alternative medical diagnosis using Bayesian arguments
Norman E. Fenton, Martin Neil
J. Biomed. Informatics2
2009 Predicting Project Velocity in XP Using a Learning Dynamic Bayesian Network Model
abstract
Bayesian networks, which can combine sparse data, prior assumptions and expert judgment into a single causal model, have already been used to build software effort prediction models. We present such a model of an extreme programming environment and show how it can learn from project data in order to make quantitative effort predictions and risk assessments without requiring any additional metrics collection program. The model's predictions are validated against a real world industrial project, with which they are in good agreement.
Peter Stewart Hearty, Norman E. Fenton, David Marquez, Martin Neil
IEEE Trans. Software Eng.4
2008 On the effectiveness of early life cycle defect prediction with Bayesian Nets
Norman E. Fenton, Martin Neil, William Marsh 0001, Peter Stewart Hearty, Lukasz Radlinski, Paul Krause
Empir. Softw. Eng.2
2007 Predicting software defects in varying development lifecycles using Bayesian nets
Norman E. Fenton, Martin Neil, William Marsh 0001, Peter Stewart Hearty, David Marquez, Paul Krause, Rajat Mishra
Inf. Softw. Technol.2
2007 Using Ranked Nodes to Model Qualitative Judgments in Bayesian Networks
abstract
Although Bayesian Nets (BNs) are increasingly being used to solve real world risk problems, their use is still constrained by the difficulty of constructing the node probability tables (NPTs). A key challenge is to construct relevant NPTs using the minimal amount of expert elicitation, recognising that it is rarely cost-effective to elicit complete sets of probability values. We describe a simple approach to defining NPTs for a large class of commonly occurring nodes (called ranked nodes). The approach is based on the doubly truncated Normal distribution with a central tendency that is invariably a type of weighted function of the parent nodes. In extensive real-world case studies we have found that this approach is sufficient for generating the NPTs of a very large class of nodes. We describe one such case study for validation purposes. The approach has been fully automated in a commercial tool, called AgenaRisk, and is thus accessible to all types of domain experts. We believe this work represents a useful contribution to BN research and technology since its application makes the difference between being able to build realistic BN models and not.
Norman E. Fenton, Martin Neil, Jose Galan Caballero
IEEE Trans. Knowl. Data Eng.2
2006 Modeling Dependable Systems using Hybrid Bayesian Networks
abstract
A hybrid Bayesian network (BN) is one that incorporates both discrete and continuous nodes. In our extensive applications of BNs for system dependability assessment the models are invariably hybrid and the need for efficient and accurate computation is paramount. We apply a new iterative algorithm that efficiently combines dynamic discretisation with robust propagation algorithms on junction tree structures to perform inference in hybrid BNs. We illustrate its use on two example dependability problems: reliability estimation and diagnosis of a faulty sensor in a temporal system. Dynamic discretisation can be used as an alternative to analytical or Monte Carlo methods with high precision and can be applied to a wide range of dependability problems.
Martin Neil, Manesh Tailor, Norman E. Fenton, David Marquez, Peter Stewart Hearty
ARES1
2006 Predicting football results using Bayesian nets and other machine learning techniques
Anito Joseph, Norman E. Fenton, Martin Neil
Knowl. Based Syst.3
2005 Automated population of causal models for improved software risk assessment
abstract
Recent work in applying causal modeling (Bayesian networks) to software engineering has resulted in improved decision support systems for software project managers. Once the causal models are built there are commercial tools that can run them. However, data to populate the models is typically entered manually and this is an impediment to their more widespread use. Hence, here we present a prototype tool for automatically extracting a range of relevant software metrics from popular project management and CASE tools. This information is used to populate Bayesian networks with the aim of providing better real world predictions of the risks associated with software costs, timescales and reliability.
Peter Stewart Hearty, Norman E. Fenton, Martin Neil, Patrick Cates
ASE3
2004 Making Resource Decisions for Software Projects
abstract
Software metrics should support managerial decision making in software projects. We explain how traditional metrics approaches, such as regression-based models for cost estimation fall short of this goal. Instead, we describe a causal model (using a Bayesian network) which incorporates empirical data, but allows it to be interpreted and supplemented using expert judgement. We show how this causal model is used in a practical decision-support tool, allowing a project manager to trade-off the resources used against the outputs (delivered functionality, quality achieved) in a software project. The model and toolset have evolved in a number of collaborative projects and hence capture significant commercial input. Extensive validation trials are taking place among partners on the EC funded project MODIST (this includes Philips, Israel Aircraft Industries and QinetiQ) and the feedback so far has been very good. The estimates are sensible and the causal modelling approach enables decision-makers to reason in a way that is not possible with other project management and resource estimation tools. To ensure wide dissemination and validation a version of the toolset with the full underlying model is being made available for free to researchers.
Norman E. Fenton, William Marsh 0001, Martin Neil, Patrick Cates, Simon Forey, Manesh Tailor
ICSE3
2001 Probabilistic Modelling for Software Quality Control
Norman E. Fenton, Paul Krause, Martin Neil
ECSQARU3
2001 Making decisions: using Bayesian nets and MCDA
Norman E. Fenton, Martin Neil
Knowl. Based Syst.2
1999 Software metrics: successes, failures and new directions
Norman E. Fenton, Martin Neil
J. Syst. Softw.2
1999 A Critique of Software Defect Prediction Models
abstract
Many organizations want to predict the number of defects (faults) in software systems, before they are deployed, to gauge the likely delivered quality and maintenance effort. To help in this numerous software metrics and statistical models have been developed, with a correspondingly large literature. We provide a critical review of this literature and the state-of-the-art. Most of the wide range of prediction models use size and complexity metrics to predict defects. Others are based on testing data, the "quality" of the development process, or take a multivariate approach. The authors of the models have often made heroic contributions to a subject otherwise bereft of empirical studies. However, there are a number of serious theoretical and practical problems in many studies. The models are weak because of their inability to cope with the, as yet, unknown relationship between defects and failures. There are fundamental statistical and data quality problems that undermine model validity. More significantly many prediction models tend to model only part of the underlying problem and seriously misspecify it. To illustrate these points the Goldilock's Conjecture, that there is an optimum module size, is used to show the considerable problems inherent in current defect prediction approaches. Careful and considered analysis of past and new results shows that the conjecture lacks support and that some models are misleading. We recommend holistic models for software defect prediction, using Bayesian belief networks, as alternative approaches to the single-issue models used at present. We also argue for research into a theory of "software decomposition" in order to test hypotheses about defect introduction and help construct a better science of software engineering.
Norman E. Fenton, Martin Neil
IEEE Trans. Software Eng.2
1998 A Strategy for Improving Safety Related Software Engineering Standards
abstract
There are many standards which are relevant for building safety- or mission-critical software systems. An effective standard is one that should help developers, assessors and users of such systems. For developers, the standard should help them build the system cost-effectively, and it should be clear what is required in order to conform to the standard. For assessors, it should be possible to objectively determine compliance to the standard. Users, and society at large, should have some assurance that a system developed to the standard has quantified risks and benefits. Unfortunately, the existing standards do not adequately fulfil any of these varied requirements. We explain why standards are the way they are, and then provide a strategy for improving them. Our approach is to evaluate standards on a number of key criteria that enable us to interpret the standard, identify its scope and check the ease with which it can be applied and checked. We also need to demonstrate that the use of a standard is likely either to deliver reliable and safe systems at an acceptable cost or to help predict reliability and safety accurately. Throughout the paper, we examine, by way of example, a specific standard for safety-critical systems (namely IEC 1508) and show how it can be improved by applying our strategy.
Norman E. Fenton, Martin Neil
IEEE Trans. Software Eng.2
1998 Lessons from Using Z to Specify a Software Tool
abstract
The authors were recently involved in the development of a COBOL parser (G. Ostrolenk et al., 1994), specified formally in Z. The type of problem tackled was well suited to a formal language. The specification process was part of a life cycle characterized by the front loading of effort in the specification stage and the inclusion of a statistical testing stage. The specification was found to be error dense and difficult to comprehend. Z was used to specify inappropriate procedural rather than declarative detail. Modularity and style problems in the Z specification made it difficult to review. In this sense, the application of formal methods was not successful. Despite these problems the estimated fault density for the product was 1.3 faults per KLOC, before delivery, which compares favorably with IBM's Cleanroom method. This was achieved, despite the low quality of the Z specification, through meticulous and effort intensive reviews. However, because the faults were in critical locations, the reliability of the product was assessed to be unacceptably low. This demonstrates the necessity of assessing reliability as well as "correctness" during system testing. Overall, the experiences reported in the paper suggest a range of important lessons for anyone contemplating the practical application of formal methods.
Martin Neil, Gary Ostrolenk, Mary Tobin, Mark Southworth
IEEE Trans. Software Eng.1
1997 Are software failures deterministic?
Martin Neil
Inf. Softw. Technol.1
1995 Metrics and Models in Software Quality Engineering, by Stephen H. Kan, Addison-Wesley, 1995 (Book Review)
Martin Neil
Softw. Test. Verification Reliab.1
1993 Data linkage maps
abstract
Abstract Most work in software metrics has been aimed at measuring the quality or cost of software. This paper gives an example of how metrics may be used to give information about how the software system is organized, specifically the data linkage between procedures in the source code. The data linkage measures capture how closely procedures are bound together by data flowing between them. The paper shows how data linkage within a program can be represented in a diagrammatic form. This approach makes use of multivariate statistics and is applied to two examples. Results are then compared with the authors' intuitions of how the example programs are organized. This approach could be used as an automatic aid to maintainers.
Martin Neil, Richard Bache
J. Softw. Maintenance Res. Pract.1
1992 Multivariate Assessment of Software Products
Martin Neil
Softw. Test. Verification Reliab.1
1990 Measures for maintenance management: A case study
abstract
Abstract Software maintenance is an important activity, but is often perceived as less challenging than software development. By following a management‐led discipline of data collection and reporting, maintenance control and visibility can be achieved. The discipline aims to promote the maintenance function within an organization and show that it is as demanding as software development. This paper describes a maintenance data collection scheme (MDCS) which was implemented at an industrial site, as a case study. This implementation resulted in: an increase in management control, promotion of maintenance to senior management, systematic recording of faults, and the successful application of a statistical model to the weekly number of maintenance incidents. An incident form was developed to record data on the effort necessary to resolve incidents, the elapsed time, status, category and severity of incidents. It turned out that maintenance incident effort was highly variable. Maintenance reports were prepared for management based on the results of data analysis. These reports addressed the following questions: Is the maintenance effort increasing or decreasing? How fast are maintenance problems being dealt with? Which systems are demanding the most effort? Statistical techniques were used to model maintenance activities. Multivariate techniques were found to be inappropriate for the type of data collected. A log‐normal distribution was fitted to the weekly number of maintenance incidents. This distribution gave a confidence interval of practical use for management.
Martin Neil, Robert J. Cole, David Slater
J. Softw. Maintenance Res. Pract.1