Stephen Swift

dblp:93/5842 · DBLP profile ↗
← Back
56ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-8918-3365ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 5 since 2021Software engineering, systems software and programming languages · 13 · 4 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 8Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 An audit of machine learning experiments on software defect prediction
abstract
Machine learning algorithms are increasingly being proposed to solve the problem of predicting defect-prone software components. In this literature, computational experiments are the primary means of evaluating and comparing learners and the credibility of findings depends critically on their experimental design and reporting. This paper audits recent software defect prediction (SDP) experiments by assessing their experimental design, analysis and reporting practices against widely accepted norms from statistics, machine learning and empirical software engineering. Our aim is to characterise the current state of practice and evaluate the reproducibility of published findings. We undertook an audit of relevant studies published from the SCOPUS database (2019-2023) focusing on their experimental design and analysis choices e.g., the outcome variables such as F-measure and the type of out of sample (OOS) validation regime, e.g., cross-validation, plus the statistical analysis and inference mechanisms. In all, we evaluated nine different study issues. This was complemented by an assessment of reproducibility using the instrument proposed by González-Barahona and Robles. Our search located approximately 1,585 experiments in SDP (2019-2023), a substantial body of work. From this, we randomly sampled 101 ( $$ \approx 6.4\%$$ ) papers, 61 journal and 40 conference papers. Almost 50% are behind ‘paywalls’. We found considerable divergence in research practice. The number of datasets used ranged 1-365, the number of learners or learner variants evaluated from 1-34 and the number of performance metrics from 1 to 9. Approximately 45% of papers made use of formal statistical inference. We detected a total of 427 issues distributed across 101 papers (median=4) with only one paper being entirely issue-free. In terms of reproducibility, experiments ranged from near perfect to lacking almost all required information. We also found two examples of tortured phrases and potential “paper mill” activity. Approaches to designing and reporting computational experiments varied greatly, but almost half the studies provided insufficient information such that reproduction would be challenging. Overall, our audit suggests that as a research community, we have considerable scope for improvement. Fortunately, many improvements should be neither difficult nor costly to achieve.
Giuseppe Destefanis, Leila Yousefi, Martin J. Shepperd, Allan Tucker, Stephen Swift, Steve Counsell, Mahir Arzoky
Empir. Softw. Eng.5
2026 ETM-F: an enriched topic modeling and filtration framework integrating ontologies and deep learning for biomedical trend analysis
Ahmad Altarawneh, Mahir Arzoky, Stephen Swift, Reem Qadan Al Fayez, Moh'd Belal Al-Zoubi, Bilal Sowan
Knowl. Inf. Syst.3
2024 Different Strokes for Different Folks: A Comparison of Developer and Tester Views on Testing
abstract
In this paper, we re-analyse the data from a previous study of Straubinger et al., which asked 284 industrial IT staff about their views on testing. In that study, as well as developers, the dedicated role of tester was included in the data - both roles were treated as the same role. In this paper, we posit that the two roles (i.e., developer and tester) are so very different that we should analyse each role separately; testers will have unique insights into testing, so separating their views and experiences from developers is important. To this end, we analyse six of the same research questions as the original study, using separate developer and tester data. Results showed that for almost every question we re-visited, testers differed in their opinions from developers, whether on the type of testing they did, measures of code quality, effort to write tests and motivation for testing.
Steve Counsell, Stephen Swift, Mahir Arzoky, Tracy Hall, Emily Winter 0001, Gareth Bennett, Thomas Shippey
SEAA2
2023 An "80-20" Approach to the Study of Coupling
abstract
In this paper, we use both afferent (i.e., incoming) and efferent (i.e., outgoing) class coupling to analyse trends in seven open-source and four industrial systems. We determine whether an 80-20 rule (or Pareto’s Law) applies to either type of coupling and whether 20% of classes contain 80% of coupling (whether incoming and outgoing). We also explore the relationship between bugs and both types of coupling. Results showed that an 80-20 rule applied overwhelmingly in the industrial systems (but not in the open-source). For the two types of coupling, the open-source systems also had a stronger, positive relationship with bugs; opposite effects were found in the industrial systems. These results highlight the need for adoption of different and impactful approaches to understanding systems which reflect on developer practice.
Steve Counsell, Stephen Swift, Amjed Tahir
SEAA2
2022 Estimating the Optimal Number of Clusters from Subsets of Ensembles
abstract
This research estimates the optimal number of clusters in a dataset using a novel ensemble technique - a preferred alternative to relying on the output of a single clustering. Combining clusterings from different algorithms can lead to a more stable and robust solution, often unattainable by any single clustering solution. Technically, we created subsets of ensembles as possible estimates; and evaluated them using a quality metric to obtain the best subset. We tested our method on publicly available datasets of varying types, sources and clustering difficulty to establish the accuracy and performance of our approach against eight standard methods. Our method outperforms all the techniques in the number of clusters estimated correctly. Due to the exhaustive nature of the initial algorithm, it is slow as the number of ensembles or the solution space increases; hence, we have provided an updated version based on the single-digit difference of Gray code that runs in linear time in terms of the subset size.
Afees Adegoke Odebode, Allan Tucker, Mahir Arzoky, Stephen Swift
DATA4
2022 Timing is Everything! A Test and Production Class View of Self-Admitted Technical Debt
Steve Counsell, Stephen Swift
SEAA2
2022 Exploring the Explicit Modelling of Bias in Machine Learning Classifiers: A Deep Multi-label ConvNet Approach
abstract
This paper addresses the problem that many machine learning classifiers make decisions based on data that are biased and can therefore result in prejudiced decisions. For example, in education (which this paper focuses on) a student may be rejected from a course based on historical decisions in the data that only exist due to historical biases in society or due to the skewed sampling of the data. Other approaches to dealing with bias in data include resampling methods (to counter imbalanced samples) and dimensionality reduction (to focus only on relevant features to the classification task). In this paper, we explore issues of modelling bias explicitly so that we can identify the types of bias and whether they are accounting for inflated predictive accuracies. In particular, we compare graphical model approaches to building classifiers, that are transparent in how they make decisions, with two forms of Deep Multi-label Convolutional Neural Networks to investigate if models can be built that maximise accuracy and minimise bias. We carry out this comparison on student entry and performance data from a higher educational institution.
Mashael Al-Luhaybi, Stephen Swift, Steve Counsell, Allan Tucker
ICMLA2
2021 Uncertainty Estimation in SARS-CoV-2 B-Cell Epitope Prediction for Vaccine Development
Bhargab Ghoshal, Biraja Ghoshal, Stephen Swift, Allan Tucker
AIME3
2021 Bayesian Deep Active Learning for Medical Image Analysis
Biraja Ghoshal, Stephen Swift, Allan Tucker
AIME2
2021 Opening the black box: Personalizing type 2 diabetes patients based on their latent phenotype and temporal associated complication rules
abstract
Abstract It is widely considered that approximately 10% of the population suffers from type 2 diabetes. Unfortunately, the impact of this disease is underestimated. Patient's mortality often occurs due to complications caused by the disease and not the disease itself. Many techniques utilized in modeling diseases are often in the form of a “black box” where the internal workings and complexities are extremely difficult to understand, both from practitioners' and patients' perspective. In this work, we address this issue and present an informative model/pattern, known as a “latent phenotype,” with an aim to capture the complexities of the associated complications' over time. We further extend this idea by using a combination of temporal association rule mining and unsupervised learning in order to find explainable subgroups of patients with more personalized prediction. Our extensive findings show how uncovering the latent phenotype aids in distinguishing the disparities among subgroups of patients based on their complications patterns. We gain insight into how best to enhance the prediction performance and reduce bias in the models applied using uncertainty in the patients' data.
Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker
Comput. Intell.2
2020 Using the Lexicon from Source Code to Determine Application Domain
abstract
Context: The vast majority of software engineering research is reported independently of the application domain: techniques and tools usage is reported without any domain context. As reported in previous research, this has not always been so: early in the computing era, the research focus was frequently application domain specific (for example, scientific and data processing).
Andrea Capiluppi, Nemitari Ajienka, Nour Ali, Mahir Arzoky, Steve Counsell, Giuseppe Destefanis, Alina Dana Miron, Bhaveet Nagaria, Rumyana Neykova, Martin J. Shepperd, Stephen Swift, Allan Tucker
EASE11
2020 Design of a Flexible, User Friendly Feature Matrix Generation System and its Application on Biomedical Datasets
abstract
Abstract The generation of a feature matrix is the first step in conducting machine learning analyses on complex data sets such as those containing DNA, RNA or protein sequences. These matrices contain information for each object which have to be identified using complex algorithms to interrogate the data. They are normally generated by combining the results of running such algorithms across various datasets from different and distributed data sources. Thus for non-computing experts the generation of such matrices prove a barrier to employing machine learning techniques. Further since datasets are becoming larger this barrier is augmented by the limitations of the single personal computer most often used by investigators to carry out such analyses. Here we propose a user friendly system to generate feature matrices in a way that is flexible, scalable and extendable. Additionally by making use of The Berkeley Open Infrastructure for Network Computing (BOINC) software, the process can be speeded up using distributed volunteer computing possible in most institutions. The system makes use of a combination of the Grid and Cloud User Support Environment (gUSE), combined with the Web Services Parallel Grid Runtime and Developer Environment Portal (WS-PGRADE) to create workflow-based science gateways that allow users to submit work to the distributed computing. This report demonstrates the use of our proposed WS-PGRADE/gUSE BOINC system to identify features to populate matrices from very large DNA sequence data repositories, however we propose that this system could be used to analyse a wide variety of feature sets including image, numerical and text data.
Mohammadmersad Ghorbani, Stephen Swift, Simon J. E. Taylor, Annette M. Payne
J. Grid Comput.2
2019 Predicting Academic Performance: A Bootstrapping Approach for Learning Dynamic Bayesian Networks
Mashael Al-Luhaybi, Leila Yousefi, Stephen Swift, Steve Counsell, Allan Tucker
AIED (1)3
2019 An Empirical Study of the AGIS Visual Field Metric and Its Seasonal Variations
abstract
The severity of the glaucoma eye disease is usually measured by the Advanced Glaucoma Intervention Studies (AGIS) metric. The metric provides a value between zero and twenty inclusive, where the former represents no evidence of glaucoma and the latter the most advanced of measurements. In a previous study by Montolio et al., the season in which the test was undertaken was shown to affect the value of eye measurements; the lowest sensitivity was found in Summer and the highest sensitivity found in Winter and Spring. In this paper, we partially replicate that study with a different set of data from 2468 patients obtained from Moorfields Eye Hospital, London. We decomposed the data according to the four seasonal dates to determine if extra sensitivity meant that patients' results improved in Winter.
Steve Counsell, Stephen Swift, Mahir Arzoky, Giuseppe Destefanis
CBMS2
2019 Opening the Black Box: Exploring Temporal Pattern of Type 2 Diabetes Complications in Patient Clustering Using Association Rules and Hidden Variable Discovery
abstract
There is a great deal of debate over the importance of explanation in AI models inferred from health data. In particular, there is a balance that needs to be made between the accuracy of complex 'deep' models such as convolutional neural networks and the transparency of models that aim to model data in a more 'human' way such as expert systems. In this paper, we explore the use of temporal association rules to validate and uncover the meaning behind discrete hidden variables that have been inferred from clinical diabetes data. We use a recently published technique based upon the IC* (Induction Causation) algorithm that limits the number of hidden variables and places them within a network structure. Here, we take the hidden variables and compare their underlying discrete states to clusters that have been generated from temporal association rules. This allows us to characterise the hidden states based upon different sequences of complications. Results are very promising, with many hidden states aligning with the discovered clusters giving us a direct interpretation.
Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker
CBMS2
2019 The Prevalence of Errors in Machine Learning Experiments
Martin J. Shepperd, Ning Li 0022, Mahir Arzoky, Andrea Capiluppi, Steve Counsell, Giuseppe Destefanis, Stephen Swift, Allan Tucker, Leila Yousefi
IDEAL (1)8
2018 Opening the Black Box: Discovering and Explaining Hidden Variables in Type 2 Diabetic Patient Modelling
Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker
BIBM2
2018 Re-visiting a Test Taxonomy with Refactoring and Defect-fix Data
abstract
In a previous empirical study by Bavota et al., multiple releases of three open-source systems reported the extent to which refactorings induced defect-fixes. In a much earlier study, van Deursen and Moonen (vD&M) provided a test taxonomy in which Fowler's seventy-two refactorings were categorized according to the post refactoring test burden of each (i.e., the changes required to unit tests after each refactoring had been undertaken). A refactoring was categorized as 'Type B' if it required no change to the original tests and 'Type E' if significant changes were necessary. In this paper, we investigate nine refactorings spread across vD&M's taxonomy and the corresponding defect-fix data provided by Bavota et al., to explore the relationship between defect-fixes due to refactoring and vD&M's taxonomy. Results showed that, in contrast to our intuition, the most defect-fix prone refactorings were of Types C and D and not, as we thought, of Type E. The 'Extract method' refactoring stood out as particularly 'defect-fix' inducing, suggesting that while it may solve one problem (i.e., in decomposing an excessively long method), it may well introduce other problems and required defect-fixes as a by-product.
Steve Counsell, Stephen Swift, Roberto Tonelli, Michele Marchesi, Michael Felderer
SEAA2
2018 Do Developers Really Worry About Refactoring Re-test? An Empirical Study of Open-Source Systems
Steve Counsell, Stephen Swift, Mahir Arzoky, Giuseppe Destefanis
PROFES2
2017 A Deconstructed Replication of a Time of Test Study Using the AGIS Metric
abstract
In medical practice, glaucoma severity is usually measured using the Advanced Glaucoma Intervention Studies (AGIS) metric. In a previous study [2], we replicated the work of Montolio et al., [5] and demonstrated that, for a larger dataset, time of day of test using the AGIS metric did make a difference to the measurement of glaucoma, supporting Montolio et als work. However, in our earlier study, we used the AGIS scores for both eyes combined. In this paper, we use the measurement from just one eye at a time. A dataset of 14389 left eye AGIS scores and the same number for the right eye from 2468 Moorfield Eye Hospital patients was used as the empirical basis. We then re-compared time of test results with those of Montolios study. Results revealed that using the values from just one eye (as opposed to both) may give a distorted picture of the AGIS scores; differences in the same time period were found between the two eyes. This may have implications for choice of sampling data and analysis of glaucoma using the AGIS metric.
Steve Counsell, Stephen Swift, Allan Tucker
CBMS2
2016 The AGIS Metric and Time of Test: A Replication Study
abstract
Visual Field (VF) tests and corresponding data are commonly used in clinical practices to manage glaucoma. The standard metric used to measure glaucoma severity is the Advanced Glaucoma Intervention Studies (AGIS) metric. We know that time of day when VF tests are applied can influence a patient's AGIS metric value; a previous study showed that this was the case for a data set of 160 patients. In this paper, we replicate that study using data from 2468 patients obtained from Moorfields Eye Hospital. This may provide further evidence and support of this phenomenon in a replication sense. Results did indeed show a tendency for the metric to be lower for early onset patients in the morning; equally, for advanced patients, the effect was less pronounced. We thus found support for the earlier work of Montolio et al. [4] and add to the body of evidence on the AGIS metric.
Steve Counsell, Stephen Swift, Allan Tucker
CBMS2
2016 Simultaneous Modelling and Clustering of Visual Field Data
abstract
This thesis was submitted for the award of Doctor of Philosophy and was awarded by Brunel University London
Mohd Zairul Mazwan Bin Jilani, Allan Tucker, Stephen Swift
CBMS3
2016 Arsonists or Firefighters? Affectiveness in Agile Software Development
abstract
In this paper, we present an analysis of more than 500 K comments from open-source repositories of software systems developed using agile methodologies. Our aim is to empirically determine how developers interact with each other under certain psychological conditions generated by politeness, sentiment and emotion expressed within developers’ comments. Developers involved in an open-source projects do not usually know each other; they mainly communicate through mailing lists, chat, and tools such as issue tracking systems. The way in which they communicate affects the development process and the productivity of the people involved in the project. We evaluated politeness, sentiment and emotions of comments posted by agile developers and studied the communication flow to understand how they interacted in the presence of impolite and negative comments (and vice versa ). Our analysis shows that “firefighters” prevail. When in presence of impolite or negative comments, the probability of the next comment being impolite or negative is 13 % and 25 %, respectively; ANGER however, has a probability of 40 % of being followed by a further ANGER comment. The result could help managers take control the development phases of a system, since social aspects can seriously affect a developer’s productivity. In a distributed agile environment this may have a particular resonance.
Marco Ortu, Giuseppe Destefanis, Steve Counsell, Stephen Swift, Roberto Tonelli, Michele Marchesi
XP4
2014 An Approach to Controlling the Runtime for Search Based Modularisation of Sequential Source Code Check-ins
Mahir Arzoky, Stephen Swift, Steve Counsell, James Cain 0002
IDA2
2014 Comparing Pre-defined Software Engineering Metrics with Free-Text for the Prediction of Code 'Ripples'
Steve Counsell, Allan Tucker, Stephen Swift, Guy Fitzgerald, Jason Peters
IDA3
2014 System performance analyses through object-oriented fault and coupling prisms
abstract
A fundamental aspect of a system's performance over time is the number of faults it generates. The relationship between the software engineering concept of "coupling" (i.e., the degree of inter-connectedness of a system's components) and faults is still a research question attracting attention and a relationship with strong implications for performance; excessive coupling is generally acknowledged to contribute to fault-proneness. In this paper, we explore the relationship between faults and coupling. Two releases from each of three open-source Eclipse projects (six releases in total) were used as an empirical basis and coupling and fault data extracted from those systems. A contrasting coupling profile between fault-free and fault-prone classes was observed and this result was statistically supported. Object-oriented (OO) classes with low values of fan-in (incoming coupling) and fan-out (outgoing coupling) appeared to support fault-free classes, while classes with high fan-out supported relatively fault-prone classes. We also considered size as an influence on fault-proneness. The study thus emphasizes the importance of minimizing coupling where possible (and particularly that of fan-out); failing to control coupling may store up problems for later in a system's life; equally, controlling class size should be a concomitant goal.
Alessandro Murgia, Roberto Tonelli, Michele Marchesi, Giulio Concas, Steve Counsell, Stephen Swift
ICPE6
2014 Improving predictive models of glaucoma severity by incorporating quality indicators
abstract
OBJECTIVE: In this paper we present an evaluation of the role of reliability indicators in glaucoma severity prediction. In particular, we investigate whether it is possible to extract useful information from tests that would be normally discarded because they are considered unreliable. METHODS: We set up a predictive modelling framework to predict glaucoma severity from visual field (VF) tests sensitivities in different reliability scenarios. Three quality indicators were considered in this study: false positives rate, false negatives rate and fixation losses. Glaucoma severity was evaluated by considering a 3-levels version of the Advanced Glaucoma Intervention Study scoring metric. A bootstrapping and class balancing technique was designed to overcome problems related to small sample size and unbalanced classes. As a classification model we selected Naïve Bayes. We also evaluated Bayesian networks to understand the relationships between the different anatomical sectors on the VF map. RESULTS: The methods were tested on a data set of 28,778 VF tests collected at Moorfields Eye Hospital between 1986 and 2010. Applying Friedman test followed by the post hoc Tukey's honestly significant difference test, we observed that the classifiers trained on any kind of test, regardless of its reliability, showed comparable performance with respect to the classifier trained only considering totally reliable tests (p-value>0.01). Moreover, we showed that different quality indicators gave different effects on prediction results. Training classifiers using tests that exceeded the fixation losses threshold did not have a deteriorating impact on classification results (p-value>0.01). On the contrary, using only tests that fail to comply with the constraint on false negatives significantly decreased the accuracy of the results (p-value<0.01). Meaningful patterns related to glaucoma evolution were also extracted. CONCLUSIONS: Results showed that classification modelling is not negatively affected by the inclusion of less reliable tests in the training process. This means that less reliable tests do not subtract useful information from a model trained using only completely reliable data. Future work will be devoted to exploring new quantitative thresholds to ensure high quality testing and low re-test rates. This could assist doctors in tuning patient follow-up and therapeutic plans, possibly slowing down disease progression.
Lucia Sacchi, Allan Tucker, Steve Counsell, David F. Garway-Heath, Stephen Swift
Artif. Intell. Medicine5
2013 The Modelling of Glaucoma Progression through the Use of Cellular Automata
Stelios Pavlidis, Stephen Swift, Allan Tucker, Steve Counsell
IDA2
2013 Modelling and analysing the dynamics of disease progression from cross-sectional studies
Stephen Swift, Allan Tucker
J. Biomed. Informatics2
2012 Use of General Purpose GPU Programming to Enhance the Classification of Leukaemia Blast Cells in Blood Smear Images
Stefan Skrobanski, Stelios Pavlidis, Waidah Ismail, Rosline Hassan, Steve Counsell, Stephen Swift
IDA6
2012 A meta-analysis of relationships between organizational characteristics and IT innovation adoption in organizations
Mumtaz Abdul Hameed, Steve Counsell, Stephen Swift
Inf. Manag.3
2012 A Constrained Evolutionary Computation Method for Detecting Controlling Regions of Cortical Networks
abstract
Controlling regions in cortical networks, which serve as key nodes to control the dynamics of networks to a desired state, can be detected by minimizing the eigenratio R and the maximum imaginary part \sigma of an extended connection matrix. Until now, optimal selection of the set of controlling regions is still an open problem and this paper represents the first attempt to include two measures of controllability into one unified framework. The detection problem of controlling regions in cortical networks is converted into a constrained optimization problem (COP), where the objective function R is minimized and \sigma is regarded as a constraint. Then, the detection of controlling regions of a weighted and directed complex network (e.g., a cortical network of a cat), is thoroughly investigated. The controlling regions of cortical networks are successfully detected by means of an improved dynamic hybrid framework (IDyHF). Our experiments verify that the proposed IDyHF outperforms two recently developed evolutionary computation methods in constrained optimization field and some traditional methods in control theory as well as graph theory. Based on the IDyHF, the controlling regions are detected in a microscopic and macroscopic way. Our results unveil the dependence of controlling regions on the number of driver nodes l and the constraint r. The controlling regions are largely selected from the regions with a large in-degree and a small out-degree. When r=+ \infty, there exists a concave shape of the mean degrees of the driver nodes, i.e., the regions with a large degree are of great importance to the control of the networks when l is small and the regions with a small degree are helpful to control the networks when l increases. When r=0, the mean degrees of the driver nodes increase as a function of l. We find that controlling \sigma is becoming more important in controlling a cortical network with increasing l. The methods and results of detecting controlling regions in this paper would promote the coordination and information consensus of various kinds of real-world complex networks including transportation networks, genetic regulatory networks, and social networks, etc.
Yang Tang 0001, Zidong Wang 0001, Huijun Gao, Stephen Swift, Jürgen Kurths
IEEE ACM Trans. Comput. Biol. Bioinform.4
2011 An integrated search-based approach for automatic testing from extended finite state machine (EFSM) models
Abdul Salam Kalaji, Robert M. Hierons, Stephen Swift
Inf. Softw. Technol.3
2010 Detecting Leukaemia (AML) Blood Cells Using Cellular Automata and Heuristic Search
Waidah Ismail, Rosline Hassan, Stephen Swift
IDA3
2010 The effect of cooling functions on ensemble clustering using simulated annealing
abstract
Simulated Annealing (SA) has been adopted by many Ensemble Clustering methods to achieve global combinational optimisation. However the performance of SA is sensitive to the settings of its parameters. Much work has been done for optimising the settings of these parameters over the last two decades , but few of them analysed the behaviour of different cooling functions for Ensemble Clustering. Our work has demonstrated that the clustering results could be invalid if we use SA for Ensemble Clustering without a good understanding of the behaviour of cooling functions. Therefore this paper aims to present the findings of how different cooling functions may affect the performance of Ensemble Clustering methods that use SA. We analyse the effect of cooling functions from three aspects: the convergence rate, the final value of the objective function, and the accuracy of results. Ten different cooling functions are tested on two Ensemble Clustering methods, and thirteen different datasets have been used for the experiments. The findings are particularly helpful for those who are interested in Ensemble Clustering methods as well as those who want to obtain a deep understanding of the behaviour of the cooling functions.
Stephen Swift, Xiaohui Liu 0001
Intell. Data Anal.2
2009 Generating Feasible Transition Paths for Testing from an Extended Finite State Machine (EFSM)
abstract
The problem of testing from an extended finite state machine (EFSM) can be expressed in terms of finding suitable paths through the EFSM and then deriving test data to follow the paths. A chosen path may be infeasible and so it is desirable to have methods that can direct the search for appropriate paths through the EFSM towards those that are likely to be feasible. However, generating feasible transition paths (FTPs) for model based testing is a challenging task and is an open research problem. This paper introduces a novel fitness metric that analyzes data flow dependence among the actions and conditions of the transitions of a path in order to estimate its feasibility. The proposed fitness metric is evaluated by being used in a genetic algorithm to guide the search for FTPs.
Abdul Salam Kalaji, Robert M. Hierons, Stephen Swift
ICST3
2009 An Application of Intelligent Data Analysis Techniques to a Large Software Engineering Dataset
James Cain 0002, Steve Counsell, Stephen Swift, Allan Tucker
IDA3
2009 Multi-Optimisation Consensus Clustering
Stephen Swift, Xiaohui Liu 0001
IDA2
2008 Refactoring Steps, Java Refactorings and Empirical Evidence
abstract
While we can determine the likely testing effort of a single refactoring through simple visual inspection, the inter-relationships between many of the seventy- two refactorings mean that a chain of refactorings and hence a chain of tests may be required for completion of each. In this paper, we establish the properties of, and the inter-relationships between, fourteen of the seventy-two refactorings described in Fowler from a testing chain perspective. We provide an empirical analysis of those refactorings and their associated testing chains. We also inform our understanding of testing effort with recourse to refactoring data from 7 Java OSS.
Steve Counsell, Stephen Swift
COMPSAC2
2007 An improved restricted growth function genetic algorithm for the consensus clustering of retinal nerve fibre data
abstract
This paper describes an extension to the Restricted Growth Function grouping Genetic Algorithm applied to the Consensus Clustering of a retinal nerve fibre layer data-set. Consensus Clustering is an optimisation based method which combines the results of a number of data clustering methods, and is used when it is unknown which clustering method is expected to perform the best. Consensus Clustering has been shown to produce results which are better than the averaged results of the input methods, but could benefit from a more efficient optimisation method. A Restricted Growth Function grouping Genetic Algorithm is a new method of grouping a number of objects into mutually exclusive subsets based upon a fitness function. This method does not suffer from degeneracy, and thus could be applied to the Consensus Clustering problem more efficiently than Simulated Annealing, the current optimisation method. Within this paper it is shown that this type of Genetic Algorithm can indeed improve the performance of Consensus Clustering, and in fact can be improved further by taking advantage of some application specific properties. These findings are demonstrated on a retinal nerve fibre layer data-set and on a synthetic data-set.
Stephen Swift, Allan Tucker, Jason Crampton, David F. Garway-Heath
GECCO1
2007 Efficiency updates for the restricted growth function GA for grouping problems
abstract
Problems that require the partitioning of a set of variables in order to compute a solution such as bin packing or line balancing are typically NP-hard. Hence, researchers have focused on producing heuristic methods for finding appropriate partitions. Many of the representations used in optimisation algorithms including those in GA methods suffer from degeneracy [2]. Furthermore, Falkenauer has found that representations with less degeneracy result in more efficient GAs with respect to grouping problems [1]. Previously we developed a new representation for grouping genetic algorithms called the Restricted Growth Function GA (RGFGA) [3]. The RGFGA effectively removes all degeneracy, resulting in a more efficient search. However, one flaw of the RGFGA is that it converges too quickly resulting in
Allan Tucker, Stephen Swift, Jason Crampton
GECCO2
2007 Mining pathway signatures from microarray data and relevant biological knowledge
Eleftherios Panteris, Stephen Swift, Annette M. Payne, Xiaohui Liu 0001
J. Biomed. Informatics2
2006 Learning short multivariate time series models through evolutionary and sparse matrix computation
Stephen Swift, Joost N. Kok, Xiaohui Liu 0001
Nat. Comput.1
2006 The interpretation and utility of three cohesion metrics for object-oriented design
abstract
The concept of cohesion in a class has been the subject of various recent empirical studies and has been measured using many different metrics. In the structured programming paradigm, the software engineering community has adopted an informal yet meaningful and understandable definition of cohesion based on the work of Yourdon and Constantine. The object-oriented (OO) paradigm has formalised various cohesion measures, but the argument over the most meaningful of those metrics continues to be debated. Yet achieving highly cohesive software is fundamental to its comprehension and thus its maintainability. In this article we subject two object-oriented cohesion metrics, CAMC and NHD, to a rigorous mathematical analysis in order to better understand and interpret them. This analysis enables us to offer substantial arguments for preferring the NHD metric to CAMC as a measure of cohesion. Furthermore, we provide a complete understanding of the behaviour of these metrics, enabling us to attach a meaning to the values calculated by the CAMC and NHD metrics. In addition, we introduce a variant of the NHD metric and demonstrate that it has several advantages over CAMC and NHD. While it may be true that a generally accepted formal and informal definition of cohesion continues to elude the OO software engineering community, there seems considerable value in being able to compare, contrast, and interpret metrics which attempt to measure the same features of software.
Steve Counsell, Stephen Swift, Jason Crampton
ACM Trans. Softw. Eng. Methodol.2
2005 ICARUS: intelligent coupon allocation for retailers using search
abstract
Many retailers run loyalty card schemes for their customers offering incentives in the form of money off coupons. The total value of the coupons depends on how much the customer has spent. This paper deals with the problem of finding the smallest set of coupons such that each possible total can be represented as the sum of a pre-defined number of coupons. A mathematical analysis of the problem leads to the development of a genetic algorithm solution. The algorithm is applied to real world data using several crossover operators and compared to well known straw-person methods. Results are promising showing that considerable time can be saved by using this method, reducing a few days worth of consultancy time to a few minutes of computation.
Stephen Swift, Amy Shi, Jason Crampton, Allan Tucker
Congress on Evolutionary Computation1
2005 An empirical study of the robustness of two module clustering fitness functions
abstract
Two of the attractions of search-based software engineering (SBSE) derive from the nature of the fitness functions used to guide the search. These have proved to be highly robust (for a variety of different search algorithms) and have yielded insight into the nature of the search space itself, shedding light upon the software engineering problem in hand.This paper aims to exploit these two benefits of SBSE in the context of search based module clustering. The paper presents empirical results which compare the robustness of two fitness functions used for software module clustering: one (MQ) used exclusively for module clustering. The other is EVM, a clustering fitness function previously applied to time series and gene expression data.The results show that both metrics are relatively robust in the presence of noise, with EVM being the more robust of the two. The results may also yield some interesting insights into the nature of software graphs.
Mark Harman, Stephen Swift, Kiarash Mahdavi
GECCO2
2005 RGFGA: An Efficient Representation and Crossover for Grouping Genetic Algorithms
abstract
There is substantial research into genetic algorithms that are used to group large numbers of objects into mutually exclusive subsets based upon some fitness function. However, nearly all methods involve degeneracy to some degree. We introduce a new representation for grouping genetic algorithms, the restricted growth function genetic algorithm, that effectively removes all degeneracy, resulting in a more efficient search. A new crossover operator is also described that exploits a measure of similarity between chromosomes in a population. Using several synthetic datasets, we compare the performance of our representation and crossover with another well known state-of-the-art GA method, a strawman optimisation method and a well-established statistical clustering algorithm, with encouraging results.
Allan Tucker, Jason Crampton, Stephen Swift
Evol. Comput.3
2005 A weighted sum validity function for clustering with a hybrid niching genetic algorithm
abstract
Clustering is inherently a difficult problem, both with respect to the construction of adequate objective functions as well as to the optimization of the objective functions. In this paper, we suggest an objective function called the Weighted Sum Validity Function (WSVF), which is a weighted sum of the several normalized cluster validity functions. Further, we propose a Hybrid Niching Genetic Algorithm (HNGA), which can be used for the optimization of the WSVF to automatically evolve the proper number of clusters as well as appropriate partitioning of the data set. Within the HNGA, a niching method is developed to preserve both the diversity of the population with respect to the number of clusters encoded in the individuals and the diversity of the subpopulation with the same number of clusters during the search. In addition, we hybridize the niching method with the k-means algorithm. In the experiments, we show the effectiveness of both the HNGA and the WSVF. In comparison with other related genetic clustering algorithms, the HNGA can consistently and efficiently converge to the best known optimum corresponding to the given data in concurrence with the convergence result. The WSVF is found generally able to improve the confidence of clustering solutions and achieve more accurate and robust results.
Weiguo Sheng 0001, Stephen Swift, Leishi Zhang, Xiaohui Liu 0001
IEEE Trans. Syst. Man Cybern. Part B2
2003 Applying Intelligent Data Analysis to Coupling Relationships in Object-Oriented Software
Steve Counsell, Xiaohui Liu 0001, Rajaa Najjar, Stephen Swift, Allan Tucker
IDA4
2002 Predicting glaucomatous visual field deterioration through short multivariate time series modelling
Stephen Swift, Xiaohui Liu 0001
Artif. Intell. Medicine1
2002 Evolutionary algorithms for grouping high dimensional Email data
Steve Counsell, Xiaohui Liu 0001, Janet McFall, Stephen Swift, Allan Tucker
Intell. Data Anal.4
2002 A framework for modelling virus gene expression data
Paul Kellam, Xiaohui Liu 0001, Nigel J. Martin 0001, Christine A. Orengo, Stephen Swift, Allan Tucker
Intell. Data Anal.5
2001 A Framework for Modelling Short, High-Dimensional Multivariate Time Series: Preliminary Results in Virus Gene Expression Data Analysis
Paul Kellam, Xiaohui Liu 0001, Nigel J. Martin 0001, Christine A. Orengo, Stephen Swift, Allan Tucker
IDA5
2001 Grouping multivariate time series variables: applications to chemical process and visual field data
Stephen Swift, Allan Tucker, Nigel J. Martin 0001, Xiaohui Liu 0001
Knowl. Based Syst.1
2001 Variable grouping in multivariate time series via correlation
abstract
The decomposition of high-dimensional multivariate time series (MTS) into a number of low-dimensional MTS is a useful but challenging task because the number of possible dependencies between variables is likely to be huge. This paper is about a systematic study of the "variable groupings" problem in MTS. In particular, we investigate different methods of utilizing the information regarding correlations among MTS variables. This type of method does not appear to have been studied before. In all, 15 methods are suggested and applied to six datasets where there are identifiable mixed groupings of MTS variables. This paper describes the general methodology, reports extensive experimental results, and concludes with useful insights on the strength and weakness of this type of grouping method.
Allan Tucker, Stephen Swift, Xiaohui Liu 0001
IEEE Trans. Syst. Man Cybern. Part B2
1999 Evolutionary Computation to Search for Strongly Correlated Variables in High-Dimensional Time-Series
Stephen Swift, Allan Tucker, Xiaohui Liu 0001
IDA1