VLDB 2026 Research / reviewers in the wild / expert
Allan Tucker
dblp:53/4404
· DBLP profile ↗
92ranked-venue papers
17as first author
22since 2021 · last 2026
0000-0001-5105-3506ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 13 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 26 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 19 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Combining Dynamic Bayesian Networks with Population Dynamics Modelling to Predict Breeding Success in Seabirds
Alan Anderson, Neda Trifonova, Beth E. Scott, Allan Tucker |
IDA | 4 |
| 2026 | Predicting and Interpolating Spatiotemporal Environmental Data: A Case Study of Groundwater Storage in Bangladesh
Anna Pazola, Mohammad Shamsudduha, Richard G. Taylor, Allan Tucker |
IDA | 4 |
| 2026 | An audit of machine learning experiments on software defect predictionabstractMachine learning algorithms are increasingly being proposed to solve the problem of predicting defect-prone software components. In this literature, computational experiments are the primary means of evaluating and comparing learners and the credibility of findings depends critically on their experimental design and reporting. This paper audits recent software defect prediction (SDP) experiments by assessing their experimental design, analysis and reporting practices against widely accepted norms from statistics, machine learning and empirical software engineering. Our aim is to characterise the current state of practice and evaluate the reproducibility of published findings. We undertook an audit of relevant studies published from the SCOPUS database (2019-2023) focusing on their experimental design and analysis choices e.g., the outcome variables such as F-measure and the type of out of sample (OOS) validation regime, e.g., cross-validation, plus the statistical analysis and inference mechanisms. In all, we evaluated nine different study issues. This was complemented by an assessment of reproducibility using the instrument proposed by González-Barahona and Robles. Our search located approximately 1,585 experiments in SDP (2019-2023), a substantial body of work. From this, we randomly sampled 101 ( $$ \approx 6.4\%$$ ) papers, 61 journal and 40 conference papers. Almost 50% are behind ‘paywalls’. We found considerable divergence in research practice. The number of datasets used ranged 1-365, the number of learners or learner variants evaluated from 1-34 and the number of performance metrics from 1 to 9. Approximately 45% of papers made use of formal statistical inference. We detected a total of 427 issues distributed across 101 papers (median=4) with only one paper being entirely issue-free. In terms of reproducibility, experiments ranged from near perfect to lacking almost all required information. We also found two examples of tortured phrases and potential “paper mill” activity. Approaches to designing and reporting computational experiments varied greatly, but almost half the studies provided insufficient information such that reproduction would be challenging. Overall, our audit suggests that as a research community, we have considerable scope for improvement. Fortunately, many improvements should be neither difficult nor costly to achieve. Giuseppe Destefanis, Leila Yousefi, Martin J. Shepperd, Allan Tucker, Stephen Swift, Steve Counsell, Mahir Arzoky |
Empir. Softw. Eng. | 4 |
| 2025 | Domain-Adversarial Neural Networks to Explore Biases in the Diagnosis of Multiple Eye Conditions from Fundus Image Data
Chiara Pullega, Arianna Dagliati, Allan Tucker |
AIME (2) | 3 |
| 2024 | Predicting Performance Drift in AI Models of Healthcare Without Ground Truth Labels
Ylenia Rotalinti, Puja Myles, Allan Tucker |
IDA (1) | 3 |
| 2023 | The Impact of Bias on Drift Detection in AI Health Software
Asal Khoshravan Azar, Barbara Draghi, Ylenia Rotalinti, Puja Myles, Allan Tucker |
AIME | 5 |
| 2023 | Creating Synthetic Geospatial Patient Data to Mimic Real Data Whilst Preserving Privacy: *2022 35th International Symposium on Computer-Based Medical Systems (CBMS)abstractSynthetic Individual-Level Geospatial Data (SIL-GSD) offers a number of advantages in Spatial Epidemiology when compared to census data or surveys conducted on regional or global levels. The use of SILGSD could bring a new dimension to the study of the patterns and causes of diseases in a particular location while minimizing the risk of patient identity disclosure, especially for rare conditions. Additionally, it could help in building and monitoring regional machine learning models, improving the quality and effectiveness of local healthcare services. Finally, SILGSD could help in controlling the spread and causes of diseases by studying disease movement across areas through the travelling patterns of populations. To our knowledge, no synthetic health records data containing synthesised geographic locations for patients has been published for research purposes so far. Therefore, in this paper we explore generating SILGSD by allocating synthetic patients to general practices (healthcare providers) in the UK using the demographics and prevalence of health conditions in each practice. The assigned general practice locations can be used as proxies for patient locations due to people being registered to their nearest practice from home. We use high-fidelity synthetic primary care patients from the Clinical Practice Research Datalink (CPRD) and allocate them to England's general practices (GPs), using the publicly available GP health conditions statistics from the Quality and Outcomes Framework (QOF). The allocation relies on similarities between patients in different locations without using real location information for the patients. We demonstrate that the Allocation Data is able to accurately mimic the real health conditions distribution in the general practices and also preserves the underlying distribution of the original primary care patients data from CPRD (Gold Standard). Dima R. Alattal, Zhenchen Wang, Puja Myles, Allan Tucker |
CBMS | 4 |
| 2023 | Privacy Assessment of Synthetic Patient DataabstractIn this paper, we quantify the privacy gain of synthetic patient data drawn from two generative models, MST and PrivBayes, which is based on real anonymized primary care patient data. This evaluation is implemented for two types of inference attacks, namely membership and attribute inference attacks using a new toolbox, TAPAS. The aim is to quantitatively evaluate the privacy gain of each attack where these two differentially private generators and different threat models are used with a focus on black-box knowledge. The evaluation that was carried out in this paper demonstrates that vulnerabilities of synthetic patient data depend on the different attack scenarios, threat models, and algorithms used to generate the synthetic patient data. It was shown empirically that although the synthetic patient data achieved high privacy gain in most attack scenarios, it does not behave uniformly against adversarial attacks, and some records and outliers remain vulnerable depending on the attack scenario. Moreover, it was shown that the PrivBayes generator is the more robust generator in comparison to MST in terms of the privacy-preservation of synthetic data. Ferdoos Hossein Nezhad, Ylenia Rotalinti, Puja Myles, Allan Tucker |
CBMS | 4 |
| 2022 | Estimating the Optimal Number of Clusters from Subsets of EnsemblesabstractThis research estimates the optimal number of clusters in a dataset using a novel ensemble technique - a preferred alternative to relying on the output of a single clustering. Combining clusterings from different algorithms can lead to a more stable and robust solution, often unattainable by any single clustering solution. Technically, we created subsets of ensembles as possible estimates; and evaluated them using a quality metric to obtain the best subset. We tested our method on publicly available datasets of varying types, sources and clustering difficulty to establish the accuracy and performance of our approach against eight standard methods. Our method outperforms all the techniques in the number of clusters estimated correctly. Due to the exhaustive nature of the initial algorithm, it is slow as the number of ensembles or the solution space increases; hence, we have provided an updated version based on the single-digit difference of Gray code that runs in linear time in terms of the subset size. Afees Adegoke Odebode, Allan Tucker, Mahir Arzoky, Stephen Swift |
DATA | 2 |
| 2022 | Exploring the Explicit Modelling of Bias in Machine Learning Classifiers: A Deep Multi-label ConvNet ApproachabstractThis paper addresses the problem that many machine learning classifiers make decisions based on data that are biased and can therefore result in prejudiced decisions. For example, in education (which this paper focuses on) a student may be rejected from a course based on historical decisions in the data that only exist due to historical biases in society or due to the skewed sampling of the data. Other approaches to dealing with bias in data include resampling methods (to counter imbalanced samples) and dimensionality reduction (to focus only on relevant features to the classification task). In this paper, we explore issues of modelling bias explicitly so that we can identify the types of bias and whether they are accounting for inflated predictive accuracies. In particular, we compare graphical model approaches to building classifiers, that are transparent in how they make decisions, with two forms of Deep Multi-label Convolutional Neural Networks to investigate if models can be built that maximise accuracy and minimise bias. We carry out this comparison on student entry and performance data from a higher educational institution. Mashael Al-Luhaybi, Stephen Swift, Steve Counsell, Allan Tucker |
ICMLA | 4 |
| 2022 | dunXai: DO-U-Net for Explainable (Multi-label) Image Classification - Applications to Biomedical Images
Toyah Overton, Allan Tucker, Tim James, Dimitar Hristozov |
IDA | 2 |
| 2022 | Identifying latent variables in Dynamic Bayesian Networks with bootstrapping applied to Type 2 Diabetes complication predictionabstractPredicting complications associated with complex disease is a challenging task given imbalanced and highly correlated disease complications along with unmeasured or latent factors. To analyse the complications associated with complex disease, this article attempts to deal with complex imbalanced clinical data, whilst determining the influence of latent variables within causal networks generated from the observation. This work proposes appropriate Intelligent Data Analysis methods for building Dynamic Bayesian networks with latent variables, applied to small-sized clinical data (a case of Type 2 Diabetes complications). First, it adopts a Time Series Bootstrapping approach to re-sample the rare complication class with a replacement with respect to the dynamics of disease progression. Then, a combination of the Induction Causation algorithm and Link Strength metric (which is called IC*LS approach) is applied on the bootstrapped data for incrementally identifying latent variables. The most highlighted contribution of this paper gained insight into the disease progression by interpreting the latent states (with respect to the associated distributions of complications). An exploration of inference methods along with confidence interval assessed the influences of these latent variables. The obtained results demonstrated an improvement in the prediction performance. Leila Yousefi, Allan Tucker |
Intell. Data Anal. | 2 |
| 2022 | New JBI policy emphasizes clinically-meaningful novel machine learning methods
Allan Tucker, Thomas George Kannampallil, Samah Jamal Fodeh, Mor Peleg |
J. Biomed. Informatics | 1 |
| 2021 | Uncertainty Estimation in SARS-CoV-2 B-Cell Epitope Prediction for Vaccine Development
Bhargab Ghoshal, Biraja Ghoshal, Stephen Swift, Allan Tucker |
AIME | 4 |
| 2021 | Bayesian Deep Active Learning for Medical Image Analysis
Biraja Ghoshal, Stephen Swift, Allan Tucker |
AIME | 3 |
| 2021 | On Cost-Sensitive Calibrated Uncertainty in Deep Learning: An application on COVID-19 detectionabstractReliable and cost-sensitive calibrated estimated uncertainty in deep learning is important in many real-world applications where safety is critical, and prediction problems are asymmetric, in the sense that different types of misclassification errors incur different costs or significant losses which may result in the loss of life in some circumstances. However, uncertainty obtained by approximate inference techniques, such as variational inference, cannot guarantee optimal predictions to represent the model error and is prone to miscalibration (and often poor calibration) due to the assumption of the constant cost of misclassification, which is not realistic in medical diagnosis. Knowing how much confidence there is in a prediction is essential for gaining clinicians' trust in the technology. Bayesian decision theory provides a principled approach for optimal decision making under uncertainty, given a utility function over actions. We propose a variational inference with Monte Carlo Drop-weights based Bayesian neural networks model, which means cost-sensitive calibrated predictive uncertainty can be estimated while minimising asymmetric cost as an expected utility function with improved accuracy. We measured bias-corrected uncertainty using Jackknife resampling technique and propose uncertainty estimation performance metrics, including risk coverage curve, which directly corresponds to well-calibrated estimated uncertainty performance. We have highlighted potential issues in commonly used performance metrics, calibration measures, the quality of the estimated uncertainty and proposed revised metrics to mitigate them. We evaluated the effectiveness of our approach using X-Ray images detecting Covid-19 to improve the reliability of computer-based diagnostics. Biraja Ghoshal, Allan Tucker |
CBMS | 2 |
| 2021 | Exploiting Clinical Staging Data to Constrain Pseudo-Time Modelling of Disease ProgressionabstractPseudo Time methods enable the construction of time-based models from non-temporal cross-sectional data. This means that temporal characteristics of disease can be inferred. However, the success of these approaches is dependent on appropriate distance metrics and labelling to guide the trajectory modelling. Clinical staging information, such as “early stage” and “advanced stage” of disease can be exploited to constrain the construction of pseudo time models to ensure more realistic trajectories are captured. In this paper we explore how clinical staging information can be used in this way on simulated data and on breast cancer data. Using the simulated data, we show how more precise estimates can be made of the underlying transition parameters in a model derived from constrained pseudo time methods, by preventing unrealistic transitions. The breast cancer pseudo time models are constrained based on uniformity of cell size, a proxy to disease staging, and this is shown to result in models that better represent the monotonically increasing symptoms over time. Seyed Erfan Sajjadi, Allan Tucker |
CBMS | 2 |
| 2021 | Evaluating a Longitudinal Synthetic Data Generator using Real World DataabstractSynthetic data offer a number of advantages over using ground truth data when working with private and personal information about individuals. Firstly, the risk of identifying individuals is reduced considerably, which enables the sharing of data for analysis amongst more organisations. Secondly, the fine tuning of synthetic datapoints to suit particular modelling and analyses could help to build more suitable models that can avoid biases found in the original ground truth data. In this paper we explore how a probabilistic synthetic data generator can be used to model data with high enough fidelity that it can be used to develop and validate state-of-the-art machine learning models. In particular, we use a Bayesian network model trained on gestational diabetes data, generated from a mobile health app collected from a number of health trusts in the UK. These data are used to train and test an established machine learning model developed by Sensyne Health using real-world data, and the resulting performance is compared to performance on ground truth data. In addition, a clinical validation is undertaken to explore if human experts can differentiate real patients from synthetic ones. We demonstrate that the Bayesian network synthetic data generator is able to mimic the ground truth closely enough to make it difficult for a human expert to distinguish between the two. We show that the data generator captures the interactions between features and the multivariate distributions close enough to enable classifiers to be inferred that imitate the key performance characteristics of models inferred from ground truth data. What is more, we demonstrate that the discovered mis-classifications found when testing using the synthetic data, are as informative as when testing using ground truth data. Zhenchen Wang, Puja Myles, Anu Jain, James L. Keidel, Roberto Liddi, Lucy Mackillop, Carmelo Velardo, Allan Tucker |
CBMS | 8 |
| 2021 | Hyperspherical Weight Uncertainty in Neural Networks
Biraja Ghoshal, Allan Tucker |
IDA | 2 |
| 2021 | Estimating uncertainty in deep learning for reporting confidence to clinicians in medical image segmentation and diseases detectionabstractAbstract Deep learning (DL), which involves powerful black box predictors, has achieved a remarkable performance in medical image analysis, such as segmentation and classification for diagnosis. However, in spite of these successes, these methods focus exclusively on improving the accuracy of point predictions without assessing the quality of their outputs. Knowing how much confidence there is in a prediction is essential for gaining clinicians' trust in the technology. In this article, we propose an uncertainty estimation framework, called MC‐DropWeights, to approximate Bayesian inference in DL by imposing a Bernoulli distribution on the incoming or outgoing weights of the model, including neurones. We demonstrate that by decomposing predictive probabilities into two main types of uncertainty, aleatoric and epistemic, using the Bayesian Residual U‐Net (BRUNet) in image segmentation. Approximation methods in Bayesian DL suffer from the “mode collapse” phenomenon in variational inference. To address this problem, we propose a model which Ensembles of Monte‐Carlo DropWeights by varying the DropWeights rate. In segmentation, we introduce a predictive uncertainty estimator, which takes the mean of the standard deviations of the class probabilities associated with every class. However, in classification, we need an alternative approach since the predictive probabilities from a forward pass through the model does not capture uncertainty. The entropy of the predictive distribution is a measure of uncertainty, but its exponential depends on sample size. The plug‐in estimate in mutual information is subject to sampling bias. We propose Jackknife resampling, to correct for sample bias, which improves estimating uncertainty quality in image classification. We demonstrate that our deep ensemble MC‐DropWeights method, using the bias‐corrected estimator produces an equally good or better result in both quantified uncertainty estimation and quality of uncertainty estimates than approximate Bayesian neural networks in practice. Biraja Ghoshal, Allan Tucker, Bal Sanghera, Wai Lup Wong |
Comput. Intell. | 2 |
| 2021 | Generating and evaluating cross-sectional synthetic electronic healthcare data: Preserving data utility and patient privacyabstractAbstract Electronic healthcare record data have been used to study risk factors of disease, treatment effectiveness and safety, and to inform healthcare service planning. There has been increasing interest in utilizing these data for new purposes such as for machine learning to develop predictive algorithms to aid diagnostic and treatment decisions. Synthetic data could potentially be an alternative to real‐world data for these purposes as well as reveal any biases in the data used for algorithm development. This article discusses the key requirements of synthetic data for multiple purposes and proposes an approach to generate and evaluate synthetic data focused on, but not limited to, cross‐sectional healthcare data. To our knowledge, this is the first article to propose a framework to generate and evaluate synthetic healthcare data with the aim of simultaneously preserving the complexities of ground truth data in the synthetic data while also ensuring privacy. We include findings and new insights from synthetic datasets modeled on both the Indian liver patient dataset and UK primary care dataset to demonstrate the application of this framework under different scenarios. Zhenchen Wang, Puja Myles, Allan Tucker |
Comput. Intell. | 3 |
| 2021 | Opening the black box: Personalizing type 2 diabetes patients based on their latent phenotype and temporal associated complication rulesabstractAbstract It is widely considered that approximately 10% of the population suffers from type 2 diabetes. Unfortunately, the impact of this disease is underestimated. Patient's mortality often occurs due to complications caused by the disease and not the disease itself. Many techniques utilized in modeling diseases are often in the form of a “black box” where the internal workings and complexities are extremely difficult to understand, both from practitioners' and patients' perspective. In this work, we address this issue and present an informative model/pattern, known as a “latent phenotype,” with an aim to capture the complexities of the associated complications' over time. We further extend this idea by using a combination of temporal association rule mining and unsupervised learning in order to find explainable subgroups of patients with more personalized prediction. Our extensive findings show how uncovering the latent phenotype aids in distinguishing the disparities among subgroups of patients based on their complications patterns. We gain insight into how best to enhance the prediction performance and reduce bias in the models applied using uncertainty in the patients' data. Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker |
Comput. Intell. | 6 |
| 2020 | Using the Lexicon from Source Code to Determine Application DomainabstractContext: The vast majority of software engineering research is reported independently of the application domain: techniques and tools usage is reported without any domain context. As reported in previous research, this has not always been so: early in the computing era, the research focus was frequently application domain specific (for example, scientific and data processing). Andrea Capiluppi, Nemitari Ajienka, Nour Ali, Mahir Arzoky, Steve Counsell, Giuseppe Destefanis, Alina Dana Miron, Bhaveet Nagaria, Rumyana Neykova, Martin J. Shepperd, Stephen Swift, Allan Tucker |
EASE | 12 |
| 2020 | Estimating Uncertainty in Deep Learning for Reporting Confidence: An Application on Cell Type Prediction in Testes Based on ProteomicsabstractMulti-label classification in deep learning is a practical yet challenging task, because class overlaps in the feature space means that each instance is associated with multiple class labels. This requires a prediction of more than one class category for each input instance. To the best of our knowledge, this is the first deep learning study which quantifies uncertainty and model interpretability in multi-label classification; as well as applying it to the problem of recognising proteins expressed in cell types in testes based on immunohistochemically stained images. Multi-label classification is achieved by thresholding the class probabilities, with the optimal thresholds adaptively determined by a grid search scheme based on Matthews correlation coefficients. We adopt MC-Dropweights to approximate Bayesian Inference in multi-label classification to evaluate the usefulness of estimating uncertainty with predictive score to avoid overconfident, incorrect predictions in decision making. Our experimental results show that the MC-Dropweights visibly improve the performance to estimate uncertainty compared to state of the art approaches. Biraja Ghoshal, Cecilia Lindskog, Allan Tucker |
IDA | 3 |
| 2020 | DO-U-Net for Segmentation and Counting - Applications to Satellite and Medical ImagesabstractMany image analysis tasks involve the automatic segmentation and counting of objects with specific characteristics. However, we find that current approaches look to either segment objects or count them through bounding boxes, and those methodologies that both segment and count struggle with co-located and overlapping objects. This restricts our capabilities when, for example, we require the area covered by particular objects as well as the number of those objects present, especially when we have a large amount of images to obtain this information for. In this paper, we address this by proposing a Dual-Output U-Net. DO-U-Net is an Encoder-Decoder style, Fully Convolutional Network (FCN) for object segmentation and counting in image processing. Our proposed architecture achieves precision and sensitivity superior to other, similar models by producing two target outputs: a segmentation mask and an edge mask. Two case studies are used to demonstrate the capabilities of DO-U-Net: locating and counting Internally Displaced People (IDP) tents in satellite imagery, and the segmentation and counting of erythrocytes in blood smears. The model was demonstrated to work with a relatively small training dataset, achieving a sensitivity of 98.69% for IDP camps of the fixed resolution, and 94.66% for a scale-invariant IDP model. DO-U-Net achieved a sensitivity of 99.07% on the erythrocytes dataset. DO-U-Net has a reduced memory footprint, allowing for training and deployment on a machine with a lower to mid-range GPU, making it accessible to a wider audience, including non-governmental organisations (NGOs) providing humanitarian aid, as well as health care organisations. Toyah Overton, Allan Tucker |
IDA | 2 |
| 2020 | Using topological data analysis and pseudo time series to infer temporal phenotypes from electronic health recordsabstractTemporal phenotyping enables clinicians to better understand observable characteristics of a disease as it progresses. Modelling disease progression that captures interactions between phenotypes is inherently challenging. Temporal models that capture change in disease over time can identify the key features that characterize disease subtypes that underpin these trajectories. These models will enable clinicians to identify early warning signs of progression in specific sub-types and therefore to make informed decisions tailored to individual patients. In this paper, we explore two approaches to building temporal phenotypes based on the topology of data: topological data analysis and pseudo time-series. Using type 2 diabetes data, we show that the topological data analysis approach is able to identify disease trajectories and that pseudo time-series can infer a state space model characterized by transitions between hidden states that represent distinct temporal phenotypes. Both approaches highlight lipid profiles as key factors in distinguishing the phenotypes. Arianna Dagliati, Nophar Geifman, Niels Peek, John H. Holmes, Lucia Sacchi, Riccardo Bellazzi, Seyed Erfan Sajjadi, Allan Tucker |
Artif. Intell. Medicine | 8 |
| 2020 | A multilevel graph approach for rainfall forecasting: A preliminary study case on London areaabstractSummary Increasing populations and rapid large‐scale urbanization has created a demand to increase the quality of life through economic development, social stability, and better quality environments. These issues are addressed in the field of Smart Cities where, through the Internet of Things, efforts are being made to support added‐value services for the administration of the city and for citizens. The continuous exchange of information inevitably produces a huge amount of data, which demands analyses of data using unconventional methods within a Big Data context. How can we properly process these data? How can we properly use these data in order to increase the competitiveness and efficiency of services, and how could they contribute to social development? Services that could be useful in this field include Early Warning Systems. Information management environments, or more generally pervasive data contexts, may be supported by context representation approaches and enhanced through adopting probabilistic approaches such as Context Dimension Tree, Ontology, and Bayesian Network. The aim of this work is to introduce and explain a methodology for merging CDTs and Ontologies, and probabilistic approach based on BNs in order to help expert users handle emergencies and provide suggestions for improving the liveability of cities for their inhabitants. Fabio Clarizia, Francesco Colace, Massimo De Santo, Marco Lombardi 0001, Francesco Pascale, Domenico Santaniello, Allan Tucker |
Concurr. Comput. Pract. Exp. | 7 |
| 2019 | Predicting Academic Performance: A Bootstrapping Approach for Learning Dynamic Bayesian Networks
Mashael Al-Luhaybi, Leila Yousefi, Stephen Swift, Steve Counsell, Allan Tucker |
AIED (1) | 5 |
| 2019 | Inferring Temporal Phenotypes with Topological Data Analysis and Pseudo Time-Series
Arianna Dagliati, Nophar Geifman, Niels Peek, John H. Holmes, Lucia Sacchi, Seyed Erfan Sajjadi, Allan Tucker |
AIME | 7 |
| 2019 | Latent Class Multi-Label Classification to Identify Subclasses of Disease for Improved PredictionabstractDisease subtyping can assist the development of precision medicine but remains a challenge in data analysis by reason of the many different methods to group individuals depending on their data. However, identification of subclasses of disease will help to produce better models which are more specific to patients and will improve prediction and interpretation of underlying characteristics of disease. This paper presents a novel algorithm that integrates latent class models with supervised learning. The new algorithm uses latent class models to cluster patients within groups that results in improved classification as well as aiding the understanding of the dissimilarities of the discovered groups. The methods are tested on data from patients with Systemic Sclerosis (SSc), a rare potentially fatal condition. Results show that the "Latent Class Multi-Label Classification Model" improves accuracy when compared with competitive similar methods. Awad Alsaid Alyousef, Svetlana I. Nihtyanova, Christopher P. Denton, Pietro Bosoni, Riccardo Bellazzi, Allan Tucker |
CBMS | 6 |
| 2019 | Retinal OCT Segmentation Using Fuzzy Region Competition and Level Set MethodsabstractOptical coherence tomography (OCT) is a noninvasive imaging modality that provides in-depth images of the retina. Properties of individual layers on OCT have become important markers for diagnosing and tracking medication of various eye diseases in current ophthalmology. Manual segmentation of OCT scans posed many challenges (errors, inconsistency), which can be addressed by automated segmentation methods. Level set method is one of the most popular methods in the literature used for this purpose. Although level set methods have a fundamental way of handling topological changes, the weak boundaries and noise in addition to inhomogeneity in OCT images make it difficult to segment the layers accurately. Inspired by the concept of region competition, we incorporate prior knowledge of the retinal structure to segment nine (9) layers of the retina. Mainly, we establish a specific region of interest, then use selected components from fuzzy C-Means for initialisation. The clustering in the initialisation stage is also used to guide the evolution through; a Mumford-Shah (MS) selective region competition force and a Hamilton-Jacobi (HJ) balloon force. The forces ensure evolution close to actual retinal boundaries. Finally, the convergence of the method is based on an improved HJ object indication function influenced by the fuzzy membership to prevent leakages at weak boundaries. Experimental results are promising based on 200 OCT images. Bashir I. Dodo, Yongmin Li 0001, Allan Tucker, Djibril Kaba, Xiaohui Liu 0001 |
CBMS | 3 |
| 2019 | Estimating Uncertainty in Deep Learning for Reporting Confidence to Clinicians when Segmenting Nuclei Image DataabstractDeep Learning, which involves powerful black box predictors, has achieved a state-of-the-art performance in medical image analysis such as segmentation and classification for diagnosis. However, in spite of these successes, these methods focus exclusively on improving the accuracy of point predictions without assessing the quality of their outputs. Knowing how much confidence there is in a prediction is essential for gaining clinicians' trust in the technology. Monte-Carlo dropout in neural networks is equivalent to a specific variational approximation in Bayesian neural networks and is simple to implement without any changes in the network architecture. It is considered state-of-the-art for estimating uncertainty. However, in classification, it does not model the predictive probabilities. This means that we are not capturing the true underlying uncertainty in the prediction. In this paper, we propose an uncertainty estimation framework for classification by decomposing predictive probabilities into two main types of uncertainty in Bayesian modelling: aleatoric and epistemic uncertainty (representing uncertainty in the quality of the data and in the model parameters, respectively). We demonstrate that the proposed uncertainty quantification framework using the Bayesian Residual U-Net (BRUNet) provides additional insight for clinicians when analysing images with help from deep learners. In addition, we demonstrate how the resulting uncertainty depends on the dropout rates using images from nuclei in divergent medical images. Biraja Ghoshal, Allan Tucker, Bal Sanghera, Wai Lup Wong |
CBMS | 2 |
| 2019 | Generating and Evaluating Synthetic UK Primary Care Data: Preserving Data Utility & Patient PrivacyabstractThere is increasing interest in the potential of synthetic data to validate and benchmark machine learning algorithms as well as reveal any biases in real-world data used for algorithm development. This paper discusses the key requirements of synthetic data for such purposes and proposes an approach to generating and evaluating synthetic data that meets these requirements. We propose a framework to generate and evaluate synthetic data with the aim of simultaneously preserving the complexities of ground truth data in the synthetic data whilst also ensuring privacy. We include as a case study, a proof-of-concept synthetic dataset modelled on UK primary care data to demonstrate the application of this framework. Zhenchen Wang, Puja Myles, Allan Tucker |
CBMS | 3 |
| 2019 | Opening the Black Box: Exploring Temporal Pattern of Type 2 Diabetes Complications in Patient Clustering Using Association Rules and Hidden Variable DiscoveryabstractThere is a great deal of debate over the importance of explanation in AI models inferred from health data. In particular, there is a balance that needs to be made between the accuracy of complex 'deep' models such as convolutional neural networks and the transparency of models that aim to model data in a more 'human' way such as expert systems. In this paper, we explore the use of temporal association rules to validate and uncover the meaning behind discrete hidden variables that have been inferred from clinical diabetes data. We use a recently published technique based upon the IC* (Induction Causation) algorithm that limits the number of hidden variables and places them within a network structure. Here, we take the hidden variables and compare their underlying discrete states to clusters that have been generated from temporal association rules. This allows us to characterise the hidden states based upon different sequences of complications. Results are very promising, with many hidden states aligning with the discovered clusters giving us a direct interpretation. Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker |
CBMS | 6 |
| 2019 | The Prevalence of Errors in Machine Learning Experiments
Martin J. Shepperd, Ning Li 0022, Mahir Arzoky, Andrea Capiluppi, Steve Counsell, Giuseppe Destefanis, Stephen Swift, Allan Tucker, Leila Yousefi |
IDEAL (1) | 9 |
| 2018 | Opening the Black Box: Discovering and Explaining Hidden Variables in Type 2 Diabetic Patient Modelling
Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker |
BIBM | 6 |
| 2018 | Predicting Disease Complications Using a Stepwise Hidden Variable Approach for Learning Dynamic Bayesian NetworksabstractPredicting Diabetes Type 2 Mellitus (T2DM) complications such as retinopathy and liver disease is still a challenge despite being a growing public health concern worldwide. This is due to the complex interactions between complications and other features, as well as between the different complications, themselves. What is more, there are likely to be many unmeasured effects that impact the disease progression of different patients. Probabilistic graphical models such as Dynamic Bayesian Networks (DBNs) have demonstrated much promise in the modeling of disease progression and they can naturally incorporate hidden (latent) variables using the EM algorithm. Unlike deep learning approaches that attempt to model complex interactions in data by using a large number of hidden variables, we adopt a different approach. We are interested in models that not only capture unmeasured effects but are also transparent in how they model data so that knowledge about disease processes can be extracted and trust in the model can be maintained by clinicians. As a result, we have developed a step-wise hidden variable structure learning process that incrementally adds hidden variables based on the IC* algorithm. To the best of our knowledge, this is the first study for classifying disease complication using a step-wise learning methodology for identifying hidden and T2DM features with a DBN structure from clinical data. Our extensive set of experiments show that the proposed method improves classification accuracy, identifying the correct number of hidden variables, and targeting their precise location within the network structure. Leila Yousefi, Allan Tucker, Mashael Al-Luhaybi, Lucia Sacchi, Riccardo Bellazzi, Luca Chiovato |
CBMS | 2 |
| 2018 | Specimens as Research Objects: Reconciliation Across Distributed Repositories to Enable Metadata PropagationabstractBotanical specimens are shared as long-term consultable research objects in a global network of specimen repositories. Multiple specimens are generated from a shared field collection event; generated specimens are then managed individually in separate repositories and independently augmented with research and management metadata which could be propagated to their duplicate peers. Establishing a data-derived network for metadata propagation will enable the reconciliation of closely related specimens which are currently dispersed, unconnected and managed independently. Following a data mining exercise applied to an aggregated dataset of 19,827,998 specimen records from 292 separate specimen repositories, 36% or 7,102,710 specimens are assessed to participate in duplication relationships, allowing the propagation of metadata among the participants in these relationships, totalling: 93,044 type citations, 1,121,865 georeferences, 1,097,168 images and 2,191,179 scientific name determinations. The results enable the creation of networks to identify which repositories could work in collaboration. Some classes of annotation (particularly those regarding scientific name determinations) represent units of scientific work: appropriate management of this data would allow the accumulation of scholarly credit to individual researchers: potential further work in this area is discussed. Nicky Nicolson, Alan Paton, Sarah Phillips, Allan Tucker |
eScience | 4 |
| 2018 | Intrusion Detection Using Transfer Learning in Machine Learning Classifiers Between Non-cloud and Cloud Datasets
Roja Ahmadi, Robert D. Macredie, Allan Tucker |
IDEAL (1) | 3 |
| 2017 | A Deconstructed Replication of a Time of Test Study Using the AGIS MetricabstractIn medical practice, glaucoma severity is usually measured using the Advanced Glaucoma Intervention Studies (AGIS) metric. In a previous study [2], we replicated the work of Montolio et al., [5] and demonstrated that, for a larger dataset, time of day of test using the AGIS metric did make a difference to the measurement of glaucoma, supporting Montolio et als work. However, in our earlier study, we used the AGIS scores for both eyes combined. In this paper, we use the measurement from just one eye at a time. A dataset of 14389 left eye AGIS scores and the same number for the right eye from 2468 Moorfield Eye Hospital patients was used as the empirical basis. We then re-compared time of test results with those of Montolios study. Results revealed that using the values from just one eye (as opposed to both) may give a distorted picture of the AGIS scores; differences in the same time period were found between the two eyes. This may have implications for choice of sampling data and analysis of glaucoma using the AGIS metric. Steve Counsell, Stephen Swift, Allan Tucker |
CBMS | 3 |
| 2017 | Predicting Comorbidities Using Resampling and Dynamic Bayesian Networks with Latent VariablesabstractComorbidities such as hypertension and lipid metabolism are often associated in diseases such as diabetes, and the early prediction of these is of great value when trying to manage progression. This is the start of a project to model multiple comorbidities in diabetes using dynamic Bayesian networks with latent variables in order to stratify patient cohorts. In this paper, we demonstrate some initial results on a dataset where the class imbalance problem poses an issue due to the rare occurrence of different individual comorbidities on a visit-by-visit basis. This is dealt with using a bootstrap technique that has been specifically designed for longitudinal data where the occurrence of the positive class occurs far less than the negative. Leila Yousefi, Lucia Sacchi, Riccardo Bellazzi, Luca Chiovato, Allan Tucker |
CBMS | 5 |
| 2017 | Identifying Novel Features from Specimen Data for the Prediction of Valuable Collection Trips
Nicky Nicolson, Allan Tucker |
IDA | 2 |
| 2017 | Updating Markov models to integrate cross-sectional and longitudinal studies
Allan Tucker, David F. Garway-Heath |
Artif. Intell. Medicine | 1 |
| 2017 | Reprint of "Updating Markov models to integrate cross-sectional and longitudinal studies"
Allan Tucker, David F. Garway-Heath |
Artif. Intell. Medicine | 1 |
| 2016 | Combining Unsupervised and Supervised Learning for Discovering Disease SubclassesabstractDiseases are often umbrella terms for many subcategories of disease. The identification of these subcategories is vital if we are to develop personalised treatments that are better focussed on individual patients. In this short paper, we explore the use of a combination of unsupervised learning to identify potential subclasses, and supervised learning to build models for better predicting a number of different health outcomes for patients that suffer from systemic sclerosis, a rare chronic connective tissue disorder - but one that shares many characteristics with other diseases. We explore a number of different algorithms for constructing models that simultaneously predict health outcomes and identify subcategories. Pietro Bosoni, Allan Tucker, Riccardo Bellazzi, Svetlana I. Nihtyanova, Christopher P. Denton |
CBMS | 2 |
| 2016 | The AGIS Metric and Time of Test: A Replication StudyabstractVisual Field (VF) tests and corresponding data are commonly used in clinical practices to manage glaucoma. The standard metric used to measure glaucoma severity is the Advanced Glaucoma Intervention Studies (AGIS) metric. We know that time of day when VF tests are applied can influence a patient's AGIS metric value; a previous study showed that this was the case for a data set of 160 patients. In this paper, we replicate that study using data from 2468 patients obtained from Moorfields Eye Hospital. This may provide further evidence and support of this phenomenon in a replication sense. Results did indeed show a tendency for the metric to be lower for early onset patients in the morning; equally, for advanced patients, the effect was less pronounced. We thus found support for the earlier work of Montolio et al. [4] and add to the body of evidence on the AGIS metric. Steve Counsell, Stephen Swift, Allan Tucker |
CBMS | 3 |
| 2016 | Simultaneous Modelling and Clustering of Visual Field DataabstractThis thesis was submitted for the award of Doctor of Philosophy and was awarded by Brunel University London Mohd Zairul Mazwan Bin Jilani, Allan Tucker, Stephen Swift |
CBMS | 2 |
| 2015 | Updating Stochastic Networks to Integrate Cross-Sectional and Longitudinal Studies
Allan Tucker |
AIME | 1 |
| 2015 | Quantifying StockTwits semantic terms' trading behavior in financial markets: An effective application of decision tree algorithmsabstractGrowing evidence is suggesting that postings on online stock forums affect stock prices, and alter investment decisions in capital markets, either because the postings contain new information or they might have predictive power to manipulate stock prices. In this paper, we propose a new intelligent trading support system based on sentiment prediction by combining text-mining techniques, feature selection and decision tree algorithms in an effort to analyze and extract semantic terms expressing a particular sentiment (sell, buy or hold) from stock-related micro-blogging messages called “StockTwits”. An attempt has been made to investigate whether the power of the collective sentiments of StockTwits might be predicted and how the changes in these predicted sentiments inform decisions on whether to sell, buy or hold the Dow Jones Industrial Average (DJIA) Index. In this paper, a filter approach of feature selection is first employed to identify the most relevant terms in tweet postings. The decision tree (DT) model is then built to determine the trading decisions of those terms or, more importantly, combinations of terms based on how they interact. Then a trading strategy based on a predetermined investment hypothesis is constructed to evaluate the profitability of the term trading decisions extracted from the DT model. The experiment results based on 122-tweet term trading (TTT) strategies achieve a promising performance and the (TTT) strategies dramatically outperform random investment strategies. Our findings also confirm that StockTwits postings contain valuable information and lead trading activities in capital markets. Alya Al Nasseri, Allan Tucker, Sergio de Cesare |
Expert Syst. Appl. | 2 |
| 2014 | Integrating Clinical Data from Cross-Sectional and Longitudinal StudiesabstractClinical trials are typically conducted over a population in order to illuminate certain characteristics of a health issue or disease process. These cross-sectional studies provide a snapshot of these disease processes over a large population but do not allow us to model the temporal nature of disease. Longitudinal studies on the other hand, are used to explore how these processes develop over time but can be expensive and time-consuming, and only cover a relatively small window within the disease process. This paper explores a technique for integrating cross-sectional and longitudinal studies to build models of disease progression. Allan Tucker |
CBMS | 2 |
| 2014 | Big Data Analysis of StockTwits to Predict Sentiments in the Stock Market
Alya Al Nasseri, Allan Tucker, Sergio de Cesare |
Discovery Science | 2 |
| 2014 | Incorporating Regime Metrics into Latent Variable Dynamic Models to Detect Early-Warning Signals of Functional Changes in Fisheries Ecology
Neda Trifonova, Daniel Duplisea, Andrew Kenny, David L. Maxwell, Allan Tucker |
Discovery Science | 5 |
| 2014 | Comparing Pre-defined Software Engineering Metrics with Free-Text for the Prediction of Code 'Ripples'
Steve Counsell, Allan Tucker, Stephen Swift, Guy Fitzgerald, Jason Peters |
IDA | 2 |
| 2014 | A Spatio-temporal Bayesian Network Approach for Revealing Functional Ecological Networks in Fisheries
Neda Trifonova, Daniel Duplisea, Andrew Kenny, Allan Tucker |
IDA | 4 |
| 2014 | Extracting Predictive Models from Marked-Up Free-Text Documents at the Royal Botanic Gardens, Kew, London
Allan Tucker, Don Kirkup |
IDA | 1 |
| 2014 | Improving predictive models of glaucoma severity by incorporating quality indicatorsabstractOBJECTIVE: In this paper we present an evaluation of the role of reliability indicators in glaucoma severity prediction. In particular, we investigate whether it is possible to extract useful information from tests that would be normally discarded because they are considered unreliable. METHODS: We set up a predictive modelling framework to predict glaucoma severity from visual field (VF) tests sensitivities in different reliability scenarios. Three quality indicators were considered in this study: false positives rate, false negatives rate and fixation losses. Glaucoma severity was evaluated by considering a 3-levels version of the Advanced Glaucoma Intervention Study scoring metric. A bootstrapping and class balancing technique was designed to overcome problems related to small sample size and unbalanced classes. As a classification model we selected Naïve Bayes. We also evaluated Bayesian networks to understand the relationships between the different anatomical sectors on the VF map. RESULTS: The methods were tested on a data set of 28,778 VF tests collected at Moorfields Eye Hospital between 1986 and 2010. Applying Friedman test followed by the post hoc Tukey's honestly significant difference test, we observed that the classifiers trained on any kind of test, regardless of its reliability, showed comparable performance with respect to the classifier trained only considering totally reliable tests (p-value>0.01). Moreover, we showed that different quality indicators gave different effects on prediction results. Training classifiers using tests that exceeded the fixation losses threshold did not have a deteriorating impact on classification results (p-value>0.01). On the contrary, using only tests that fail to comply with the constraint on false negatives significantly decreased the accuracy of the results (p-value<0.01). Meaningful patterns related to glaucoma evolution were also extracted. CONCLUSIONS: Results showed that classification modelling is not negatively affected by the inclusion of less reliable tests in the training process. This means that less reliable tests do not subtract useful information from a model trained using only completely reliable data. Future work will be devoted to exploring new quantitative thresholds to ensure high quality testing and low re-test rates. This could assist doctors in tuning patient follow-up and therapeutic plans, possibly slowing down disease progression. Lucia Sacchi, Allan Tucker, Steve Counsell, David F. Garway-Heath, Stephen Swift |
Artif. Intell. Medicine | 2 |
| 2014 | Exploring Early Glaucoma and the Visual Field Test: Classification and Clustering Using Bayesian NetworksabstractBayesian networks (BNs) are probabilistic models used for classification and clustering in several fields. Their ability to deal with unobserved variables and to integrate data and expert knowledge make them an appropriate technique for modeling eye functionality measurements in glaucoma. In this study, a set of BNs is used to simultaneously perform classification of early glaucoma and cluster data into different stages of disease. A novel learning algorithm that combines clustering and quasi-greedy search is also proposed. The classification performances of the models are evaluated on an independent dataset, while the clusters are compared to K-means, previous publications, and direct knowledge. The use of clustering and structure learning enabled the exploration of the visual field patterns of the disease while obtaining good results both on pre- (50% sensitivity at 90% specificity) and post- (85% sensitivity at 90% specificity) diagnosis data. Clusters obtained were insightful and in conformity with consolidated knowledge in the field. Stefano Ceccon, David F. Garway-Heath, David P. Crabb, Allan Tucker |
IEEE J. Biomed. Health Informatics | 4 |
| 2013 | Integrating Multiple Studies of Wheat Microarray Data to Identify Treatment-Specific Regulatory Networks
Valeria Bo, Artem Lysenko, Mansoor A. S. Saqi, Dimah Z. Habash, Allan Tucker |
IDA | 5 |
| 2013 | The Modelling of Glaucoma Progression through the Use of Cellular Automata
Stelios Pavlidis, Stephen Swift, Allan Tucker, Steve Counsell |
IDA | 3 |
| 2013 | Modelling and analysing the dynamics of disease progression from cross-sectional studies
Stephen Swift, Allan Tucker |
J. Biomed. Informatics | 3 |
| 2011 | The Dynamic Stage Bayesian Network: Identifying and Modelling Key Stages in a Temporal Process
Stefano Ceccon, David F. Garway-Heath, David P. Crabb, Allan Tucker |
IDA | 4 |
| 2011 | Automatic Layout Design Solution
Fadratul Hafinaz Hassan, Allan Tucker |
IDA | 2 |
| 2011 | Integrating Marine Species Biomass Data by Modelling Functional Knowledge
Allan Tucker, Daniel Duplisea |
IDA | 1 |
| 2011 | Interspecies Translation of Disease Networks Increases Robustness and Predictive AccuracyabstractGene regulatory networks give important insights into the mechanisms underlying physiology and pathophysiology. The derivation of gene regulatory networks from high-throughput expression data via machine learning strategies is problematic as the reliability of these models is often compromised by limited and highly variable samples, heterogeneity in transcript isoforms, noise, and other artifacts. Here, we develop a novel algorithm, dubbed Dandelion, in which we construct and train intraspecies Bayesian networks that are translated and assessed on independent test sets from other species in a reiterative procedure. The interspecies disease networks are subjected to multi-layers of analysis and evaluation, leading to the identification of the most consistent relationships within the network structure. In this study, we demonstrate the performance of our algorithms on datasets from animal models of oculopharyngeal muscular dystrophy (OPMD) and patient materials. We show that the interspecies network of genes coding for the proteasome provide highly accurate predictions on gene expression levels and disease phenotype. Moreover, the cross-species translation increases the stability and robustness of these networks. Unlike existing modeling approaches, our algorithms do not require assumptions on notoriously difficult one-to-one mapping of protein orthologues or alternative transcripts and can deal with missing data. We show that the identified key components of the OPMD disease network can be confirmed in an unseen and independent disease model. This study presents a state-of-the-art strategy in constructing interspecies disease networks that provide crucial information on regulatory relationships among genes, leading to better understanding of the disease molecular mechanisms. Seyed Yahya Anvar, Allan Tucker, Veronica Vinciotti, Andrea Venema, Gert-Jan B. van Ommen, Silvere M. van der Maarel, Vered Raz, Peter A. C. 't Hoen |
PLoS Comput. Biol. | 2 |
| 2010 | Using Cellular Automata Pedestrian Flow Statistics with Heuristic Search to Automatically Design Spatial LayoutabstractThe spatial layout of public space has an enormous impact upon the ease with which people can move. For example, a well designed public building such as a train station or hospital should allow for the smooth flow of a large number of people. In extreme cases, design can be so poor that during emergency evacuations people have been crushed to death. It is vital that design takes into account the smooth flow of pedestrians. In this paper we make an initial exploration in using pedestrian flow combined with heuristic search to assist in the automatic design for spatial layout planning. Using pedestrian simulations, the activity of crowds can be used to study the consequences of different spatial layouts. In this paper, a two-way pedestrian flow system is simulated and heuristic search techniques (hill climbing and simulated annealing) are used to find feasible spatial layouts based upon the generated statistics with promising results. Fadratul Hafinaz Hassan, Allan Tucker |
ICTAI (2) | 2 |
| 2010 | The identification of informative genes from multiple datasets with increasing complexityabstractBACKGROUND: In microarray data analysis, factors such as data quality, biological variation, and the increasingly multi-layered nature of more complex biological systems complicates the modelling of regulatory networks that can represent and capture the interactions among genes. We believe that the use of multiple datasets derived from related biological systems leads to more robust models. Therefore, we developed a novel framework for modelling regulatory networks that involves training and evaluation on independent datasets. Our approach includes the following steps: (1) ordering the datasets based on their level of noise and informativeness; (2) selection of a Bayesian classifier with an appropriate level of complexity by evaluation of predictive performance on independent data sets; (3) comparing the different gene selections and the influence of increasing the model complexity; (4) functional analysis of the informative genes. RESULTS: In this paper, we identify the most appropriate model complexity using cross-validation and independent test set validation for predicting gene expression in three published datasets related to myogenesis and muscle differentiation. Furthermore, we demonstrate that models trained on simpler datasets can be used to identify interactions among genes and select the most informative. We also show that these models can explain the myogenesis-related genes (genes of interest) significantly better than others (P < 0.004) since the improvement in their rankings is much more pronounced. Finally, after further evaluating our results on synthetic datasets, we show that our approach outperforms a concordance method by Lai et al. in identifying informative genes from multiple datasets with increasing complexity whilst additionally modelling the interaction between genes. CONCLUSIONS: We show that Bayesian networks derived from simpler controlled systems have better performance than those trained on datasets from more complex biological systems. Further, we present that highly predictive and consistent genes, from the pool of differentially expressed genes, across independent datasets are more likely to be fundamentally involved in the biological process under study. We conclude that networks trained on simpler controlled systems, such as in vitro experiments, can be used to model and capture interactions among genes in more complex datasets, such as in vivo experiments, where these interactions would otherwise be concealed by a multitude of other ongoing events. Seyed Yahya Anvar, Peter A. C. 't Hoen, Allan Tucker |
BMC Bioinform. | 3 |
| 2010 | A niching genetic k-means algorithm and its applications to gene expression data
Weiguo Sheng 0001, Allan Tucker, Xiaohui Liu 0001 |
Soft Comput. | 2 |
| 2010 | The pseudotemporal bootstrap for predicting glaucoma from cross-sectional visual field dataabstractProgressive loss of the field of vision is characteristic of a number of eye diseases such as glaucoma, a leading cause of irreversible blindness in the world. Recently, there has been an explosion in the amount of data being stored on patients who suffer from visual deterioration, including visual field (VF) test, retinal image, and frequent intraocular pressure measurements. Like the progression of many biological and medical processes, VF progression is inherently temporal in nature. However, many datasets associated with the study of such processes are often cross sectional and the time dimension is not measured due to the expensive nature of such studies. In this paper, we address this issue by developing a method to build artificial time series, which we call pseudo time series from cross-sectional data. This involves building trajectories through all of the data that can then, in turn, be used to build temporal models for forecasting (which would otherwise be impossible without longitudinal data). Glaucoma, like many diseases, is a family of conditions and it is, therefore, likely that there will be a number of key trajectories that are important in understanding the disease. In order to deal with such situations, we extend the idea of pseudo time series by using resampling techniques to build multiple sequences prior to model building. This approach naturally handles outliers and multiple possible disease trajectories. We demonstrate some key properties of our approach on synthetic data and present very promising results on VF data for predicting glaucoma. Allan Tucker, David F. Garway-Heath |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2009 | An Application of Intelligent Data Analysis Techniques to a Large Software Engineering Dataset
James Cain 0002, Steve Counsell, Stephen Swift, Allan Tucker |
IDA | 4 |
| 2009 | Selecting and Weighting Data for Building Consensus Gene Regulatory Networks
Emma Steele, Allan Tucker |
IDA | 2 |
| 2009 | Literature-based priors for gene regulatory networksabstractMOTIVATION: The use of prior knowledge to improve gene regulatory network modelling has often been proposed. In this article we present the first research on the massive incorporation of prior knowledge from literature for Bayesian network learning of gene networks. As the publication rate of scientific papers grows, updating online databases, which have been proposed as potential prior knowledge in past research, becomes increasingly challenging. The novelty of our approach lies in the use of gene-pair association scores that describe the overlap in the contexts in which the genes are mentioned, generated from a large database of scientific literature, harnessing the information contained in a huge number of documents into a simple, clear format. RESULTS: We present a method to transform such literature-based gene association scores to network prior probabilities, and apply it to learn gene sub-networks for yeast, Escherichia coli and Human organisms. We also investigate the effect of weighting the influence of the prior knowledge. Our findings show that literature-based priors can improve both the number of true regulatory interactions present in the network and the accuracy of expression value prediction on genes, in comparison to a network learnt solely from expression data. Networks learnt with priors also show an improved biological interpretation, with identified subnetworks that coincide with known biological pathways. Emma Steele, Allan Tucker, Peter A. C. 't Hoen, Martijn J. Schuemie |
Bioinform. | 2 |
| 2008 | Consensus and Meta-analysis regulatory networks for combining multiple microarray gene expression datasets
Emma Steele, Allan Tucker |
J. Biomed. Informatics | 2 |
| 2007 | An improved restricted growth function genetic algorithm for the consensus clustering of retinal nerve fibre dataabstractThis paper describes an extension to the Restricted Growth Function grouping Genetic Algorithm applied to the Consensus Clustering of a retinal nerve fibre layer data-set. Consensus Clustering is an optimisation based method which combines the results of a number of data clustering methods, and is used when it is unknown which clustering method is expected to perform the best. Consensus Clustering has been shown to produce results which are better than the averaged results of the input methods, but could benefit from a more efficient optimisation method. A Restricted Growth Function grouping Genetic Algorithm is a new method of grouping a number of objects into mutually exclusive subsets based upon a fitness function. This method does not suffer from degeneracy, and thus could be applied to the Consensus Clustering problem more efficiently than Simulated Annealing, the current optimisation method. Within this paper it is shown that this type of Genetic Algorithm can indeed improve the performance of Consensus Clustering, and in fact can be improved further by taking advantage of some application specific properties. These findings are demonstrated on a retinal nerve fibre layer data-set and on a synthetic data-set. Stephen Swift, Allan Tucker, Jason Crampton, David F. Garway-Heath |
GECCO | 2 |
| 2007 | Efficiency updates for the restricted growth function GA for grouping problemsabstractProblems that require the partitioning of a set of variables in order to compute a solution such as bin packing or line balancing are typically NP-hard. Hence, researchers have focused on producing heuristic methods for finding appropriate partitions. Many of the representations used in optimisation algorithms including those in GA methods suffer from degeneracy [2]. Furthermore, Falkenauer has found that representations with less degeneracy result in more efficient GAs with respect to grouping problems [1]. Previously we developed a new representation for grouping genetic algorithms called the Restricted Growth Function GA (RGFGA) [3]. The RGFGA effectively removes all degeneracy, resulting in a more efficient search. However, one flaw of the RGFGA is that it converges too quickly resulting in Allan Tucker, Stephen Swift, Jason Crampton |
GECCO | 1 |
| 2007 | Making Time: Pseudo Time-Series for the Temporal Analysis of Cross Section Data
Emma Peeling, Allan Tucker |
IDA | 2 |
| 2006 | Temporal Bayesian classifiers for modelling muscular dystrophy expression data
Allan Tucker, Peter A. C. 't Hoen, Veronica Vinciotti, Xiaohui Liu 0001 |
Intell. Data Anal. | 1 |
| 2005 | ICARUS: intelligent coupon allocation for retailers using searchabstractMany retailers run loyalty card schemes for their customers offering incentives in the form of money off coupons. The total value of the coupons depends on how much the customer has spent. This paper deals with the problem of finding the smallest set of coupons such that each possible total can be represented as the sum of a pre-defined number of coupons. A mathematical analysis of the problem leads to the development of a genetic algorithm solution. The algorithm is applied to real world data using several crossover operators and compared to well known straw-person methods. Results are promising showing that considerable time can be saved by using this method, reducing a few days worth of consultancy time to a few minutes of computation. Stephen Swift, Amy Shi, Jason Crampton, Allan Tucker |
Congress on Evolutionary Computation | 4 |
| 2005 | Bayesian Network Classifiers for Time-Series Microarray Data
Allan Tucker, Veronica Vinciotti, Peter A. C. 't Hoen, Xiaohui Liu 0001 |
IDA | 1 |
| 2005 | A spatio-temporal Bayesian network classifier for understanding visual field deterioration
Allan Tucker, Veronica Vinciotti, Xiaohui Liu 0001, David F. Garway-Heath |
Artif. Intell. Medicine | 1 |
| 2005 | RGFGA: An Efficient Representation and Crossover for Grouping Genetic AlgorithmsabstractThere is substantial research into genetic algorithms that are used to group large numbers of objects into mutually exclusive subsets based upon some fitness function. However, nearly all methods involve degeneracy to some degree. We introduce a new representation for grouping genetic algorithms, the restricted growth function genetic algorithm, that effectively removes all degeneracy, resulting in a more efficient search. A new crossover operator is also described that exploits a measure of similarity between chromosomes in a population. Using several synthetic datasets, we compare the performance of our representation and crossover with another well known state-of-the-art GA method, a strawman optimisation method and a well-established statistical clustering algorithm, with encouraging results. Allan Tucker, Jason Crampton, Stephen Swift |
Evol. Comput. | 1 |
| 2004 | Clustering with Niching Genetic K-means Algorithm
Weiguo Sheng 0001, Allan Tucker, Xiaohui Liu 0001 |
GECCO (2) | 2 |
| 2004 | A Bayesian network approach to explaining time series with changing structure
Allan Tucker |
Intell. Data Anal. | 1 |
| 2003 | Spatial Operators for Evolving Dynamic Bayesian Networks from Spatio-temporal Data
Allan Tucker, Xiaohui Liu 0001, David F. Garway-Heath |
GECCO | 1 |
| 2003 | Applying Intelligent Data Analysis to Coupling Relationships in Object-Oriented Software
Steve Counsell, Xiaohui Liu 0001, Rajaa Najjar, Stephen Swift, Allan Tucker |
IDA | 5 |
| 2003 | Learning Dynamic Bayesian Networks from Multivariate Time Series with Changing Dependencies
Allan Tucker, Xiaohui Liu 0001 |
IDA | 1 |
| 2002 | Evolutionary algorithms for grouping high dimensional Email data
Steve Counsell, Xiaohui Liu 0001, Janet McFall, Stephen Swift, Allan Tucker |
Intell. Data Anal. | 5 |
| 2002 | A framework for modelling virus gene expression data
Paul Kellam, Xiaohui Liu 0001, Nigel J. Martin 0001, Christine A. Orengo, Stephen Swift, Allan Tucker |
Intell. Data Anal. | 6 |
| 2001 | A Framework for Modelling Short, High-Dimensional Multivariate Time Series: Preliminary Results in Virus Gene Expression Data Analysis
Paul Kellam, Xiaohui Liu 0001, Nigel J. Martin 0001, Christine A. Orengo, Stephen Swift, Allan Tucker |
IDA | 6 |
| 2001 | Evolutionary learning of dynamic probabilistic models with large time lagsabstractIn this paper, we explore the automatic explanation of multivariate time series (MTS) through learning dynamic Bayesian networks (DBNs). We have developed an evolutionary algorithm which exploits certain characteristics of MTS in order to generate good networks as quickly as possible. We compare this algorithm to other standard learning algorithms that have traditionally been used for static Bayesian networks but are adapted for DBNs in this paper. These are extensively tested on both synthetic and real-world MTS for various aspects of efficiency and accuracy. By proposing a simple representation scheme, an efficient learning methodology, and several useful heuristics, we have found that the proposed method is more efficient for learning DBNs from MTS with large time lags, especially in time-demanding situations. © 2001 John Wiley & Sons, Inc. Allan Tucker, Xiaohui Liu 0001, Andrew Ogden-Swift |
Int. J. Intell. Syst. | 1 |
| 2001 | Grouping multivariate time series variables: applications to chemical process and visual field data
Stephen Swift, Allan Tucker, Nigel J. Martin 0001, Xiaohui Liu 0001 |
Knowl. Based Syst. | 2 |
| 2001 | Variable grouping in multivariate time series via correlationabstractThe decomposition of high-dimensional multivariate time series (MTS) into a number of low-dimensional MTS is a useful but challenging task because the number of possible dependencies between variables is likely to be huge. This paper is about a systematic study of the "variable groupings" problem in MTS. In particular, we investigate different methods of utilizing the information regarding correlations among MTS variables. This type of method does not appear to have been studied before. In all, 15 methods are suggested and applied to six datasets where there are identifiable mixed groupings of MTS variables. This paper describes the general methodology, reports extensive experimental results, and concludes with useful insights on the strength and weakness of this type of grouping method. Allan Tucker, Stephen Swift, Xiaohui Liu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1999 | Evolutionary Computation to Search for Strongly Correlated Variables in High-Dimensional Time-Series
Stephen Swift, Allan Tucker, Xiaohui Liu 0001 |
IDA | 2 |