Colin Campbell

dblp:48/2217 · DBLP profile ↗
← Back
48ranked-venue papers
11as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorComputer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2024 DrivR-Base: a feature extraction toolkit for variant effect prediction model construction
abstract
MOTIVATION: Recent advancements in sequencing technologies have led to the discovery of numerous variants in the human genome. However, understanding their precise roles in diseases remains challenging due to their complex functional mechanisms. Various methodologies have emerged to predict the pathogenic significance of these genetic variants. Typically, these methods employ an integrative approach, leveraging diverse data sources that provide important insights into genomic function. Despite the abundance of publicly available data sources and databases, the process of navigating, extracting, and pre-processing features for machine learning models can be highly challenging and time-consuming. Furthermore, researchers often invest substantial effort in feature extraction, only to later discover that these features lack informativeness. RESULTS: In this article, we introduce DrivR-Base, an innovative resource that efficiently extracts and integrates molecular information (features) related to single nucleotide variants. These features encompass information about the genomic positions and the associated protein positions of a variant. They are derived from a wide array of databases and tools, including structural properties obtained from AlphaFold, regulatory information sourced from ENCODE, and predicted variant consequences from Variant Effect Predictor. DrivR-Base is easily deployable via a Docker container to ensure reproducibility and ease of access across diverse computational environments. The resulting features can be used as input for machine learning models designed to predict the pathogenic impact of human genome variants in disease. Moreover, these feature sets have applications beyond this, including haploinsufficiency prediction and the development of drug repurposing tools. We describe the resource's development, practical applications, and potential for future expansion and enhancement. AVAILABILITY AND IMPLEMENTATION: DrivR-Base source code is available at https://github.com/amyfrancis97/DrivR-Base.
Amy Francis, Colin Campbell, Tom R. Gaunt
Bioinform.2
2022 Little rewards, big changes: Using exercise analytics to motivate sustainable changes in physical activity
Kirk Plangger, Colin Campbell, Karen Robson, Matteo Montecchi
Inf. Manag.2
2022 An IoT Edge Computing Framework Using Cordova Accessor Host
abstract
The Internet of Things (IoT) is a rapidly growing system of physical sensors and connected devices, enabling advanced information gathering, interpretation, and monitoring. The realization of a versatile IoT edge computing framework will accelerate seamless integration of the cyber-world with new physical IoT devices, and will fundamentally change and empower the way humans interact with the world. While there are many cloud-based IoT computing frameworks, they cannot support the needs of IoT applications that require local processing and guarantee of consumer’s privacy. This article presents experimentation with the opensource plug-and-play IoT middleware, called Cordova Accesor Host. We demonstrated that Cordova Accessor Host supports the essential ingredients of the composition and reusability of IoT services using the accessor as the basic building block and adopting an accessor-module-plugin design pattern. The portability is demonstrated by using the same accessor for collecting sensor data from radically different IoT devices such as, wearables (e.g., smartwatches) and microcontrollers (e.g., Arduino). Our energy profiling experiments show that IoT services deployed using the Cordova Accessor Host consume around 35% less battery power than the same IoT services deployed in the native Android operating system.
Anne H. H. Ngu, Jesuloluwa S. Eyitayo, Guowei Yang 0001, Colin Campbell, Quan Z. Sheng, Jianyuan Ni
IEEE Internet Things J.4
2022 Whole community invasions and the integration of novel ecosystems
abstract
The impact of invasion by a single non-native species on the function and structure of ecological communities can be significant, and the effects can become more drastic-and harder to predict-when multiple species invade as a group. Here we modify a dynamic Boolean model of plant-pollinator community assembly to consider the invasion of native communities by multiple invasive species that are selected either randomly or such that the invaders constitute a stable community. We show that, compared to random invasion, whole community invasion leads to final stable communities (where the initial process of species turnover has given way to a static or near-static set of species in the community) including both native and non-native species that are larger, more likely to retain native species, and which experience smaller changes to the topological measures of nestedness and connectance. We consider the relationship between the prevalence of mutualistic interactions among native and invasive species in the final stable communities and demonstrate that mutualistic interactions may act as a buffer against significant disruptions to the native community.
Colin Campbell, Laura Russo, Réka Albert, Angus Buckling, Katriona Shea
PLoS Comput. Biol.1
2021 Prediction of driver variants in the cancer genome via machine learning methodologies
abstract
Sequencing technologies have led to the identification of many variants in the human genome which could act as disease-drivers. As a consequence, a variety of bioinformatics tools have been proposed for predicting which variants may drive disease, and which may be causatively neutral. After briefly reviewing generic tools, we focus on a subset of these methods specifically geared toward predicting which variants in the human cancer genome may act as enablers of unregulated cell proliferation. We consider the resultant view of the cancer genome indicated by these predictors and discuss ways in which these types of prediction tools may be progressed by further research.
Mark F. Rogers, Tom R. Gaunt, Colin Campbell
Briefings Bioinform.3
2021 CScape-somatic: distinguishing driver and passenger point mutations in the cancer genome
abstract
Bioinformatics (2020) doi:10.1093/bioinformatics/btaa242 In the originally published article, a funding acknowledgement was missing. This should read: “Financial Support: The Integrative Epidemiology Unit is supported by the Medical Research Council (MC_UU_00011/4) and the University of Bristol, and we also acknowledge funding from the Cancer Research UK Integrative Cancer Epidemiology Programme (C18281/A19169).” instead of: “Financial Support: none declared.” These details have been corrected online.
Mark F. Rogers, Tom R. Gaunt, Colin Campbell
Bioinform.3
2020 CScape-somatic: distinguishing driver and passenger point mutations in the cancer genome
abstract
MOTIVATION: Next-generation sequencing technologies have accelerated the discovery of single nucleotide variants in the human genome, stimulating the development of predictors for classifying which of these variants are likely functional in disease, and which neutral. Recently, we proposed CScape, a method for discriminating between cancer driver mutations and presumed benign variants. For the neutral class, this method relied on benign germline variants found in the 1000 Genomes Project database. Discrimination could, therefore, be influenced by the distinction of germline versus somatic, rather than neutral versus disease driver. This motivates this article in which we consider predictive discrimination between recurrent and rare somatic single point mutations based solely on using cancer data, and the distinction between these two somatic classes and germline single point mutations. RESULTS: For somatic point mutations in coding and non-coding regions of the genome, we propose CScape-somatic, an integrative classifier for predictively discriminating between recurrent and rare variants in the human cancer genome. In this study, we use purely cancer genome data and investigate the distinction between minimal occurrence and significantly recurrent somatic single point mutations in the human cancer genome. We show that this type of predictive distinction can give novel insight, and may deliver more meaningful prediction in both coding and non-coding regions of the cancer genome. Tested on somatic mutations, CScape-somatic outperforms alternative methods, reaching 74% balanced accuracy in coding regions and 69% in non-coding regions, whereas even higher accuracy may be achieved using thresholds to isolate high-confidence predictions. AVAILABILITY AND IMPLEMENTATION: Predictions and software are available at http://CScape-somatic.biocompute.org.uk/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mark F. Rogers, Tom R. Gaunt, Colin Campbell
Bioinform.3
2018 FATHMM-XF: accurate prediction of pathogenic point mutations via extended features
abstract
Summary: We present FATHMM-XF, a method for predicting pathogenic point mutations in the human genome. Drawing on an extensive feature set, FATHMM-XF outperforms competitors on benchmark tests, particularly in non-coding regions where the majority of pathogenic mutations are likely to be found. Availability and implementation: The FATHMM-XF web server is available at http://fathmm.biocompute.org.uk/fathmm-xf/, and as tracks on the Genome Tolerance Browser: http://gtb.biocompute.org.uk. Predictions are provided for human genome version GRCh37/hg19. The data used for this project can be downloaded from: http://fathmm.biocompute.org.uk/fathmm-xf/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Mark F. Rogers, Hashem A. Shihab, Matthew E. Mort, David N. Cooper, Tom R. Gaunt, Colin Campbell
Bioinform.6
2017 HIPred: an integrative approach to predicting haploinsufficient genes
abstract
MOTIVATION: A major cause of autosomal dominant disease is haploinsufficiency, whereby a single copy of a gene is not sufficient to maintain the normal function of the gene. A large proportion of existing methods for predicting haploinsufficiency incorporate biological networks, e.g. protein-protein interaction networks that have recently been shown to introduce study bias. As a result, these methods tend to perform best on well-studied genes, but underperform on less studied genes. The advent of large genome sequencing consortia, such as the 1000 genomes project, NHLBI Exome Sequencing Project and the Exome Aggregation Consortium creates an urgent need for unbiased haploinsufficiency prediction methods. RESULTS: Here, we describe a machine learning approach, called HIPred, that integrates genomic and evolutionary information from ENSEMBL, with functional annotations from the Encyclopaedia of DNA Elements consortium and the NIH Roadmap Epigenomics Project to predict haploinsufficiency, without the study bias described earlier. We benchmark HIPred using several datasets and show that our unbiased method performs as well as, and in most cases, outperforms existing biased algorithms. AVAILABILITY AND IMPLEMENTATION: HIPred scores for all gene identifiers are available at: https://github.com/HAShihab/HIPred . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hashem A. Shihab, Mark F. Rogers, Colin Campbell, Tom R. Gaunt
Bioinform.3
2017 An integrative approach to predicting the functional effects of small indels in non-coding regions of the human genome
abstract
BACKGROUND: Small insertions and deletions (indels) have a significant influence in human disease and, in terms of frequency, they are second only to single nucleotide variants as pathogenic mutations. As the majority of mutations associated with complex traits are located outside the exome, it is crucial to investigate the potential pathogenic impact of indels in non-coding regions of the human genome. RESULTS: We present FATHMM-indel, an integrative approach to predict the functional effect, pathogenic or neutral, of indels in non-coding regions of the human genome. Our method exploits various genomic annotations in addition to sequence data. When validated on benchmark data, FATHMM-indel significantly outperforms CADD and GAVIN, state of the art models in assessing the pathogenic impact of non-coding variants. FATHMM-indel is available via a web server at indels.biocompute.org.uk. CONCLUSIONS: FATHMM-indel can accurately predict the functional impact and prioritise small indels throughout the whole non-coding genome.
Michael Ferlaino, Mark F. Rogers, Hashem A. Shihab, Matthew E. Mort, David N. Cooper, Tom R. Gaunt, Colin Campbell
BMC Bioinform.7
2017 GTB - an online genome tolerance browser
abstract
BACKGROUND: Accurate methods capable of predicting the impact of single nucleotide variants (SNVs) are assuming ever increasing importance. There exists a plethora of in silico algorithms designed to help identify and prioritize SNVs across the human genome for further investigation. However, no tool exists to visualize the predicted tolerance of the genome to mutation, or the similarities between these methods. RESULTS: We present the Genome Tolerance Browser (GTB, http://gtb.biocompute.org.uk ): an online genome browser for visualizing the predicted tolerance of the genome to mutation. The server summarizes several in silico prediction algorithms and conservation scores: including 13 genome-wide prediction algorithms and conservation scores, 12 non-synonymous prediction algorithms and four cancer-specific algorithms. CONCLUSION: The GTB enables users to visualize the similarities and differences between several prediction algorithms and to upload their own data as additional tracks; thereby facilitating the rapid identification of potential regions of interest.
Hashem A. Shihab, Mark F. Rogers, Michael Ferlaino, Colin Campbell, Tom R. Gaunt
BMC Bioinform.4
2015 Sequential data selection for predicting the pathogenic effects of sequence variation
abstract
Recent improvements in sequencing technologies provide unprecedented opportunities to investigate the role of genetic variation in human disease. In previous work we have proposed a machine learning approach to predicting whether single nucleotide variants (SNVs) are functional or neutral in human disease. Many data sources from the Encyclopaedia of DNA Elements (ENCODE) may be relevant to this problem. To integrate these data sources, we applied integrative multiple kernel learning (MKL) that weights each source according to its relevance. Using an MKL optimization that yields sparse weights, we were able to eliminate the least informative data sources from our model. However, when selecting from a wide assortment of data sources, we have found that MKL may not be an efficient method for eliminating uninformative sources. Many data sources related to the human genome are incomplete: this can reduce dramatically the data available for training and the proportion of novel predictions that exploit all relevant sources. Here we introduce a greedy sequential selection method that assesses data sources in a structured fashion prior to MKL weight optimization. This method allows us to eliminate a majority of uninformative data sources prior to assigning kernel weights. When we use this method with our coding-region predictor, we select just five kernels for our final model, yielding increased accuracy over our previous model. In addition, by reducing the amount of data required for novel predictions, we are able to increase by five fold our model's coverage for new predictions.
Mark F. Rogers, Colin Campbell, Hashem A. Shihab, Tom R. Gaunt, Matthew E. Mort, David N. Cooper
BIBM2
2015 An integrative approach to predicting the functional effects of non-coding and coding sequence variation
abstract
Abstract Motivation: Technological advances have enabled the identification of an increasingly large spectrum of single nucleotide variants within the human genome, many of which may be associated with monogenic disease or complex traits. Here, we propose an integrative approach, named FATHMM-MKL, to predict the functional consequences of both coding and non-coding sequence variants. Our method utilizes various genomic annotations, which have recently become available, and learns to weight the significance of each component annotation source. Results: We show that our method outperforms current state-of-the-art algorithms, CADD and GWAVA, when predicting the functional consequences of non-coding variants. In addition, FATHMM-MKL is comparable to the best of these algorithms when predicting the impact of coding variants. The method includes a confidence measure to rank order predictions. Availability and implementation: The FATHMM-MKL webserver is available at: http://fathmm.biocompute.org.uk Contact: [email protected] or [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Hashem A. Shihab, Mark F. Rogers, Julian Gough, Matthew E. Mort, David N. Cooper, Ian N. M. Day, Tom R. Gaunt, Colin Campbell
Bioinform.8
2015 Texture classification using feature selection and kernel-based techniques
Carlos Fernandez-Lozano, José Antonio Seoane Fernández, Marcos Gestal, Tom R. Gaunt, Julián Dorado, Colin Campbell
Soft Comput.6
2014 A Random Forest proximity matrix as a new measure for gene annotation
José Antonio Seoane Fernández, Ian N. M. Day, Juan P. Casas, Colin Campbell, Tom R. Gaunt
ESANN4
2014 A pathway-based data integration framework for prediction of disease progression
abstract
MOTIVATION: Within medical research there is an increasing trend toward deriving multiple types of data from the same individual. The most effective prognostic prediction methods should use all available data, as this maximizes the amount of information used. In this article, we consider a variety of learning strategies to boost prediction performance based on the use of all available data. IMPLEMENTATION: We consider data integration via the use of multiple kernel learning supervised learning methods. We propose a scheme in which feature selection by statistical score is performed separately per data type and by pathway membership. We further consider the introduction of a confidence measure for the class assignment, both to remove some ambiguously labeled datapoints from the training data and to implement a cautious classifier that only makes predictions when the associated confidence is high. RESULTS: We use the METABRIC dataset for breast cancer, with prediction of survival at 2000 days from diagnosis. Predictive accuracy is improved by using kernels that exclusively use those genes, as features, which are known members of particular pathways. We show that yet further improvements can be made by using a range of additional kernels based on clinical covariates such as Estrogen Receptor (ER) status. Using this range of measures to improve prediction performance, we show that the test accuracy on new instances is nearly 80%, though predictions are only made on 69.2% of the patient cohort. AVAILABILITY: https://github.com/jseoane/FSMKL CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
José Antonio Seoane Fernández, Ian N. M. Day, Tom R. Gaunt, Colin Campbell
Bioinform.4
2014 Canonical Correlation Analysis for Gene-Based Pleiotropy Discovery
abstract
Genome-wide association studies have identified a wealth of genetic variants involved in complex traits and multifactorial diseases. There is now considerable interest in testing variants for association with multiple phenotypes (pleiotropy) and for testing multiple variants for association with a single phenotype (gene-based association tests). Such approaches can increase statistical power by combining evidence for association over multiple phenotypes or genetic variants respectively. Canonical Correlation Analysis (CCA) measures the correlation between two sets of multidimensional variables, and thus offers the potential to combine these two approaches. To apply CCA, we must restrict the number of attributes relative to the number of samples. Hence we consider modules of genetic variation that can comprise a gene, a pathway or another biologically relevant grouping, and/or a set of phenotypes. In order to do this, we use an attribute selection strategy based on a binary genetic algorithm. Applied to a UK-based prospective cohort study of 4286 women (the British Women's Heart and Health Study), we find improved statistical power in the detection of previously reported genetic associations, and identify a number of novel pleiotropic associations between genetic variants and phenotypes. New discoveries include gene-based association of NSF with triglyceride levels and several genes (ACSM3, ERI2, IL18RAP, IL23RAP and NRG1) with left ventricular hypertrophy phenotypes. In multiple-phenotype analyses we find association of NRG1 with left ventricular hypertrophy phenotypes, fibrinogen and urea and pleiotropic relationships of F7 and F10 with Factor VII, Factor IX and cholesterol levels.
José Antonio Seoane Fernández, Colin Campbell, Ian N. M. Day, Juan P. Casas, Tom R. Gaunt
PLoS Comput. Biol.2
2011 Generalized sparse metric learning with relative comparisons
Kaizhu Huang, Yiming Ying, Colin Campbell
Knowl. Inf. Syst.3
2010 Rademacher Chaos Complexities for Learning the Kernel Problem
abstract
We develop a novel generalization bound for learning the kernel problem. First, we show that the generalization analysis of the kernel learning problem reduces to investigation of the suprema of the Rademacher chaos process of order 2 over candidate kernels, which we refer to as Rademacher chaos complexity. Next, we show how to estimate the empirical Rademacher chaos complexity by well-established metric entropy integrals and pseudo-dimension of the set of candidate kernels. Our new methodology mainly depends on the principal theory of U-processes and entropy integrals. Finally, we establish satisfactory excess generalization bounds and misclassification error rates for learning gaussian kernels and general radial basis kernels.
Yiming Ying, Colin Campbell
Neural Comput.2
2009 Generalization Bounds for Learning the Kernel Problem
Yiming Ying, Colin Campbell
COLT2
2009 A Variational Approach to Semi-Supervised Clustering
Yiming Ying, Colin Campbell
ESANN3
2009 GSML: A Unified Framework for Sparse Metric Learning
abstract
There has been significant recent interest in sparse metric learning (SML) in which we simultaneously learn both a good distance metric and a low-dimensional representation. Unfortunately, the performance of existing sparse metric learning approaches is usually limited because the authors assumed certain problem relaxations or they target the SML objective indirectly. In this paper, we propose a Generalized Sparse Metric Learning method (GSML). This novel framework offers a unified view for understanding many of the popular sparse metric learning algorithms including the Sparse Metric Learning framework proposed, the Large Margin Nearest Neighbor (LMNN), and the D-ranking Vector Machine (D-ranking VM). Moreover, GSML also establishes a close relationship with the Pairwise Support Vector Machine. Furthermore, the proposed framework is capable of extending many current non-sparse metric learning models such as Relevant Vector Machine (RCA) and a state-of-the-art method proposed into their sparse versions. We present the detailed framework, provide theoretical justifications, build various connections with other models, and propose a practical iterative optimization method, making the framework both theoretically important and practically scalable for medium or large datasets. A series of experiments show that the proposed approach can outperform previous methods in terms of both test accuracy and dimension reduction, on six real-world benchmark datasets.
Kaizhu Huang, Yiming Ying, Colin Campbell
ICDM3
2009 Supervised Self-taught Learning: Actively transferring knowledge from unlabeled data
abstract
We consider the task of Self-taught Learning (STL) from unlabeled data. In contrast to semi-supervised learning, which requires unlabeled data to have the same set of class labels as labeled data, STL can transfer knowledge from different types of unlabeled data. STL uses a three-step strategy: (1) learning high-level representations from unlabeled data only, (2) re-constructing the labeled data via such representations and (3) building a classifier over the re-constructed labeled data. However, the high-level representations which are exclusively determined by the unlabeled data, may be inappropriate or even misleading for the latter classifier training process. In this paper, we propose a novel Supervised Self-taught Learning (SSTL) framework that successfully integrates the three isolated steps of STL into a single optimization problem. Benefiting from the interaction between the classifier optimization and the process of choosing high-level representations, the proposed model is able to select those discriminative representations which are more appropriate for classification. One important feature of our novel framework is that the final optimization can be iteratively solved with convergence guaranteed. We evaluate our novel framework on various data sets. The experimental results show that the proposed SSTL can outperform STL and traditional supervised learning methods in certain instances.
Kaizhu Huang, Zenglin Xu, Irwin King, Michael R. Lyu, Colin Campbell
IJCNN5
2009 Analysis of SVM with Indefinite Kernels
abstract
The recent introduction of indefinite SVM by Luss and dAspremont [15] has effectively demonstrated SVM classification with a non-positive semi-definite kernel (indefinite kernel). This paper studies the properties of the objective function introduced there. In particular, we show that the objective function is continuously differentiable and its gradient can be explicitly computed. Indeed, we further show that its gradient is Lipschitz continuous. The main idea behind our analysis is that the objective function is smoothed by the penalty term, in its saddle (min-max) representation, measuring the distance between the indefinite kernel matrix and the proxy positive semi-definite one. Our elementary result greatly facilitates the application of gradient-based algorithms. Based on our analysis, we further develop Nesterovs smooth optimization approach [16,17] for indefinite SVM which has an optimal convergence rate for smooth problems. Experiments on various benchmark datasets validate our analysis and demonstrate the efficiency of our proposed algorithms.
Yiming Ying, Colin Campbell, Mark A. Girolami
NIPS2
2009 Sparse Metric Learning via Smooth Optimization
abstract
In this paper we study the problem of learning a low-dimensional (sparse) distance matrix. We propose a novel metric learning model which can simultaneously conduct dimension reduction and learn a distance matrix. The sparse representation involves a mixed-norm regularization which is non-convex. We then show that it can be equivalently formulated as a convex saddle (min-max) problem. From this saddle representation, we develop an efficient smooth optimization approach for sparse metric learning although the learning model is based on a non-differential loss function. This smooth optimization approach has an optimal convergence rate of $O(1 /\ell^2)$ for smooth problems where $\ell$ is the iteration number. Finally, we run experiments to validate the effectiveness and efficiency of our sparse metric learning model on various datasets.
Yiming Ying, Kaizhu Huang, Colin Campbell
NIPS3
2009 Enhanced protein fold recognition through a novel data integration approach
abstract
BACKGROUND: Protein fold recognition is a key step in protein three-dimensional (3D) structure discovery. There are multiple fold discriminatory data sources which use physicochemical and structural properties as well as further data sources derived from local sequence alignments. This raises the issue of finding the most efficient method for combining these different informative data sources and exploring their relative significance for protein fold classification. Kernel methods have been extensively used for biological data analysis. They can incorporate separate fold discriminatory features into kernel matrices which encode the similarity between samples in their respective data sources. RESULTS: In this paper we consider the problem of integrating multiple data sources using a kernel-based approach. We propose a novel information-theoretic approach based on a Kullback-Leibler (KL) divergence between the output kernel matrix and the input kernel matrix so as to integrate heterogeneous data sources. One of the most appealing properties of this approach is that it can easily cope with multi-class classification and multi-task learning by an appropriate choice of the output kernel matrix. Based on the position of the output and input kernel matrices in the KL-divergence objective, there are two formulations which we respectively refer to as MKLdiv-dc and MKLdiv-conv. We propose to efficiently solve MKLdiv-dc by a difference of convex (DC) programming method and MKLdiv-conv by a projected gradient descent algorithm. The effectiveness of the proposed approaches is evaluated on a benchmark dataset for protein fold recognition and a yeast protein function prediction problem. CONCLUSION: Our proposed methods MKLdiv-dc and MKLdiv-conv are able to achieve state-of-the-art performance on the SCOP PDB-40D benchmark dataset for protein fold prediction and provide useful insights into the relative significance of informative data sources. In particular, MKLdiv-dc further improves the fold discrimination accuracy to 75.19% which is a more than 5% improvement over competitive Bayesian probabilistic and SVM margin-based kernel learning methods. Furthermore, we report a competitive performance on the yeast protein function prediction problem.
Yiming Ying, Kaizhu Huang, Colin Campbell
BMC Bioinform.3
2008 Learning Coordinate Gradients with Multi-Task Kernels
Yiming Ying, Colin Campbell
COLT2
2008 Inferring Sparse Kernel Combinations and Relevance Vectors: An Application to Subcellular Localization of Proteins
abstract
In this paper, we introduce two new formulations for multi-class multi-kernel relevance vector machines (m-RVMs) that explicitly lead to sparse solutions, both in samples and in number of kernels. This enables their application to large-scale multi-feature multinomial classification problems where there is an abundance of training samples, classes and feature spaces. The proposed methods are based on an expectation-maximization (EM) framework employing a multinomial probit likelihood and explicit pruning of non-relevant training samples. We demonstrate the methods on a low-dimensional artificial dataset. We then demonstrate the accuracy and sparsity of the method when applied to the challenging bioinformatics task of predicting protein subcellular localization.
Theodoros Damoulas, Yiming Ying, Mark A. Girolami, Colin Campbell
ICMLA4
2007 Composition of Model Programs
Margus Veanes, Colin Campbell, Wolfram Schulte
FORTE2
2007 State Isomorphism in Model Programs with Abstract Data Structures
Margus Veanes, Juhan P. Ernits, Colin Campbell
FORTE3
2005 Testing Concurrent Object-Oriented Systems with Spec Explorer
Colin Campbell, Wolfgang Grieskamp, Lev Nachmanson, Wolfram Schulte, Nikolai Tillmann, Margus Veanes
FM1
2005 Online testing with model programs
abstract
Online testing is a technique in which test derivation from a model program and test execution are combined into a single algorithm. We describe a practical online testing algorithm that is implemented in the model-based testing tool developed at Microsoft Research called Spec Explorer. Spec Explorer is being used daily by several Microsoft product groups. Model programs in Spec Explorer are written in the high level specification languages AsmL or Spec\#. We view model programs as implicit definitions of interface automata. The conformance relation between a model and an implementation under test is formalized in terms of refinement between interface automata. Testing then amounts to a game between the test tool and the implementation under test.
Margus Veanes, Colin Campbell, Wolfram Schulte, Nikolai Tillmann
ESEC/SIGSOFT FSE2
2005 The Latent Process Decomposition of cDNA Microarray Data Sets
abstract
We present a new computational technique (a software implementation, data sets, and supplementary information are available at http://www.enm.bris.ac.uk/lpd/) which enables the probabilistic analysis of cDNA microarray data and we demonstrate its effectiveness in identifying features of biomedical importance. A hierarchical Bayesian model, called Latent Process Decomposition (LPD), is introduced in which each sample in the data set is represented as a combinatorial mixture over a finite set of latent processes, which are expected to correspond to biological processes. Parameters in the model are estimated using efficient variational methods. This type of probabilistic model is most appropriate for the interpretation of measurement data generated by cDNA microarray technology. For determining informative substructure in such data sets, the proposed model has several important advantages over the standard use of dendrograms. First, the ability to objectively assess the optimal number of sample clusters. Second, the ability to represent samples and gene expression levels using a common set of latent variables (dendrograms cluster samples and gene expression values separately which amounts to two distinct reduced space representations). Third, in constrast to standard cluster models, observations are not assigned to a single cluster and, thus, for example, gene expression levels are modeled via combinations of the latent processes identified by the algorithm. We show this new method compares favorably with alternative cluster analysis methods. To illustrate its potential, we apply the proposed technique to several microarray data sets for cancer. For these data sets it successfully decomposes the data into known subtypes and indicates possible further taxonomic subdivision in addition to highlighting, in a wholly unsupervised manner, the importance of certain genes which are known to be medically significant. To illustrate its wider applicability, we also illustrate its performance on a microarray data set for yeast.
Simon Rogers, Mark A. Girolami, Colin Campbell, Rainer Breitling
IEEE ACM Trans. Comput. Biol. Bioinform.3
2003 New Analytical Techniques for the Interpretation of Microarray Data
abstract
Colin Campbell; New analytical techniques for the interpretation of microarray data, Bioinformatics, Volume 19, Issue 9, 12 June 2003, Pages 1045, https://doi.o
Colin Campbell
Bioinform.1
2003 Special issue on support vector machines
Colin Campbell, Chih-Jen Lin, S. Sathiya Keerthi, V. David Sánchez A.
Neurocomputing1
2002 Kernel methods: a survey of current techniques
Colin Campbell
Neurocomputing1
2002 Editorial: Kernel Methods: Current Research and Future Directions
Nello Cristianini, Colin Campbell, Christopher J. C. Burges
Mach. Learn.2
2001 Bayes Point Machines
Ralf Herbrich, Thore Graepel, Colin Campbell
J. Mach. Learn. Res.3
2000 Algorithmic approaches to training Support Vector Machines: a survey
Colin Campbell
ESANN1
2000 Robust Bayes Point Machines
Ralf Herbrich, Thore Graepel, Colin Campbell
ESANN3
2000 Query Learning with Large Margin Classifiers
Colin Campbell, Nello Cristianini, Alexander J. Smola
ICML1
2000 A Linear Programming Approach to Novelty Detection
abstract
Novelty detection involves modeling the normal behaviour of a sys(cid:173) tem hence enabling detection of any divergence from normality. It has potential applications in many areas such as detection of ma(cid:173) chine damage or highlighting abnormal features in medical data. One approach is to build a hypothesis estimating the support of the normal data i.e. constructing a function which is positive in the region where the data is located and negative elsewhere. Recently kernel methods have been proposed for estimating the support of a distribution and they have performed well in practice - training involves solution of a quadratic programming problem. In this pa(cid:173) per we propose a simpler kernel method for estimating the support based on linear programming. The method is easy to implement and can learn large datasets rapidly. We demonstrate the method on medical and fault detection datasets. 1
Colin Campbell, Kristin P. Bennett
NIPS1
1999 A multiplicative updating algorithm for training support vector machine
Nello Cristianini, Colin Campbell, John Shawe-Taylor
ESANN2
1998 The Kernel-Adatron Algorithm: A Fast and Simple Learning Procedure for Support Vector Machines
Thilo-Thomas Frieß, Nello Cristianini, Colin Campbell
ICML3
1998 Dynamically Adapting Kernels in Support Vector Machines
Nello Cristianini, Colin Campbell, John Shawe-Taylor
NIPS2
1995 Constructing feed-forward neural networks for binary classification tasks
Colin Campbell, Conrado J. Pérez Vicente
ESANN1
1995 The target switch algorithm: a constructive learning procedure for feed-forward neural networks
abstract
We propose an efficient procedure for constructing and training a feed-forward neural network. The network can perform binary classification for binary or analogue input data. We show that the procedure can also be used to construct feedforward neural networks with binary-valued weights. Neural networks with binary-valued weights are potentially straightforward to implement using microelectronic or optical devices and they can also exhibit good generalization.
Colin Campbell, Conrado J. Pérez Vicente
Neural Comput.1
1993 Neural Networks with Sign-Constrained Weights
Colin Campbell
Neural Comput. Appl.1