Christina S. Leslie

dblp:14/3641 · DBLP profile ↗
← Back
26ranked-venue papers
3as first author
0since 2021 · last 2016
0000-0002-4571-5910ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 19 · 1 first-authorArtificial intelligence and machine learning · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
12 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
4 papers
Learning theory · 55% Kernel, tree and ensemble methods · 45%
Databases, data mining, and information retrieval
2 papers
Machine learning and data management · 64% Data mining · 36%

Topics — the 25 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › protein function prediction
protein classification
0.352007
Multi-class Protein Classification Using Adaptive Codes · J. Mach. Learn. Res. 2007
Semi-supervised protein classification using cluster kernels · Bioinform. 2005
Mismatch string kernels for discriminative protein classification · Bioinform. 2004
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection
0.232009
RANKPROP: a web server for protein remote homology detection · Bioinform. 2009
Mismatch string kernels for discriminative protein classification · Bioinform. 2004
Mismatch String Kernels for SVM Protein Classification · NIPS 2002
Bioinformatics and computational biology
protein function prediction
0.222010
iDBPs: a web server for the identification of DNA binding proteins · Bioinform. 2010
Motif-based protein ranking by network propagation · Bioinform. 2005
Bioinformatics and computational biology › protein function prediction › protein classification
DNA-binding protein prediction
0.112010
iDBPs: a web server for the identification of DNA binding proteins · Bioinform. 2010
Bioinformatics and computational biology
protein structure analysis
0.112010
iDBPs: a web server for the identification of DNA binding proteins · Bioinform. 2010
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.132005
Semi-supervised protein classification using cluster kernels · Bioinform. 2005
Semi-supervised Protein Classification Using Cluster Kernels · NIPS 2003
Mismatch String Kernels for SVM Protein Classification · NIPS 2002
Bioinformatics and computational biology
protein structure prediction
0.122005
Multi-class protein fold recognition using adaptive codes · ICML 2005
Protein backbone angle prediction with machine learning approaches · Bioinform. 2004
Bioinformatics and computational biology › structural biology
protein structure and function
0.112009
RANKPROP: a web server for protein remote homology detection · Bioinform. 2009
Machine learning › Learning theory › classification › multiclass classification
error-correcting output codes
0.112007
Multi-class Protein Classification Using Adaptive Codes · J. Mach. Learn. Res. 2007
Machine learning › Learning theory › classification
multiclass classification
0.112007
Multi-class Protein Classification Using Adaptive Codes · J. Mach. Learn. Res. 2007
Bioinformatics and computational biology › genomics
computational genomics
0.112005
Motif Discovery Through Predictive Modeling of Gene Regulation · RECOMB 2005
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition
0.112005
Multi-class protein fold recognition using adaptive codes · ICML 2005
Bioinformatics and computational biology › gene regulation
gene regulation analysis
0.112005
Motif Discovery Through Predictive Modeling of Gene Regulation · RECOMB 2005
Bioinformatics and computational biology › sequence analysis
homology detection
0.112005
Motif-based protein ranking by network propagation · Bioinform. 2005
Bioinformatics and computational biology › sequence analysis
motif discovery
0.112005
Motif Discovery Through Predictive Modeling of Gene Regulation · RECOMB 2005
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network analysis
0.112005
Motif-based protein ranking by network propagation · Bioinform. 2005
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network propagation
0.112005
Motif-based protein ranking by network propagation · Bioinform. 2005
Data mining › predictive modeling › classification
multiclass classification
0.112005
Multi-class protein fold recognition using adaptive codes · ICML 2005
Bioinformatics and computational biology › protein structure prediction › protein structural feature prediction
backbone torsion angle prediction
0.012004
Protein backbone angle prediction with machine learning approaches · Bioinform. 2004
Bioinformatics and computational biology
protein sequence analysis
0.012004
Fast String Kernels using Inexact Matching for Protein Sequences · J. Mach. Learn. Res. 2004
Bioinformatics and computational biology › kernel methods
string kernel
0.012004
Fast String Kernels using Inexact Matching for Protein Sequences · J. Mach. Learn. Res. 2004
Machine learning and data management
kernel methods
0.012004
Fast String Kernels using Inexact Matching for Protein Sequences · J. Mach. Learn. Res. 2004
Machine learning and data management › kernel methods
string kernels
0.012004
Fast String Kernels using Inexact Matching for Protein Sequences · J. Mach. Learn. Res. 2004
Bioinformatics and computational biology › biological network › network biology
protein similarity network
0.012009
RANKPROP: a web server for protein remote homology detection · Bioinform. 2009
Machine learning › Kernel, tree and ensemble methods › kernel methods › structured kernel
string kernel
0.012002
Mismatch String Kernels for SVM Protein Classification · NIPS 2002

Methods — techniques the papers use, named apart from their topics

PSI-BLAST · 0.2cluster kernels · 0.2support vector machine · 0.2network propagation · 0.1mismatch tree · 0.1random forest · 0.1evolutionary profile · 0.1electrostatic potential · 0.1string kernels · 0.1string kernel · 0.1rankprop · 0.1adaptive codes · 0.1semi-supervised learning · 0.1one-vs-all classifiers · 0.1nearest neighbor · 0.1inexact matching · 0.0
YearPublicationVenuePosition
2016 Learning to Predict miRNA-mRNA Interactions from AGO CLIP Sequencing and CLASH Data
abstract
Recent technologies like AGO CLIP sequencing and CLASH enable direct transcriptome-wide identification of AGO binding and miRNA target sites, but the most widely used miRNA target prediction algorithms do not exploit these data. Here we use discriminative learning on AGO CLIP and CLASH interactions to train a novel miRNA target prediction model. Our method combines two SVM classifiers, one to predict miRNA-mRNA duplexes and a second to learn a binding model of AGO's local UTR sequence preferences and positional bias in 3'UTR isoforms. The duplex SVM model enables the prediction of non-canonical target sites and more accurately resolves miRNA interactions from AGO CLIP data than previous methods. The binding model is trained using a multi-task strategy to learn context-specific and common AGO sequence preferences. The duplex and common AGO binding models together outperform existing miRNA target prediction algorithms on held-out binding data. Open source code is available at https://bitbucket.org/leslielab/chimiric.
Yuheng Lu, Christina S. Leslie
PLoS Comput. Biol.2
2015 SeqGL Identifies Context-Dependent Binding Signals in Genome-Wide Regulatory Element Maps
abstract
Genome-wide maps of transcription factor (TF) occupancy and regions of open chromatin implicitly contain DNA sequence signals for multiple factors. We present SeqGL, a novel de novo motif discovery algorithm to identify multiple TF sequence signals from ChIP-, DNase-, and ATAC-seq profiles. SeqGL trains a discriminative model using a k-mer feature representation together with group lasso regularization to extract a collection of sequence signals that distinguish peak sequences from flanking regions. Benchmarked on over 100 ChIP-seq experiments, SeqGL outperformed traditional motif discovery tools in discriminative accuracy. Furthermore, SeqGL can be naturally used with multitask learning to identify genomic and cell-type context determinants of TF binding. SeqGL successfully scales to the large multiplicity of sequence signals in DNase- or ATAC-seq maps. In particular, SeqGL was able to identify a number of ChIP-seq validated sequence signals that were not found by traditional motif discovery algorithms. Thus compared to widely used motif discovery algorithms, SeqGL demonstrates both greater discriminative accuracy and higher sensitivity for detecting the DNA sequence signals underlying regulatory element maps. SeqGL is available at http://cbio.mskcc.org/public/Leslie/SeqGL/.
Manu Setty, Christina S. Leslie
PLoS Comput. Biol.2
2011 Detecting Remote Evolutionary Relationships among Proteins by Large-Scale Semantic Embedding
abstract
Virtually every molecular biologist has searched a protein or DNA sequence database to find sequences that are evolutionarily related to a given query. Pairwise sequence comparison methods--i.e., measures of similarity between query and target sequences--provide the engine for sequence database search and have been the subject of 30 years of computational research. For the difficult problem of detecting remote evolutionary relationships between protein sequences, the most successful pairwise comparison methods involve building local models (e.g., profile hidden Markov models) of protein sequences. However, recent work in massive data domains like web search and natural language processing demonstrate the advantage of exploiting the global structure of the data space. Motivated by this work, we present a large-scale algorithm called ProtEmbed, which learns an embedding of protein sequences into a low-dimensional "semantic space." Evolutionarily related proteins are embedded in close proximity, and additional pieces of evidence, such as 3D structural similarity or class labels, can be incorporated into the learning process. We find that ProtEmbed achieves superior accuracy to widely used pairwise sequence methods like PSI-BLAST and HHSearch for remote homology detection; it also outperforms our previous RankProp algorithm, which incorporates global structure in the form of a protein similarity network. Finally, the ProtEmbed embedding space can be visualized, both at the global level and local to a given query, yielding intuition about the structure of protein sequence space.
Iain Melvin, Jason Weston, William Stafford Noble, Christina S. Leslie
PLoS Comput. Biol.4
2010 iDBPs: a web server for the identification of DNA binding proteins
abstract
SUMMARY: The iDBPs server uses the three-dimensional (3D) structure of a query protein to predict whether it binds DNA. First, the algorithm predicts the functional region of the protein based on its evolutionary profile; the assumption is that large clusters of conserved residues are good markers of functional regions. Next, various characteristics of the predicted functional region as well as global features of the protein are calculated, such as the average surface electrostatic potential, the dipole moment and cluster-based amino acid conservation patterns. Finally, a random forests classifier is used to predict whether the query protein is likely to bind DNA and to estimate the prediction confidence. We have trained and tested the classifier on various datasets and shown that it outperformed related methods. On a dataset that reflects the fraction of DNA binding proteins (DBPs) in a proteome, the area under the ROC curve was 0.90. The application of the server to an updated version of the N-Func database, which contains proteins of unknown function with solved 3D-structure, suggested new putative DBPs for experimental studies. AVAILABILITY: http://idbps.tau.ac.il/
Guy Nimrod, Maya Schushan, András Szilágyi 0002, Christina S. Leslie, Nir Ben-Tal
Bioinform.4
2010 High Resolution Models of Transcription Factor-DNA Affinities Improve In Vitro and In Vivo Binding Predictions
abstract
Accurately modeling the DNA sequence preferences of transcription factors (TFs), and using these models to predict in vivo genomic binding sites for TFs, are key pieces in deciphering the regulatory code. These efforts have been frustrated by the limited availability and accuracy of TF binding site motifs, usually represented as position-specific scoring matrices (PSSMs), which may match large numbers of sites and produce an unreliable list of target genes. Recently, protein binding microarray (PBM) experiments have emerged as a new source of high resolution data on in vitro TF binding specificities. PBM data has been analyzed either by estimating PSSMs or via rank statistics on probe intensities, so that individual sequence patterns are assigned enrichment scores (E-scores). This representation is informative but unwieldy because every TF is assigned a list of thousands of scored sequence patterns. Meanwhile, high-resolution in vivo TF occupancy data from ChIP-seq experiments is also increasingly available. We have developed a flexible discriminative framework for learning TF binding preferences from high resolution in vitro and in vivo data. We first trained support vector regression (SVR) models on PBM data to learn the mapping from probe sequences to binding intensities. We used a novel -mer based string kernel called the di-mismatch kernel to represent probe sequence similarities. The SVR models are more compact than E-scores, more expressive than PSSMs, and can be readily used to scan genomics regions to predict in vivo occupancy. Using a large data set of yeast and mouse TFs, we found that our SVR models can better predict probe intensity than the E-score method or PBM-derived PSSMs. Moreover, by using SVRs to score yeast, mouse, and human genomic regions, we were better able to predict genomic occupancy as measured by ChIP-chip and ChIP-seq experiments. Finally, we found that by training kernel-based models directly on ChIP-seq data, we greatly improved in vivo occupancy prediction, and by comparing a TF's in vitro and in vivo models, we could identify cofactors and disambiguate direct and indirect binding.
Phaedra Agius, Aaron Arvey, William Stafford Noble, Christina S. Leslie
PLoS Comput. Biol.5
2010 Learning "graph-mer" Motifs that Predict Gene Expression Trajectories in Development
abstract
A key problem in understanding transcriptional regulatory networks is deciphering what cis regulatory logic is encoded in gene promoter sequences and how this sequence information maps to expression. A typical computational approach to this problem involves clustering genes by their expression profiles and then searching for overrepresented motifs in the promoter sequences of genes in a cluster. However, genes with similar expression profiles may be controlled by distinct regulatory programs. Moreover, if many gene expression profiles in a data set are highly correlated, as in the case of whole organism developmental time series, it may be difficult to resolve fine-grained clusters in the first place. We present a predictive framework for modeling the natural flow of information, from promoter sequence to expression, to learn cis regulatory motifs and characterize gene expression patterns in developmental time courses. We introduce a cluster-free algorithm based on a graph-regularized version of partial least squares (PLS) regression to learn sequence patterns--represented by graphs of k-mers, or "graph-mers"--that predict gene expression trajectories. Applying the approach to wildtype germline development in Caenorhabditis elegans, we found that the first and second latent PLS factors mapped to expression profiles for oocyte and sperm genes, respectively. We extracted both known and novel motifs from the graph-mers associated to these germline-specific patterns, including novel CG-rich motifs specific to oocyte genes. We found evidence supporting the functional relevance of these putative regulatory elements through analysis of positional bias, motif conservation and in situ gene expression. This study demonstrates that our regression model can learn biologically meaningful latent structure and identify potentially functional motifs from subtle developmental time course expression data.
Xuejing Li, Casandra Panea, Chris Wiggins 0001, Valerie Reinke, Christina S. Leslie
PLoS Comput. Biol.5
2009 RANKPROP: a web server for protein remote homology detection
abstract
UNLABELLED: We present a large-scale implementation of the Rankprop protein homology ranking algorithm in the form of an openly accessible web server. We use the NRDB40 PSI-BLAST all-versus-all protein similarity network of 1.1 million proteins to construct the graph for the Rankprop algorithm, whereas previously, results were only reported for a database of 108 000 proteins. We also describe two algorithmic improvements to the original algorithm, including propagation from multiple homologs of the query and better normalization of ranking scores, that lead to higher accuracy and to scores with a probabilistic interpretation. AVAILABILITY: The Rankprop web server and source code are available at http://rankprop.gs.washington.edu
Iain Melvin, Jason Weston, Christina S. Leslie, William Stafford Noble
Bioinform.3
2008 Combining classifiers for improved classification of proteins from sequence or structure
abstract
BACKGROUND: Predicting a protein's structural or functional class from its amino acid sequence or structure is a fundamental problem in computational biology. Recently, there has been considerable interest in using discriminative learning algorithms, in particular support vector machines (SVMs), for classification of proteins. However, because sufficiently many positive examples are required to train such classifiers, all SVM-based methods are hampered by limited coverage. RESULTS: In this study, we develop a hybrid machine learning approach for classifying proteins, and we apply the method to the problem of assigning proteins to structural categories based on their sequences or their 3D structures. The method combines a full-coverage but lower accuracy nearest neighbor method with higher accuracy but reduced coverage multiclass SVMs to produce a full coverage classifier with overall improved accuracy. The hybrid approach is based on the simple idea of "punting" from one method to another using a learned threshold. CONCLUSION: In cross-validated experiments on the SCOP hierarchy, the hybrid methods consistently outperform the individual component methods at all levels of coverage. Code and data sets are available at http://noble.gs.washington.edu/proj/sabretooth.
Iain Melvin, Jason Weston, Christina S. Leslie, William Stafford Noble
BMC Bioinform.3
2008 A Predictive Model of the Oxygen and Heme Regulatory Network in Yeast
abstract
Deciphering gene regulatory mechanisms through the analysis of high-throughput expression data is a challenging computational problem. Previous computational studies have used large expression datasets in order to resolve fine patterns of coexpression, producing clusters or modules of potentially coregulated genes. These methods typically examine promoter sequence information, such as DNA motifs or transcription factor occupancy data, in a separate step after clustering. We needed an alternative and more integrative approach to study the oxygen regulatory network in Saccharomyces cerevisiae using a small dataset of perturbation experiments. Mechanisms of oxygen sensing and regulation underlie many physiological and pathological processes, and only a handful of oxygen regulators have been identified in previous studies. We used a new machine learning algorithm called MEDUSA to uncover detailed information about the oxygen regulatory network using genome-wide expression changes in response to perturbations in the levels of oxygen, heme, Hap1, and Co2+. MEDUSA integrates mRNA expression, promoter sequence, and ChIP-chip occupancy data to learn a model that accurately predicts the differential expression of target genes in held-out data. We used a novel margin-based score to extract significant condition-specific regulators and assemble a global map of the oxygen sensing and regulatory network. This network includes both known oxygen and heme regulators, such as Hap1, Mga2, Hap4, and Upc2, as well as many new candidate regulators. MEDUSA also identified many DNA motifs that are consistent with previous experimentally identified transcription factor binding sites. Because MEDUSA's regulatory program associates regulators to target genes through their promoter sequences, we directly tested the predicted regulators for OLE1, a gene specifically induced under hypoxia, by experimental analysis of the activity of its promoter. In each case, deletion of the candidate regulator resulted in the predicted effect on promoter activity, confirming that several novel regulators identified by MEDUSA are indeed involved in oxygen regulation. MEDUSA can reveal important information from a small dataset and generate testable hypotheses for further experimental analysis. Supplemental data are included.
Anshul Kundaje, Xiantong Xin, Changgui Lan, Steve Lianoglou, Mei Zhou, Li Zhang 0014, Christina S. Leslie
PLoS Comput. Biol.7
2007 NIPS workshop on New Problems and Methods in Computational Biology
abstract
The field of computational biology has seen dramatic growth over the past few years, both in terms of available data, scientific questions and challenges for learning and inference. These new types of scientific and clinical problems require the development of novel supervised and unsupervised learning approaches. In particular, the field is characterized by a diversity of heterogeneous data. The human genome sequence is accompanied by real-valued gene expression data, functional annotation of genes, genotyping information, a graph of interacting proteins, a set of equations describing the dynamics of a system, localization of proteins in a cell, a phylogenetic tree relating species, natural language text in the form of papers describing experiments, partial models that provide priors, and numerous other data sources. This supplementary issue consists of seven peer-reviewed papers based on the NIPS Workshop on New Problems and Methods in Computational Biology held at Whistler, British Columbia, Canada on December 8, 2006. The Neural Information Processing Systems Conference is the premier scientific meeting on neural computation, with session topics spanning artificial intelligence, learning theory, neuroscience, etc. The goal of this workshop was to present emerging problems and machine learning techniques in computational biology, with a particular emphasis on methods for computational learning from heterogeneous data. We received 37 extended abstract submissions, from which 13 were selected for oral presentation. The current supplement contains seven papers based on a subset of the 13 extended abstracts. Submitted manuscripts were rigorously reviewed by at least two referees. The quality of each paper was evaluated with respect to its contribution to biology as well as the novelty of the machine learning methods employed.
Gal Chechik, Christina S. Leslie, William Stafford Noble, Gunnar Rätsch, Quaid Morris, Koji Tsuda
BMC Bioinform.2
2007 SVM-Fold: a tool for discriminative multi-class protein fold and superfamily recognition
abstract
BACKGROUND: Predicting a protein's structural class from its amino acid sequence is a fundamental problem in computational biology. Much recent work has focused on developing new representations for protein sequences, called string kernels, for use with support vector machine (SVM) classifiers. However, while some of these approaches exhibit state-of-the-art performance at the binary protein classification problem, i.e. discriminating between a particular protein class and all other classes, few of these studies have addressed the real problem of multi-class superfamily or fold recognition. Moreover, there are only limited software tools and systems for SVM-based protein classification available to the bioinformatics community. RESULTS: We present a new multi-class SVM-based protein fold and superfamily recognition system and web server called SVM-Fold, which can be found at http://svm-fold.c2b2.columbia.edu. Our system uses an efficient implementation of a state-of-the-art string kernel for sequence profiles, called the profile kernel, where the underlying feature representation is a histogram of inexact matching k-mer frequencies. We also employ a novel machine learning approach to solve the difficult multi-class problem of classifying a sequence of amino acids into one of many known protein structural classes. Binary one-vs-the-rest SVM classifiers that are trained to recognize individual structural classes yield prediction scores that are not comparable, so that standard "one-vs-all" classification fails to perform well. Moreover, SVMs for classes at different levels of the protein structural hierarchy may make useful predictions, but one-vs-all does not try to combine these multiple predictions. To deal with these problems, our method learns relative weights between one-vs-the-rest classifiers and encodes information about the protein structural hierarchy for multi-class prediction. In large-scale benchmark results based on the SCOP database, our code weighting approach significantly improves on the standard one-vs-all method for both the superfamily and fold prediction in the remote homology setting and on the fold recognition problem. Moreover, our code weight learning algorithm strongly outperforms nearest-neighbor methods based on PSI-BLAST in terms of prediction accuracy on every structure classification problem we consider. CONCLUSION: By combining state-of-the-art SVM kernel methods with a novel multi-class algorithm, the SVM-Fold system delivers efficient and accurate protein fold and superfamily recognition.
Iain Melvin, Eugene Ie, Rui Kuang, Jason Weston, William Stafford Noble, Christina S. Leslie
BMC Bioinform.6
2007 Multi-class Protein Classification Using Adaptive Codes
Iain Melvin, Eugene Ie, Jason Weston, William Stafford Noble, Christina S. Leslie
J. Mach. Learn. Res.5
2006 A classification-based framework for predicting and analyzing gene regulatory response
abstract
BACKGROUND: We have recently introduced a predictive framework for studying gene transcriptional regulation in simpler organisms using a novel supervised learning algorithm called GeneClass. GeneClass is motivated by the hypothesis that in model organisms such as Saccharomyces cerevisiae, we can learn a decision rule for predicting whether a gene is up- or down-regulated in a particular microarray experiment based on the presence of binding site subsequences ("motifs") in the gene's regulatory region and the expression levels of regulators such as transcription factors in the experiment ("parents"). GeneClass formulates the learning task as a classification problem--predicting +1 and -1 labels corresponding to up- and down-regulation beyond the levels of biological and measurement noise in microarray measurements. Using the Adaboost algorithm, GeneClass learns a prediction function in the form of an alternating decision tree, a margin-based generalization of a decision tree. METHODS: In the current work, we introduce a new, robust version of the GeneClass algorithm that increases stability and computational efficiency, yielding a more scalable and reliable predictive model. The improved stability of the prediction tree enables us to introduce a detailed post-processing framework for biological interpretation, including individual and group target gene analysis to reveal condition-specific regulation programs and to suggest signaling pathways. Robust GeneClass uses a novel stabilized variant of boosting that allows a set of correlated features, rather than single features, to be included at nodes of the tree; in this way, biologically important features that are correlated with the single best feature are retained rather than decorrelated and lost in the next round of boosting. Other computational developments include fast matrix computation of the loss function for all features, allowing scalability to large datasets, and the use of abstaining weak rules, which results in a more shallow and interpretable tree. We also show how to incorporate genome-wide protein-DNA binding data from ChIP chip experiments into the GeneClass algorithm, and we use an improved noise model for gene expression data. RESULTS: Using the improved scalability of Robust GeneClass, we present larger scale experiments on a yeast environmental stress dataset, training and testing on all genes and using a comprehensive set of potential regulators. We demonstrate the improved stability of the features in the learned prediction tree, and we show the utility of the post-processing framework by analyzing two groups of genes in yeast--the protein chaperones and a set of putative targets of the Nrg1 and Nrg2 transcription factors--and suggesting novel hypotheses about their transcriptional and post-transcriptional regulation. Detailed results and Robust GeneClass source code is available for download from http://www.cs.columbia.edu/compbio/robust-geneclass.
Anshul Kundaje, Manuel Middendorf, Mihir Shah, Chris Wiggins 0001, Yoav Freund, Christina S. Leslie
BMC Bioinform.6
2006 Protein Ranking by Semi-Supervised Network Propagation
abstract
BACKGROUND: Biologists regularly search DNA or protein databases for sequences that share an evolutionary or functional relationship with a given query sequence. Traditional search methods, such as BLAST and PSI-BLAST, focus on detecting statistically significant pairwise sequence alignments and often miss more subtle sequence similarity. Recent work in the machine learning community has shown that exploiting the global structure of the network defined by these pairwise similarities can help detect more remote relationships than a purely local measure. METHODS: We review RankProp, a ranking algorithm that exploits the global network structure of similarity relationships among proteins in a database by performing a diffusion operation on a protein similarity network with weighted edges. The original RankProp algorithm is unsupervised. Here, we describe a semi-supervised version of the algorithm that uses labeled examples. Three possible ways of incorporating label information are considered: (i) as a validation set for model selection, (ii) to learn a new network, by choosing which transfer function to use for a given query, and (iii) to estimate edge weights, which measure the probability of inferring structural similarity. RESULTS: Benchmarked on a human-curated database of protein structures, the original RankProp algorithm provides significant improvement over local network search algorithms such as PSI-BLAST. Furthermore, we show here that labeled data can be used to learn a network without any need for estimating parameters of the transfer function, and that diffusion on this learned network produces better results than the original RankProp algorithm with a fixed network. CONCLUSION: In order to gain maximal information from a network, labeled and unlabeled data should be used to extract both local and global structure.
Jason Weston, Rui Kuang, Christina S. Leslie, William Stafford Noble
BMC Bioinform.3
2005 Multi-class protein fold recognition using adaptive codes
abstract
We develop a novel multi-class classification method based on output codes for the problem of classifying a sequence of amino acids into one of many known protein structural classes, called folds. Our method learns relative weights between one-vs-all classifiers and encodes information about the protein structural hierarchy for multi-class prediction. Our code weighting approach significantly improves on the standard one-vs-all method for the fold recognition problem. In order to compare against widely used methods in protein sequence analysis, we also test nearest neighbor approaches based on the PSI-BLAST algorithm. Our code weight learning algorithm strongly outperforms these PSI-BLAST methods on every structure recognition problem we consider. 1.
Eugene Ie, Jason Weston, William Stafford Noble, Christina S. Leslie
ICML4
2005 Motif Discovery Through Predictive Modeling of Gene Regulation
Manuel Middendorf, Anshul Kundaje, Mihir Shah, Yoav Freund, Chris Wiggins 0001, Christina S. Leslie
RECOMB6
2005 Motif-based protein ranking by network propagation
abstract
MOTIVATION: Sequence similarity often suggests evolutionary relationships between protein sequences that can be important for inferring similarity of structure or function. The most widely-used pairwise sequence comparison algorithms for homology detection, such as BLAST and PSI-BLAST, often fail to detect less conserved remotely-related targets. RESULTS: In this paper, we propose a new general graph-based propagation algorithm called MotifProp to detect more subtle similarity relationships than pairwise comparison methods. MotifProp is based on a protein-motif network, in which edges connect proteins and the k-mer based motif features that they contain. We show that our new motif-based propagation algorithm can improve the ranking results over a base algorithm, such as PSI-BLAST, that is used to initialize the ranking. Despite the complex structure of the protein-motif network, MotifProp can be easily interpreted using the top-ranked motifs and motif-rich regions induced by the propagation, both of which are helpful for discovering conserved structural components in remote homologies.
Rui Kuang, Jason Weston, William Stafford Noble, Christina S. Leslie
Bioinform.4
2005 Semi-supervised protein classification using cluster kernels
abstract
MOTIVATION: Building an accurate protein classification system depends critically upon choosing a good representation of the input sequences of amino acids. Recent work using string kernels for protein data has achieved state-of-the-art classification performance. However, such representations are based only on labeled data--examples with known 3D structures, organized into structural classes--whereas in practice, unlabeled data are far more plentiful. RESULTS: In this work, we develop simple and scalable cluster kernel techniques for incorporating unlabeled data into the representation of protein sequences. We show that our methods greatly improve the classification performance of string kernels and outperform standard approaches for using unlabeled data, such as adding close homologs of the positive examples to the training data. We achieve equal or superior performance to previously presented cluster kernel methods and at the same time achieving far greater computational efficiency. AVAILABILITY: Source code is available at www.kyb.tuebingen.mpg.de/bs/people/weston/semiprot. The Spider matlab package is available at www.kyb.tuebingen.mpg.de/bs/people/spider. SUPPLEMENTARY INFORMATION: www.kyb.tuebingen.mpg.de/bs/people/weston/semiprot.
Jason Weston, Christina S. Leslie, Eugene Ie, Dengyong Zhou, André Elisseeff, William Stafford Noble
Bioinform.2
2005 Combining Sequence and Time Series Expression Data to Learn Transcriptional Modules
abstract
Our goal is to cluster genes into transcriptional modules--sets of genes where similarity in expression is explained by common regulatory mechanisms at the transcriptional level. We want to learn modules from both time series gene expression data and genome-wide motif data that are now readily available for organisms such as S. cereviseae as a result of prior computational studies or experimental results. We present a generative probabilistic model for combining regulatory sequence and time series expression data to cluster genes into coherent transcriptional modules. Starting with a set of motifs representing known or putative regulatory elements (transcription factor binding sites) and the counts of occurrences of these motifs in each gene's promoter region, together with a time series expression profile for each gene, the learning algorithm uses expectation maximization to learn module assignments based on both types of data. We also present a technique based on the Jensen-Shannon entropy contributions of motifs in the learned model for associating the most significant motifs to each module. Thus, the algorithm gives a global approach for associating sets of regulatory elements to "modules" of genes with similar time series expression profiles. The model for expression data exploits our prior belief of smooth dependence on time by using statistical splines and is suitable for typical time course data sets with relatively few experiments. Moreover, the model is sufficiently interpretable that we can understand how both sequence data and expression data contribute to the cluster assignments, and how to interpolate between the two data sources. We present experimental results on the yeast cell cycle to validate our method and find that our combined expression and motif clustering algorithm discovers modules with both coherent expression and similar motif patterns, including binding motifs associated to known cell cycle transcription factors.
Anshul Kundaje, Manuel Middendorf, Chris Wiggins 0001, Christina S. Leslie
IEEE ACM Trans. Comput. Biol. Bioinform.5
2004 Protein backbone angle prediction with machine learning approaches
abstract
MOTIVATION: Protein backbone torsion angle prediction provides useful local structural information that goes beyond conventional three-state (alpha, beta and coil) secondary structure predictions. Accurate prediction of protein backbone torsion angles will substantially improve modeling procedures for local structures of protein sequence segments, especially in modeling loop conformations that do not form regular structures as in alpha-helices or beta-strands. RESULTS: We have devised two novel automated methods in protein backbone conformational state prediction: one method is based on support vector machines (SVMs); the other method combines a standard feed-forward back-propagation artificial neural network (NN) with a local structure-based sequence profile database (LSBSP1). Extensive benchmark experiments demonstrate that both methods have improved the prediction accuracy rate over the previously published methods for conformation state prediction when using an alphabet of three or four states. AVAILABILITY: LSBSP1 and the NN algorithm have been implemented in PrISM.1, which is available from www.columbia.edu/~ay1/. SUPPLEMENTARY INFORMATION: Supplementary data for the SVM method can be downloaded from the Website www.cs.columbia.edu/compbio/backbone.
Rui Kuang, Christina S. Leslie, An-Suei Yang
Bioinform.2
2004 Mismatch string kernels for discriminative protein classification
abstract
MOTIVATION: Classification of proteins sequences into functional and structural families based on sequence homology is a central problem in computational biology. Discriminative supervised machine learning approaches provide good performance, but simplicity and computational efficiency of training and prediction are also important concerns. RESULTS: We introduce a class of string kernels, called mismatch kernels, for use with support vector machines (SVMs) in a discriminative approach to the problem of protein classification and remote homology detection. These kernels measure sequence similarity based on shared occurrences of fixed-length patterns in the data, allowing for mutations between patterns. Thus, the kernels provide a biologically well-motivated way to compare protein sequences without relying on family-based generative models such as hidden Markov models. We compute the kernels efficiently using a mismatch tree data structure, allowing us to calculate the contributions of all patterns occurring in the data in one pass while traversing the tree. When used with an SVM, the kernels enable fast prediction on test sequences. We report experiments on two benchmark SCOP datasets, where we show that the mismatch kernel used with an SVM classifier performs competitively with state-of-the-art methods for homology detection, particularly when very few training examples are available. Examination of the highest-weighted patterns learned by the SVM classifier recovers biologically important motifs in protein families and superfamilies.
Christina S. Leslie, Eleazar Eskin, Adiel Cohen, Jason Weston, William Stafford Noble
Bioinform.1
2004 Fast String Kernels using Inexact Matching for Protein Sequences
Christina S. Leslie, Rui Kuang
J. Mach. Learn. Res.1
2003 Semi-supervised Protein Classification Using Cluster Kernels
abstract
A key issue in supervised protein classification is the representation of in- put sequences of amino acids. Recent work using string kernels for pro- tein data has achieved state-of-the-art classification performance. How- ever, such representations are based only on labeled data — examples with known 3D structures, organized into structural classes — while in practice, unlabeled data is far more plentiful. In this work, we de- velop simple and scalable cluster kernel techniques for incorporating un- labeled data into the representation of protein sequences. We show that our methods greatly improve the classification performance of string ker- nels and outperform standard approaches for using unlabeled data, such as adding close homologs of the positive examples to the training data. We achieve equal or superior performance to previously presented cluster kernel methods while achieving far greater computational efficiency.
Jason Weston, Christina S. Leslie, Dengyong Zhou, André Elisseeff, William Stafford Noble
NIPS2
2002 A Kernel Approach for Learning from almost Orthogonal Patterns
Bernhard Schölkopf, Jason Weston, Eleazar Eskin, Christina S. Leslie, William Stafford Noble
ECML4
2002 Mismatch String Kernels for SVM Protein Classification
abstract
We introduce a class of string kernels, called mismatch kernels, for use with support vector machines (SVMs) in a discriminative approach to the protein classification problem. These kernels measure sequence sim- ilarity based on shared occurrences of  -length subsequences, counted with up to mismatches, and do not rely on any generative model for the positive training sequences. We compute the kernels efficiently using a mismatch tree data structure and report experiments on a benchmark SCOP dataset, where we show that the mismatch kernel used with an SVM classifier performs as well as the Fisher kernel, the most success- ful method for remote homology detection, while achieving considerable computational savings.
Christina S. Leslie, Eleazar Eskin, Jason Weston, William Stafford Noble
NIPS1
2002 A Kernel Approach for Learning from Almost Orthogonal Patterns
Bernhard Schölkopf, Jason Weston, Eleazar Eskin, Christina S. Leslie, William Stafford Noble
PKDD4