Mikael Bodén

dblp:30/5501 · DBLP profile ↗
← Back
48ranked-venue papers
16as first author
5since 2021 · last 2025
0000-0003-3548-268XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 33 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 15 · 11 first-author
YearPublicationVenuePosition
2025 Do protein language models learn phylogeny?
abstract
Deep machine learning demonstrates a capacity to uncover evolutionary relationships directly from protein sequences, in effect internalising notions inherent to classical phylogenetic tree inference. We connect these two paradigms by assessing the capacity of protein-based language models (pLMs) to discern phylogenetic relationships without being explicitly trained to do so. We evaluate ESM2, ProtTrans, and MSA-Transformer relative to classical phylogenetic methods, while also considering sequence insertions and deletions (indels) across 114 Pfam datasets. The largest ESM2 model tends to outperform other pLMs (including the multimodal ESM3) by recovering phylogenetic relationships among homologous protein sequences in both low- and high-gap settings. pLMs agree with conventional phylogenetic methods in general, but more so for protein families with fewer implied indels, highlighting indels as a key factor differentiating classical phylogenetics from pLMs. We find that pLMs preferentially capture broader as opposed to finer evolutionary relationships within a specific protein family, where ESM2 has a sweet spot for highly divergent sequences, at remote distance. Less than 10% of neurons are sufficient to broadly recapitulate classical phylogenetic distances; when used in isolation, the difference between the paradigms is further diminished. We show these neurons are polysemantic, shared among different homologous families but never fully overlapping. We highlight the potential of ESM2 as a complementary tool for phylogenetic analysis, especially when extending to remote homologs that are difficult to align and imply complex histories of insertions and deletions. Implementations of analyses are available at https://github.com/santule/pLMEvo.
Sanjana Tule, Gabriel Foley, Mikael Bodén
Briefings Bioinform.3
2025 TRIAGE: an R package for regulatory gene analysis
abstract
Regulatory genes are critical determinants of cellular responses in development and disease, but standard RNA sequencing (RNA-seq) analysis workflows, such as differential expression analysis, have significant limitations in revealing the regulatory basis of cell identity and function. To address this challenge, we present the TRIAGE R package, a toolkit specifically designed to analyze regulatory elements in both bulk and single-cell RNA-seq datasets. The package is built upon TRIAGE methods, which leverage consortium-level H3K27me3 data to enrich for cell-type-specific regulatory regions. It facilitates the construction of efficient and adaptable pipelines for transcriptomic data analysis and visualization, with a focus on revealing regulatory gene networks. We demonstrate the utility of the TRIAGE R package using three independent transcriptomic datasets, showcasing its integration into standard analysis workflows for examining regulatory mechanisms across diverse biological contexts. The TRIAGE R package is available on GitHub at https://github.com/palpant-comp/TRIAGE_R_Package.
Qiongyi Zhao, Woo Jun Shim, Yuliangzi Sun, Enakshi Sinniah, Sophie Shen, Mikael Bodén, Nathan J. Palpant
Briefings Bioinform.6
2024 Optimal phylogenetic reconstruction of insertion and deletion events
abstract
MOTIVATION: Insertions and deletions (indels) influence the genetic code in fundamentally distinct ways from substitutions, significantly impacting gene product structure and function. Despite their influence, the evolutionary history of indels is often neglected in phylogenetic tree inference and ancestral sequence reconstruction, hindering efforts to comprehend biological diversity determinants and engineer variants for medical and industrial applications. RESULTS: We frame determining the optimal history of indel events as a single Mixed-Integer Programming (MIP) problem, across all branch points in a phylogenetic tree adhering to topological constraints, and all sites implied by a given set of aligned, extant sequences. By disentangling the impact on ancestral sequences at each branch point, this approach identifies the minimal indel events that jointly explain the diversity in sequences mapped to the tips of that tree. MIP can recover alternate optimal indel histories, if available. We evaluated MIP for indel inference on a dataset comprising 15 real phylogenetic trees associated with protein families ranging from 165 to 2000 extant sequences, and on 60 synthetic trees at comparable scales of data and reflecting realistic rates of mutation. Across relevant metrics, MIP outperformed alternative parsimony-based approaches and reported the fewest indel events, on par or below their occurrence in synthetic datasets. MIP offers a rational justification for indel patterns in extant sequences; importantly, it uniquely identifies global optima on complex protein data sets without making unrealistic assumptions of independence or evolutionary underpinnings, promising a deeper understanding of molecular evolution and aiding novel protein design. AVAILABILITY AND IMPLEMENTATION: The implementation is available via GitHub at https://github.com/santule/indelmip.
Sanjana Tule, Gabriel Foley, Chongting Zhao, Mikael Bodén
Bioinform.5
2023 Cytocipher determines significantly different populations of cells in single-cell RNA-seq data
abstract
MOTIVATION: Identification of cell types using single-cell RNA-seq is revolutionizing the study of multicellular organisms. However, typical single-cell RNA-seq analysis often involves post hoc manual curation to ensure clusters are transcriptionally distinct, which is time-consuming, error-prone, and irreproducible. RESULTS: To overcome these obstacles, we developed Cytocipher, a bioinformatics method and scverse compatible software package that statistically determines significant clusters. Application of Cytocipher to normal tissue, development, disease, and large-scale atlas data reveals the broad applicability and power of Cytocipher to generate biological insights in numerous contexts. This included the identification of cell types not previously described in the datasets analysed, such as CD8+ T cell subtypes in human peripheral blood mononuclear cells; cell lineage intermediate states during mouse pancreas development; and subpopulations of luminal epithelial cells over-represented in prostate cancer. Cytocipher also scales to large datasets with high-test performance, as shown by application to the Tabula Sapiens Atlas representing >480 000 cells. Cytocipher is a novel and generalizable method that statistically determines transcriptionally distinct and programmatically reproducible clusters from single-cell data. AVAILABILITY AND IMPLEMENTATION: The software version used for this manuscript has been deposited on Zenodo (https://doi.org/10.5281/zenodo.8089546), and is also available via github (https://github.com/BradBalderson/Cytocipher).
Brad Balderson, Michael Piper, Stefan Thor, Mikael Bodén
Bioinform.4
2022 Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP)
abstract
Ancestral sequence reconstruction is a technique that is gaining widespread use in molecular evolution studies and protein engineering. Accurate reconstruction requires the ability to handle appropriately large numbers of sequences, as well as insertion and deletion (indel) events, but available approaches exhibit limitations. To address these limitations, we developed Graphical Representation of Ancestral Sequence Predictions (GRASP), which efficiently implements maximum likelihood methods to enable the inference of ancestors of families with more than 10,000 members. GRASP implements partial order graphs (POGs) to represent and infer insertion and deletion events across ancestors, enabling the identification of building blocks for protein engineering. To validate the capacity to engineer novel proteins from realistic data, we predicted ancestor sequences across three distinct enzyme families: glucose-methanol-choline (GMC) oxidoreductases, cytochromes P450, and dihydroxy/sugar acid dehydratases (DHAD). All tested ancestors demonstrated enzymatic activity. Our study demonstrates the ability of GRASP (1) to support large data sets over 10,000 sequences and (2) to employ insertions and deletions to identify building blocks for engineering biologically active ancestors, by exploring variation over evolutionary time.
Gabriel Foley, Ariane Mora, Connie M. Ross, Scott Bottoms, Leander Sützl, Marnie L. Lamprecht, Julian Zaugg, Alexandra Essebier, Brad Balderson, Rhys Newell, Raine E. S. Thomson, Bostjan Kobe, Ross T. Barnard, Luke Guddat, Gerhard Schenk, Jörg Carsten, Yosephine Gumulya, Burkhard Rost, Dietmar Haltrich, Volker Sieber, Elizabeth M. J. Gillam, Mikael Bodén
PLoS Comput. Biol.22
2020 T-Gene: improved target gene prediction
abstract
MOTIVATION: Identifying the genes regulated by a given transcription factor (TF) (its 'target genes') is a key step in developing a comprehensive understanding of gene regulation. Previously, we developed a method (CisMapper) for predicting the target genes of a TF based solely on the correlation between a histone modification at the TF's binding site and the expression of the gene across a set of tissues or cell lines. That approach is limited to organisms for which extensive histone and expression data are available, and does not explicitly incorporate the genomic distance between the TF and the gene. RESULTS: We present the T-Gene algorithm, which overcomes these limitations. It can be used to predict which genes are most likely to be regulated by a TF, and which of the TF's binding sites are most likely involved in regulating particular genes. T-Gene calculates a novel score that combines distance and histone/expression correlation, and we show that this score accurately predicts when a regulatory element bound by a TF is in contact with a gene's promoter, achieving median precision above 60%. T-Gene is easy to use via its web server or as a command-line tool, and can also make accurate predictions (median precision above 40%) based on distance alone when extensive histone/expression data is not available for the organism. T-Gene provides an estimate of the statistical significance of each of its predictions. AVAILABILITY AND IMPLEMENTATION: The T-Gene web server, source code, histone/expression data and genome annotation files are provided at http://meme-suite.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Timothy O'Connor, Charles E. Grant, Mikael Bodén, Timothy L. Bailey
Bioinform.3
2019 Detailed prediction of protein sub-nuclear localization
abstract
BACKGROUND: Sub-nuclear structures or locations are associated with various nuclear processes. Proteins localized in these substructures are important to understand the interior nuclear mechanisms. Despite advances in high-throughput methods, experimental protein annotations remain limited. Predictions of cellular compartments have become very accurate, largely at the expense of leaving out substructures inside the nucleus making a fine-grained analysis impossible. RESULTS: Here, we present a new method (LocNuclei) that predicts nuclear substructures from sequence alone. LocNuclei used a string-based Profile Kernel with Support Vector Machines (SVMs). It distinguishes sub-nuclear localization in 13 distinct substructures and distinguishes between nuclear proteins confined to the nucleus and those that are also native to other compartments (traveler proteins). High performance was achieved by implicitly leveraging a large biological knowledge-base in creating predictions by homology-based inference through BLAST. Using this approach, the performance reached AUC = 0.70-0.74 and Q13 = 59-65%. Travelling proteins (nucleus and other) were identified at Q2 = 70-74%. A Gene Ontology (GO) analysis of the enrichment of biological processes revealed that the predicted sub-nuclear compartments matched the expected functionality. Analysis of protein-protein interactions (PPI) show that formation of compartments and functionality of proteins in these compartments highly rely on interactions between proteins. This suggested that the LocNuclei predictions carry important information about function. The source code and data sets are available through GitHub: https://github.com/Rostlab/LocNuclei . CONCLUSIONS: LocNuclei predicts subnuclear compartments and traveler proteins accurately. These predictions carry important information about functionality and PPIs.
Maria Littmann, Tatyana Goldberg, Sebastian Seitz, Mikael Bodén, Burkhard Rost
BMC Bioinform.4
2019 Correction to: Detailed prediction of protein sub-nuclear localization
abstract
Following publication of the original article [1], the author reported that an incorrect figure has been published as Figure 2. The correct Figure 2 is shown below.
Maria Littmann, Tatyana Goldberg, Sebastian Seitz, Mikael Bodén, Burkhard Rost
BMC Bioinform.4
2018 SCRAM: a pipeline for fast index-free small RNA read alignment and visualization
abstract
Summary: Small RNAs play key roles in gene regulation, defense against viral pathogens and maintenance of genome stability, though many aspects of their biogenesis and function remain to be elucidated. SCRAM (Small Complementary RNA Mapper) is a novel, simple-to-use short read aligner and visualization suite that enhances exploration of small RNA datasets. Availability and implementation: The SCRAM pipeline is implemented in Go and Python, and is freely available under MIT license. Source code, multiplatform binaries and a Docker image can be accessed via https://sfletc.github.io/scram/. Supplementary information: Supplementary data are available at Bioinformatics online.
Stephen J. Fletcher, Mikael Bodén, Neena Mitter, Bernard J. Carroll
Bioinform.2
2017 PhosphoPICK-SNP: quantifying the effect of amino acid variants on protein phosphorylation
abstract
MOTIVATION: Genome-wide association studies are identifying single nucleotide variants (SNVs) linked to various diseases, however the functional effect caused by these variants is often unknown. One potential functional effect, the loss or gain of protein phosphorylation sites, can be induced through variations in key amino acids that disrupt or introduce valid kinase binding patterns. Current methods for predicting the effect of SNVs on phosphorylation operate on the sequence content of reference and variant proteins. However, consideration of the amino acid sequence alone is insufficient for predicting phosphorylation change, as context factors determine kinase-substrate selection. RESULTS: We present here a method for quantifying the effect of SNVs on protein phosphorylation through an integrated system of motif analysis and context-based assessment of kinase targets. By predicting the effect that known variants across the proteome have on phosphorylation, we are able to use this background of proteome-wide variant effects to quantify the significance of novel variants for modifying phosphorylation. We validate our method on a manually curated set of phosphorylation change-causing variants from the primary literature, showing that the method predicts known examples of phosphorylation change at high levels of specificity. We apply our approach to data-sets of variants in phosphorylation site regions, showing that variants causing predicted phosphorylation loss are over-represented among disease-associated variants. AVAILABILITY AND IMPLEMENTATION: The method is freely available as a web-service at the website http://bioinf.scmb.uq.edu.au/phosphopick/snp. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ralph Patrick, Bostjan Kobe, Kim-Anh Lê Cao, Mikael Bodén
Bioinform.4
2015 Sequencing technologies and tools for short tandem repeat variation detection
abstract
Short tandem repeats are highly polymorphic and associated with a wide range of phenotypic variation, some of which cause neurodegenerative disease in humans. With advances in high-throughput sequencing technologies, there are novel opportunities to study genetic variation. While available sequencing technologies and bioinformatics tools provide options for mining high-throughput sequencing data, their suitability for analysis of repeat variation is an open question, with tools for quantifying variability in repetitive sequence still in their infancy. We present here a comprehensive survey and empirical evaluation of current sequencing technologies and bioinformatics tools in all stages of an analysis pipeline. While there is not one optimal pipeline to suit all circumstances, we find that the choice of alignment and repeat genotyping tools greatly impacts the accuracy and efficiency by which short tandem repeat variation can be detected. We further note that to detect variation relevant to many repeat diseases, it is essential to choose technologies that offer either long read-lengths or paired-end sequencing, coupled with specific genotyping tools.
Minh Duc Cao, Sureshkumar Balasubramanian, Mikael Bodén
Briefings Bioinform.3
2015 PhosphoPICK: modelling cellular context to map kinase-substrate phosphorylation events
abstract
MOTIVATION: The determinants of kinase-substrate phosphorylation can be found both in the substrate sequence and the surrounding cellular context. Cell cycle progression, interactions with mediating proteins and even prior phosphorylation events are necessary for kinases to maintain substrate specificity. While much work has focussed on the use of sequence-based methods to predict phosphorylation sites, there has been very little work invested into the application of systems biology to understand phosphorylation. Lack of specificity in many kinase substrate binding motifs means that sequence methods for predicting kinase binding sites are susceptible to high false-positive rates. RESULTS: We present here a model that takes into account protein-protein interaction information, and protein abundance data across the cell cycle to predict kinase substrates for 59 human kinases that are representative of important biological pathways. The model shows high accuracy for substrate prediction (with an average AUC of 0.86) across the 59 kinases tested. When using the model to complement sequence-based kinase-specific phosphorylation site prediction, we found that the additional information increased prediction performance for most comparisons made, particularly on kinases from the CMGC family. We then used our model to identify functional overlaps between predicted CDK2 substrates and targets from the E2F family of transcription factors. Our results demonstrate that a model harnessing context data can account for the short-falls in sequence information and provide a robust description of the cellular events that regulate protein phosphorylation. AVAILABILITY AND IMPLEMENTATION: The method is freely available online as a web server at the website http://bioinf.scmb.uq.edu.au/phosphopick. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ralph Patrick, Kim-Anh Lê Cao, Bostjan Kobe, Mikael Bodén
Bioinform.4
2014 Dimensionality reduction and topographic mapping of binary tensors
Jakub Mazgut, Peter Tiño, Mikael Bodén, Hong Yan 0001
Pattern Anal. Appl.3
2013 DLocalMotif: a discriminative approach for discovering local motifs in protein sequences
abstract
MOTIVATION: Local motifs are patterns of DNA or protein sequences that occur within a sequence interval relative to a biologically defined anchor or landmark. Current protein motif discovery methods do not adequately consider such constraints to identify biologically significant motifs that are only weakly over-represented but spatially confined. Using negatives, i.e. sequences known to not contain a local motif, can further increase the specificity of their discovery. RESULTS: This article introduces the method DLocalMotif that makes use of positional information and negative data for local motif discovery in protein sequences. DLocalMotif combines three scoring functions, measuring degrees of motif over-representation, entropy and spatial confinement, specifically designed to discriminatively exploit the availability of negative data. The method is shown to outperform current methods that use only a subset of these motif characteristics. We apply the method to several biological datasets. The analysis of peroxisomal targeting signals uncovers several novel motifs that occur immediately upstream of the dominant peroxisomal targeting signal-1 signal. The analysis of proline-tyrosine nuclear localization signals uncovers multiple novel motifs that overlap with C2H2 zinc finger domains. We also evaluate the method on classical nuclear localization signals and endoplasmic reticulum retention signals and find that DLocalMotif successfully recovers biologically relevant sequence properties. AVAILABILITY: http://bioinf.scmb.uq.edu.au/dlocalmotif/
Ahmed M. Mehdi, Muhammad Shoaib B. Sehgal, Bostjan Kobe, Timothy L. Bailey, Mikael Bodén
Bioinform.5
2013 PREDIVAC: CD4+ T-cell epitope prediction for vaccine design that covers 95% of HLA class II DR protein diversity
abstract
BACKGROUND: CD4+ T-cell epitopes play a crucial role in eliciting vigorous protective immune responses during peptide (epitope)-based vaccination. The prediction of these epitopes focuses on the peptide binding process by MHC class II proteins. The ability to account for MHC class II polymorphism is critical for epitope-based vaccine design tools, as different allelic variants can have different peptide repertoires. In addition, the specificity of CD4+ T-cells is often directed to a very limited set of immunodominant peptides in pathogen proteins. The ability to predict what epitopes are most likely to dominate an immune response remains a challenge. RESULTS: We developed the computational tool Predivac to predict CD4+ T-cell epitopes. Predivac can make predictions for 95% of all MHC class II protein variants (allotypes), a substantial advance over other available methods. Predivac bases its prediction on the concept of specificity-determining residues. The performance of the method was assessed both for high-affinity HLA class II peptide binding and CD4+ T-cell epitope prediction. In terms of epitope prediction, Predivac outperformed three available pan-specific approaches (delivering the highest specificity). A central finding was the high accuracy delivered by the method in the identification of immunodominant and promiscuous CD4+ T-cell epitopes, which play an essential role in epitope-based vaccine design. CONCLUSIONS: The comprehensive HLA class II allele coverage along with the high specificity in identifying immunodominant CD4+ T-cell epitopes makes Predivac a valuable tool to aid epitope-based vaccine design in the context of a genetically heterogeneous human population.The tool is available at: http://predivac.biosci.uq.edu.au/.
Patricio Oyarzún, Jonathan J. Ellis, Mikael Bodén, Bostjan Kobe
BMC Bioinform.3
2011 Sorting the nuclear proteome
abstract
MOTIVATION: Quantitative experimental analyses of the nuclear interior reveal a morphologically structured yet dynamic mix of membraneless compartments. Major nuclear events depend on the functional integrity and timely assembly of these intra-nuclear compartments. Yet, unknown drivers of protein mobility ensure that they are in the right place at the time when they are needed. RESULTS: This study investigates determinants of associations between eight intra-nuclear compartments and their proteins in heterogeneous genome-wide data. We develop a model based on a range of candidate determinants, capable of mapping the intra-nuclear organization of proteins. The model integrates protein interactions, protein domains, post-translational modification sites and protein sequence data. The predictions of our model are accurate with a mean AUC (over all compartments) of 0.71. We present a complete map of the association of 3567 mouse nuclear proteins with intra-nuclear compartments. Each decision is explained in terms of essential interactions and domains, and qualified with a false discovery assessment. Using this resource, we uncover the collective role of transcription factors in each of the compartments. We create diagrams illustrating the outcomes of a Gene Ontology enrichment analysis. Associated with an extensive range of transcription factors, the analysis suggests that PML bodies coordinate regulatory immune responses.
Denis C. Bauer, Kai Willadsen, Fabian A. Buske, Kim-Anh Lê Cao, Timothy L. Bailey, Graham Dellaire, Mikael Bodén
Bioinform.7
2011 A probabilistic model of nuclear import of proteins
abstract
MOTIVATION: Nucleo-cytoplasmic trafficking of proteins is a core regulatory process that sustains the integrity of the nuclear space of eukaryotic cells via an interplay between numerous factors. Despite progress on experimentally characterizing a number of nuclear localization signals, their presence alone remains an unreliable indicator of actual translocation. RESULTS: This article introduces a probabilistic model that explicitly recognizes a variety of nuclear localization signals, and integrates relevant amino acid sequence and interaction data for any candidate nuclear protein. In particular, we develop and incorporate scoring functions based on distinct classes of classical nuclear localization signals. Our empirical results show that the model accurately predicts whether a protein is imported into the nucleus, surpassing the classification accuracy of similar predictors when evaluated on the mouse and yeast proteomes (area under the receiver operator characteristic curve of 0.84 and 0.80, respectively). The model also predicts the sequence position of a nuclear localization signal and whether it interacts with importin-α. AVAILABILITY: http://pprowler.itee.uq.edu.au/NucImport
Ahmed M. Mehdi, Muhammad Shoaib B. Sehgal, Bostjan Kobe, Timothy L. Bailey, Mikael Bodén
Bioinform.5
2010 Multilinear Decomposition and Topographic Mapping of Binary Tensors
Jakub Mazgut, Peter Tiño, Mikael Bodén, Hong Yan 0001
ICANN (1)3
2010 Assigning roles to DNA regulatory motifs using comparative genomics
abstract
MOTIVATION: Transcription factors (TFs) are crucial during the lifetime of the cell. Their functional roles are defined by the genes they regulate. Uncovering these roles not only sheds light on the TF at hand but puts it into the context of the complete regulatory network. RESULTS: Here, we present an alignment- and threshold-free comparative genomics approach for assigning functional roles to DNA regulatory motifs. We incorporate our approach into the Gomo algorithm, a computational tool for detecting associations between a user-specified DNA regulatory motif [expressed as a position weight matrix (PWM)] and Gene Ontology (GO) terms. Incorporating multiple species into the analysis significantly improves Gomo's ability to identify GO terms associated with the regulatory targets of TFs. Including three comparative species in the process of predicting TF roles in Saccharomyces cerevisiae and Homo sapiens increases the number of significant predictions by 75 and 200%, respectively. The predicted GO terms are also more specific, yielding deeper biological insight into the role of the TF. Adjusting motif (binding) affinity scores for individual sequence composition proves to be essential for avoiding false positive associations. We describe a novel DNA sequence-scoring algorithm that compensates a thermodynamic measure of DNA-binding affinity for individual sequence base composition. GOMO's prediction accuracy proves to be relatively insensitive to how promoters are defined. Because GOMO uses a threshold-free form of gene set analysis, there are no free parameters to tune. Biologists can investigate the potential roles of DNA regulatory motifs of interest using GOMO via the web (http://meme.nbcr.net).
Fabian A. Buske, Mikael Bodén, Denis C. Bauer, Timothy L. Bailey
Bioinform.2
2010 The value of position-specific priors in motif discovery using MEME
abstract
BACKGROUND: Position-specific priors have been shown to be a flexible and elegant way to extend the power of Gibbs sampler-based motif discovery algorithms. Information of many types-including sequence conservation, nucleosome positioning, and negative examples-can be converted into a prior over the location of motif sites, which then guides the sequence motif discovery algorithm. This approach has been shown to confer many of the benefits of conservation-based and discriminative motif discovery approaches on Gibbs sampler-based motif discovery methods, but has not previously been studied with methods based on expectation maximization (EM). RESULTS: We extend the popular EM-based MEME algorithm to utilize position-specific priors and demonstrate their effectiveness for discovering transcription factor (TF) motifs in yeast and mouse DNA sequences. Utilizing a discriminative, conservation-based prior dramatically improves MEME's ability to discover motifs in 156 yeast TF ChIP-chip datasets, more than doubling the number of datasets where it finds the correct motif. On these datasets, MEME using the prior has a higher success rate than eight other conservation-based motif discovery approaches. We also show that the same type of prior improves the accuracy of motifs discovered by MEME in mouse TF ChIP-seq data, and that the motifs tend to be of slightly higher quality those found by a Gibbs sampling algorithm using the same prior. CONCLUSIONS: We conclude that using position-specific priors can substantially increase the power of EM-based motif discovery algorithms such as MEME algorithm.
Timothy L. Bailey, Mikael Bodén, Tom Whitington, Philip Machanick
BMC Bioinform.2
2010 Predicting SUMOylation sites in developmental transcription factors of Drosophila melanogaster
Denis C. Bauer, Fabian A. Buske, Timothy L. Bailey, Mikael Bodén
Neurocomputing4
2010 Using Gaussian Process with Test Rejection to Detect T-Cell Epitopes in Pathogen Genomes
abstract
A major challenge in the development of peptide-based vaccines is finding the right immunogenic element, with efficient and long-lasting immunization effects, from large potential targets encoded by pathogen genomes. Computer models are convenient tools for scanning pathogen genomes to preselect candidate immunogenic peptides for experimental validation. Current methods predict many false positives resulting from a low prevalence of true positives. We develop a test reject method based on the prediction uncertainty estimates determined by Gaussian process regression. This method filters false positives among predicted epitopes from a pathogen genome. The performance of stand-alone Gaussian process regression is compared to other state-of-the-art methods using cross validation on 11 benchmark data sets. The results show that the Gaussian process method has the same accuracy as the top performing algorithms. The combination of Gaussian process regression with the proposed test reject method is used to detect true epitopes from the Vaccinia virus genome. The test rejection increases the prediction accuracy by reducing the number of false positives without sacrificing the method's sensitivity. We show that the Gaussian process in combination with test rejection is an effective method for prediction of T-cell epitopes in large and diverse pathogen genomes, where false positives are of concern.
Liwen You, Vladimir Brusic, Marcus Gallagher, Mikael Bodén
IEEE ACM Trans. Comput. Biol. Bioinform.4
2009 Detecting sequence and structure homology via an integrative kernel: A case-study in recognizing enzymes
abstract
Sequence and structure are complementary pieces of information that can be used to infer protein function. We study and compare sequence, structure and sequence-structure integrative kernels to recognize proteins with enzymatic function. Using a support-vector machine, we show that kernels that combine sequence and structure information typically perform better (AUC 0.73) at this task than kernels that exploit either type of information exclusively. We find that the feature space of structure kernels complements that of sequence kernels, making both sources of similarity more accessible to kernel methods.
Isye Arieshanti, Mikael Bodén, Stefan Maetschke, Fabian A. Buske
CIBCB2
2009 It's about time: Signal recognition in staged models of protein translocation
Fabian A. Buske, Stefan Maetschke, Mikael Bodén
Pattern Recognit.3
2008 Predicting Nucleolar Proteins Using Support-Vector Machines
Mikael Bodén
APBC1
2007 A Comparison of Sequence Kernels for Localization Prediction of Transmembrane Proteins
abstract
We applied support vector machines to the prediction of the subcellular localization of transmembrane proteins, and compared the performance of different sequence kernels on this task. More specifically we measured prediction accuracy, computation time, number of kernel evaluations and number of support vectors for the spectrum, the full spectrum, the wildcard, the mismatch, the local-alignment and the residue-coupling kernel. The local-alignment achieved the highest prediction accuracy, with a Matthews correlation coefficient of 0.51, closely followed by the mismatch kernel. However, the local-alignment kernel was also the most time consuming kernel and seven times slower than the mismatch kernel. The spectrum kernel was the fastest kernel but linked to the highest number of support vectors and kernel evaluations. The residue-coupling kernel showed the lowest number of support vectors and kernel evaluations. No correlation between the number of support vectors and prediction accuracy could be observed. A localization predictor (TMPLoc) has been made available at http://pprowler.itee.uq.edu.au/TMPLoc
Stefan Maetschke, Marcus Gallagher, Mikael Bodén
CIBCB3
2006 Evolving discriminative motifs for recognizing proteins imported to the peroxisome via the PTS2 pathway
abstract
Peroxisomes are small subcellular compartments that utilize proteins manufactured in the cytoplasm. Proteins use one of two peroxisomal import pathways. This paper presents a simple evolutionary search for a motif that describes the signal used by one of the two pathways: PTS2. The evolved motif has a discriminative accuracy exceeding previously manually curated motifs and can be used to screen genomic data for putative peroxisomal proteins.
Mikael Bodén, John Hawkins
IEEE Congress on Evolutionary Computation1
2006 Identifying sequence regions undergoing conformational change via predicted continuum secondary structure
abstract
MOTIVATION: Conformational flexibility is essential to the function of many proteins, e.g. catalytic activity. To assist efforts in determining and exploring the functional properties of a protein, it is desirable to automatically identify regions that are prone to undergo conformational changes. It was recently shown that a probabilistic predictor of continuum secondary structure is more accurate than categorical predictors for structurally ambivalent sequence regions, suggesting that such models are suited to characterize protein flexibility. RESULTS: We develop a computational method for identifying regions that are prone to conformational change directly from the amino acid sequence. The method uses the entropy of the probabilistic output of an 8-class continuum secondary structure predictor. Results for 171 unique amino acid sequences with well-characterized variable structure (identified in the 'Macromolecular movements database') indicate that the method is highly sensitive at identifying flexible protein regions, but false positives remain a problem. The method can be used to explore conformational flexibility of proteins (including hypothetical or synthetic ones) whose structure is yet to be determined experimentally. AVAILABILITY: The predictor, sequence data and supplementary studies are available at http://pprowler.itee.uq.edu.au/sspred/ and are free for academic use.
Mikael Bodén, Timothy L. Bailey
Bioinform.1
2006 STAR: predicting recombination sites from amino acid sequence
abstract
Abstract Background Designing novel proteins with site-directed recombination has enormous prospects. By locating effective recombination sites for swapping sequence parts, the probability that hybrid sequences have the desired properties is increased dramatically. The prohibitive requirements for applying current tools led us to investigate machine learning to assist in finding useful recombination sites from amino acid sequence alone. Results We present STAR, Site Targeted Amino acid Recombination predictor, which produces a score indicating the structural disruption caused by recombination, for each position in an amino acid sequence. Example predictions contrasted with those of alternative tools, illustrate STAR'S utility to assist in determining useful recombination sites. Overall, the correlation coefficient between the output of the experimentally validated protein design algorithm SCHEMA and the prediction of STAR is very high (0.89). Conclusion STAR allows the user to explore useful recombination sites in amino acid sequences with unknown structure and unknown evolutionary origin. The predictor service is available from http://pprowler.itee.uq.edu.au/star .
Denis C. Bauer, Mikael Bodén, Ricarda Thier, Elizabeth M. J. Gillam
BMC Bioinform.2
2006 Prediction of protein continuum secondary structure with probabilistic models based on NMR solved structures
abstract
BACKGROUND: The structure of proteins may change as a result of the inherent flexibility of some protein regions. We develop and explore probabilistic machine learning methods for predicting a continuum secondary structure, i.e. assigning probabilities to the conformational states of a residue. We train our methods using data derived from high-quality NMR models. RESULTS: Several probabilistic models not only successfully estimate the continuum secondary structure, but also provide a categorical output on par with models directly trained on categorical data. Importantly, models trained on the continuum secondary structure are also better than their categorical counterparts at identifying the conformational state for structurally ambivalent residues. CONCLUSION: Cascaded probabilistic neural networks trained on the continuum secondary structure exhibit better accuracy in structurally ambivalent regions of proteins, while sustaining an overall classification accuracy on par with standard, categorical prediction methods.
Mikael Bodén, Timothy L. Bailey
BMC Bioinform.1
2005 Detecting residues in targeting peptides
Mikael Bodén, John Hawkins
APBC1
2005 BLOMAP: An encoding of amino acids which improves signal peptide cleavage site prediction
Stefan Maetschke, Michael W. Towsey, Mikael Bodén
APBC3
2005 Predicting Structural Disruption of Proteins Caused by Crossover (based on Amino Acid Sequence)
Denis C. Bauer, Mikael Bodén, Ricarda Thier
CIBCB2
2005 Predicting Peroxisomal Proteins
John Hawkins, Mikael Bodén
CIBCB2
2005 Heuristic Algorithm for Computing Reversal Distance with MultiGene Families via Binary Integer Programming
Jakkarin Suksawatchon, Chidchanok Lursinsap, Mikael Bodén
CIBCB3
2005 Exploiting Sequence Dependencies in the Prediction of Peroxisomal Proteins
Mark Wakabayashi, John Hawkins, Stefan Maetschke, Mikael Bodén
IDEAL4
2005 Prediction of subcellular localization using sequence-biased recurrent networks
abstract
MOTIVATION: Targeting peptides direct nascent proteins to their specific subcellular compartment. Knowledge of targeting signals enables informed drug design and reliable annotation of gene products. However, due to the low similarity of such sequences and the dynamical nature of the sorting process, the computational prediction of subcellular localization of proteins is challenging. RESULTS: We contrast the use of feed forward models as employed by the popular TargetP/SignalP predictors with a sequence-biased recurrent network model. The models are evaluated in terms of performance at the residue level and at the sequence level, and demonstrate that recurrent networks improve the overall prediction performance. Compared to the original results reported for TargetP, an ensemble of the tested models increases the accuracy by 6 and 5% on non-plant and plant data, respectively. AVAILABILITY: The Protein Prowler incorporating the recurrent network predictor described in this paper is available online at http://pprowler.imb.uq.edu.au/
Mikael Bodén, John Hawkins
Bioinform.1
2005 The Applicability of Recurrent Neural Networks for Biological Sequence Analysis
abstract
Selection of machine learning techniques requires a certain sensitivity to the requirements of the problem. In particular, the problem can be made more tractable by deliberately using algorithms that are biased toward solutions of the requisite kind. In this paper, we argue that recurrent neural networks have a natural bias toward a problem domain of which biological sequence analysis tasks are a subset. We use experiments with synthetic data to illustrate this bias. We then demonstrate that this bias can be exploitable using a data set of protein sequences containing several classes of subcellular localization targeting peptides. The results show that, compared with feed forward, recurrent neural networks will generally perform better on sequence analysis tasks. Furthermore, as the patterns within the sequence become more ambiguous, the choice of specific recurrent architecture becomes more critical.
John Hawkins, Mikael Bodén
IEEE ACM Trans. Comput. Biol. Bioinform.2
2005 Improved access to sequential motifs: a note on the architectural bias of recurrent networks
abstract
For many biological sequence problems the available data occupies only sparse regions of the problem space. To use machine learning effectively for the analysis of sparse data we must employ architectures with an appropriate bias. By experimentation we show that the bias of recurrent neural networks--recently analyzed by Tino et al. and Hammer and Tino--offers superior access to motifs (sequential patterns) compared to the, in bioinformatics, standardly used feedforward neural networks.
Mikael Bodén, John Hawkins
IEEE Trans. Neural Networks1
2004 Generalization by symbolic abstraction in cascaded recurrent networks
Mikael Bodén
Neurocomputing1
2003 Using evolutionary noise to improve prediction of rapidly evolving targeting peptides
abstract
Targeting peptides are responsible for directing proteins to the appropriate subcellular location. As a group of biological sequences, targeting peptides are evolving at a relatively high rate and exhibit diversity. We investigate if evolutionary noise-simulated mutation at the molecular level-improves target classification for a neural network predictor. Comparison with the well-known TargetP prediction service illustrates some advantages of the approach. Specifically, classification of signal peptides, which exhibit an extremely high rate of evolution, is improved.
Mikael Bodén
IEEE Congress on Evolutionary Computation1
2003 Learning the Dynamics of Embedded Clauses
Mikael Bodén, Alan Blair 0001
Appl. Intell.1
2002 Generalization by structural properties from sparse nested symbolic data
Mikael Bodén
ESANN1
2002 On learning context-free and context-sensitive languages
abstract
The long short-term memory (LSTM) is not the only neural network which learns a context sensitive language. Second-order sequential cascaded networks (SCNs) are able to induce means from a finite fragment of a context-sensitive language for processing strings outside the training set. The dynamical behavior of the SCN is qualitatively distinct from that observed in LSTM networks. Differences in performance and dynamics are discussed.
Mikael Bodén, Janet Wiles
IEEE Trans. Neural Networks1
2000 Evolving context-free language predictors
Mikael Bodén, Henrik Jacobsson, Tom Ziemke
GECCO1
2000 Semantic systematicity and context in connectionist networks
abstract
A bstract . Fodor and Pylyshyn argued that connectionist models could not be used to exhibit and explain a phenomenon that they termed systematicity, and which they explained by possession of composition syntax and semantics for mental representations and structure sensitivity of mental processes. This inability of connectionist models, they argued, was particularly serious since it meant that these models could not be used as alternative models to classical symbolic models to explain cognition. In this paper, a connectionist model is used to identify some properties which collectively show that connectionist networks supply means for accomplishing a stronger version ofsystematicity than Fodor and Pylyshyn opted for. It is argued that 'context-dependent systematicity' is achievable within a connectionist framework. The arguments put forward rest on a particular formulation of content and context of connectionist representation, firmly and technically based on connectionist primitives in a learning environment. The perspective is motivated by the fundamental differences between the connectionist and classical architectures, in terms of prerequisites, lower-level functionality and inherent constraints. The claim is supported by a set of experiments using a connectionist architecture that demonstrates both an ability of enforcing, what Fodor and Pylyshyn term systematic and nonsystematic processing using a single mechanism, and how novel items can be handled without prior classification. The claim relies on extended learning feedback which enforces representational context dependence.
Mikael Bodén, Lars Niklasson
Connect. Sci.1
2000 Context-free and context-sensitive dynamics in recurrent neural networks
abstract
Continuous-valued recurrent neural networks can learn mechanisms for processing context-free languages. The dynamics of such networks is usually based on damped oscillation around fixed points in state space and requires that the dynamical components are arranged in certain ways. It is shown that qualitatively similar dynamics with similar constraints hold for anbncn , a context-sensitive language. The additional difficulty with anbncn , compared with the context-free language anbn , consists of 'counting up' and 'counting down' letters simultaneously. The network solution is to oscillate in two principal dimensions, one for counting up and one for counting down. This study focuses on the dynamics employed by the sequential cascaded network, in contrast to the simple recurrent network, and the use of backpropagation through time. Found solutions generalize well beyond training data, however, learning is not reliable. The contribution of this study lies in demonstrating how the dynamics in recurrent neural networks that process context-free languages can also be employed in processing some context-sensitive languages (traditionally thought of as requiring additional computation resources). This continuity of mechanism between language classes contributes to our understanding of neural networks in modelling language learning and processing.
Mikael Bodén, Janet Wiles
Connect. Sci.1
1996 A Connectionist Variation of Inheritance
Mikael Bodén
ICANN1