Søren Brunak

dblp:83/546 · DBLP profile ↗
← Back
40ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0003-0316-5866ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 2 since 2021Artificial intelligence and machine learning · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
22 papers
Bioinformatics and computational biology · 92% Environmental and earth informatics · 8%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › knowledge representation in biology
biomedical ontology
0.812024
Lifestyle factors in the biomedical literature: an ontology and comprehensive resources for named entity recognition · Bioinform. 2024
Bioinformatics and computational biology
biomedical text mining
0.812024
Lifestyle factors in the biomedical literature: an ontology and comprehensive resources for named entity recognition · Bioinform. 2024
Bioinformatics and computational biology › biomedical text mining
named entity recognition
0.812024
Lifestyle factors in the biomedical literature: an ontology and comprehensive resources for named entity recognition · Bioinform. 2024
Environmental and earth informatics
adverse outcome pathway
0.412019
sAOP: linking chemical stressors to adverse outcomes pathway networks · Bioinform. 2019
Bioinformatics and computational biology › drug discovery › computational drug discovery
protein-ligand interaction prediction
0.412019
A generic deep convolutional neural network framework for prediction of receptor-ligand interactions - NetPhosPan: application to kinase phosphorylation prediction · Bioinform. 2019
Bioinformatics and computational biology › genomics
toxicogenomics
0.412019
sAOP: linking chemical stressors to adverse outcomes pathway networks · Bioinform. 2019
Bioinformatics and computational biology
comparative genomics
0.112011
The strength of intron donor splice sites in human genes displays a bell-shaped pattern · Bioinform. 2011
Bioinformatics and computational biology › RNA biology › RNA analysis
RNA bioinformatics
0.112011
The strength of intron donor splice sites in human genes displays a bell-shaped pattern · Bioinform. 2011
Bioinformatics and computational biology
immunoinformatics
0.122007
Modeling the adaptive immune system: predictions and simulations · Bioinform. 2007
Improved prediction of MHC class I and class II epitopes using a novel Gibbs sampling approach · Bioinform. 2004
Bioinformatics and computational biology › drug discovery
compound prioritization
0.112019
sAOP: linking chemical stressors to adverse outcomes pathway networks · Bioinform. 2019
Bioinformatics and computational biology
protein structure prediction
0.142000
Assessing the accuracy of prediction algorithms for classification: an overview · Bioinform. 2000
Matching Protein b-Sheet Partners by Feedforward and Recurrent Neural Networks · ISMB 2000
Exploiting the past and the future in protein secondary structure prediction · Bioinform. 1999
Bioinformatics and computational biology › immunoinformatics
epitope prediction
0.112007
Modeling the adaptive immune system: predictions and simulations · Bioinform. 2007
Bioinformatics and computational biology › immunoinformatics
immune system modeling
0.112007
Modeling the adaptive immune system: predictions and simulations · Bioinform. 2007
Bioinformatics and computational biology › gene expression analysis › periodic gene detection
cell-cycle gene identification
0.112005
Comparison of computational methods for the identification of cell cycle-regulated genes · Bioinform. 2005
Bioinformatics and computational biology › molecular informatics
cheminformatics
0.112005
Prediction methods and databases within chemoinformatics: emphasis on drugs and drug candidates · Bioinform. 2005
Bioinformatics and computational biology
gene expression analysis
0.112005
Comparison of computational methods for the identification of cell cycle-regulated genes · Bioinform. 2005
Bioinformatics and computational biology › protein structure prediction
secondary structure prediction
0.122000
Assessing the accuracy of prediction algorithms for classification: an overview · Bioinform. 2000
Exploiting the past and the future in protein secondary structure prediction · Bioinform. 1999
Bioinformatics and computational biology › sequence analysis › motif discovery
sequence motif discovery
0.012004
Improved prediction of MHC class I and class II epitopes using a novel Gibbs sampling approach · Bioinform. 2004
Bioinformatics and computational biology › immunoinformatics
TCR-epitope binding prediction
0.012004
Improved prediction of MHC class I and class II epitopes using a novel Gibbs sampling approach · Bioinform. 2004
Bioinformatics and computational biology
protein function prediction
0.012003
Prediction of human protein function according to Gene Ontology categories · Bioinform. 2003
Bioinformatics and computational biology
sequence analysis
0.021999
MatrixPlot: visualizing sequence constraints · Bioinform. 1999
Characterization of Prokaryotic and Eukaryotic Promoters Using Hidden Markov Models · ISMB 1996
Bioinformatics and computational biology › structural biology
DNA structure
0.012001
Flexibility of the genetic code with respect to DNA structure · Bioinform. 2001
Bioinformatics and computational biology › molecular evolution
genetic code evolution
0.012001
Flexibility of the genetic code with respect to DNA structure · Bioinform. 2001
Bioinformatics and computational biology › molecular evolution › evolutionary bioinformatics › evolutionary genomics
genome evolution
0.012001
Flexibility of the genetic code with respect to DNA structure · Bioinform. 2001
Bioinformatics and computational biology › genomics
repetitive DNA analysis
0.011999
Structural basis for triplet repeat disorders: a computational analysis · Bioinform. 1999
Bioinformatics and computational biology › immunoinformatics
MHC binding prediction
0.012007
Modeling the adaptive immune system: predictions and simulations · Bioinform. 2007
Bioinformatics and computational biology › structural bioinformatics › nucleic acid structure analysis
RNA structural alignment
0.011997
Displaying the information contents of structural RNA alignments: the structure logos · Comput. Appl. Biosci. 1997
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA structure analysis
0.011997
Displaying the information contents of structural RNA alignments: the structure logos · Comput. Appl. Biosci. 1997
Bioinformatics and computational biology › biological data visualization
sequence logo
0.011997
Displaying the information contents of structural RNA alignments: the structure logos · Comput. Appl. Biosci. 1997
Bioinformatics and computational biology
drug discovery
0.012005
Prediction methods and databases within chemoinformatics: emphasis on drugs and drug candidates · Bioinform. 2005

Methods — techniques the papers use, named apart from their topics

transformer-based NER · 0.8dictionary-based NER · 0.8toxcast data integration · 0.4network mapping · 0.4deep convolutional neural network · 0.4statistical analysis of splice site strengths · 0.1statistical benchmarking · 0.1permutation testing · 0.1gibbs sampling · 0.0ROC analysis · 0.0multiple sequence alignment · 0.0ensemble methods · 0.0bidirectional recurrent neural network · 0.0
YearPublicationVenuePosition
2024 Lifestyle factors in the biomedical literature: an ontology and comprehensive resources for named entity recognition
abstract
MOTIVATION: Despite lifestyle factors (LSFs) being increasingly acknowledged in shaping individual health trajectories, particularly in chronic diseases, they have still not been systematically described in the biomedical literature. This is in part because no named entity recognition (NER) system exists, which can comprehensively detect all types of LSFs in text. The task is challenging due to their inherent diversity, lack of a comprehensive LSF classification for dictionary-based NER, and lack of a corpus for deep learning-based NER. RESULTS: We present a novel lifestyle factor ontology (LSFO), which we used to develop a dictionary-based system for recognition and normalization of LSFs. Additionally, we introduce a manually annotated corpus for LSFs (LSF200) suitable for training and evaluation of NER systems, and use it to train a transformer-based system. Evaluating the performance of both NER systems on the corpus revealed an F-score of 64% for the dictionary-based system and 76% for the transformer-based system. Large-scale application of these systems on PubMed abstracts and PMC Open Access articles identified over 300 million mentions of LSF in the biomedical literature. AVAILABILITY AND IMPLEMENTATION: LSFO, the annotated LSF200 corpus, and the detected LSFs in PubMed and PMC-OA articles using both NER systems, are available under open licenses via the following GitHub repository: https://github.com/EsmaeilNourani/LSFO-expansion. This repository contains links to two associated GitHub repositories and a Zenodo project related to the study. LSFO is also available at BioPortal: https://bioportal.bioontology.org/ontologies/LSFO.
Esmaeil Nourani, Mikaela Koutrouli, Yijia Xie, Danai Vagiaki, Sampo Pyysalo, Katerina C. Nastou, Søren Brunak, Lars Juhl Jensen
Bioinform.7
2023 BALDR: A Web-based platform for informed comparison and prioritization of biomarker candidates for type 2 diabetes mellitus
abstract
Novel biomarkers are key to addressing the ongoing pandemic of type 2 diabetes mellitus. While new technologies have improved the potential of identifying such biomarkers, at the same time there is an increasing need for informed prioritization to ensure efficient downstream verification. We have built BALDR, an automated pipeline for biomarker comparison and prioritization in the context of diabetes. BALDR includes protein, gene, and disease data from major public repositories, text-mining data, and human and mouse experimental data from the IMI2 RHAPSODY consortium. These data are provided as easy-to-read figures and tables enabling direct comparison of up to 20 biomarker candidates for diabetes through the public website https://baldr.cpr.ku.dk.
Agnete T. Lundgaard, Frédéric Burdet, Troels Siggaard, David Westergaard, Danai Vagiaki, Lisa Cantwell, Timo Röder, Dorte Vistisen, Thomas Sparsø, Giuseppe N. Giordano, Mark Ibberson, Karina Banasik, Søren Brunak
PLoS Comput. Biol.13
2020 Alcoholic liver disease: A registry view on comorbidities and disease prediction
abstract
Alcoholic-related liver disease (ALD) is the cause of more than half of all liver-related deaths. Sustained excess drinking causes fatty liver and alcohol-related steatohepatitis, which may progress to alcoholic liver fibrosis (ALF) and eventually to alcohol-related liver cirrhosis (ALC). Unfortunately, it is difficult to identify patients with early-stage ALD, as these are largely asymptomatic. Consequently, the majority of ALD patients are only diagnosed by the time ALD has reached decompensated cirrhosis, a symptomatic phase marked by the development of complications as bleeding and ascites. The main goal of this study is to discover relevant upstream diagnoses helping to understand the development of ALD, and to highlight meaningful downstream diagnoses that represent its progression to liver failure. Here, we use data from the Danish health registries covering the entire population of Denmark during nineteen years (1996-2014), to examine if it is possible to identify patients likely to develop ALF or ALC based on their past medical history. To this end, we explore a knowledge discovery approach by using high-dimensional statistical and machine learning techniques to extract and analyze data from the Danish National Patient Registry. Consistent with the late diagnoses of ALD, we find that ALC is the most common form of ALD in the registry data and that ALC patients have a strong over-representation of diagnoses associated with liver dysfunction. By contrast, we identify a small number of patients diagnosed with ALF who appear to be much less sick than those with ALC. We perform a matched case-control study using the group of patients with ALC as cases and their matched patients with non-ALD as controls. Machine learning models (SVM, RF, LightGBM and NaiveBayes) trained and tested on the set of ALC patients achieve a high performance for data classification (AUC = 0.89). When testing the same trained models on the small set of ALF patients, their performance unsurprisingly drops a lot (AUC = 0.67 for NaiveBayes). The statistical and machine learning results underscore small groups of upstream and downstream comorbidities that accurately detect ALC patients and show promise in prediction of ALF. Some of these groups are conditions either caused by alcohol or caused by malnutrition associated with alcohol-overuse. Others are comorbidities either related to trauma and life-style or to complications to cirrhosis, such as oesophageal varices. Our findings highlight the potential of this approach to uncover knowledge in registry data related to ALD.
Dhouha Grissa, Ditlev Nytoft Rasmussen, Aleksander Krag, Søren Brunak, Lars Juhl Jensen
PLoS Comput. Biol.4
2019 sAOP: linking chemical stressors to adverse outcomes pathway networks
abstract
MOTIVATION: Adverse outcome pathway (AOP) is a toxicological concept proposed to provide a mechanistic representation of biological perturbation over different layers of biological organization. Although AOPs are by definition chemical-agnostic, many chemical stressors can putatively interfere with one or several AOPs and such information would be relevant for regulatory decision-making. RESULTS: With the recent development of AOPs networks aiming to facilitate the identification of interactions among AOPs, we developed a stressor-AOP network (sAOP). Using the 'cytotoxitiy burst' (CTB) approach, we mapped bioactive compounds from the ToxCast data to a list of AOPs reported in AOP-Wiki database. With this analysis, a variety of relevant connections between chemicals and AOP components can be identified suggesting multiple effects not observed in the simplified 'one-biological perturbation to one-adverse outcome' model. The results may assist in the prioritization of chemicals to assess risk-based evaluations in the context of human health. AVAILABILITY AND IMPLEMENTATION: sAOP is available at http://saop.cpr.ku.dk. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alejandro Aguayo-Orozco, Karine Audouze, Troels Siggaard, Robert Barouki, Søren Brunak, Olivier Taboureau
Bioinform.5
2019 A generic deep convolutional neural network framework for prediction of receptor-ligand interactions - NetPhosPan: application to kinase phosphorylation prediction
abstract
MOTIVATION: Understanding the specificity of protein receptor-ligand interactions is pivotal for our comprehension of biological mechanisms and systems. Receptor protein families often have a certain level of sequence diversity that converges into fewer conserved protein structures, allowing the exertion of well-defined functions. T and B cell receptors of the immune system and protein kinases that control the dynamic behaviour and decision processes in eukaryotic cells by catalysing phosphorylation represent prime examples. Driven by the large sequence diversity, the receptors within such protein families are often found to share specificities although divergent at the sequence level. This observation has led to the notion that prediction models of such systems are most effectively handled in a receptor-specific manner. RESULTS: We show that this approach in many cases is suboptimal, and describe an alternative improved framework for generating models with pan-receptor-predictive power for receptor protein families. The framework is based on deep artificial neural networks and integrates information from individual receptors into a single pan-receptor model, leveraging information across multiple receptor-specific datasets allowing predictions of the receptor specificity for all members of a given protein family including those described by limited or no ligand data. The approach was applied to the protein kinase superfamily, leading to the method NetPhosPan. The method was extensively validated and benchmarked against state-of-the-art prediction methods and was found to have unprecedented performance in particularly for kinase domains characterized by limited or no experimental data. AVAILABILITY AND IMPLEMENTATION: The method is freely available to non-commercial users and can be downloaded at http://www.cbs.dtu.dk/services/NetPhospan-1.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Emilio Fenoy, José M. G. Izarzugaza, Vanessa Isabell Jurtz, Søren Brunak, Morten Nielsen 0001
Bioinform.4
2018 Benchmarking the HLA typing performance of Polysolver and Optitype in 50 Danish parental trios
abstract
BACKGROUND: The adaptive immune response intrinsically depends on hypervariable human leukocyte antigen (HLA) genes. Concomitantly, correct HLA phenotyping is crucial for successful donor-patient matching in organ transplantation. The cost and technical limitations of current laboratory techniques, together with advances in next-generation sequencing (NGS) methodologies, have increased the need for precise computational typing methods. RESULTS: We tested two widespread HLA typing methods using high quality full genome sequencing data from 150 individuals in 50 family trios from the Genome Denmark project. First, we computed descendant accuracies assessing the agreement in the inheritance of alleles from parents to offspring. Second, we compared the locus-specific homozygosity rates as well as the allele frequencies; and we compared those to the observed values in related populations. We provide guidelines for testing the accuracy of HLA typing methods by comparing family information, which is independent of the availability of curated alleles. CONCLUSIONS: Although current computational methods for HLA typing generally provide satisfactory results, our benchmark - using data with ultra-high sequencing depth - demonstrates the incompleteness of current reference databases, and highlights the importance of providing genomic databases addressing current sequencing standards, a problem yet to be resolved before benefiting fully from personalised medicine approaches HLA phenotyping is essential.
Maria Luisa Matey-Hernandez, Lasse Maretty, Jacob Malte Jensen, Bent Petersen, Jonas Andreas Sibbesen, Siyang Liu 0007, Palle Villesen, Laurits Skov, Kirstine Belling, Christian Theil Have, José M. G. Izarzugaza, Marie Grosjean, Jette Bork-Jensen, Jakob Grove, Thomas D. Als, Shujia Huang, Yuqi Chang, Weijian Ye, Junhua Rao, Xiaosen Guo, Jihua Sun, Hongzhi Cao, John van Beusekom, Thomas Espeseth, Esben N. Flindt, Rune Møllegaard Friborg, Anders E. Halager, Stephanie Le Hellard, Christina M. Hultman, Francesco Lescai, Shengting Li, Ole Lund, Peter Løngreen, Thomas Mailund, Ole Mors, Christian N. S. Pedersen, Thomas Sicheritz-Pontén, Patrick F. Sullivan, Ali Syed, David Westergaard, Rachita Yadav, Torben Hansen, Anders Krogh, Lars Bolund, Thorkild I. A. Sørensen, Oluf Pedersen, Ramneek Gupta, Simon Rasmussen, Søren Besenbacher, Anders D. Børglum, Jun Wang 0004, Hans Eiberg, Karsten Kristiansen, Mikkel H. Schierup, Søren Brunak
BMC Bioinform.59
2018 A comprehensive and quantitative comparison of text-mining in 15 million full-text articles versus their corresponding abstracts
abstract
Across academia and industry, text mining has become a popular strategy for keeping up with the rapid growth of the scientific literature. Text mining of the scientific literature has mostly been carried out on collections of abstracts, due to their availability. Here we present an analysis of 15 million English scientific full-text articles published during the period 1823-2016. We describe the development in article length and publication sub-topics during these nearly 250 years. We showcase the potential of text mining by extracting published protein-protein, disease-gene, and protein subcellular associations using a named entity recognition system, and quantitatively report on their accuracy using gold standard benchmark data sets. We subsequently compare the findings to corresponding results obtained on 16.5 million abstracts included in MEDLINE and show that text mining of full-text articles consistently outperforms using abstracts only.
David Westergaard, Hans Henrik Stærfeldt, Christian Tønsberg, Lars Juhl Jensen, Søren Brunak
PLoS Comput. Biol.5
2015 Finding Cervical Cancer Symptoms in Swedish Clinical Text using a Machine Learning Approach and NegEx
Rebecka Weegar, Maria Kvist, Karin Sundström, Søren Brunak, Hercules Dalianis
AMIA4
2014 Facilitating the use of large-scale biological data and tools in the era of translational bioinformatics
abstract
As both the amount of generated biological data and the processing compute power increase, computational experimentation is no longer the exclusivity of bioinformaticians, but it is moving across all biomedical domains. For bioinformatics to realize its translational potential, domain experts need access to user-friendly solutions to navigate, integrate and extract information out of biological databases, as well as to combine tools and data resources in bioinformatics workflows. In this review, we present services that assist biomedical scientists in incorporating bioinformatics tools into their research. We review recent applications of Cytoscape, BioGPS and DAVID for data visualization, integration and functional enrichment. Moreover, we illustrate the use of Taverna, Kepler, GenePattern, and Galaxy as open-access workbenches for bioinformatics workflows. Finally, we mention services that facilitate the integration of biomedical ontologies and bioinformatics tools in computational workflows.
Irene Kouskoumvekaki, Nour Shublaq, Søren Brunak
Briefings Bioinform.3
2014 Compass: A hybrid method for clinical and biobank data mining
Konrad Krysiak-Baltyn, Thomas Nordahl Petersen, Karine Audouze, Niels Jørgensen, Lars Ängquist, Søren Brunak
J. Biomed. Informatics6
2013 Dictionary construction and identification of possible adverse drug events in Danish clinical narrative text
abstract
OBJECTIVE: Drugs have tremendous potential to cure and relieve disease, but the risk of unintended effects is always present. Healthcare providers increasingly record data in electronic patient records (EPRs), in which we aim to identify possible adverse events (AEs) and, specifically, possible adverse drug events (ADEs). MATERIALS AND METHODS: Based on the undesirable effects section from the summary of product characteristics (SPC) of 7446 drugs, we have built a Danish ADE dictionary. Starting from this dictionary we have developed a pipeline for identifying possible ADEs in unstructured clinical narrative text. We use a named entity recognition (NER) tagger to identify dictionary matches in the text and post-coordination rules to construct ADE compound terms. Finally, we apply post-processing rules and filters to handle, for example, negations and sentences about subjects other than the patient. Moreover, this method allows synonyms to be identified and anatomical location descriptions can be merged to allow appropriate grouping of effects in the same location. RESULTS: The method identified 1 970 731 (35 477 unique) possible ADEs in a large corpus of 6011 psychiatric hospital patient records. Validation was performed through manual inspection of possible ADEs, resulting in precision of 89% and recall of 75%. DISCUSSION: The presented dictionary-building method could be used to construct other ADE dictionaries. The complication of compound words in Germanic languages was addressed. Additionally, the synonym and anatomical location collapse improve the method. CONCLUSIONS: The developed dictionary and method can be used to identify possible ADEs in Danish clinical narratives.
Robert Eriksson, Peter Bjødstrup Jensen, Sune Pletscher-Frankild, Lars Juhl Jensen, Søren Brunak
J. Am. Medical Informatics Assoc.5
2011 The strength of intron donor splice sites in human genes displays a bell-shaped pattern
abstract
MOTIVATION: The gene concept has recently changed from the classical one protein notion into a much more diverse picture, where overlapping or fused transcripts, alternative transcription initiation, and genes within genes, add to the complexity generated by alternative splicing. Increased understanding of the mechanisms controlling pre-mRNA splicing is thus important for a wide range of aspects relating to gene expression. RESULTS: We have discovered a convex gene delineating pattern in the strength of 5' intron splice sites. When comparing the strengths of > 18,000 intron containing Human genes, we found that when analysing them separately according to the number of introns they contain, initial splice sites were always stronger on average than subsequent ones, and that a similar reversed trend exist towards the terminal gene part. The convex pattern is strongest for genes with up to 10 introns. Interestingly, when analysing the intron containing gene pool from mouse consisting of >15,000 genes, we found the convex pattern to be conserved despite > 75 million years of evolutionary divergence between the two organisms. We also analysed an interesting, novel class of chimeric genes which during spliceosome assembly are fused and in tandem are transcribed and spliced into a single mature mRNA sequence. In their splice site patterns, these genes individually seem to deviate from the convex pattern, offering a possible rationale behind their fusion into a single transcript.
Rasmus Wernersson, Søren Brunak
Bioinform.3
2011 Consistent metagenes from cancer expression profiles yield agent specific predictors of chemotherapy response
abstract
BACKGROUND: Genome scale expression profiling of human tumor samples is likely to yield improved cancer treatment decisions. However, identification of clinically predictive or prognostic classifiers can be challenging when a large number of genes are measured in a small number of tumors. RESULTS: We describe an unsupervised method to extract robust, consistent metagenes from multiple analogous data sets. We applied this method to expression profiles from five "double negative breast cancer" (DNBC) (not expressing ESR1 or HER2) cohorts and derived four metagenes. We assessed these metagenes in four similar but independent cohorts and found strong associations between three of the metagenes and agent-specific response to neoadjuvant therapy. Furthermore, we applied the method to ovarian and early stage lung cancer, two tumor types that lack reliable predictors of outcome, and found that the metagenes yield predictors of survival for both. CONCLUSIONS: These results suggest that the use of multiple data sets to derive potential biomarkers can filter out data set-specific noise and can increase the efficiency in identifying clinically accurate biomarkers.
Aron C. Eklund, Nicolai J. Birkbak, Christine Desmedt, Benjamin Haibe-Kains, Christos P. Sotiriou, W. Fraser Symmans, Lajos Pusztai, Søren Brunak, Andrea L. Richardson, Zoltan Szallasi
BMC Bioinform.9
2011 Using Electronic Patient Records to Discover Disease Correlations and Stratify Patient Cohorts
abstract
Electronic patient records remain a rather unexplored, but potentially rich data source for discovering correlations between diseases. We describe a general approach for gathering phenotypic descriptions of patients from medical records in a systematic and non-cohort dependent manner. By extracting phenotype information from the free-text in such records we demonstrate that we can extend the information contained in the structured record data, and use it for producing fine-grained patient stratification and disease co-occurrence statistics. The approach uses a dictionary based on the International Classification of Disease ontology and is therefore in principle language independent. As a use case we show how records from a Danish psychiatric hospital lead to the identification of disease correlations, which subsequently can be mapped to systems biology frameworks.
Francisco S. Roque, Peter Bjødstrup Jensen, Henriette Schmock, Marlene Dalgaard, Massimo Andreatta, Thomas Folkmann Hansen, Karen Søeby, Søren Bredkjær, Anders Juul, Thomas Werge, Lars Juhl Jensen, Søren Brunak
PLoS Comput. Biol.12
2010 Deciphering Diseases and Biological Targets for Environmental Chemicals using Toxicogenomics Networks
abstract
Exposure to environmental chemicals and drugs may have a negative effect on human health. A better understanding of the molecular mechanism of such compounds is needed to determine the risk. We present a high confidence human protein-protein association network built upon the integration of chemical toxicology and systems biology. This computational systems chemical biology model reveals uncharacterized connections between compounds and diseases, thus predicting which compounds may be risk factors for human health. Additionally, the network can be used to identify unexpected potential associations between chemicals and proteins. Examples are shown for chemicals associated with breast cancer, lung cancer and necrosis, and potential protein targets for di-ethylhexyl-phthalate, 2,3,7,8-tetrachlorodibenzo-p-dioxin, pirinixic acid and permethrine. The chemical-protein associations are supported through recent published studies, which illustrate the power of our approach that integrates toxicogenomics data with other data types.
Karine Audouze, Agnieszka Sierakowska Juncker, Francisco S. Roque, Konrad Krysiak-Baltyn, Nils Weinhold, Olivier Taboureau, Thomas Skøt Jensen, Søren Brunak
PLoS Comput. Biol.8
2009 ImmunoGrid, an integrative environment for large-scale simulation of the immune system for vaccine discovery, design and optimization
abstract
Vaccine research is a combinatorial science requiring computational analysis of vaccine components, formulations and optimization. We have developed a framework that combines computational tools for the study of immune function and vaccine development. This framework, named ImmunoGrid combines conceptual models of the immune system, models of antigen processing and presentation, system-level models of the immune system, Grid computing, and database technology to facilitate discovery, formulation and optimization of vaccines. ImmunoGrid modules share common conceptual models and ontologies. The ImmunoGrid portal offers access to educational simulators where previously defined cases can be displayed, and to research simulators that allow the development of new, or tuning of existing, computational models. The portal is accessible at .
Francesco Pappalardo 0001, Mark D. Halling-Brown, Nicolas Rapin, Ping Zhang 0008, Davide Alemani, Andrew P. J. Emerson, Paola Paci, Patrice Duroux, Marzio Pennisi, Arianna Palladini, Olivo Miotto, Daniel Churchill, Elda Rossi, Adrian J. Shepherd, David S. Moss, Filippo Castiglione, Massimo Bernaschi, Marie-Paule Lefranc, Søren Brunak, Santo Motta, Pierluigi Lollini, Kaye E. Basford, Vladimir Brusic
Briefings Bioinform.19
2007 Modeling the adaptive immune system: predictions and simulations
abstract
MOTIVATION: Immunological bioinformatics methods are applicable to a broad range of scientific areas. The specifics of how and where they might be implemented have recently been reviewed in the literature. However, the background and concerns for selecting between the different available methods have so far not been adequately covered. SUMMARY: Before using predictions systems, it is necessary to not only understand how the methods are constructed but also their strength and limitations. The prediction systems in humoral epitope discovery are still in their infancy, but have reached a reasonable level of predictive strength. In cellular immunology, MHC class I binding predictions are now very strong and cover most of the known HLA specificities. These systems work well for epitope discovery, and predictions of the MHC class I pathway have been further improved by integration with state-of-the-art prediction tools for proteasomal cleavage and TAP binding. By comparison, class II MHC binding predictions have not developed to a comparable accuracy level, but new tools have emerged that deliver significantly improved predictions not only in terms of accuracy, but also in MHC specificity coverage. Simulation systems and mathematical modeling are also now beginning to reach a level where these methods will be able to answer more complex immunological questions.
Claus Lundegaard, Ole Lund, Can Kesmir, Søren Brunak, Morten Nielsen 0001
Bioinform.4
2006 ISMB 2006
Goran Neshich, Philip E. Bourne, Søren Brunak
PLoS Comput. Biol.3
2005 Prediction methods and databases within chemoinformatics: emphasis on drugs and drug candidates
abstract
MOTIVATION: To gather information about available databases and chemoinformatics methods for prediction of properties relevant to the drug discovery and optimization process. RESULTS: We present an overview of the most important databases with 2-dimensional and 3-dimensional structural information about drugs and drug candidates, and of databases with relevant properties. Access to experimental data and numerical methods for selecting and utilizing these data is crucial for developing accurate predictive in silico models. Many interesting predictive methods for classifying the suitability of chemical compounds as potential drugs, as well as for predicting their physico-chemical and ADMET properties have been proposed in recent years. These methods are discussed, and some possible future directions in this rapidly developing field are described.
Svava Ósk Jónsdóttir, Flemming Steen Jørgensen, Søren Brunak
Bioinform.3
2005 Comparison of computational methods for the identification of cell cycle-regulated genes
abstract
MOTIVATION: DNA microarrays have been used extensively to study the cell cycle transcription programme in a number of model organisms. The Saccharomyces cerevisiae data in particular have been subjected to a wide range of bioinformatics analysis methods, aimed at identifying the correct and complete set of periodically expressed genes. RESULTS: Here, we provide the first thorough benchmark of such methods, surprisingly revealing that most new and more mathematically advanced methods actually perform worse than the analysis published with the original microarray data sets. We show that this loss of accuracy specifically affects methods that only model the shape of the expression profile without taking into account the magnitude of regulation. We present a simple permutation-based method that performs better than most existing methods.
Ulrik de Lichtenberg, Lars Juhl Jensen, Anders Fausbøll, Thomas Skøt Jensen, Peer Bork, Søren Brunak
Bioinform.6
2005 Prediction of twin-arginine signal peptides
abstract
BACKGROUND: Proteins carrying twin-arginine (Tat) signal peptides are exported into the periplasmic compartment or extracellular environment independently of the classical Sec-dependent translocation pathway. To complement other methods for classical signal peptide prediction we here present a publicly available method, TatP, for prediction of bacterial Tat signal peptides. RESULTS: We have retrieved sequence data for Tat substrates in order to train a computational method for discrimination of Sec and Tat signal peptides. The TatP method is able to positively classify 91% of 35 known Tat signal peptides and 84% of the annotated cleavage sites of these Tat signal peptides were correctly predicted. This method generates far less false positive predictions on various datasets than using simple pattern matching. Moreover, on the same datasets TatP generates less false positive predictions than a complementary rule based prediction method. CONCLUSION: The method developed here is able to discriminate Tat signal peptides from cytoplasmic proteins carrying a similar motif, as well as from Sec signal peptides, with high accuracy. The method allows filtering of input sequences based on Perl syntax regular expressions, whereas hydrophobicity discrimination of Tat- and Sec-signal peptides is carried out by an artificial neural network. A potential cleavage site of the predicted Tat signal peptide is also reported. The TatP prediction server is available as a public web server at http://www.cbs.dtu.dk/services/TatP/.
Jannick Dyrløv Bendtsen, Henrik Nielsen, David Widdick, Tracy Palmer, Søren Brunak
BMC Bioinform.5
2004 Improved prediction of MHC class I and class II epitopes using a novel Gibbs sampling approach
abstract
MOTIVATION: Prediction of which peptides will bind a specific major histocompatibility complex (MHC) constitutes an important step in identifying potential T-cell epitopes suitable as vaccine candidates. MHC class II binding peptides have a broad length distribution complicating such predictions. Thus, identifying the correct alignment is a crucial part of identifying the core of an MHC class II binding motif. In this context, we wish to describe a novel Gibbs motif sampler method ideally suited for recognizing such weak sequence motifs. The method is based on the Gibbs sampling method, and it incorporates novel features optimized for the task of recognizing the binding motif of MHC classes I and II. The method locates the binding motif in a set of sequences and characterizes the motif in terms of a weight-matrix. Subsequently, the weight-matrix can be applied to identifying effectively potential MHC binding peptides and to guiding the process of rational vaccine design. RESULTS: We apply the motif sampler method to the complex problem of MHC class II binding. The input to the method is amino acid peptide sequences extracted from the public databases of SYFPEITHI and MHCPEP and known to bind to the MHC class II complex HLA-DR4(B1*0401). Prior identification of information-rich (anchor) positions in the binding motif is shown to improve the predictive performance of the Gibbs sampler. Similarly, a consensus solution obtained from an ensemble average over suboptimal solutions is shown to outperform the use of a single optimal solution. In a large-scale benchmark calculation, the performance is quantified using relative operating characteristics curve (ROC) plots and we make a detailed comparison of the performance with that of both the TEPITOPE method and a weight-matrix derived using the conventional alignment algorithm of ClustalW. The calculation demonstrates that the predictive performance of the Gibbs sampler is higher than that of ClustalW and in most cases also higher than that of the TEPITOPE method.
Morten Nielsen 0001, Claus Lundegaard, Peder Worning, Christina Sylvester-Hvid, Kasper Lamberth, Søren Buus, Søren Brunak, Ole Lund
Bioinform.7
2004 Coronavirus 3CLpro proteinase cleavage sites: Possible relevance to SARS virus pathology
abstract
BACKGROUND: Despite the passing of more than a year since the first outbreak of Severe Acute Respiratory Syndrome (SARS), efficient counter-measures are still few and many believe that reappearance of SARS, or a similar disease caused by a coronavirus, is not unlikely. For other virus families like the picornaviruses it is known that pathology is related to proteolytic cleavage of host proteins by viral proteinases. Furthermore, several studies indicate that virus proliferation can be arrested using specific proteinase inhibitors supporting the belief that proteinases are indeed important during infection. Prompted by this, we set out to analyse and predict cleavage by the coronavirus main proteinase using computational methods. RESULTS: We retrieved sequence data on seven fully sequenced coronaviruses and identified the main 3CL proteinase cleavage sites in polyproteins using alignments. A neural network was trained to recognise the cleavage sites in the genomes obtaining a sensitivity of 87.0% and a specificity of 99.0%. Several proteins known to be cleaved by other viruses were submitted to prediction as well as proteins suspected relevant in coronavirus pathology. Cleavage sites were predicted in proteins such as the cystic fibrosis transmembrane conductance regulator (CFTR), transcription factors CREB-RP and OCT-1, and components of the ubiquitin pathway. CONCLUSIONS: Our prediction method NetCorona predicts coronavirus cleavage sites with high specificity and several potential cleavage candidates were identified which might be important to elucidate coronavirus pathology. Furthermore, the method might assist in design of proteinase inhibitors for treatment of SARS and possible future diseases caused by coronaviruses. It is made available for public use at our website: http://www.cbs.dtu.dk/services/NetCorona/.
Lars Kiemer, Ole Lund, Søren Brunak, Nikolaj Blom
BMC Bioinform.3
2003 Prediction of human protein function according to Gene Ontology categories
abstract
MOTIVATION: The human genome project has led to the discovery of many human protein coding genes which were previously unknown. As a large fraction of these are functionally uncharacterized, it is of interest to develop methods for predicting their molecular function from sequence. RESULTS: We have developed a method for prediction of protein function for a subset of classes from the Gene Ontology classification scheme. This subset includes several pharmaceutically interesting categories-transcription factors, receptors, ion channels, stress and immune response proteins, hormones and growth factors can all be predicted. Although the method relies on protein sequences as the sole input, it does not rely on sequence similarity, but instead on sequence derived protein features such as predicted post translational modifications (PTMs), protein sorting signals and physical/chemical properties calculated from the amino acid composition. This allows for prediction of the function for orphan proteins where no homologs can be found. Using this method we propose two novel receptors in the human genome, and further demonstrate chromosomal clustering of related proteins.
Lars Juhl Jensen, Ramneek Gupta, Hans Henrik Stærfeldt, Søren Brunak
Bioinform.4
2003 Selecting Informative Data for Developing Peptide-MHC Binding Predictors Using a Query by Committee Approach
abstract
Strategies for selecting informative data points for training prediction algorithms are important, particularly when data points are difficult and costly to obtain. A Query by Committee (QBC) training strategy for selecting new data points uses the disagreement between a committee of different algorithms to suggest new data points, which most rationally complement existing data, that is, they are the most informative data points. In order to evaluate this QBC approach on a real-world problem, we compared strategies for selecting new data points. We trained neural network algorithms to obtain methods to predict the binding affinity of peptides binding to the MHC class I molecule, HLA-A2. We show that the QBC strategy leads to a higher performance than a baseline strategy where new data points are selected at random from a pool of available data. Most peptides bind HLA-A2 with a low affinity, and as expected using a strategy of selecting peptides that are predicted to have high binding affinities also lead to more accurate predictors than the base line strategy. The QBC value is shown to correlate with the measured binding affinity. This demonstrates that the different predictors can easily learn if a peptide will fail to bind, but often conflict in predicting if a peptide binds. Using a carefully constructed computational setup, we demonstrate that selecting peptides with a high QBC performs better than low QBC peptides independently from binding affinity. When predictors are trained on a very limited set of data they cannot be expected to disagree in a meaningful way and we find a data limit below which the QBC strategy fails. Finally, it should be noted that data selection strategies similar to those used here might be of use in other settings in which generation of more data is a costly process.
Jens Kaae Christensen, Kasper Lamberth, Morten Nielsen 0001, Claus Lundegaard, Peder Worning, Sanne Lise Lauemøller, Søren Buus, Søren Brunak, Ole Lund
Neural Comput.8
2002 Neural network predicts sequence of TP53 gene based on DNA chip
abstract
UNLABELLED: We have trained an artificial neural network to predict the sequence of the human TP53 tumor suppressor gene based on a p53 GeneChip. The trained neural network uses as input the fluorescence intensities of DNA hybridized to oligonucleotides on the surface of the chip and makes between zero and four errors in the predicted 1300 bp sequence when tested on wild-type TP53 sequence. AVAILABILITY: The trained neural network is available for academic use by contacting [email protected]
Jeppe S. Spicker, Friedrik Wikman, Ming-Lan Lu, Carlos Cordon-Cardo, Christopher T. Workman, Torben F. Ørntoft, Søren Brunak, Steen Knudsen
Bioinform.7
2001 Flexibility of the genetic code with respect to DNA structure
abstract
MOTIVATION: The primary function of DNA is to carry genetic information through the genetic code. DNA, however, contains a variety of other signals related, for instance, to reading frame, codon bias, pairwise codon bias, splice sites and transcription regulation, nucleosome positioning and DNA structure. Here we study the relationship between the genetic code and DNA structure and address two questions. First, to which degree does the degeneracy of the genetic code and the acceptable amino acid substitution patterns allow for the superimposition of DNA structural signals to protein coding sequences? Second, is the origin or evolution of the genetic code likely to have been constrained by DNA structure? RESULTS: We develop an index for code flexibility with respect to DNA structure. Using five different di- or tri-nucleotide models of sequence-dependent DNA structure, we show that the standard genetic code provides a fair level of flexibility at the level of broad amino acid categories. Thus the code generally allows for the superimposition of any structural signal on any protein-coding sequence, through amino acid substitution. The flexibility observed at the level of single amino acids allows only for the superimposition of punctual and loosely positioned signals to conserved amino acid sequences. The degree of flexibility of the genetic code is low or average with respect to several classes of alternative codes. This result is consistent with the view that DNA structure is not likely to have played a significant role in the origin and evolution of the genetic code.
Pierre-François Baisnée, Pierre Baldi, Søren Brunak, Anders Gorm Pedersen
Bioinform.3
2000 Matching Protein b-Sheet Partners by Feedforward and Recurrent Neural Networks
Pierre Baldi, Gianluca Pollastri, Claus A. F. Andersen, Søren Brunak
ISMB4
2000 Assessing the accuracy of prediction algorithms for classification: an overview
abstract
Abstract 4 Also at the Department of Biological Sciences, University of California, Irvine, USA, to whom all correspondence should be addressed. We provide a unified overview of methods that currently are widely used to assess the accuracy of prediction algorithms, from raw percentages, quadratic error measures and other distances, and correlation coefficients, and to information theoretic measures such as relative entropy and mutual information. We briefly discuss the advantages and disadvantages of each approach. For classification tasks, we derive new learning algorithms for the design of prediction systems by directly optimising the correlation coefficient. We observe and prove several results relating sensitivity and specificity of optimal systems. While the principles are general, we illustrate the applicability on specific problems such as protein secondary structure and signal peptide prediction. Contact: [email protected]
Pierre Baldi, Søren Brunak, Yves Chauvin, Claus A. F. Andersen, Henrik Nielsen
Bioinform.2
1999 Using Sequence Motifs for Enhanced Neural Network Prediction of Protein Distance Constraints
Jan Gorodkin, Ole Lund, Claus A. F. Andersen, Søren Brunak
ISMB4
1999 Structural basis for triplet repeat disorders: a computational analysis
abstract
MOTIVATION: Over a dozen major degenerative disorders, including myotonic distrophy, Huntington's disease and fragile X syndrome, result from unstable expansions of particular trinucleotides. Remarkably, only some of all the possible triplets, namely CAG/CTG, CGG/CCG and GAA/TTC, have been associated with the known pathological expansions. This raises some basic questions at the DNA level. Why do particular triplets seem to be singled out? What is the mechanism for their expansion and how does it depend on the triplet itself? Could other triplets or longer repeats be involved in other diseases? RESULTS: Using several different computational models of DNA structure, we show that the triplets involved in the pathological repeats generally fall into extreme classes. Thus, CAG/CTG repeats are particularly flexible, whereas GCC, CGG and GAA repeats appear to display both flexible and rigid (but curved) characteristics depending on the method of analysis. The fact that (1) trinucleotide repeats often become increasingly unstable when they exceed a length of approximately 50 repeats, and (2) repeated 12-mers display a similar increase in instability above 13 repeats, together suggest that approximately 150 bp is a general threshold length for repeat instability. Since this is about the length of DNA wrapped up in a single nucleosome core particle, we speculate that chromatin structure may play an important role in the expansion mechanism. We furthermore suggest that expansion of a dodecamer repeat, which we predict to have very high flexibility, may play a role in the pathogenesis of the neurodegenerative disorder multiple system atrophy (MSA). CONTACT: [email protected], [email protected], [email protected], [email protected].
Pierre Baldi, Søren Brunak, Yves Chauvin, Anders Gorm Pedersen
Bioinform.2
1999 Exploiting the past and the future in protein secondary structure prediction
abstract
MOTIVATION: Predicting the secondary structure of a protein (alpha-helix, beta-sheet, coil) is an important step towards elucidating its three-dimensional structure, as well as its function. Presently, the best predictors are based on machine learning approaches, in particular neural network architectures with a fixed, and relatively short, input window of amino acids, centered at the prediction site. Although a fixed small window avoids overfitting problems, it does not permit capturing variable long-rang information. RESULTS: We introduce a family of novel architectures which can learn to make predictions based on variable ranges of dependencies. These architectures extend recurrent neural networks, introducing non-causal bidirectional dynamics to capture both upstream and downstream information. The prediction algorithm is completed by the use of mixtures of estimators that leverage evolutionary information, expressed in terms of multiple alignments, both at the input and output levels. While our system currently achieves an overall performance close to 76% correct prediction--at least comparable to the best existing systems--the main emphasis here is on the development of new algorithmic ideas. AVAILABILITY: The executable program for predicting protein secondary structure is available from the authors free of charge. CONTACT: [email protected], [email protected], [email protected], [email protected].
Pierre Baldi, Søren Brunak, Paolo Frasconi, Giovanni Soda, Gianluca Pollastri
Bioinform.2
1999 MatrixPlot: visualizing sequence constraints
abstract
UNLABELLED: MatrixPlot is a program for making high-quality matrix plots, such as mutual information plots of sequence alignments and distance matrices of sequences with known three-dimensional coordinates. The user can add information about the sequences (e.g. a sequence logo profile) along the edges of the plot, as well as zoom in on any region in the plot. AVAILABILITY: MatrixPlot can be obtained on request, and can also be accessed online at http://www. cbs.dtu.dk/services/MatrixPlot. CONTACT: [email protected]
Jan Gorodkin, Hans Henrik Stærfeldt, Ole Lund, Søren Brunak
Bioinform.4
1998 Computational Applications of DNA Structural Scales
Pierre Baldi, Søren Brunak, Yves Chauvin, Anders Gorm Pedersen
ISMB2
1997 Displaying the information contents of structural RNA alignments: the structure logos
abstract
MOTIVATION: We extend the standard 'Sequence Logo' method of Schneider and Stevens (Nucleic Acids Res., 18, 6097-6100, 1990) to incorporate prior frequencies on the bases, allow for gaps in the alignments, and indicate the mutual information of base-paired regions in RNA. RESULTS: Given an alignment of RNA sequences with the base pairings indicated, the program will calculate the information at each position, including the mutual information of the base pairs, and display the results in a 'Structure Logo'. Alignments without base pairing can also be displayed in a 'Sequence Logo', but still allowing gaps and incorporating prior frequencies if desired. AVAILABILITY: The code is available from, and an Internet server can be used to run the program at, http://www.cbs.dtu.dk/gorodkin/appl/slogo. html.
Jan Gorodkin, Laurie J. Heyer, Søren Brunak, Gary D. Stormo
Comput. Appl. Biosci.3
1997 A Neural Network Method for Identification of Prokaryotic and Eukaryotic Signal Peptides and Prediction of their Cleavage Sites
abstract
We have developed a new method for the identification of signal peptides and their cleavage sites based on neural networks trained on separate sets of prokaryotic and eukaryotic sequences. The method performs significantly better than previous prediction schemes, and can easily be applied to genome-wide data sets. Discrimination between cleaved signal peptides and uncleaved N-terminal signal-anchor sequences is also possible, though with lower precision. Predictions can be made on a publicly available WWW server: http://www.cbs.dtu.dk/services/SignalP/.
Henrik Nielsen, Jacob Engelbrecht, Søren Brunak, Gunnar von Heijne
Int. J. Neural Syst.3
1996 Characterization of Prokaryotic and Eukaryotic Promoters Using Hidden Markov Models
Anders Gorm Pedersen, Pierre Baldi, Søren Brunak, Yves Chauvin
ISMB3
1995 Periodic Sequence Patterns in Human Exons
Pierre Baldi, Søren Brunak, Yves Chauvin, Jacob Engelbrecht, Anders Krogh
ISMB2
1993 Hidden Markov Models for Human Genes
Pierre Baldi, Søren Brunak, Yves Chauvin, Jacob Engelbrecht, Anders Krogh
NIPS2
1990 A Novel Approach to Prediction of the 3-Dimensional Structures
Henrik Fredholm, Henrik Bohr, Jakob Bohr, Søren Brunak, Rodney M. J. Cotterill, Benny Lautrup, Steffen B. Petersen
NIPS4