EDBT 2026 Demo / reviewers in the wild / expert
Lars Juhl Jensen
dblp:06/6699
· DBLP profile ↗
36ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-7885-715XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 34 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SPACE: STRING proteins as complementary embeddingsabstractMOTIVATION: Representation learning has revolutionized sequence-based prediction of protein function and subcellular localization. Protein networks are an important source of information complementary to sequences, but the use of protein networks has proven to be challenging in the context of machine learning, especially in a cross-species setting. RESULTS: We leveraged the STRING database of protein networks and orthology relations for 1322 eukaryotes to generate network-based cross-species protein embeddings. We did this by first creating species-specific network embeddings and subsequently aligning them based on orthology relations to facilitate direct cross-species comparisons. We show that these aligned network embeddings ensure consistency across species without sacrificing quality compared to species-specific network embeddings. We also show that the aligned network embeddings are complementary to sequence embedding techniques, despite the use of sequence-based orthology relations in the alignment process. Finally, we validated the embeddings by using them for two well-established tasks: subcellular localization prediction and protein function prediction. Training logistic regression classifiers on aligned network embeddings and sequence embeddings improved the accuracy over using sequence alone, reaching performance numbers close to state-of-the-art deep-learning methods. AVAILABILITY AND IMPLEMENTATION: The source code and scripts for generating the network-based cross-species protein embeddings are available at https://github.com/deweihu96/SPACE. Precomputed network embeddings and sequence embeddings for all eukaryotic proteins are included in STRING version 12.0 (https://string-db.org/cgi/download). Dewei Hu, Damian Szklarczyk, Christian von Mering, Lars Juhl Jensen |
Bioinform. | 4 |
| 2024 | Graph databases in systems biology: a systematic reviewabstractGraph databases are becoming increasingly popular across scientific disciplines, being highly suitable for storing and connecting complex heterogeneous data. In systems biology, they are used as a backend solution for biological data repositories, ontologies, networks, pathways, and knowledge graph databases. In this review, we analyse all publications using or mentioning graph databases retrieved from PubMed and PubMed Central full-text search, focusing on the top 16 available graph databases, Publications are categorized according to their domain and application, focusing on pathway and network biology and relevant ontologies and tools. We detail different approaches and highlight the advantages of outstanding resources, such as UniProtKB, Disease Ontology, and Reactome, which provide graph-based solutions. We discuss ongoing efforts of the systems biology community to standardize and harmonize knowledge graph creation and the maintenance of integrated resources. Outlining prospects, including the use of graph databases as a way of communication between biological data repositories, we conclude that efficient design, querying, and maintenance of graph databases will be key for knowledge generation in systems biology and other research fields with heterogeneous data. Ilya Mazein, Adrien Rougny, Alexander Mazein, Ron Henkel, Lea Gütebier, Lea Michaelis, Marek Ostaszewski, Reinhard Schneider 0002, Venkata P. Satagopam, Lars Juhl Jensen, Dagmar Waltemath, Judith A. H. Wodke, Irina Balaur |
Briefings Bioinform. | 10 |
| 2024 | FAVA: high-quality functional association networks inferred from scRNA-seq and proteomics dataabstractMOTIVATION: Protein networks are commonly used for understanding how proteins interact. However, they are typically biased by data availability, favoring well-studied proteins with more interactions. To uncover functions of understudied proteins, we must use data that are not affected by this literature bias, such as single-cell RNA-seq and proteomics. Due to data sparseness and redundancy, functional association analysis becomes complex. RESULTS: To address this, we have developed FAVA (Functional Associations using Variational Autoencoders), which compresses high-dimensional data into a low-dimensional space. FAVA infers networks from high-dimensional omics data with much higher accuracy than existing methods, across a diverse collection of real as well as simulated datasets. FAVA can process large datasets with over 0.5 million conditions and has predicted 4210 interactions between 1039 understudied proteins. Our findings showcase FAVA's capability to offer novel perspectives on protein interactions. FAVA functions within the scverse ecosystem, employing AnnData as its input source. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, and tutorials for FAVA are accessible on GitHub at https://github.com/mikelkou/fava. FAVA can also be installed and used via pip/PyPI as well as via the scverse ecosystem https://github.com/scverse/ecosystem-packages/tree/main/packages/favapy. Mikaela Koutrouli, Katerina C. Nastou, Pau Piera Líndez, Robbin Bouwmeester, Simon Rasmussen, Lennart Martens, Lars Juhl Jensen |
Bioinform. | 7 |
| 2024 | ECCB2024: The 23rd European Conference on Computational BiologyabstractThis volume of Bioinformatics includes the proceedings papers of the 23rd European Conference on Computational Biology (ECCB2024) to be held in Turku, Finland, from 16 September to 20 September 2024, under the theme Data and Algorithms for Health and Science. More information on the ECCB2024 conference is available at the conference website https://eccb2024.fi/. ECCB is one of the main international conferences in the field of computational biology and bioinformatics together with the Intelligent Systems for Molecular Biology (ISMB) and the Research in Computational Molecular Biology (RECOMB). It is held jointly with the ISMB conference in odd-numbered years and independently in even-numbered years. ECCB attracts scientists and industry professionals from diverse disciplines, including mathematics, statistics, computer science, biology, and medicine. Rapid technological advancements enable life scientists to gain increasingly detailed insights in complex biological systems. These improvements in measurement technologies, however, introduce new challenges for data interpretation and necessitate the development of advanced computational techniques to manage the complexity and volume of the data. Consequently, the field of computational biology is rapidly evolving with new algorithms, software, and databases. The ECCB2024 conference showcases cutting-edge developments in systems biology, artificial intelligence, single-cell and spatial technologies, data integration, and more, addressing the growing demand for sophisticated algorithms to enhance the analysis of large-scale biological and biomedical datasets. This is highlighted by the most frequent keywords accompanying all the accepted submissions in ECCB2024 (tutorials/workshops, proceedings, highlight talks, posters), with the most popular keywords including ‘machine learning’, ‘deep learning’, ‘single-cell rnaseq’, and ‘multiomics’ (Fig. 1). Most frequent keywords among all accepted submissions in ECCB2024. The barplot shows the frequencies of top keywords appearing in at least ten submissions, while the word cloud illustrates the relative frequencies of all keywords appearing in at least five submissions. The ECCB2024 edition features five keynote lectures by distinguished speakers: Sarah Teichmann (Cambridge Stem Cell Institute, University of Cambridge, UK), Peer Bork (EMBL—European Molecular Biology Laboratory, Germany), Ileana Cristea (Princeton University, US), Jussi Taipale (University of Cambridge, UK), and Fabian Theis (Helmholtz Munich Computational Health Center, Germany). In addition, ECCB2024 hosts a scientific debate on data sharing and privacy protection by Melissa Haendel (University of North Carolina, US) and Yves Moreau (KU Leuven, Belgium). The ECCB2024 conference covers a wide range of topics, focusing on methodological advancements in computational biology as well as innovative application of computational techniques to life sciences and medicine. To provide a cohesive overview of recent scientific progress, the conference presentations are organized under six broad themes: (i) Genomes, (ii) Proteins, (iii) Systems biology and multiomics, (iv) Single-cell omics, (v) Microbiomes and planetary health, and (vi) Digital health. The proceedings talks present new scientific contributions, while the highlight talks showcase already published cutting-edge science in computational biology and related further developments. Poster presentations provide an opportunity for the participants to discuss their recent work with other researchers in the field. Workshops and tutorials on specialized topics prior to the main conference program are platforms to share practical experiences and learn new skills. In addition to scientific presentation tracks, ECCB2024 has a separate ELIXIR track, overseen by ELIXIR, focusing on advancements in infrastructure and services within ELIXIR nodes in support of the expert groups known as ‘ELIXIR Communities’. ELIXIR is a distributed pan-European life science infrastructure for biological data that coordinates, integrates, and sustains bioinformatics resources across its member states. This enables academic and industry users to access data, tools, standards, computing, and training services for life science research. ELIXIR has selected ECCB as a primary dissemination platform, serving as a co-organizing sponsor. Apart from Community activities the track will also introduce three scientific areas as outlined in the 2024–2028 ELIXIR Scientific Programme (https://elixir-europe.org), enabling scientists to access and analyse life science data across ‘Cellular and Molecular Research’, ‘Biodiversity, food security & pathogens’, and ‘Human data & translational research’. ECCB2024 received an impressive number of 200 submissions of full manuscripts for the proceedings call, highlighting the growing interest and engagement within the community. These submissions were organized under the six conference themes. Each submission was subjected to a peer-review process, with at least two reviews per manuscript, managed by the members of the ECCB2024 Programme Committee. The Programme Committee was chaired by Laura Elo (University of Turku, Finland) and each conference theme had an area chair (Table 1). The Proceedings Review Committee had over 100 reviewers. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. The review process focused on the impact, reproducibility, and scientific quality of the submitted research, as well as its relevance, interest, and value for the ECCB2024 audience. Upon completion of the review, the Program Committee chairs selected 24 papers to be included in the ECCB2024 proceedings, with an acceptance rate of 12% across the different areas (Table 1). All proceedings papers are published open access in this Proceedings issue of September 2024 of Bioinformatics. The ECCB2024 call for Highlight Talks invited presentations of studies recently published in scientific journals since 1 March 2023, or accepted for publication. The existence of further developments related to the published paper was also positively considered. The call received a total of 212 proposals, which were ranked according to their relevance and impact on computational biology, as well as their suitability to be presented to the large and diverse audience. Ultimately, the Program Committee selected 23% of the submissions for presentation as a Highlight talk at the conference. ECCB2024 hosts two poster sessions where researchers can introduce and discuss their work under the conference themes. In total, nearly 700 posters were accepted to be presented. The poster submissions were reviewed by the Posters Committee: Sini Junttila, Asta Laiho, Tomi Suomi, and Sampsa Hautaniemi (late posters). In addition, the ELIXIR track selected posters to be presented at ECCB2024. Overall, the ECCB2024 tracks attracted participation of presenting authors from 48 countries. The countries with the highest number of presenting authors were Germany, Finland, the USA, the United Kingdom, and Spain, jointly contributing to half of the overall number (Fig. 2). Following these were Turkey, France, and China, each contributing 4%–5% of the presenting authors. Proportion of countries where the presenting author of each accepted submission had their primary affiliation. Countries with the proportion below 2% were grouped under ‘Other’. Exhibitor booths are open throughout the conference, including ELIXIR, ISCB, Oxford University Press, Royal Society Publishing and eLife Sciences Publications Ltd, as well as the two organizing institutions: University of Turku and CSC—IT Center for Science. The booths showcase the latest scientific literature in computational biology and bioinformatics, data stewardship, as well as new developments in hardware, software, and technology. In addition, the Visit Turku Archipelago (tourist office) is also represented. Before the main ECCB2024 conference, a satellite meeting, seven workshops, and nine tutorials are organized. The satellite meeting is organized by the Student Council of the International Society for Computational Biology (ISCB) by and for early-stage researchers. This 8th European Student Council Symposium (ESCS) continues the successful collaboration between ESCS and ECCB, building the future of research in computational biology. The 16 workshops/tutorials preceding the ECCB2024 main conference were selected out of a total of 24 proposals by the workshops and tutorials chair Bengt Persson (Uppsala University, Sweden), each running for half a day. The workshops foster discussions and exchange of ideas on a range of specialized or emerging topics in computational biology and provide opportunities to share practical experiences. The tutorials provide participants with lectures and hands-on training to enable learning about new areas of computational biology or important established topics. Following the tradition of previous ECCB conferences, the ISCB and ECCB sponsored a number of travel fellowships for students and postdoctoral fellows. These fellowships were primarily awarded to members presenting talks and those from low- or middle-income countries, facilitating their participation in the conference. A total of 68 applications were received from scientists across 26 countries. After a careful review, the ECCB2024 Organizing Committee in collaboration with the ISCB and ECCB awarded 10 and 18 fellowships, respectively. The ECCB2024 Code of Conduct is designed to ensure a safe and respectful environment for all attendees, outlining clear standards of behaviour to foster an inclusive conference atmosphere. In addition, the code provides specific procedures for attendees to follow if they feel these standards have been violated. This ensures that all participants have access to support and resources to address any concerns, reinforcing the commitment of ECCB2024 to uphold the highest ethical and professional standards during the conference. The ECCB2024 Organizing Committee is strongly committed to promote gender equity throughout the conference planning and execution. In particular, an important effort was made to ensure that the review panels included a balanced composition of female and male professionals across the world. In addition, three out of the seven distinguished keynote speakers are women. The conference program also features a collaborative workshop with the Bioinfo4Women initiative, discussing sex and gender bias in artificial intelligence. We would like to thank everyone who has contributed to the success of ECCB2024, ensuring it meets high standards of excellence. Special recognition is due to the Program Committee, including theme area chairs and reviewers, whose critical and dedicated efforts have been fundamental to the conference organization in selecting excellent tutorials, workshops, manuscripts, talks, and posters. We are equally thankful to the ECCB steering committee for their invaluable support and advice, especially ECCB2022 organizers for providing detailed information about the organization of the previous ECCB conference. The collaboration with the ISCB has been crucial in offering travel fellowships and promoting ECCB2024 globally, for which we are deeply grateful. Our gratitude also goes to all our financial sponsors, including our co-organizing sponsor ELIXIR. We also acknowledge the Oxford University Press production team for their work on the ECCB2024 Proceedings issue. Many individuals have played significant roles in the local organization, often exceeding their responsibilities, and we are immensely thankful for their dedication. Lastly, the conference would not be what it is without the diverse participants from around the world, whose scientific contributions, presentations, and discussions enrich ECCB2024. Thank you all for being there and for allowing us to enjoy science at ECCB2024 in Turku! None declared. This paper was published as part of a supplement financially supported by ECCB2024. All data are incorporated into the article. Additional information is available at the conference website https://eccb2024.fi. Anu Kukkonen-Macchi, Sampsa Hautaniemi, Katharina F. Heil, Merja Heinäniemi, Lars Juhl Jensen, Sini Junttila, Lukas Käll, Asta Laiho, Peter Maccallum, Matti Nykter, Bengt Persson, Tomi Suomi, Tim Van Den Bossche, Tommi H. Nyrönen, Laura Elo |
Bioinform. | 5 |
| 2024 | STRING-ing together protein complexes: corpus and methods for extracting physical protein interactions from the biomedical literatureabstractMOTIVATION: Understanding biological processes relies heavily on curated knowledge of physical interactions between proteins. Yet, a notable gap remains between the information stored in databases of curated knowledge and the plethora of interactions documented in the scientific literature. RESULTS: To bridge this gap, we introduce ComplexTome, a manually annotated corpus designed to facilitate the development of text-mining methods for the extraction of complex formation relationships among biomedical entities targeting the downstream semantics of the physical interaction subnetwork of the STRING database. This corpus comprises 1287 documents with ∼3500 relationships. We train a novel relation extraction model on this corpus and find that it can highly reliably identify physical protein interactions (F1-score = 82.8%). We additionally enhance the model's capabilities through unsupervised trigger word detection and apply it to extract relations and trigger words for these relations from all open publications in the domain literature. This information has been fully integrated into the latest version of the STRING database. AVAILABILITY AND IMPLEMENTATION: We provide the corpus, code, and all results produced by the large-scale runs of our systems biomedical on literature via Zenodo https://doi.org/10.5281/zenodo.8139716, Github https://github.com/farmeh/ComplexTome_extraction, and the latest version of STRING database https://string-db.org/. Farrokh Mehryary, Katerina C. Nastou, Tomoko Ohta, Lars Juhl Jensen, Sampo Pyysalo |
Bioinform. | 4 |
| 2024 | Improving dictionary-based named entity recognition with deep learningabstractMOTIVATION: Dictionary-based named entity recognition (NER) allows terms to be detected in a corpus and normalized to biomedical databases and ontologies. However, adaptation to different entity types requires new high-quality dictionaries and associated lists of blocked names for each type. The latter are so far created by identifying cases that cause many false positives through manual inspection of individual names, a process that scales poorly. RESULTS: In this work, we aim to improve block list s by automatically identifying names to block, based on the context in which they appear. By comparing results of three well-established biomedical NER methods, we generated a dataset of over 12.5 million text spans where the methods agree on the boundaries and type of entity tagged. These were used to generate positive and negative examples of contexts for four entity types (genes, diseases, species, and chemicals), which were used to train a Transformer-based model (BioBERT) to perform entity type classification. Application of the best model (F1-score = 96.7%) allowed us to generate a list of problematic names that should be blocked. Introducing this into our system doubled the size of the previous list of corpus-wide blocked names. In addition, we generated a document-specific list that allows ambiguous names to be blocked in specific documents. These changes boosted text mining precision by ∼5.5% on average, and over 8.5% for chemical and 7.5% for gene names, positively affecting several biological databases utilizing this NER system, like the STRING database, with only a minor drop in recall (0.6%). AVAILABILITY AND IMPLEMENTATION: All resources are available through Zenodo https://doi.org/10.5281/zenodo.11243139 and GitHub https://doi.org/10.5281/zenodo.10289360. Katerina C. Nastou, Mikaela Koutrouli, Sampo Pyysalo, Lars Juhl Jensen |
Bioinform. | 4 |
| 2024 | Lifestyle factors in the biomedical literature: an ontology and comprehensive resources for named entity recognitionabstractMOTIVATION: Despite lifestyle factors (LSFs) being increasingly acknowledged in shaping individual health trajectories, particularly in chronic diseases, they have still not been systematically described in the biomedical literature. This is in part because no named entity recognition (NER) system exists, which can comprehensively detect all types of LSFs in text. The task is challenging due to their inherent diversity, lack of a comprehensive LSF classification for dictionary-based NER, and lack of a corpus for deep learning-based NER. RESULTS: We present a novel lifestyle factor ontology (LSFO), which we used to develop a dictionary-based system for recognition and normalization of LSFs. Additionally, we introduce a manually annotated corpus for LSFs (LSF200) suitable for training and evaluation of NER systems, and use it to train a transformer-based system. Evaluating the performance of both NER systems on the corpus revealed an F-score of 64% for the dictionary-based system and 76% for the transformer-based system. Large-scale application of these systems on PubMed abstracts and PMC Open Access articles identified over 300 million mentions of LSF in the biomedical literature. AVAILABILITY AND IMPLEMENTATION: LSFO, the annotated LSF200 corpus, and the detected LSFs in PubMed and PMC-OA articles using both NER systems, are available under open licenses via the following GitHub repository: https://github.com/EsmaeilNourani/LSFO-expansion. This repository contains links to two associated GitHub repositories and a Zenodo project related to the study. LSFO is also available at BioPortal: https://bioportal.bioontology.org/ontologies/LSFO. Esmaeil Nourani, Mikaela Koutrouli, Yijia Xie, Danai Vagiaki, Sampo Pyysalo, Katerina C. Nastou, Søren Brunak, Lars Juhl Jensen |
Bioinform. | 8 |
| 2023 | S1000: a better taxonomic name corpus for biomedical information extractionabstractMOTIVATION: The recognition of mentions of species names in text is a critically important task for biomedical text mining. While deep learning-based methods have made great advances in many named entity recognition tasks, results for species name recognition remain poor. We hypothesize that this is primarily due to the lack of appropriate corpora. RESULTS: We introduce the S1000 corpus, a comprehensive manual re-annotation and extension of the S800 corpus. We demonstrate that S1000 makes highly accurate recognition of species names possible (F-score =93.1%), both for deep learning and dictionary-based methods. AVAILABILITY AND IMPLEMENTATION: All resources introduced in this study are available under open licenses from https://jensenlab.org/resources/s1000/. The webpage contains links to a Zenodo project and three GitHub repositories associated with the study. Jouni Luoma, Katerina C. Nastou, Tomoko Ohta, Harttu Toivonen, Evangelos Pafilis, Lars Juhl Jensen, Sampo Pyysalo |
Bioinform. | 6 |
| 2022 | Identifying the genes impacted by cell proliferation in proteomics and transcriptomics studiesabstractHypothesis-free high-throughput profiling allows relative quantification of thousands of proteins or transcripts across samples and thereby identification of differentially expressed genes. It is used in many biological contexts to characterize differences between cell lines and tissues, identify drug mode of action or drivers of drug resistance, among others. Changes in gene expression can also be due to confounding factors that were not accounted for in the experimental plan, such as change in cell proliferation. We combined the analysis of 1,076 and 1,040 cell lines in five proteomics and three transcriptomics data sets to identify 157 genes that correlate with cell proliferation rates. These include actors in DNA replication and mitosis, and genes periodically expressed during the cell cycle. This signature of cell proliferation is a valuable resource when analyzing high-throughput data showing changes in proliferation across conditions. We show how to use this resource to help in interpretation of in vitro drug screens and tumor samples. It informs on differences of cell proliferation rates between conditions where such information is not directly available. The signature genes also highlight which hits in a screen may be due to proliferation changes; this can either contribute to biological interpretation or help focus on experiment-specific regulation events otherwise buried in the statistical analysis. Marie Locard-Paulet, Oana Palasca, Lars Juhl Jensen |
PLoS Comput. Biol. | 3 |
| 2021 | TIGA: target illumination GWAS analyticsabstractMOTIVATION: Genome-wide association studies can reveal important genotype-phenotype associations; however, data quality and interpretability issues must be addressed. For drug discovery scientists seeking to prioritize targets based on the available evidence, these issues go beyond the single study. RESULTS: Here, we describe rational ranking, filtering and interpretation of inferred gene-trait associations and data aggregation across studies by leveraging existing curation and harmonization efforts. Each gene-trait association is evaluated for confidence, with scores derived solely from aggregated statistics, linking a protein-coding gene and phenotype. We propose a method for assessing confidence in gene-trait associations from evidence aggregated across studies, including a bibliometric assessment of scientific consensus based on the iCite relative citation ratio, and meanRank scores, to aggregate multivariate evidence.This method, intended for drug target hypothesis generation, scoring and ranking, has been implemented as an analytical pipeline, available as open source, with public datasets of results, and a web application designed for usability by drug discovery scientists. AVAILABILITY AND IMPLEMENTATION: Web application, datasets and source code via https://unmtid-shinyapps.net/tiga/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jeremy J. Yang, Dhouha Grissa, Christophe G. Lambert, Cristian Bologa, Stephen L. Mathias, Anna Waller, David J. Wild 0001, Lars Juhl Jensen, Tudor I. Oprea |
Bioinform. | 8 |
| 2020 | CoCoScore: context-aware co-occurrence scoring for text mining applications using distant supervisionabstractMOTIVATION: Information extraction by mining the scientific literature is key to uncovering relations between biomedical entities. Most existing approaches based on natural language processing extract relations from single sentence-level co-mentions, ignoring co-occurrence statistics over the whole corpus. Existing approaches counting entity co-occurrences ignore the textual context of each co-occurrence. RESULTS: We propose a novel corpus-wide co-occurrence scoring approach to relation extraction that takes the textual context of each co-mention into account. Our method, called CoCoScore, scores the certainty of stating an association for each sentence that co-mentions two entities. CoCoScore is trained using distant supervision based on a gold-standard set of associations between entities of interest. Instead of requiring a manually annotated training corpus, co-mentions are labeled as positives/negatives according to their presence/absence in the gold standard. We show that CoCoScore outperforms previous approaches in identifying human disease-gene and tissue-gene associations as well as in identifying physical and functional protein-protein associations in different species. CoCoScore is a versatile text mining tool to uncover pairwise associations via co-occurrence mining, within and beyond biomedical applications. AVAILABILITY AND IMPLEMENTATION: CoCoScore is available at: https://github.com/JungeAlexander/cocoscore. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander Junge, Lars Juhl Jensen |
Bioinform. | 2 |
| 2020 | Alcoholic liver disease: A registry view on comorbidities and disease predictionabstractAlcoholic-related liver disease (ALD) is the cause of more than half of all liver-related deaths. Sustained excess drinking causes fatty liver and alcohol-related steatohepatitis, which may progress to alcoholic liver fibrosis (ALF) and eventually to alcohol-related liver cirrhosis (ALC). Unfortunately, it is difficult to identify patients with early-stage ALD, as these are largely asymptomatic. Consequently, the majority of ALD patients are only diagnosed by the time ALD has reached decompensated cirrhosis, a symptomatic phase marked by the development of complications as bleeding and ascites. The main goal of this study is to discover relevant upstream diagnoses helping to understand the development of ALD, and to highlight meaningful downstream diagnoses that represent its progression to liver failure. Here, we use data from the Danish health registries covering the entire population of Denmark during nineteen years (1996-2014), to examine if it is possible to identify patients likely to develop ALF or ALC based on their past medical history. To this end, we explore a knowledge discovery approach by using high-dimensional statistical and machine learning techniques to extract and analyze data from the Danish National Patient Registry. Consistent with the late diagnoses of ALD, we find that ALC is the most common form of ALD in the registry data and that ALC patients have a strong over-representation of diagnoses associated with liver dysfunction. By contrast, we identify a small number of patients diagnosed with ALF who appear to be much less sick than those with ALC. We perform a matched case-control study using the group of patients with ALC as cases and their matched patients with non-ALD as controls. Machine learning models (SVM, RF, LightGBM and NaiveBayes) trained and tested on the set of ALC patients achieve a high performance for data classification (AUC = 0.89). When testing the same trained models on the small set of ALF patients, their performance unsurprisingly drops a lot (AUC = 0.67 for NaiveBayes). The statistical and machine learning results underscore small groups of upstream and downstream comorbidities that accurately detect ALC patients and show promise in prediction of ALF. Some of these groups are conditions either caused by alcohol or caused by malnutrition associated with alcohol-overuse. Others are comorbidities either related to trauma and life-style or to complications to cirrhosis, such as oesophageal varices. Our findings highlight the potential of this approach to uncover knowledge in registry data related to ALD. Dhouha Grissa, Ditlev Nytoft Rasmussen, Aleksander Krag, Søren Brunak, Lars Juhl Jensen |
PLoS Comput. Biol. | 5 |
| 2019 | Inferring disease-associated long non-coding RNAs using genome-wide tissue expression profilesabstractMOTIVATION: Long non-coding RNAs (lncRNAs) are important regulators in wide variety of biological processes, which are linked to many diseases. Compared to protein-coding genes (PCGs), the association between diseases and lncRNAs is still not well studied. Thus, inferring disease-associated lncRNAs on a genome-wide scale has become imperative. RESULTS: In this study, we propose a machine learning-based method, DislncRF, which infers disease-associated lncRNAs on a genome-wide scale based on tissue expression profiles. DislncRF uses random forest models trained on expression profiles of known disease-associated PCGs across human tissues to extract general patterns between expression profiles and diseases. These models are then applied to score associations between lncRNAs and diseases. DislncRF was benchmarked against a gold standard dataset and compared to other methods. The results show that DislncRF yields promising performance and outperforms the existing methods. The utility of DislncRF is further substantiated on two diseases in which we find that top scoring candidates are supported by literature or independent datasets. AVAILABILITY AND IMPLEMENTATION: https://github.com/xypan1232/DislncRF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaoyong Pan, Lars Juhl Jensen, Jan Gorodkin |
Bioinform. | 2 |
| 2019 | ProtFus: A Comprehensive Method Characterizing Protein-Protein Interactions of Fusion ProteinsabstractTailored therapy aims to cure cancer patients effectively and safely, based on the complex interactions between patients' genomic features, disease pathology and drug metabolism. Thus, the continual increase in scientific literature drives the need for efficient methods of data mining to improve the extraction of useful information from texts based on patients' genomic features. An important application of text mining to tailored therapy in cancer encompasses the use of mutations and cancer fusion genes as moieties that change patients' cellular networks to develop cancer, and also affect drug metabolism. Fusion proteins, which are derived from the slippage of two parental genes, are produced in cancer by chromosomal aberrations and trans-splicing. Given that the two parental proteins for predicted fusion proteins are known, we used our previously developed method for identifying chimeric protein-protein interactions (ChiPPIs) associated with the fusion proteins. Here, we present a validation approach that receives fusion proteins of interest, predicts their cellular network alterations by ChiPPI and validates them by our new method, ProtFus, using an online literature search. This process resulted in a set of 358 fusion proteins and their corresponding protein interactions, as a training set for a Naïve Bayes classifier, to identify predicted fusion proteins that have reliable evidence in the literature and that were confirmed experimentally. Next, for a test group of 1817 fusion proteins, we were able to identify from the literature 2908 PPIs in total, across 18 cancer types. The described method, ProtFus, can be used for screening the literature to identify unique cases of fusion proteins and their PPIs, as means of studying alterations of protein networks in cancers. Availability: http://protfus.md.biu.ac.il/. Somnath Tagore, Alessandro Gorohovski, Lars Juhl Jensen, Milana Frenkel-Morgenstern |
PLoS Comput. Biol. | 3 |
| 2018 | LocText: relation extraction of protein localizations to assist database curationabstractBACKGROUND: The subcellular localization of a protein is an important aspect of its function. However, the experimental annotation of locations is not even complete for well-studied model organisms. Text mining might aid database curators to add experimental annotations from the scientific literature. Existing extraction methods have difficulties to distinguish relationships between proteins and cellular locations co-mentioned in the same sentence. RESULTS: LocText was created as a new method to extract protein locations from abstracts and full texts. LocText learned patterns from syntax parse trees and was trained and evaluated on a newly improved LocTextCorpus. Combined with an automatic named-entity recognizer, LocText achieved high precision (P = 86%±4). After completing development, we mined the latest research publications for three organisms: human (Homo sapiens), budding yeast (Saccharomyces cerevisiae), and thale cress (Arabidopsis thaliana). Examining 60 novel, text-mined annotations, we found that 65% (human), 85% (yeast), and 80% (cress) were correct. Of all validated annotations, 40% were completely novel, i.e. did neither appear in the annotations nor the text descriptions of Swiss-Prot. CONCLUSIONS: LocText provides a cost-effective, semi-automated workflow to assist database curators in identifying novel protein localization annotations. The annotations suggested through text-mining would be verified by experts to guarantee high-quality standards of manually-curated databases such as Swiss-Prot. Juan Miguel Cejuela, Shrikant Vinchurkar, Tatyana Goldberg, Madhukar Sollepura Prabhu Shankar, Ashish Baghudana, Aleksandar Bojchevski, Carsten Uhlig, André Ofner, Pandu Raharja-Liu, Lars Juhl Jensen, Burkhard Rost |
BMC Bioinform. | 10 |
| 2018 | A comprehensive and quantitative comparison of text-mining in 15 million full-text articles versus their corresponding abstractsabstractAcross academia and industry, text mining has become a popular strategy for keeping up with the rapid growth of the scientific literature. Text mining of the scientific literature has mostly been carried out on collections of abstracts, due to their availability. Here we present an analysis of 15 million English scientific full-text articles published during the period 1823-2016. We describe the development in article length and publication sub-topics during these nearly 250 years. We showcase the potential of text mining by extracting published protein-protein, disease-gene, and protein subcellular associations using a named entity recognition system, and quantitatively report on their accuracy using gold standard benchmark data sets. We subsequently compare the findings to corresponding results obtained on 16.5 million abstracts included in MEDLINE and show that text mining of full-text articles consistently outperforms using abstracts only. David Westergaard, Hans Henrik Stærfeldt, Christian Tønsberg, Lars Juhl Jensen, Søren Brunak |
PLoS Comput. Biol. | 4 |
| 2017 | TIN-X: target importance and novelty explorerabstractMOTIVATION: The increasing amount of peer-reviewed manuscripts requires the development of specific mining tools to facilitate the visual exploration of evidence linking diseases and proteins. RESULTS: We developed TIN-X, the Target Importance and Novelty eXplorer, to visualize the association between proteins and diseases, based on text mining data processed from scientific literature. In the current implementation, TIN-X supports exploration of data for G-protein coupled receptors, kinases, ion channels, and nuclear receptors. TIN-X supports browsing and navigating across proteins and diseases based on ontology classes, and displays a scatter plot with two proposed new bibliometric statistics: Importance and Novelty. AVAILABILITY AND IMPLEMENTATION: http://www.newdrugtargets.org. CONTACT: [email protected]. Daniel Cannon, Jeremy J. Yang, Stephen L. Mathias, Oleg Ursu, Subramani Mani, Anna Waller, Stephan C. Schürer, Lars Juhl Jensen, Larry A. Sklar, Cristian Bologa, Tudor I. Oprea |
Bioinform. | 8 |
| 2016 | SVD-phy: improved prediction of protein functional associations through singular value decomposition of phylogenetic profilesabstractUNLABELLED: A successful approach for predicting functional associations between non-homologous genes is to compare their phylogenetic distributions. We have devised a phylogenetic profiling algorithm, SVD-Phy, which uses truncated singular value decomposition to address the problem of uninformative profiles giving rise to false positive predictions. Benchmarking the algorithm against the KEGG pathway database, we found that it has substantially improved performance over existing phylogenetic profiling methods. AVAILABILITY AND IMPLEMENTATION: The software is available under the open-source BSD license at https://bitbucket.org/andrea/svd-phy CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andrea Franceschini, Jianyi Lin, Christian von Mering, Lars Juhl Jensen |
Bioinform. | 4 |
| 2015 | ENVIRONMENTS and EOL: identification of Environment Ontology terms in text and the annotation of the Encyclopedia of LifeabstractUNLABELLED: The association of organisms to their environments is a key issue in exploring biodiversity patterns. This knowledge has traditionally been scattered, but textual descriptions of taxa and their habitats are now being consolidated in centralized resources. However, structured annotations are needed to facilitate large-scale analyses. Therefore, we developed ENVIRONMENTS, a fast dictionary-based tagger capable of identifying Environment Ontology (ENVO) terms in text. We evaluate the accuracy of the tagger on a new manually curated corpus of 600 Encyclopedia of Life (EOL) species pages. We use the tagger to associate taxa with environments by tagging EOL text content monthly, and integrate the results into the EOL to disseminate them to a broad audience of users. AVAILABILITY AND IMPLEMENTATION: The software and the corpus are available under the open-source BSD and the CC-BY-NC-SA 3.0 licenses, respectively, at http://environments.hcmr.gr. Evangelos Pafilis, Sune Pletscher-Frankild, Julia Schnetzer, Lucia Fanini, Sarah Faulwetter, Christina Pavloudi, Katerina Vasileiadou, Patrick Leary, Jennifer Hammock, Katja Schulz, Cynthia Sims Parr, Christos Arvanitidis, Lars Juhl Jensen |
Bioinform. | 13 |
| 2014 | Protein-driven inference of miRNA-disease associationsabstractMOTIVATION: MicroRNAs (miRNAs) are a highly abundant class of non-coding RNA genes involved in cellular regulation and thus also diseases. Despite miRNAs being important disease factors, miRNA-disease associations remain low in number and of variable reliability. Furthermore, existing databases and prediction methods do not explicitly facilitate forming hypotheses about the possible molecular causes of the association, thereby making the path to experimental follow-up longer. RESULTS: Here we present miRPD in which miRNA-Protein-Disease associations are explicitly inferred. Besides linking miRNAs to diseases, it directly suggests the underlying proteins involved, which can be used to form hypotheses that can be experimentally tested. The inference of miRNAs and diseases is made by coupling known and predicted miRNA-protein associations with protein-disease associations text mined from the literature. We present scoring schemes that allow us to rank miRNA-disease associations inferred from both curated and predicted miRNA targets by reliability and thereby to create high- and medium-confidence sets of associations. Analyzing these, we find statistically significant enrichment for proteins involved in pathways related to cancer and type I diabetes mellitus, suggesting either a literature bias or a genuine biological trend. We show by example how the associations can be used to extract proteins for disease hypothesis. AVAILABILITY AND IMPLEMENTATION: All datasets, software and a searchable Web site are available at http://mirpd.jensenlab.org. Søren Mørk, Sune Pletscher-Frankild, Albert Pallejà, Jan Gorodkin, Lars Juhl Jensen |
Bioinform. | 5 |
| 2013 | OnTheFly 2.0: A tool for automatic annotation of files and biological information extractionabstractRetrieving all of the necessary information from databases about bioentities mentioned in an article is not a trivial or an easy task. Following the daily literature about a specific biological topic and collecting all the necessary information about the bioentities mentioned in the literature manually is tedious and time consuming. OnTheFly 2.0 is a web application mainly designed for non-computer experts which aims to automate data collection and knowledge extraction from biological literature in a user friendly and efficient way. OnTheFly 2.0 is able to extract bioentities from individual articles such as text, Microsoft Word, Excel and PDF files. With a simple drag-and-drop motion, the text of a document is extensively parsed for bioentities such as protein/gene names and chemical compound names. Utilizing high quality data integration platforms, OnTheFly allows the generation of informative summaries, interaction networks and at-a-glance popup windows containing knowledge related to the bioentities found in documents. OnTheFly 2.0 provides a concise application to automate the extraction of bioentities hidden in various documents and is offered as a web based application. It can be found at: http://onthefly.embl.de, http://onthefly.med.uoc.gr or http://onthefly.hcmr.gr. Evangelos Pafilis, Georgios A. Pavlopoulos, Venkata P. Satagopam, Nikolas Papanikolaou, Heiko Horn, Christos Arvanitidis, Lars Juhl Jensen, Reinhard Schneider 0002 |
BIBE | 7 |
| 2013 | Are graph databases ready for bioinformatics?abstractAbstract Contact: [email protected] Christian Theil Have, Lars Juhl Jensen |
Bioinform. | 2 |
| 2013 | Dictionary construction and identification of possible adverse drug events in Danish clinical narrative textabstractOBJECTIVE: Drugs have tremendous potential to cure and relieve disease, but the risk of unintended effects is always present. Healthcare providers increasingly record data in electronic patient records (EPRs), in which we aim to identify possible adverse events (AEs) and, specifically, possible adverse drug events (ADEs). MATERIALS AND METHODS: Based on the undesirable effects section from the summary of product characteristics (SPC) of 7446 drugs, we have built a Danish ADE dictionary. Starting from this dictionary we have developed a pipeline for identifying possible ADEs in unstructured clinical narrative text. We use a named entity recognition (NER) tagger to identify dictionary matches in the text and post-coordination rules to construct ADE compound terms. Finally, we apply post-processing rules and filters to handle, for example, negations and sentences about subjects other than the patient. Moreover, this method allows synonyms to be identified and anatomical location descriptions can be merged to allow appropriate grouping of effects in the same location. RESULTS: The method identified 1 970 731 (35 477 unique) possible ADEs in a large corpus of 6011 psychiatric hospital patient records. Validation was performed through manual inspection of possible ADEs, resulting in precision of 89% and recall of 75%. DISCUSSION: The presented dictionary-building method could be used to construct other ADE dictionaries. The complication of compound words in Germanic languages was addressed. Additionally, the synonym and anatomical location collapse improve the method. CONCLUSIONS: The developed dictionary and method can be used to identify possible ADEs in Danish clinical narratives. Robert Eriksson, Peter Bjødstrup Jensen, Sune Pletscher-Frankild, Lars Juhl Jensen, Søren Brunak |
J. Am. Medical Informatics Assoc. | 4 |
| 2011 | The rise and fall of supervised machine learning techniquesabstractMachine learning is of immense importance in bioinformatics and biomedical science more generally (Larrañaga et al., 2006; Tarca et al., 2007). In particular, supervised machine learning has been used to great effect in numerous bioinformatics prediction methods. Through many years of editing and reviewing manuscripts, we noticed that some supervised machine learning techniques seem to be gaining in popularity while others seemed, at least to our eyes, to be looking ‘unfashionable’. We were motivated to create a league table of machine learning techniques to learn what is hot and what is not in the machine learning field. In this editorial, we only include those that we considered major league and leave analysis of the minor league methods as an exercise for the interested reader. To create our league table, we created a list of supervised machine learning techniques commonly used in bioinformatics and their common synonyms, plural forms and abbreviations. We then searched this list against the PubMed titles and abstracts to identify the number of papers published per year for each machine learning technique. To match as many papers as possible, searches were case insensitive and allowed for variation in hyphenation. To our surprise, the artificial neural network (ANN) is not only the dominant league leader in 2011 but has been in this position since at least the 1970s (see Fig. 1). However, in recent years the usage of support vector machines (SVMs) grew tremendously, and we predict that SVMs will challenge ANNs for the dominant position in the coming decade. Since 2007 the number of publications using ANNs has decreased by 21%, which we hypothesize may be directly attributed to researchers increasingly using SVMs in place of ANNs. SVMs caught up with and overtook Markov models in 2004 to gain second spot in our machine learning league. The growth of supervised machine learning methods in PubMed. As for the question of ‘what is hot?’, one can see that Random forests are a rapidly growing method with not a single mention of them before 2003 and now a total of 407 papers published to date. We were hoping to find techniques that were not so hot and perhaps going out of fashion. The results show that none of the major league methods has gone out of fashion, but we do see moderate decreases in the use of both ANNs and Markov models in the literature. We were also curious to find out if certain machine learning techniques were used in combination with each other. To investigate this, we looked at what machine learning methods are co-mentioned in articles (See Fig. 2). For all pairs of methods from the Supervised Machine Learning Top-5, we counted the number of abstracts that mention both methods and normalized the counts with the number of co-occurrences that would be expected by chance (based on the frequencies with which the methods are mentioned over the years). The strongest correlation (185 times higher than random expectation) is seen between decision trees and random forests, which is to be expected as random forests are ensembles of decision trees. Apart from this, the next strongest correlation (88 times higher than random expectation) is found between the two newest methods on the list, namely SVMs and random forests. We hypothesize that this is due to many researchers using these algorithms through machine learning frameworks such as Weka (Frank et al., 2004), which allows many different algorithms to easily be applied to the same dataset. Heatmap showing the co-occurrence of machine learning techniques within articles. Applications of supervised machine learning methodology continue to grow in the biomedical literature. Despite new methods growing in usage, for example support vector machines and random forests, we see little evidence that any widely adopted methods are falling out of use. Conflict of Interest: none declared. Lars Juhl Jensen, Alex Bateman |
Bioinform. | 1 |
| 2011 | Ten Simple Rules for Getting Help from Online Scientific CommunitiesabstractInternational audience Giovanni Marco Dall'Olio, Jacopo Marino, Michael Schubert, Kevin L. Keys, Melanie I. Stefan, Colin S. Gillespie, Pierre Poulain, Khader Shameer, Robert Sugar, Brandon M. Invergo, Lars Juhl Jensen, Jaume Bertranpetit, Hafid Laayouni |
PLoS Comput. Biol. | 11 |
| 2011 | BioStar: An Online Question & Answer Resource for the Bioinformatics CommunityabstractInternational audience Laurence D. Parnell, Pierre Lindenbaum, Khader Shameer, Giovanni Marco Dall'Olio, Daniel C. Swan, Lars Juhl Jensen, Simon J. Cockell, Brent S. Pedersen, Mary E. Mangan, Christopher A. Miller 0002, István Albert |
PLoS Comput. Biol. | 6 |
| 2011 | Using Electronic Patient Records to Discover Disease Correlations and Stratify Patient CohortsabstractElectronic patient records remain a rather unexplored, but potentially rich data source for discovering correlations between diseases. We describe a general approach for gathering phenotypic descriptions of patients from medical records in a systematic and non-cohort dependent manner. By extracting phenotype information from the free-text in such records we demonstrate that we can extend the information contained in the structured record data, and use it for producing fine-grained patient stratification and disease co-occurrence statistics. The approach uses a dictionary based on the International Classification of Disease ontology and is therefore in principle language independent. As a use case we show how records from a Danish psychiatric hospital lead to the identification of disease correlations, which subsequently can be mapped to systems biology frameworks. Francisco S. Roque, Peter Bjødstrup Jensen, Henriette Schmock, Marlene Dalgaard, Massimo Andreatta, Thomas Folkmann Hansen, Karen Søeby, Søren Bredkjær, Anders Juul, Thomas Werge, Lars Juhl Jensen, Søren Brunak |
PLoS Comput. Biol. | 11 |
| 2010 | Drug-Induced Regulation of Target ExpressionabstractDrug perturbations of human cells lead to complex responses upon target binding. One of the known mechanisms is a (positive or negative) feedback loop that adjusts the expression level of the respective target protein. To quantify this mechanism systems-wide in an unbiased way, drug-induced differential expression of drug target mRNA was examined in three cell lines using the Connectivity Map. To overcome various biases in this valuable resource, we have developed a computational normalization and scoring procedure that is applicable to gene expression recording upon heterogeneous drug treatments. In 1290 drug-target relations, corresponding to 466 drugs acting on 167 drug targets studied, 8% of the targets are subject to regulation at the mRNA level. We confirmed systematically that in particular G-protein coupled receptors, when serving as known targets, are regulated upon drug treatment. We further newly identified drug-induced differential regulation of Lanosterol 14-alpha demethylase, Endoplasmin, DNA topoisomerase 2-alpha and Calmodulin 1. The feedback regulation in these and other targets is likely to be relevant for the success or failure of the molecular intervention. Murat Iskar, Monica Campillos, Michael Kuhn 0004, Lars Juhl Jensen, Vera van Noort, Peer Bork |
PLoS Comput. Biol. | 4 |
| 2010 | Reflect: A practical approach to web semantics
Seán I. O'Donoghue, Heiko Horn, Evangelos Pafilis, Sven Haag, Michael Kuhn 0004, Venkata P. Satagopam, Reinhard Schneider 0002, Lars Juhl Jensen |
J. Web Semant. | 8 |
| 2009 | Microblogging the ISMB: A New Approach to Conference ReportingabstractMicroblogging platforms and other tools for videos, podcasts, and virtual environments provide an untapped potential for science conferences. Our experiment using FriendFeed to cover ISMB 2008 was educational and surprisingly successful. We found that it enhanced our note-taking skills, allowed us to compile notes from parallel sessions, attracted wider interest from non-attendees, and, in addition to the “live” aspect, generated a permanent archive of the meeting. ISMB/ECCB 2009 will be held in Stockholm. We look forward to the new developments in Web usage by scientists that are sure to emerge between now and then. We also anticipate new and exciting ways to report from Stockholm as it happens; perhaps the ISMB/ECCB 2009 Web site will look something like this: http://www.bork.embl.de/̃jensen/ismb2008/keynotes.php.html? Neil F. W. Saunders, Pedro Beltrão, Lars Juhl Jensen, Daniel Jurczak, Roland Krause, Michael Kuhn 0004, Shirley Wu |
PLoS Comput. Biol. | 3 |
| 2006 | Extraction of regulatory gene/protein networks from MedlineabstractMOTIVATION: We have previously developed a rule-based approach for extracting information on the regulation of gene expression in yeast. The biomedical literature, however, contains information on several other equally important regulatory mechanisms, in particular phosphorylation, which we now expanded for our rule-based system also to extract. RESULTS: This paper presents new results for extraction of relational information from biomedical text. We have improved our system, STRING-IE, to capture both new types of linguistic constructs as well as new types of biological information [i.e. (de-)phosphorylation]. The precision remains stable with a slight increase in recall. From almost one million PubMed abstracts related to four model organisms, we manage to extract regulatory networks and binary phosphorylations comprising 3,319 relation chunks. The accuracy is 83-90% and 86-95% for gene expression and (de-)phosphorylation relations, respectively. To achieve this, we made use of an organism-specific resource of gene/protein names considerably larger than those used in most other biology related information extraction approaches. These names were included in the lexicon when retraining the part-of-speech (POS) tagger on the GENIA corpus. For the domain in question, an accuracy of 96.4% was attained on POS tags. It should be noted that the rules were developed for yeast and successfully applied to both abstracts and full-text articles related to other organisms with comparable accuracy. AVAILABILITY: The revised GENIA corpus, the POS tagger, the extraction rules and the full sets of extracted relations are available from http://www.bork.embl.de/Docu/STRING-IE Jasmin Saric, Lars Juhl Jensen, Rossitza Ouzounova, Isabel Rojas, Peer Bork |
Bioinform. | 2 |
| 2005 | Comparison of computational methods for the identification of cell cycle-regulated genesabstractMOTIVATION: DNA microarrays have been used extensively to study the cell cycle transcription programme in a number of model organisms. The Saccharomyces cerevisiae data in particular have been subjected to a wide range of bioinformatics analysis methods, aimed at identifying the correct and complete set of periodically expressed genes. RESULTS: Here, we provide the first thorough benchmark of such methods, surprisingly revealing that most new and more mathematically advanced methods actually perform worse than the analysis published with the original microarray data sets. We show that this loss of accuracy specifically affects methods that only model the shape of the expression profile without taking into account the magnitude of regulation. We present a simple permutation-based method that performs better than most existing methods. Ulrik de Lichtenberg, Lars Juhl Jensen, Anders Fausbøll, Thomas Skøt Jensen, Peer Bork, Søren Brunak |
Bioinform. | 2 |
| 2005 | Extraction of Transcript Diversity from Scientific LiteratureabstractTranscript diversity generated by alternative splicing and associated mechanisms contributes heavily to the functional complexity of biological systems. The numerous examples of the mechanisms and functional implications of these events are scattered throughout the scientific literature. Thus, it is crucial to have a tool that can automatically extract the relevant facts and collect them in a knowledge base that can aid the interpretation of data from high-throughput methods. We have developed and applied a composite text-mining method for extracting information on transcript diversity from the entire MEDLINE database in order to create a database of genes with alternative transcripts. It contains information on tissue specificity, number of isoforms, causative mechanisms, functional implications, and experimental methods used for detection. We have mined this resource to identify 959 instances of tissue-specific splicing. Our results in combination with those from EST-based methods suggest that alternative splicing is the preferred mechanism for generating transcript diversity in the nervous system. We provide new annotations for 1,860 genes with the potential for generating transcript diversity. We assign the MeSH term "alternative splicing" to 1,536 additional abstracts in the MEDLINE database and suggest new MeSH terms for other events. We have successfully extracted information about transcript diversity and semiautomatically generated a database, LSAT, that can provide a quantitative understanding of the mechanisms behind tissue-specific gene expression. LSAT (Literature Support for Alternative Transcripts) is publicly available at http://www.bork.embl.de/LSAT/. Parantu K. Shah, Lars Juhl Jensen, Stéphanie Boué, Peer Bork |
PLoS Comput. Biol. | 2 |
| 2004 | Extracting Regulatory Gene Expression Networks From PubmedabstractWe present an approach using syntacto-semantic rules for the extraction of relational information from biomedical abstracts. The results show that by overcoming the hurdle of technical terminology, high precision results can be achieved. From abstracts related to baker's yeast, we manage to extract a regulatory network comprised of 441 pairwise relations from 58,664 abstracts with an accuracy of 83-90%. To achieve this, we made use of a resource of gene/protein names considerably larger than those used in most other biology related information extraction approaches. This list of names was included in the lexicon of our retrained part-of-speech tagger for use on molecular biology abstracts. For the domain in question an accuracy of 93.6-97.7% was attained on POS-tags. The method is easily adapted to other organisms than yeast, allowing us to extract many more biologically relevant relations. Jasmin Saric, Lars Juhl Jensen, Peer Bork, Rossitza Ouzounova, Isabel Rojas |
ACL | 2 |
| 2003 | Prediction of human protein function according to Gene Ontology categoriesabstractMOTIVATION: The human genome project has led to the discovery of many human protein coding genes which were previously unknown. As a large fraction of these are functionally uncharacterized, it is of interest to develop methods for predicting their molecular function from sequence. RESULTS: We have developed a method for prediction of protein function for a subset of classes from the Gene Ontology classification scheme. This subset includes several pharmaceutically interesting categories-transcription factors, receptors, ion channels, stress and immune response proteins, hormones and growth factors can all be predicted. Although the method relies on protein sequences as the sole input, it does not rely on sequence similarity, but instead on sequence derived protein features such as predicted post translational modifications (PTMs), protein sorting signals and physical/chemical properties calculated from the amino acid composition. This allows for prediction of the function for orphan proteins where no homologs can be found. Using this method we propose two novel receptors in the human genome, and further demonstrate chromosomal clustering of related proteins. Lars Juhl Jensen, Ramneek Gupta, Hans Henrik Stærfeldt, Søren Brunak |
Bioinform. | 1 |
| 2000 | Automatic discovery of regulatory patterns in promoter regions based on whole cell expression data and functional annotationabstractMOTIVATION: The whole genomes submitted to GenBank contain valuable information about the function of genes as well as the upstream sequences and whole cell expression provides valuable information on gene regulation. To utilize these large amounts of data for a biological understanding of the regulation of gene expression, new automatic methods for pattern finding are needed. RESULTS: Two word-analysis algorithms for automatic discovery of regulatory sequence elements have been developed. We show that sequence patterns correlated to whole cell expression data can be found using Kolmogorov-Smirnov tests on the raw data, thereby eliminating the need for clustering co-regulated genes. Regulatory elements have also been identified by systematic calculations of the significance of correlations between words found in the functional annotation of genes and DNA words occurring in their promoter regions. Application of these algorithms to the Saccharomyces cerevisiae genome and publicly available DNA array data sets revealed a highly conserved 9-mer occurring in the upstream regions of genes coding for proteasomal subunits. Several other putative and known regulatory elements were also found. AVAILABILITY: Upon request. Lars Juhl Jensen, Steen Knudsen |
Bioinform. | 1 |