Björn A. Grüning

dblp:00/9446 · DBLP profile ↗
← Back
22ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-3079-6586ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 14 since 2021
YearPublicationVenuePosition
2026 Ten common misconceptions about Galaxy (and why they are wrong!)
abstract
Galaxy is a widely used open-source platform for accessible, reproducible, transparent and scalable data analysis in the life sciences and beyond. Despite its growing adoption across domains, several misconceptions persist about its scope, usability, scalability and relevance to academia and industry. In this manuscript, we identify and address 10 common misconceptions about Galaxy, ranging from the belief that it is limited to genomics, lacks scalability, or is only useful for teaching, to doubts about its ability to support secure data analysis or maintain high software quality as a free and open-source project. We refute each misconception with present evidence based on Galaxy's technical features, real-world use cases, user communities and governance structures. We show that Galaxy is a mature and versatile platform capable of supporting cutting-edge scientific research, education and even clinical workflows across a wide variety of disciplines. By clarifying existing misconceptions, we aim to help researchers, educators, developers and decision-makers better appreciate Galaxy's capabilities and potential within their fields.
Wendi Bacon, Bérénice Batut, Sanjay Kumar Srikakulam, Paul F. Zierep, Anthony Bretaudeau, Björn A. Grüning, Gildas Le Corguillé, Helge Hecht, Hans-Rudolf Hotz, Beatriz Serrano-Solano
PLoS Comput. Biol.6
2024 DIMet: an open-source tool for differential analysis of targeted isotope-labeled metabolomics data
abstract
MOTIVATION: Many diseases, such as cancer, are characterized by an alteration of cellular metabolism allowing cells to adapt to changes in the microenvironment. Stable isotope-resolved metabolomics (SIRM) and downstream data analyses are widely used techniques for unraveling cells' metabolic activity to understand the altered functioning of metabolic pathways in the diseased state. While a number of bioinformatic solutions exist for the differential analysis of SIRM data, there is currently no available resource providing a comprehensive toolbox. RESULTS: In this work, we present DIMet, a one-stop comprehensive tool for differential analysis of targeted tracer data. DIMet accepts metabolite total abundances, isotopologue contributions, and isotopic mean enrichment, and supports differential comparison (pairwise and multi-group), time-series analyses, and labeling profile comparison. Moreover, it integrates transcriptomics and targeted metabolomics data through network-based metabolograms. We illustrate the use of DIMet in real SIRM datasets obtained from Glioblastoma P3 cell-line samples. DIMet is open-source, and is readily available for routine downstream analysis of isotope-labeled targeted metabolomics data, as it can be used both in the command line interface or as a complete toolkit in the public Galaxy Europe and Workfow4Metabolomics web platforms. AVAILABILITY AND IMPLEMENTATION: DIMet is freely available at https://github.com/cbib/DIMet, and through https://usegalaxy.eu and https://workflow4metabolomics.usegalaxy.fr. All the datasets are available at Zenodo https://zenodo.org/records/10925786.
Johanna Galvis, Joris Guyon, Benjamin Dartigues, Helge Hecht, Björn A. Grüning, Florian Specque, Hayssam Soueidan, Slim Karkar, Thomas Daubon, Macha Nikolski
Bioinform.5
2023 Fast and accurate genome-wide predictions and structural modeling of protein-protein interactions using Galaxy
abstract
BACKGROUND: Protein-protein interactions play a crucial role in almost all cellular processes. Identifying interacting proteins reveals insight into living organisms and yields novel drug targets for disease treatment. Here, we present a publicly available, automated pipeline to predict genome-wide protein-protein interactions and produce high-quality multimeric structural models. RESULTS: Application of our method to the Human and Yeast genomes yield protein-protein interaction networks similar in quality to common experimental methods. We identified and modeled Human proteins likely to interact with the papain-like protease of SARS-CoV2's non-structural protein 3. We also produced models of SARS-CoV2's spike protein (S) interacting with myelin-oligodendrocyte glycoprotein receptor and dipeptidyl peptidase-4. CONCLUSIONS: The presented method is capable of confidently identifying interactions while providing high-quality multimeric structural models for experimental validation. The interactome modeling pipeline is available at usegalaxy.org and usegalaxy.eu.
Aysam Guerler, Dannon Baker, Marius van den Beek, Björn A. Grüning, Dave Bouvier, Nathan Coraor, Stephen D. Shank, Jordan D. Zehr, Michael C. Schatz, Anton Nekrutenko
BMC Bioinform.4
2023 Transformer-based tool recommendation system in Galaxy
abstract
BACKGROUND: Galaxy is a web-based open-source platform for scientific analyses. Researchers use thousands of high-quality tools and workflows for their respective analyses in Galaxy. Tool recommender system predicts a collection of tools that can be used to extend an analysis. In this work, a tool recommender system is developed by training a transformer on workflows available on Galaxy Europe and its performance is compared to other neural networks such as recurrent, convolutional and dense neural networks. RESULTS: The transformer neural network achieves two times faster convergence, has significantly lower model usage (model reconstruction and prediction) time and shows a better generalisation that goes beyond training workflows than the older tool recommender system created using RNN in Galaxy. In addition, the transformer also outperforms CNN and DNN on several key indicators. It achieves a faster convergence time, lower model usage time, and higher quality tool recommendations than CNN. Compared to DNN, it converges faster to a higher precision@k metric (approximately 0.98 by transformer compared to approximately 0.9 by DNN) and shows higher quality tool recommendations. CONCLUSION: Our work shows a novel usage of transformers to recommend tools for extending scientific workflows. A more robust tool recommendation model, created using a transformer, having significantly lower usage time than RNN and CNN, higher precision@k than DNN, and higher quality tool recommendations than all three neural networks, will benefit researchers in creating scientifically significant workflows and exploratory data analysis in Galaxy. Additionally, the ability to train faster than all three neural networks imparts more scalability for training on larger datasets consisting of millions of tool sequences. Open-source scripts to create the recommendation model are available under MIT licence at https://github.com/anuprulez/galaxy_tool_recommendation_transformers.
Björn A. Grüning, Rolf Backofen
BMC Bioinform.2
2023 Galaxy Training: A powerful framework for teaching!
abstract
There is an ongoing explosion of scientific datasets being generated, brought on by recent technological advances in many areas of the natural sciences. As a result, the life sciences have become increasingly computational in nature, and bioinformatics has taken on a central role in research studies. However, basic computational skills, data analysis, and stewardship are still rarely taught in life science educational programs, resulting in a skills gap in many of the researchers tasked with analysing these big datasets. In order to address this skills gap and empower researchers to perform their own data analyses, the Galaxy Training Network (GTN) has previously developed the Galaxy Training Platform (https://training.galaxyproject.org), an open access, community-driven framework for the collection of FAIR (Findable, Accessible, Interoperable, Reusable) training materials for data analysis utilizing the user-friendly Galaxy framework as its primary data analysis platform. Since its inception, this training platform has thrived, with the number of tutorials and contributors growing rapidly, and the range of topics extending beyond life sciences to include topics such as climatology, cheminformatics, and machine learning. While initially aimed at supporting researchers directly, the GTN framework has proven to be an invaluable resource for educators as well. We have focused our efforts in recent years on adding increased support for this growing community of instructors. New features have been added to facilitate the use of the materials in a classroom setting, simplifying the contribution flow for new materials, and have added a set of train-the-trainer lessons. Here, we present the latest developments in the GTN project, aimed at facilitating the use of the Galaxy Training materials by educators, and its usage in different learning environments.
Saskia D. Hiltemann, Helena Rasche, Simon L. Gladman, Hans-Rudolf Hotz, Delphine Larivière, Daniel J. Blankenberg, Pratik D. Jagtap, Thomas Wollmann, Anthony Bretaudeau, Nadia Goué, Timothy J. Griffin, Coline Royaux, Yvan Le Bras, Subina P. Mehta, Anna Syme, Frederik Coppens, Bert Droesbeke, Nicola Soranzo, Wendi Bacon, Fotis E. Psomopoulos, Cristóbal Gallardo-Alba, Melanie Christine Föll, Matthias Fahrner, Maria A. Doyle, Beatriz Serrano-Solano, Anne Fouilloux, Peter van Heusden, Wolfgang Maier 0003, Dave Clements, Florian Heyl, Björn A. Grüning, Bérénice Batut
PLoS Comput. Biol.33
2022 Ten simple rules for making a software tool workflow-ready
Paul Brack, Peter Crowther, Stian Soiland-Reyes, Stuart Owen, Douglas Lowe, Alan R. Williams, Quentin Groom, Mathias Dillen, Frederik Coppens, Björn A. Grüning, Ignacio Eguinoa, Philip Ewels, Carole A. Goble
PLoS Comput. Biol.10
2021 pyGenomeTracks: reproducible plots for multivariate genomic datasets
abstract
MOTIVATION: Generating publication ready plots to display multiple genomic tracks can pose a serious challenge. Making desirable and accurate figures requires considerable effort. This is usually done by hand or using a vector graphic software. RESULTS: pyGenomeTracks (PGT) is a modular plotting tool that easily combines multiple tracks. It enables a reproducible and standardized generation of highly customizable and publication ready images. AVAILABILITY AND IMPLEMENTATION: PGT is available through a graphical interface on https://usegalaxy.eu and through the command line. It is provided on conda via the bioconda channel, on pip and it is openly developed on github: https://github.com/deeptools/pyGenomeTracks. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lucille Lopez-Delisle, Leily Rabbani, Joachim Wolff, Vivek Bhardwaj 0002, Rolf Backofen, Björn A. Grüning, Fidel Ramírez, Thomas Manke
Bioinform.6
2021 A SARS-CoV-2 sequence submission tool for the European Nucleotide Archive
abstract
SUMMARY: Many aspects of the global response to the COVID-19 pandemic are enabled by the fast and open publication of SARS-CoV-2 genetic sequence data. The European Nucleotide Archive (ENA) is the European recommended open repository for genetic sequences. In this work, we present a tool for submitting raw sequencing reads of SARS-CoV-2 to ENA. The tool features a single-step submission process, a graphical user interface, tabular-formatted metadata and the possibility to remove human reads prior to submission. A Galaxy wrap of the tool allows users with little or no bioinformatics knowledge to do bulk sequencing read submissions. The tool is also packed in a Docker container to ease deployment. AVAILABILITY AND IMPLEMENTATION: CLI ENA upload tool is available at github.com/usegalaxy-eu/ena-upload-cli (DOI 10.5281/zenodo.4537621); Galaxy ENA upload tool at toolshed.g2.bx.psu.edu/view/iuc/ena_upload/382518f24d6d and github.com/galaxyproject/tools-iuc/tree/master/tools/ena_upload (development); and ENA upload Galaxy container at github.com/ELIXIR-Belgium/ena-upload-container (DOI 10.5281/zenodo.4730785).
Miguel Roncoroni, Bert Droesbeke, Ignacio Eguinoa, Kim De Ruyck, Flora D'anna, Dilmurat Yusuf, Björn A. Grüning, Rolf Backofen, Frederik Coppens
Bioinform.7
2021 Scool: a new data storage format for single-cell Hi-C data
abstract
Bioinformatics (2020) doi: 10.1093/bioinformatics/btaa924 In the originally published version of this manuscript, requested author amendments to Figure 1. were inadvertently omitted prior to publishing. This error has now been corrected.
Joachim Wolff, Nezar Abdennur, Rolf Backofen, Björn A. Grüning
Bioinform.4
2021 Scool: a new data storage format for single-cell Hi-C data
abstract
MOTIVATION: Single-cell Hi-C research currently lacks an efficient, easy to use and shareable data storage format. Recent studies have used a variety of sub-optimal solutions: publishing raw data only, text-based interaction matrices, or reusing established Hi-C storage formats for single interaction matrices. These approaches are storage and pre-processing intensive, require long labour time and are often error-prone. RESULTS: The single-cell cooler file format (scool) provides an efficient, user-friendly and storage-saving approach for single-cell Hi-C data. It is a flavour of the established cooler format and guarantees stable API support. AVAILABILITY AND IMPLEMENTATION: The single-cell cooler format is part of the cooler file format as of API version 0.8.9. It is available via pip, conda and github: https://github.com/mirnylab/cooler. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Joachim Wolff, Nezar Abdennur, Rolf Backofen, Björn A. Grüning
Bioinform.4
2021 Robust and efficient single-cell Hi-C clustering with approximate k-nearest neighbor graphs
abstract
MOTIVATION: Hi-C technology provides insights into the 3D organization of the chromatin, and the single-cell Hi-C method enables researchers to gain knowledge about the chromatin state in individual cell levels. Single-cell Hi-C interaction matrices are high dimensional and very sparse. To cluster thousands of single-cell Hi-C interaction matrices, they are flattened and compiled into one matrix. Depending on the resolution, this matrix can have a few million or even billions of features; therefore, computations can be memory intensive. We present a single-cell Hi-C clustering approach using an approximate nearest neighbors method based on locality-sensitive hashing to reduce the dimensions and the computational resources. RESULTS: The presented method can process a 10 kb single-cell Hi-C dataset with 2600 cells and needs 40 GB of memory, while competitive approaches are not computable even with 1 TB of memory. It can be shown that the differentiation of the cells by their chromatin folding properties and, therefore, the quality of the clustering of single-cell Hi-C data is advantageous compared to competitive algorithms. AVAILABILITY AND IMPLEMENTATION: The presented clustering algorithm is part of the scHiCExplorer, is available on Github https://github.com/joachimwolff/scHiCExplorer, and as a conda package via the bioconda channel. The approximate nearest neighbors implementation is available via https://github.com/joachimwolff/sparse-neighbors-search and as a conda package via the bioconda channel. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Joachim Wolff, Rolf Backofen, Björn A. Grüning
Bioinform.3
2021 A constructivist-based proposal for bioinformatics teaching practices during lockdown
abstract
The Coronavirus Disease 2019 (COVID-19) outbreaks have caused universities all across the globe to close their campuses and forced them to initiate online teaching. This article reviews the pedagogical foundations for developing effective distance education practices, starting from the assumption that promoting autonomous thinking is an essential element to guarantee full citizenship in a democracy and for moral decision-making in situations of rapid change, which has become a pressing need in the context of a pandemic. In addition, the main obstacles related to this new context are identified, and solutions are proposed according to the existing bibliography in learning sciences.
Cristóbal Gallardo-Alba, Björn A. Grüning, Beatriz Serrano-Solano
PLoS Comput. Biol.2
2021 Galaxy-ML: An accessible, reproducible, and scalable machine learning toolkit for biomedicine
abstract
Supervised machine learning is an essential but difficult to use approach in biomedical data analysis. The Galaxy-ML toolkit (https://galaxyproject.org/community/machine-learning/) makes supervised machine learning more accessible to biomedical scientists by enabling them to perform end-to-end reproducible machine learning analyses at large scale using only a web browser. Galaxy-ML extends Galaxy (https://galaxyproject.org), a biomedical computational workbench used by tens of thousands of scientists across the world, with a suite of tools for all aspects of supervised machine learning.
Qiang Gu, Simon Bray, Allison L. Creason, Alireza Khanteymoori, Vahid Jalili, Björn A. Grüning, Jeremy Goecks
PLoS Comput. Biol.7
2021 Fostering accessible online education using Galaxy as an e-learning platform
abstract
The COVID-19 pandemic is shifting teaching to an online setting all over the world. The Galaxy framework facilitates the online learning process and makes it accessible by providing a library of high-quality community-curated training materials, enabling easy access to data and tools, and facilitates sharing achievements and progress between students and instructors. By combining Galaxy with robust communication channels, effective instruction can be designed inclusively, regardless of the students' environments.
Beatriz Serrano-Solano, Melanie Christine Föll, Cristóbal Gallardo-Alba, Anika Erxleben-Eggenhofer, Helena Rasche, Saskia D. Hiltemann, Matthias Fahrner, Mark J. Dunning, Marcel H. Schulz, Beáta Scholtz, Dave Clements, Anton Nekrutenko, Bérénice Batut, Björn A. Grüning
PLoS Comput. Biol.14
2020 GLASSgo in Galaxy: high-throughput, reproducible and easy-to-integrate prediction of sRNA homologs
abstract
MOTIVATION: The correct prediction of bacterial sRNA homologs is a prerequisite for many downstream analyses based on comparative genomics, but it is frequently challenging due to the short length and distinct heterogeneity of such homologs. GLobal Automatic Small RNA Search go (GLASSgo) is an efficient tool for the prediction of sRNA homologs from a single input query. To make the algorithm available to a broader community, we offer a Docker container along with a free-access web service. For non-computer scientists, the web service provides a user-friendly interface. However, capabilities were lacking so far for batch processing, version control and direct interaction with compatible software applications as a workflow management system can provide. RESULTS: Here, we present GLASSgo 1.5.2, an updated version that is fully incorporated into the workflow management system Galaxy. The improved version contains a new feature for extracting the upstream regions, allowing the search for conserved promoter elements. Additionally, it supports the use of accession numbers instead of the outdated GI numbers, which widens the applicability of the tool. AVAILABILITY AND IMPLEMENTATION: GLASSgo is available at https://github.com/lotts/GLASSgo/ under the MIT license and is accompanied by instruction and application data. Furthermore, it can be installed into any Galaxy instance using the Galaxy ToolShed.
Richard A. Schäfer, Steffen Lott, Jens Georg, Björn A. Grüning, Wolfgang R. Hess, Björn Voß
Bioinform.4
2019 Parkour LIMS: high-quality sample preparation in next generation sequencing
abstract
MOTIVATION: This paper presents Parkour, a software package for sample processing and quality management of next generation sequencing data and samples. RESULTS: Starting with user requests, Parkour allows tracking and assessing samples based on predefined quality criteria through different stages of the sample preparation workflow. Ideally suited for academic core laboratories, the software aims to maximize efficiency and reduce turnaround time by intelligent sample grouping and a clear assignment of staff to work units. Tools for automated invoicing, interactive statistics on facility usage and simple report generation minimize administrative tasks. Provided as a web application, Parkour is a convenient tool for both deep sequencing service users and laboratory personal. A set of web APIs allow coordinated information sharing with local and remote bioinformaticians. The flexible structure allows workflow customization and simple addition of new features as well as the expansion to other domains. AVAILABILITY AND IMPLEMENTATION: The code and documentation are available at https://github.com/maxplanck-ie/parkour. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Evgeny Anatskiy, Devon Patrick Ryan, Björn A. Grüning, Laura Arrigoni, Thomas Manke, Ulrike Bönisch
Bioinform.3
2019 Biomolecular Reaction and Interaction Dynamics Global Environment (BRIDGE)
abstract
MOTIVATION: The pathway from genomics through proteomics and onto a molecular description of biochemical processes makes the discovery of drugs and biomaterials possible. A research framework common to genomics and proteomics is needed to conduct biomolecular simulations that will connect biological data to the dynamic molecular mechanisms of enzymes and proteins. Novice biomolecular modelers are faced with the daunting task of complex setups and a myriad of possible choices preventing their use of molecular simulations and their ability to conduct reliable and reproducible computations that can be shared with collaborators and verified for procedural accuracy. RESULTS: We present the foundations of Biomolecular Reaction and Interaction Dynamics Global Environment (BRIDGE) developed on the Galaxy platform that makes possible fundamental molecular dynamics of proteins through workflows and pipelines via commonly used packages, such as NAMD, GROMACS and CHARMM. BRIDGE can be used to set up and simulate biological macromolecules, perform conformational analysis from trajectory data and conduct data analytics of large scale protein motions using statistical rigor. We illustrate the basic BRIDGE simulation and analytics capabilities on a previously reported CBH1 protein simulation. AVAILABILITY AND IMPLEMENTATION: Publicly available at https://github.com/scientificomputing/BRIDGE and https://usegalaxy.eu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tharindu Senapathi, Simon Bray, Christopher B. Barnett, Björn A. Grüning, Kevin J. Naidoo
Bioinform.4
2017 BioContainers: an open-source and community-driven framework for software standardization
abstract
MOTIVATION: BioContainers (biocontainers.pro) is an open-source and community-driven framework which provides platform independent executable environments for bioinformatics software. BioContainers allows labs of all sizes to easily install bioinformatics software, maintain multiple versions of the same software and combine tools into powerful analysis pipelines. BioContainers is based on popular open-source projects Docker and rkt frameworks, that allow software to be installed and executed under an isolated and controlled environment. Also, it provides infrastructure and basic guidelines to create, manage and distribute bioinformatics containers with a special focus on omics technologies. These containers can be integrated into more comprehensive bioinformatics pipelines and different architectures (local desktop, cloud environments or HPC clusters). AVAILABILITY AND IMPLEMENTATION: The software is freely available at github.com/BioContainers/. CONTACT: [email protected].
Felipe da Veiga Leprevost, Björn A. Grüning, Saulo Aflitos, Hannes L. Röst, Julian Uszkoreit, Harald Barsnes, Marc Vaudel, Pablo A. Moreno, Laurent Gatto, Jonas Weber, Mingze Bai, Rafael C. Jiménez, Timo Sachsenberg, Julianus Pfeuffer, Roberto Vera, Johannes Griss, Alexey I. Nesvizhskii, Yasset Pérez-Riverol
Bioinform.2
2017 Jupyter and Galaxy: Easing entry barriers into complex data analyses for biomedical researchers
abstract
What does it take to convert a heap of sequencing data into a publishable result? First, common tools are employed to reduce primary data (sequencing reads) to a form suitable for further analyses (i.e., the list of variable sites). The subsequent exploratory stage is much more ad hoc and requires the development of custom scripts and pipelines, making it problematic for biomedical researchers. Here, we describe a hybrid platform combining common analysis pathways with the ability to explore data interactively. It aims to fully encompass and simplify the "raw data-to-publication" pathway and make it reproducible.
Björn A. Grüning, Eric Rasche, Boris Rebolledo-Jaramillo, Carl Eberhard, Torsten Houwaart, John Chilton, Nathan Coraor, Rolf Backofen, James Taylor 0001, Anton Nekrutenko
PLoS Comput. Biol.1
2014 PyWATER: a PyMOL plug-in to find conserved water molecules in proteins by clustering
abstract
SUMMARY: Conserved water molecules play a crucial role in protein structure, stabilization of secondary structure, protein activity, flexibility and ligand binding. Clustering of water molecules in superimposed protein structures, obtained by X-ray crystallography at high resolution, is an established method to identify consensus water molecules in all known protein structures of the same family. PyWATER is an easy-to-use PyMOL plug-in and identifies conserved water molecules in the protein structure of interest. PyWATER can be installed via the user interface of PyMOL. No programming or command-line knowledge is required for its use. AVAILABILITY AND IMPLEMENTATION: PyWATER and a tutorial are available at https://github.com/hiteshpatel379/PyWATER. PyMOL is available at http://www.pymol.org/ or http://sourceforge.net/projects/pymol/. CONTACT: [email protected].
Hitesh Patel, Björn A. Grüning, Stefan Günther, Irmgard Merfort
Bioinform.2
2012 Mining and evaluation of molecular relationships in literature
abstract
MOTIVATION: Specific information on newly discovered proteins is often difficult to find in literature. Particularly if only sequences and no common names of proteins or genes are available, preceding sequence similarity searches can be crucial for the process of information collection. In drug research, it is important to know whether a small molecule targets only one specific protein or whether similar or homologous proteins are also influenced that may account for possible side effects. RESULTS: prolific (protein-literature investigation for interacting compounds) provides a one-step solution to investigate available information on given protein names, sequences, similar proteins or sequences on the gene level. Co-occurrences of UniProtKB/Swiss-Prot proteins and PubChem compounds in all PubMed abstracts are retrievable. Concise 'heat-maps' and tables display frequencies of co-occurrences. They provide links to processed literature with highlighted found protein and compound synonyms. Evaluation with manually curated drug-protein relationships showed that up to 69% could be discovered by automatic text-processing. Examples are presented to demonstrate the capabilities of prolific. AVAILABILITY: The web-application is available at http://prolific.pharmaceutical-bioinformatics.de and a web service at http://www.pharmaceutical-bioinformatics.de/prolific/soap/prolific.wsdl. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Christian Senger, Björn A. Grüning, Anika Erxleben-Eggenhofer, Kersten Döring, Hitesh Patel, Stephan Flemming, Irmgard Merfort, Stefan Günther
Bioinform.2
2011 Compounds In Literature (CIL): screening for compounds and relatives in PubMed
abstract
SUMMARY: Searching for certain compounds in literature can be an elaborate task, with many compounds having several different synonyms. Often, only the structure is known but not its name. Furthermore, rarely investigated compounds may not be described in the available literature at all. In such cases, preceding searches for described similar compounds facilitate literature mining. Highlighted names of proteins in selected texts may further accelerate the time-consuming process of literary research. Compounds In Literature (CIL) provides a web interface to automatically find names, structures, and similar structures in over 28 million compounds of PubChem and more than 18 million citations provided by the PubMed service. CIL's pre-calculated database contains more than 56 million parent compound-abstract relations. Found compounds, relatives and abstracts are related to proteins in a concise 'heat map'-like overview. Compounds and proteins are highlighted in their respective abstracts, and are provided with links to PubChem and UniProt. AVAILABILITY: An easy-to-use web interface with detailed descriptions, help and statistics is available from http://cil.pharmaceutical-bioinformatics.de. CONTACT: [email protected].
Björn A. Grüning, Christian Senger, Anika Erxleben-Eggenhofer, Stephan Flemming, Stefan Günther
Bioinform.1