Nicola J. Mulder

dblp:20/6780 · DBLP profile ↗
← Back
45ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-4905-0941ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 44 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 A scalable HPC framework for bioinformatics in resource-limited settings: design principles, implementation, and sustainability from the UVRI experience
abstract
MOTIVATION: Building and sustaining High-Performance Computing (HPC) infrastructure for bioinformatics research in resource-limited settings presents significant technical, financial and operational challenges. Institutions in low-and middle-income regions often face constraints such as limited technical expertise, unstable infrastructure and restricted funding which can hinder the deployment of large-scale computational platforms necessary for modern genomics and bioinformatics analyses. RESULTS: We present a scalable and modular HPC framework developed at the Uganda Virus Research Institute (UVRI) to support large-scale genomics and other omics data analyses in resource-limited settings. The framework integrates open-source HPC management tools, infrastructure automation, and reproducible configuration management to enable reliable deployment and maintenance. Optimized storage and networking configurations combined with a phased capacity-building strategy support high-throughput genomic workflows while strengthening local technical expertise. From our implementation experience, we derive ten practical design and operational rules that provide a transferable methodology for establishing and sustaining in-house HPC infrastructure. These rules emphasize strategic investment in human capacity, structured planning, leveraging collaborations, adoption of open-source technologies and service management practices to improve operational resilience and long-term sustainability. AVAILABILITY: The design principles, automation strategies and implementation guidelines described in this work are applicable to institutions seeking to establish sustainable HPC resources for bioinformatics research in resource-constrained environments.
Edward Lukyamuzi, Timothy W. Kimbowa, Alfred Ssekagiri, Ronald Galiwango, Grace Kebirungi, Mugume T. Atwine, Mike Nsubuga, Suresh Maslamoney, Sumir Panji, Nicola J. Mulder, Daudi Jjingo, Jonathan K. Kayondo
Bioinform.10
2025 Ten simple rules for building and maintaining sustainable high-performance computing infrastructure for research in resource-limited settings
abstract
[No abstract]
Ronald Galiwango, Christopher J. Whalen, Grace Kebirungi, Mugume T. Atwine, Rodgers Kimera, Alfred Ssekagiri, Timothy W. Kimbowa, Edward Lukyamuzi, Mike Nsubuga, Lloyd Ssentongo, Henry Mutegeki, John M. Fonner, Frank Würthwein, Ari Berman, Laura B. Okalebo, Meghan Coakley McCarthy, Victor S. Kramer, Mariam Quiñones, Phillip Cruz, Darrell E. Hurt, Maria Y. Giovanni, Nicola J. Mulder, Michael Tartakovsky, Jonathan K. Kayondo, Daudi Jjingo
PLoS Comput. Biol.22
2023 MetaNovo: An open-source pipeline for probabilistic peptide discovery in complex metaproteomic datasets
abstract
BACKGROUND: Microbiome research is providing important new insights into the metabolic interactions of complex microbial ecosystems involved in fields as diverse as the pathogenesis of human diseases, agriculture and climate change. Poor correlations typically observed between RNA and protein expression datasets make it hard to accurately infer microbial protein synthesis from metagenomic data. Additionally, mass spectrometry-based metaproteomic analyses typically rely on focused search sequence databases based on prior knowledge for protein identification that may not represent all the proteins present in a set of samples. Metagenomic 16S rRNA sequencing only targets the bacterial component, while whole genome sequencing is at best an indirect measure of expressed proteomes. Here we describe a novel approach, MetaNovo, that combines existing open-source software tools to perform scalable de novo sequence tag matching with a novel algorithm for probabilistic optimization of the entire UniProt knowledgebase to create tailored sequence databases for target-decoy searches directly at the proteome level, enabling metaproteomic analyses without prior expectation of sample composition or metagenomic data generation and compatible with standard downstream analysis pipelines. RESULTS: We compared MetaNovo to published results from the MetaPro-IQ pipeline on 8 human mucosal-luminal interface samples, with comparable numbers of peptide and protein identifications, many shared peptide sequences and a similar bacterial taxonomic distribution compared to that found using a matched metagenome sequence database-but simultaneously identified many more non-bacterial peptides than the previous approaches. MetaNovo was also benchmarked on samples of known microbial composition against matched metagenomic and whole genomic sequence database workflows, yielding many more MS/MS identifications for the expected taxa, with improved taxonomic representation, while also highlighting previously described genome sequencing quality concerns for one of the organisms, and identifying an experimental sample contaminant without prior expectation. CONCLUSIONS: By estimating taxonomic and peptide level information directly on microbiome samples from tandem mass spectrometry data, MetaNovo enables the simultaneous identification of peptides from all domains of life in metaproteome samples, bypassing the need for curated sequence databases to search. We show that the MetaNovo approach to mass spectrometry metaproteomics is more accurate than current gold standard approaches of tailored or matched genomic sequence database searches, can identify sample contaminants without prior expectation and yields insights into previously unidentified metaproteomic signals, building on the potential for complex mass spectrometry metaproteomic data to speak for itself.
Matthys G. Potgieter, Andrew J. M. Nel, Suereta Fortuin, Shaun Garnett, Jerome M. Wendoh, David L. Tabb, Nicola J. Mulder, Jonathan M. Blackburn
PLoS Comput. Biol.7
2022 Correction: Ten simple rules for organizing a bioinformatics training course in low- and middle-income countries
abstract
[This corrects the article DOI: 10.1371/journal.pcbi.1009218.].
Benjamin L. Moore, Patricia Carvajal López, Paballo Abel Chauke, Marco Cristancho, Victoria Dominguez Del Angel, Selene L. Fernandez-Valverde, Amel Ghouila, Piraveen Gopalasingam, Fatma Z. Guerfali, Alice Matimba, Sarah L. Morgan, Guilherme C. Oliveira 0001, Verena Ras, Javier De Las Rivas, Nicola J. Mulder
PLoS Comput. Biol.16
2021 Simulation of African and non-African low and high coverage whole genome sequence data to assess variant calling approaches
abstract
Current variant calling (VC) approaches have been designed to leverage populations of long-range haplotypes and were benchmarked using populations of European descent, whereas most genetic diversity is found in non-European such as Africa populations. Working with these genetically diverse populations, VC tools may produce false positive and false negative results, which may produce misleading conclusions in prioritization of mutations, clinical relevancy and actionability of genes. The most prominent question is which tool or pipeline has a high rate of sensitivity and precision when analysing African data with either low or high sequence coverage, given the high genetic diversity and heterogeneity of this data. Here, a total of 100 synthetic Whole Genome Sequencing (WGS) samples, mimicking the genetics profile of African and European subjects for different specific coverage levels (high/low), have been generated to assess the performance of nine different VC tools on these contrasting datasets. The performances of these tools were assessed in false positive and false negative call rates by comparing the simulated golden variants to the variants identified by each VC tool. Combining our results on sensitivity and positive predictive value (PPV), VarDict [PPV = 0.999 and Matthews correlation coefficient (MCC) = 0.832] and BCFtools (PPV = 0.999 and MCC = 0.813) perform best when using African population data on high and low coverage data. Overall, current VC tools produce high false positive and false negative rates when analysing African compared with European data. This highlights the need for development of VC approaches with high sensitivity and precision tailored for populations characterized by high genetic variations and low linkage disequilibrium.
Shatha Alosaimi, Noëlle van Biljon, Denis Awany, Prisca K. Thami, Joel Defo, Jacquiline W. Mugo, Christian D. Bope, Gaston K. Mazandu, Nicola J. Mulder, Emile R. Chimusa
Briefings Bioinform.9
2021 Reviewing and assessing existing meta-analysis models and tools
abstract
Over the past few years, meta-analysis has become popular among biomedical researchers for detecting biomarkers across multiple cohort studies with increased predictive power. Combining datasets from different sources increases sample size, thus overcoming the issue related to limited sample size from each individual study and boosting the predictive power. This leads to an increased likelihood of more accurately predicting differentially expressed genes/proteins or significant biomarkers underlying the biological condition of interest. Currently, several meta-analysis methods and tools exist, each having its own strengths and limitations. In this paper, we survey existing meta-analysis methods, and assess the performance of different methods based on results from different datasets as well as assessment from prior knowledge of each method. This provides a reference summary of meta-analysis models and tools, which helps to guide end-users on the choice of appropriate models or tools for given types of datasets and enables developers to consider current advances when planning the development of new meta-analysis models and more practical integrative tools.
Funmilayo L. Makinde, Milaine S. S. Tchamga, James Jafali, Segun A. Fatumo, Emile R. Chimusa, Nicola J. Mulder, Gaston K. Mazandu
Briefings Bioinform.6
2021 IHP-PING - generating integrated human protein-protein interaction networks on-the-fly
abstract
Advances in high-throughput sequencing technologies have resulted in an exponential growth of publicly accessible biological datasets. In the 'big data' driven 'post-genomic' context, much work is being done to explore human protein-protein interactions (PPIs) for a systems level based analysis to uncover useful signals and gain more insights to advance current knowledge and answer specific biological and health questions. These PPIs are experimentally or computationally predicted, stored in different online databases and some of PPI resources are updated regularly. As with many biological datasets, such regular updates continuously render older PPI datasets potentially outdated. Moreover, while many of these interactions are shared between these online resources, each resource includes its own identified PPIs and none of these databases exhaustively contains all existing human PPI maps. In this context, it is essential to enable the integration of or combining interaction datasets from different resources, to generate a PPI map with increased coverage and confidence. To allow researchers to produce an integrated human PPI datasets in real-time, we introduce the integrated human protein-protein interaction network generator (IHP-PING) tool. IHP-PING is a flexible python package which generates a human PPI network from freely available online resources. This tool extracts and integrates heterogeneous PPI datasets to generate a unified PPI network, which is stored locally for further applications.
Gaston K. Mazandu, Christopher Hooper, Kenneth Opap, Funmilayo L. Makinde, Victoria Nembaware, Nicholas E. Thomford, Emile R. Chimusa, Ambroise Wonkam, Nicola J. Mulder
Briefings Bioinform.9
2021 Data Management Plans in the genomics research revolution of Africa: Challenges and recommendations
abstract
Drafting and writing a data management plan (DMP) is increasingly seen as a key part of the academic research process. A DMP is a document that describes how a researcher will collect, document, describe, share, and preserve the data that will be generated as part of a research project. The DMP illustrates the importance of utilizing best practices through all stages of working with data while ensuring accessibility, quality, and longevity of the data. The benefits of writing a DMP include compliance with funder and institutional mandates; making research more transparent (for reproduction and validation purposes); and FAIR (findable, accessible, interoperable, reusable); protecting data subjects and compliance with the General Data Protection Regulation (GDPR) and/or local data protection policies. In this review, we highlight the importance of a DMP in modern biomedical research, explaining both the rationale and current best practices associated with DMPs. In addition, we outline various funders' requirements concerning DMPs and discuss open-source tools that facilitate the development and implementation of a DMP. Finally, we discuss DMPs in the context of African research, and the considerations that need to be made in this regard.
Faisal M. Fadlelmola, Lyndon Zass, Melek Chaouch, Chaimae Samtal, Verena Ras, Judit Kumuthini, Sumir Panji, Nicola J. Mulder
J. Biomed. Informatics8
2021 Ten simple rules for developing bioinformatics capacity at an academic institution
abstract
Bioinformatics is an applied interdisciplinary field whose primary purpose is to develop and deploy computational techniques to store, organize, and aid in the analysis and interpretation of large-scale data obtained from biological systems.While rooted in the analysis of nucleotide and protein sequences, it now encompasses techniques targeting multiple data acquisition modalities and seeks to comprehend the functioning of biological systems at many different levels.Bioinformaticians need to be cognizant of diverse scientific fields: basic and molecular biology, genetics, mathematics, statistics, and computer science at a minimum, thus requiring a thoroughly interdisciplinary set of skills to successfully carry out their duties.Due to the growing importance of bioinformatics in enabling modern biomedical research, programs and core facilities have been established in most academic institutions in the developed world over the last 30 years.At present, there are relatively few research and higher education institutions in low-and middle-income countries (LMICs) that have incorporated bioinformatics into their academic curricula or are hosting bioinformatics research or service groups [1,2].There are many reasons for this, including the relative lack of research projects requiring computational analysis, lack of local expertise, lack of required infrastructure, lack of appropriate internal financial support, and the resistance of many academic institutions to the incorporation of new fields of teaching and research [3].However, there is an increasing recognition in some African countries, and other LMICs, that bioinformatics needs to be developed as an independent discipline, both in the academic curricula and in research portfolios [4,5].H3ABioNet is a Pan-African network of bioinformaticians, funded by the United States National Institutes of Health (NIH), that aims to provide technical and scientific support to the Human Heredity and Health in Africa (H3Africa) genomics research program [6].As part of its mission, H3ABioNet has developed training and capacity building projects ranging from introductory courses to advanced workshops on bioinformatics and related topics, and from
Shaun Aron, C. Victor Jongeneel, Paballo Abel Chauke, Melek Chaouch, Judit Kumuthini, Lyndon Zass, Fouzia Radouani, Samar Kamal Kassim, Faisal M. Fadlelmola, Nicola J. Mulder
PLoS Comput. Biol.10
2021 Ten simple rules for organizing a bioinformatics training course in low- and middle-income countries
abstract
ntroductionBioinformatics training is required at every stage of a scientist's research career.Continual bioinformatics training allows exposure to an ever-changing and growing repertoire of techniques and databases, and so biologists, computational scientists, and healthcare practitioners are all seeking learning opportunities in the use of computational resources and tools designed for data storage, retrieval, and analysis.TAU : PleasecheckwhethertheeditstothesentenceThereareabundantopportunitiesforaccessing:::areco here are abundant opportunities for accessing bioinformatics training for scientists in high-income countries (HICs), with well-equipped facilities and participants and trainers requiring minimal travel and financial costs alongside a range of general advice for developing short bioinformatics training courses [1-3].However, regionally targeted bioinformatics training in low-and middle-income countries (LAU : Pleasenotethatlow À middleincomecountrieshasbeenc MICs) often requires more extensive local and external support, organization, and travel.Due to the limited expertise in bioinformatics in LMICs in general, most bioinformatics training requires a fair amount of collaboration with experts beyond the local community, country, or region.A common model of training, used as the basis of this article, includes a local host collaborating with local, regional, and international experts gathering to train local or regional participants.Recently, there has been a growth of capacity strengthening initiatives in LMICs, such as the Pan African Bioinformatics Network for Human Heredity and Health in Africa (H3ABi-oNet) Initiative [4-6], the Capacity Building for Bioinformatics in Latin America (CABANA) Project [7], the Asia Pacific BioInformatics Network (APBioNet) [8], and the Wellcome Connecting Science Courses and Conferences program [9].One of the important strands of these initiatives is a drive to organize and deliver valuable bioinformatics training, but organizing
Benjamin L. Moore, Patricia Carvajal López, Paballo Abel Chauke, Marco Cristancho, Victoria Dominguez Del Angel, Selene L. Fernandez-Valverde, Amel Ghouila, Piraveen Gopalasingam, Fatma Z. Guerfali, Alice Matimba, Sarah L. Morgan, Guilherme C. Oliveira 0001, Verena Ras, Javier De Las Rivas, Nicola J. Mulder
PLoS Comput. Biol.16
2021 Using a multiple-delivery-mode training approach to develop local capacity and infrastructure for advanced bioinformatics in Africa
abstract
With more microbiome studies being conducted by African-based research groups, there is an increasing demand for knowledge and skills in the design and analysis of microbiome studies and data. However, high-quality bioinformatics courses are often impeded by differences in computational environments, complicated software stacks, numerous dependencies, and versions of bioinformatics tools along with a lack of local computational infrastructure and expertise. To address this, H3ABioNet developed a 16S rRNA Microbiome Intermediate Bioinformatics Training course, extending its remote classroom model. The course was developed alongside experienced microbiome researchers, bioinformaticians, and systems administrators, who identified key topics to address. Development of containerised workflows has previously been undertaken by H3ABioNet, and Singularity containers were used here to enable the deployment of a standard replicable software stack across different hosting sites. The pilot ran successfully in 2019 across 23 sites registered in 11 African countries, with more than 200 participants formally enrolled and 106 volunteer staff for onsite support. The pulling, running, and testing of the containers, software, and analyses on various clusters were performed prior to the start of the course by hosting classrooms. The containers allowed the replication of analyses and results across all participating classrooms running a cluster and remained available posttraining ensuring analyses could be repeated on real data. Participants thus received the opportunity to analyse their own data, while local staff were trained and supported by experienced experts, increasing local capacity for ongoing research support. This provides a model for delivering topic-specific bioinformatics courses across Africa and other remote/low-resourced regions which overcomes barriers such as inadequate infrastructures, geographical distance, and access to expertise and educational materials.
Verena Ras, Gerrit Botha, Shaun Aron, Katie Lennard, Imane Allali, Shantelle Claassen-Weitz, Kilaza Samson Mwaikono, Dane Kennedy, Jessica R. Holmes, Gloria Rendon, Sumir Panji, Christopher J. Fields, Nicola J. Mulder
PLoS Comput. Biol.13
2020 FRANC: a unified framework for multi-way local ancestry deconvolution with high density SNP data
abstract
Abstract Several thousand genomes have been completed with millions of variants identified in the human deoxyribonucleic acid sequences. These genomic variations, especially those introduced by admixture, significantly contribute to a remarkable phenotypic variability with medical and/or evolutionary implications. Elucidating local ancestry estimates is necessary for a better understanding of genomic variation patterns throughout modern human evolution and adaptive processes, and consequences in human heredity and health. However, existing local ancestry deconvolution tools are accessible as individual scripts, each requiring input and producing output in its own complex format. This limits the user’s ability to retrieve local ancestry estimates. We introduce a unified framework for multi-way local ancestry inference, FRANC, integrating eight existing state-of-the-art local ancestry deconvolution tools. FRANC is an adaptable, expandable and portable tool that manipulates tool-specific inputs, deconvolutes ancestry and standardizes tool-specific results. To facilitate both medical and population genetics studies, FRANC requires convenient and easy to manipulate input files and allows users to choose output formats to ease their use in further potential local ancestry deconvolution applications.
Ephifania Geza, Nicola J. Mulder, Emile R. Chimusa, Gaston K. Mazandu
Briefings Bioinform.2
2019 A comprehensive survey of models for dissecting local ancestry deconvolution in human genome
abstract
Over the past decade, studies of admixed populations have increasingly gained interest in both medical and population genetics. These studies have so far shed light on the patterns of genetic variation throughout modern human evolution and have improved our understanding of the demographics and adaptive processes of human populations. To date, there exist about 20 methods or tools to deconvolve local ancestry. These methods have merits and drawbacks in estimating local ancestry in multiway admixed populations. In this article, we survey existing ancestry deconvolution methods, with special emphasis on multiway admixture, and compare these methods based on simulation results reported by different studies, computational approaches used, including mathematical and statistical models, and biological challenges related to each method. This should orient users on the choice of an appropriate method or tool for given population admixture characteristics and update researchers on current advances, challenges and opportunities behind existing ancestry deconvolution methods.
Ephifania Geza, Jacquiline W. Mugo, Nicola J. Mulder, Ambroise Wonkam, Emile R. Chimusa, Gaston K. Mazandu
Briefings Bioinform.3
2019 GenGraph: a python module for the simple generation and manipulation of genome graphs
abstract
BACKGROUND: As sequencing technology improves, the concept of a single reference genome is becoming increasingly restricting. In the case of Mycobacterium tuberculosis, one must often choose between using a genome that is closely related to the isolate, or one that is annotated in detail. One promising solution to this problem is through the graph based representation of collections of genomes as a single genome graph. Though there are currently a handful of tools that can create genome graphs and have demonstrated the advantages of this new paradigm, there still exists a need for flexible tools that can be used by researchers to overcome challenges in genomics studies. RESULTS: We present GenGraph, a Python toolkit and accompanying modules that use existing multiple sequence alignment tools to create genome graphs. Python is one of the most popular coding languages for the biological sciences, and by providing these tools, GenGraph makes it easier to experiment and develop new tools that utilise genome graphs. The conceptual model used is highly intuitive, and as much as possible the graph structure represents the biological relationship between the genomes. This design means that users will quickly be able to start creating genome graphs and using them in their own projects. We outline the methods used in the generation of the graphs, and give some examples of how the created graphs may be used. GenGraph utilises existing file formats and methods in the generation of these graphs, allowing graphs to be visualised and imported with widely used applications, including Cytoscape, R, and Java Script. CONCLUSIONS: GenGraph provides a set of tools for generating graph based representations of sets of sequences with a simple conceptual model, written in the widely used coding language Python, and publicly available on Github.
Jon Mitchell Ambler, Shandukani Mulaudzi, Nicola J. Mulder
BMC Bioinform.3
2019 The H3ABioNet helpdesk: an online bioinformatics resource, enhancing Africa's capacity for genomics research
abstract
BACKGROUND: Currently, formal mechanisms for bioinformatics support are limited. The H3Africa Bioinformatics Network has implemented a public and freely available Helpdesk (HD), which provides generic bioinformatics support to researchers through an online ticketing platform. The following article reports on the H3ABioNet HD (H3A-HD)'s development, outlining its design, management, usage and evaluation framework, as well as the lessons learned through implementation. RESULTS: The H3A-HD evaluated using automatically generated usage logs, user feedback and qualitative ticket evaluation. Evaluation revealed that communication methods, ticketing strategies and the technical platforms used are some of the primary factors which may influence the effectivity of HD. CONCLUSION: To continuously improve the H3A-HD services, the resource should be regularly monitored and evaluated. The H3A-HD design, implementation and evaluation framework could be easily adapted for use by interested stakeholders within the Bioinformatics community and beyond.
Judit Kumuthini, Lyndon Zass, Sumir Panji, Samson P. Salifu, Jonathan K. Kayondo, Victoria Nembaware, Mamana Mbiyavanga, Ajayi Olabode, Ali Kishk, Gordon Wells, Nicola J. Mulder
BMC Bioinform.11
2019 Ten simple rules for organizing a webinar series
abstract
International audience
Faisal M. Fadlelmola, Sumir Panji, Azza E. Ahmed, Amel Ghouila, Wisdom A. Akurugu, Jean-Baka Domelevo Entfellner, Oussema Souiai, Nicola J. Mulder
PLoS Comput. Biol.8
2018 Large-scale data-driven integrative framework for extracting essential targets and processes from disease-associated gene data sets
abstract
Populations worldwide currently face several public health challenges, including growing prevalence of infections and the emergence of new pathogenic organisms. The cost and risk associated with drug development make the development of new drugs for several diseases, especially orphan or rare diseases, unappealing to the pharmaceutical industry. Proof of drug safety and efficacy is required before market approval, and rigorous testing makes the drug development process slow, expensive and frequently result in failure. This failure is often because of the use of irrelevant targets identified in the early steps of the drug discovery process, suggesting that target identification and validation are cornerstones for the success of drug discovery and development. Here, we present a large-scale data-driven integrative computational framework to extract essential targets and processes from an existing disease-associated data set and enhance target selection by leveraging drug-target-disease association at the systems level. We applied this framework to tuberculosis and Ebola virus diseases combining heterogeneous data from multiple sources, including protein-protein functional interaction, functional annotation and pharmaceutical data sets. Results obtained demonstrate the effectiveness of the pipeline, leading to the extraction of essential drug targets and to the rational use of existing approved drugs. This provides an opportunity to move toward optimal target-based strategies for screening available drugs and for drug discovery. There is potential for this model to bridge the gap in the production of orphan disease therapies, offering a systematic approach to predict new uses for existing drugs, thereby harnessing their full therapeutic potential.
Gaston K. Mazandu, Emile R. Chimusa, Kayleigh Rutherford, Elsa-Gayle Zekeng, Zoe Z. Gebremariam, Maryam Y. Onifade, Nicola J. Mulder
Briefings Bioinform.7
2018 Developing reproducible bioinformatics analysis workflows for heterogeneous computing environments to support African genomics
abstract
BACKGROUND: The Pan-African bioinformatics network, H3ABioNet, comprises 27 research institutions in 17 African countries. H3ABioNet is part of the Human Health and Heredity in Africa program (H3Africa), an African-led research consortium funded by the US National Institutes of Health and the UK Wellcome Trust, aimed at using genomics to study and improve the health of Africans. A key role of H3ABioNet is to support H3Africa projects by building bioinformatics infrastructure such as portable and reproducible bioinformatics workflows for use on heterogeneous African computing environments. Processing and analysis of genomic data is an example of a big data application requiring complex interdependent data analysis workflows. Such bioinformatics workflows take the primary and secondary input data through several computationally-intensive processing steps using different software packages, where some of the outputs form inputs for other steps. Implementing scalable, reproducible, portable and easy-to-use workflows is particularly challenging. RESULTS: H3ABioNet has built four workflows to support (1) the calling of variants from high-throughput sequencing data; (2) the analysis of microbial populations from 16S rDNA sequence data; (3) genotyping and genome-wide association studies; and (4) single nucleotide polymorphism imputation. A week-long hackathon was organized in August 2016 with participants from six African bioinformatics groups, and US and European collaborators. Two of the workflows are built using the Common Workflow Language framework (CWL) and two using Nextflow. All the workflows are containerized for improved portability and reproducibility using Docker, and are publicly available for use by members of the H3Africa consortium and the international research community. CONCLUSION: The H3ABioNet workflows have been implemented in view of offering ease of use for the end user and high levels of reproducibility and portability, all while following modern state of the art bioinformatics data processing protocols. The H3ABioNet workflows will service the H3Africa consortium projects and are currently in use. All four workflows are also publicly available for research scientists worldwide to use and adapt for their respective needs. The H3ABioNet workflows will help develop bioinformatics capacity and assist genomics research within Africa and serve to increase the scientific output of H3Africa and its Pan-African Bioinformatics Network.
Shakuntala Baichoo, Yassine Souilmi, Sumir Panji, Gerrit Botha, Ayton Meintjes, Scott Hazelhurst, Hocine Bendou, Eugene de Beste, Phelelani T. Mpangase, Oussema Souiai, Mustafa Alghali, Long Yi, Brian D. O'Connor, Michael R. Crusoe, Don Armstrong, Shaun Aron, Fourie Joubert, Azza E. Ahmed, Mamana Mbiyavanga, Peter van Heusden, Lerato E. Magosi, Jennie Zermeno, Liudmila S. Mainzer, Faisal M. Fadlelmola, C. Victor Jongeneel, Nicola J. Mulder
BMC Bioinform.26
2018 The development and application of bioinformatics core competencies to improve bioinformatics training and education
abstract
Bioinformatics is recognized as part of the essential knowledge base of numerous career paths in biomedical research and healthcare. However, there is little agreement in the field over what that knowledge entails or how best to provide it. These disagreements are compounded by the wide range of populations in need of bioinformatics training, with divergent prior backgrounds and intended application areas. The Curriculum Task Force of the International Society of Computational Biology (ISCB) Education Committee has sought to provide a framework for training needs and curricula in terms of a set of bioinformatics core competencies that cut across many user personas and training programs. The initial competencies developed based on surveys of employers and training programs have since been refined through a multiyear process of community engagement. This report describes the current status of the competencies and presents a series of use cases illustrating how they are being applied in diverse training contexts. These use cases are intended to demonstrate how others can make use of the competencies and engage in the process of their continuing refinement and application. The report concludes with a consideration of remaining challenges and future plans.
Nicola J. Mulder, Russell Schwartz, Michelle D. Brazas, Catherine Brooksbank, Bruno A. Gaëta, Sarah L. Morgan, Mark A. Pauley, Anne G. Rosenwald, Gabriella Rustici, Michael L. Sierk, Tandy J. Warnow, Lonnie R. Welch
PLoS Comput. Biol.1
2018 Strategies and opportunities for promoting bioinformatics in Zimbabwe
abstract
Overview of bioinformatics in AfricaThe establishment of the South African National Bioinformatics Institute in South Africa in the 1990s heralded the development of bioinformatics on the continent [9].Countries such as Kenya and Nigeria established pockets of high-quality bioinformatics teams soon after.However, most African research institutions still lagged behind.The introduction of bioinformatics
Ryman Shoko, Justen Manasa, Mcebisi Maphosa, Joshua Mbanga, Reagan Mudziwapasi, Victoria Nembaware, Walter T. Sanyika, Tawanda Tinago, Zedias Chikwambi, Cephas Mawere, Alice Matimba, Grace Mugumbate, Jonathan Mufandaedza, Nicola J. Mulder, Hugh Patterton
PLoS Comput. Biol.14
2017 Gene Ontology semantic similarity tools: survey on features and challenges for biological knowledge discovery
abstract
Gene Ontology (GO) semantic similarity tools enable retrieval of semantic similarity scores, which incorporate biological knowledge embedded in the GO structure for comparing or classifying different proteins or list of proteins based on their GO annotations. This facilitates a better understanding of biological phenomena underlying the corresponding experiment and enables the identification of processes pertinent to different biological conditions. Currently, about 14 tools are available, which may play an important role in improving protein analyses at the functional level using different GO semantic similarity measures. Here we survey these tools to provide a comprehensive view of the challenges and advances made in this area to avoid redundant effort in developing features that already exist, or implementing ideas already proven to be obsolete in the context of GO. This helps researchers, tool developers, as well as end users, understand the underlying semantic similarity measures implemented through knowledge of pertinent features of, and issues related to, a particular tool. This should empower users to make appropriate choices for their biological applications and ensure effective knowledge discovery based on GO annotations.
Gaston K. Mazandu, Emile R. Chimusa, Nicola J. Mulder
Briefings Bioinform.3
2017 A multi-scenario genome-wide medical population genetics simulation framework
abstract
MOTIVATION: Recent technological advances in high-throughput sequencing and genotyping have facilitated an improved understanding of genomic structure and disease-associated genetic factors. In this context, simulation models can play a critical role in revealing various evolutionary and demographic effects on genomic variation, enabling researchers to assess existing and design novel analytical approaches. Although various simulation frameworks have been suggested, they do not account for natural selection in admixture processes. Most are tailored to a single chromosome or a genomic region, very few capture large-scale genomic data, and most are not accessible for genomic communities. RESULTS: Here we develop a multi-scenario genome-wide medical population genetics simulation framework called 'FractalSIM'. FractalSIM has the capability to accurately mimic and generate genome-wide data under various genetic models on genetic diversity, genomic variation affecting diseases and DNA sequence patterns of admixed and/or homogeneous populations. Moreover, the framework accounts for natural selection in both homogeneous and admixture processes. The outputs of FractalSIM have been assessed using popular tools, and the results demonstrated its capability to accurately mimic real scenarios. They can be used to evaluate the performance of a range of genomic tools from ancestry inference to genome-wide association studies. AVAILABILITY AND IMPLEMENTATION: The FractalSIM package is available at http://www.cbio.uct.ac.za/FractalSIM. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jacquiline W. Mugo, Ephifania Geza, Joel Defo, Samar S. M. Elsheikh, Gaston K. Mazandu, Nicola J. Mulder, Emile R. Chimusa
Bioinform.6
2017 Ten simple rules for forming a scientific professional society
abstract
Starting a professional society is not something that should be entered into lightly: it requires work and dedication that can detract from your research projects and other career objectives [15]. It certainly should not be attempted on your own. But there are many potential benefits and rewards in terms of promoting the profile of your discipline (which, in turn, can affect your grant success), boosting your own profile, developing useful management and leadership skills, finding mentors, and forming essential contacts and partnerships, as science is becoming increasingly collaborative. A successful society will be a source of lifelong learning and new ideas, will open up career opportunities for students and investigators, and will provide a much stronger voice for your discipline than an isolated scientist.
Bruno A. Gaëta, Javier De Las Rivas, Paul Horton, Pieter Meysman, Nicola J. Mulder, Paolo Romano 0001, Lonnie R. Welch
PLoS Comput. Biol.5
2017 Designing a course model for distance-based online bioinformatics training in Africa: The H3ABioNet experience
abstract
Africa is not unique in its need for basic bioinformatics training for individuals from a diverse range of academic backgrounds. However, particular logistical challenges in Africa, most notably access to bioinformatics expertise and internet stability, must be addressed in order to meet this need on the continent. H3ABioNet (www.h3abionet.org), the Pan African Bioinformatics Network for H3Africa, has therefore developed an innovative, free-of-charge "Introduction to Bioinformatics" course, taking these challenges into account as part of its educational efforts to provide on-site training and develop local expertise inside its network. A multiple-delivery-mode learning model was selected for this 3-month course in order to increase access to (mostly) African, expert bioinformatics trainers. The content of the course was developed to include a range of fundamental bioinformatics topics at the introductory level. For the first iteration of the course (2016), classrooms with a total of 364 enrolled participants were hosted at 20 institutions across 10 African countries. To ensure that classroom success did not depend on stable internet, trainers pre-recorded their lectures, and classrooms downloaded and watched these locally during biweekly contact sessions. The trainers were available via video conferencing to take questions during contact sessions, as well as via online "question and discussion" forums outside of contact session time. This learning model, developed for a resource-limited setting, could easily be adapted to other settings.
Kim T. Gurwitz, Shaun Aron, Sumir Panji, Suresh Maslamoney, Pedro L. Fernandes, David P. Judge, Amel Ghouila, Jean-Baka Domelevo Entfellner, Fatma Z. Guerfali, Colleen Saunders, Ahmed Mansour Alzohairy, Samson P. Salifu, Rehab Ahmed, Ruben Cloete, Jonathan K. Kayondo, Deogratius Ssemwanga, Nicola J. Mulder
PLoS Comput. Biol.17
2017 Assessing computational genomics skills: Our experience in the H3ABioNet African bioinformatics network
abstract
The H3ABioNet pan-African bioinformatics network, which is funded to support the Human Heredity and Health in Africa (H3Africa) program, has developed node-assessment exercises to gauge the ability of its participating research and service groups to analyze typical genome-wide datasets being generated by H3Africa research groups. We describe a framework for the assessment of computational genomics analysis skills, which includes standard operating procedures, training and test datasets, and a process for administering the exercise. We present the experiences of 3 research groups that have taken the exercise and the impact on their ability to manage complex projects. Finally, we discuss the reasons why many H3ABioNet nodes have declined so far to participate and potential strategies to encourage them to do so.
C. Victor Jongeneel, Ovokeraye Achinike-Oduaran, Ezekiel F. Adebiyi, Marion O. Adebiyi, Seun Adeyemi, Bola Akanle, Shaun Aron, Efejiro Ashano, Hocine Bendou, Gerrit Botha, Emile R. Chimusa, Ananyo Choudhury, Ravikiran Donthu, Jenny Drnevich, Oluwadamilare Falola, Christopher J. Fields, Scott Hazelhurst, Liesl Hendry, Itunuoluwa Isewon, Radhika S. Khetani, Judit Kumuthini, Magambo Phillip Kimuda, Lerato E. Magosi, Liudmila S. Mainzer, Suresh Maslamoney, Mamana Mbiyavanga, Ayton Meintjes, Danny Mugutso, Phelelani T. Mpangase, Richard Munthali, Victoria Nembaware, Andrew Ndhlovu, Trust Odia, Adaobi Okafor, Olaleye Oladipo, Sumir Panji, Venesa Pillay, Gloria Rendon, Dhriti Sengupta, Nicola J. Mulder
PLoS Comput. Biol.40
2016 ancGWAS: a post genome-wide association study method for interaction, pathway and ancestry analysis in homogeneous and admixed populations
abstract
MOTIVATION: Despite numerous successful Genome-wide Association Studies (GWAS), detecting variants that have low disease risk still poses a challenge. GWAS may miss disease genes with weak genetic effects or strong epistatic effects due to the single-marker testing approach commonly used. GWAS may thus generate false negative or inconclusive results, suggesting the need for novel methods to combine effects of single nucleotide polymorphisms within a gene to increase the likelihood of fully characterizing the susceptibility gene. RESULTS: We developed ancGWAS, an algebraic graph-based centrality measure that accounts for linkage disequilibrium in identifying significant disease sub-networks by integrating the association signal from GWAS data sets into the human protein-protein interaction (PPI) network. We validated ancGWAS using an association study result from a breast cancer data set and the simulation of interactive disease loci in the simulation of a complex admixed population, as well as pathway-based GWAS simulation. This new approach holds promise for deconvoluting the interactions between genes underlying the pathogenesis of complex diseases. Results obtained yield a novel central breast cancer sub-network of the human interactome implicated in the proteoglycan syndecan-mediated signaling events pathway which is known to play a major role in mesenchymal tumor cell proliferation, thus providing further insights into breast cancer pathogenesis. AVAILABILITY AND IMPLEMENTATION: The ancGWAS package and documents are available at http://www.cbio.uct.ac.za/~emile/software.html.
Emile R. Chimusa, Mamana Mbiyavanga, Gaston K. Mazandu, Nicola J. Mulder
Bioinform.4
2016 A-DaGO-Fun: an adaptable Gene Ontology semantic similarity-based functional analysis tool
abstract
SUMMARY: Gene Ontology (GO) semantic similarity measures are being used for biological knowledge discovery based on GO annotations by integrating biological information contained in the GO structure into data analyses. To empower users to quickly compute, manipulate and explore these measures, we introduce A-DaGO-Fun (ADaptable Gene Ontology semantic similarity-based Functional analysis). It is a portable software package integrating all known GO information content-based semantic similarity measures and relevant biological applications associated with these measures. A-DaGO-Fun has the advantage not only of handling datasets from the current high-throughput genome-wide applications, but also allowing users to choose the most relevant semantic similarity approach for their biological applications and to adapt a given module to their needs. AVAILABILITY AND IMPLEMENTATION: A-DaGO-Fun is freely available to the research community at http://web.cbio.uct.ac.za/ITGOM/adagofun. It is implemented in Linux using Python under free software (GNU General Public Licence). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gaston K. Mazandu, Emile R. Chimusa, Mamana Mbiyavanga, Nicola J. Mulder
Bioinform.4
2016 The Development of Computational Biology in South Africa: Successes Achieved and Lessons Learnt
abstract
Bioinformatics is now a critical skill in many research and commercial environments as biological data are increasing in both size and complexity. South African researchers recognized this need in the mid-1990s and responded by working with the government as well as international bodies to develop initiatives to build bioinformatics capacity in the country. Significant injections of support from these bodies provided a springboard for the establishment of computational biology units at multiple universities throughout the country, which took on teaching, basic research and support roles. Several challenges were encountered, for example with unreliability of funding, lack of skills, and lack of infrastructure. However, the bioinformatics community worked together to overcome these, and South Africa is now arguably the leading country in bioinformatics on the African continent. Here we discuss how the discipline developed in the country, highlighting the challenges, successes, and lessons learnt.
Nicola J. Mulder, Alan Christoffels, Tulio de Oliveira, Junaid Gamieldien, Scott Hazelhurst, Fourie Joubert, Judit Kumuthini, Ché S. Pillay, Jacky L. Snoep, Özlem Tastan Bishop, Nicki Tiffin
PLoS Comput. Biol.1
2016 Applying, Evaluating and Refining Bioinformatics Core Competencies (An Update from the Curriculum Task Force of ISCB's Education Committee)
abstract
The Curriculum Task Force (CTF) of ISCB’s Education Committee seeks to define curricular guidelines for those who educate or train bioinformatics professionals at all career stages. A recent report of the CTF [1] presented a draft set of bioinformatics core competencies, derived from the results of surveys of (1) core facility directors, (2) career opportunities, and (3) existing curricula. Since the publication of its 2014 report, the CTF has focused on the application of the guidelines in varied contexts to identify areas where refinement is needed. As a first step, the task force held an open meeting at the ISMB conference in July 2014. The ideas discussed at the meeting spawned four working groups (WGs), which focus on (i) defining core competencies for specific types and levels of bioinformatics training, (ii) mapping the curriculum guidelines and competencies to existing materials in order to identify the need for development of new materials, and (iii) identifying where revision of the guidelines may be valuable. The CTF is engaging the ISCB community through open WG meetings at ISCB’s official conferences. Thus far, the WGs have convened at the ISCB Great Lakes Bioinformatics Conference (Purdue University, May 2015) and at the ISMB/ECCB Conference (Dublin, Ireland, July 2015). Additionally, the CTF held a workshop at the Annual General Meeting of the Global Organization of Bioinformatics Learning, Education and Training (Cape Town, South Africa, November 2015). Specifically, the draft competencies have been employed in a wide range of activities and contexts (see Table 1 and [2–11]), including the development of new curricula, the analysis of existing curricula, and the creation of new roles involving bioinformatics. These activities have resulted in the identification of several areas where refinement would be useful: Table 1 Summary of the activities of the ISCB Curriculum Task Force. Identify different levels or phases of competency. It would be helpful to define different phases of competency development, or different levels of competency appropriate for distinct roles. Define competency profiles for disciplines that don’t fit into our current silos. Bioengineering provides an illustrative example of a discipline that requires core competency in bioinformatics but does not fit into our current categories. There are almost certainly others. It would be helpful if we could provide some guidance on how to produce ‘hybrid’ competency profiles, perhaps borrowing some competencies from the TF’s core set and others from different disciplines. The LifeTrain initiative (www.lifetrain.eu) [2, 3] is collecting competency profiles for a range of disciplines of relevance to the biomedical sciences and may provide a useful resource kit for this. Broaden the scope of the competency profiles in response to cutting-edge and emerging research. Current areas requiring improvement include incorporating competencies that capture a fundamental understanding of the biological principles central to analyzing biomolecular data, and broadening the user WG to include applications beyond medicine. Provide guidance on the evidence required to assess whether someone has acquired each competency. For undergraduate, Master’s and PhD programs, learning outcomes for each competency, perhaps with examples of appropriate means of assessment, would be valuable. For established professionals who need to assimilate competencies into their working lives, a different approach may be required (such as keeping a portfolio to capture evidence of competency); the CTF should seek guidance from relevant professional bodies, especially in regulated professions such as healthcare. Provide indicative course content or examples of programs that map to the competency requirements. We do not wish to prescribe what course providers should teach or how they should teach it; however, if a course provider is designing a course to meet a specific competency requirement, it may be helpful to find examples of other programs that do this successfully. One way of achieving this is by mapping existing training content to the TF’s competencies. Another way might be to provide an indication, perhaps based on several courses, of the course content that would meet the competency requirements. This would give course providers the freedom to build their own course syllabi without having to reinvent the wheel. Initiatives to collect examples of Creative Commons (or otherwise reusable) course materials will provide an extremely valuable bank of training materials that could be mapped to the core competencies.
Lonnie R. Welch, Catherine Brooksbank, Russell Schwartz, Sarah L. Morgan, Bruno A. Gaëta, Alastair M. Kilpatrick, Daniel Mietchen, Benjamin L. Moore, Nicola J. Mulder, Mark A. Pauley, William R. Pearson, Predrag Radivojac, Naomi Rosenberg, Anne G. Rosenwald, Gabriella Rustici, Tandy J. Warnow
PLoS Comput. Biol.9
2015 Bioinformatics Education - Perspectives and Challenges out of Africa
abstract
The discipline of bioinformatics has developed rapidly since the complete sequencing of the first genomes in the 1990s. The development of many high-throughput techniques during the last decades has ensured that bioinformatics has grown into a discipline that overlaps with, and is required for, the modern practice of virtually every field in the life sciences. This has placed a scientific premium on the availability of skilled bioinformaticians, a qualification that is extremely scarce on the African continent. The reasons for this are numerous, although the absence of a skilled bioinformatician at academic institutions to initiate a training process and build sustained capacity seems to be a common African shortcoming. This dearth of bioinformatics expertise has had a knock-on effect on the establishment of many modern high-throughput projects at African institutes, including the comprehensive and systematic analysis of genomes from African populations, which are among the most genetically diverse anywhere on the planet. Recent funding initiatives from the National Institutes of Health and the Wellcome Trust are aimed at ameliorating this shortcoming. In this paper, we discuss the problems that have limited the establishment of the bioinformatics field in Africa, as well as propose specific actions that will help with the education and training of bioinformaticians on the continent. This is an absolute requirement in anticipation of a boom in high-throughput approaches to human health issues unique to data from African populations.
Özlem Tastan Bishop, Ezekiel F. Adebiyi, Ahmed M. Alzohairy, Dean Everett, Kais Ghedira, Amel Ghouila, Judit Kumuthini, Nicola J. Mulder, Sumir Panji, Hugh-George Patterton
Briefings Bioinform.8
2015 GOBLET: The Global Organisation for Bioinformatics Learning, Education and Training
abstract
In recent years, high-throughput technologies have brought big data to the life sciences. The march of progress has been rapid, leaving in its wake a demand for courses in data analysis, data stewardship, computing fundamentals, etc., a need that universities have not yet been able to satisfy--paradoxically, many are actually closing "niche" bioinformatics courses at a time of critical need. The impact of this is being felt across continents, as many students and early-stage researchers are being left without appropriate skills to manage, analyse, and interpret their data with confidence. This situation has galvanised a group of scientists to address the problems on an international scale. For the first time, bioinformatics educators and trainers across the globe have come together to address common needs, rising above institutional and international boundaries to cooperate in sharing bioinformatics training expertise, experience, and resources, aiming to put ad hoc training practices on a more professional footing for the benefit of all.
Terri K. Attwood, Erik Bongcam-Rudloff, Michelle D. Brazas, Manuel Corpas, Pascale Gaudet, Fran Lewitter, Nicola J. Mulder, Patricia M. Palagi, Maria Victoria Schneider, Celia W. G. van Gelder
PLoS Comput. Biol.7
2015 A Quick Guide for Building a Successful Bioinformatics Community
abstract
"Scientific community" refers to a group of people collaborating together on scientific-research-related activities who also share common goals, interests, and values. Such communities play a key role in many bioinformatics activities. Communities may be linked to a specific location or institute, or involve people working at many different institutions and locations. Education and training is typically an important component of these communities, providing a valuable context in which to develop skills and expertise, while also strengthening links and relationships within the community. Scientific communities facilitate: (i) the exchange and development of ideas and expertise; (ii) career development; (iii) coordinated funding activities; (iv) interactions and engagement with professionals from other fields; and (v) other activities beneficial to individual participants, communities, and the scientific field as a whole. It is thus beneficial at many different levels to understand the general features of successful, high-impact bioinformatics communities; how individual participants can contribute to the success of these communities; and the role of education and training within these communities. We present here a quick guide to building and maintaining a successful, high-impact bioinformatics community, along with an overview of the general benefits of participating in such communities. This article grew out of contributions made by organizers, presenters, panelists, and other participants of the ISMB/ECCB 2013 workshop "The 'How To Guide' for Establishing a Successful Bioinformatics Network" at the 21st Annual International Conference on Intelligent Systems for Molecular Biology (ISMB) and the 12th European Conference on Computational Biology (ECCB).
Aidan Budd, Manuel Corpas, Michelle D. Brazas, Jonathan C. Fuller, Jeremy Goecks, Nicola J. Mulder, Magali Michaut, B. F. Francis Ouellette, Aleksandra Pawlik, Niklas Blomberg
PLoS Comput. Biol.6
2014 A web-based protein interaction network visualizer
abstract
BACKGROUND: Interaction between proteins is one of the most important mechanisms in the execution of cellular functions. The study of these interactions has provided insight into the functioning of an organism's processes. As of October 2013, Homo sapiens had over 170000 Protein-Protein interactions (PPI) registered in the Interologous Interaction Database, which is only one of the many public resources where protein interactions can be accessed. These numbers exemplify the volume of data that research on the topic has generated. Visualization of large data sets is a well known strategy to make sense of information, and protein interaction data is no exception. There are several tools that allow the exploration of this data, providing different methods to visualize protein network interactions. However, there is still no native web tool that allows this data to be explored interactively online. RESULTS: Given the advances that web technologies have made recently it is time to bring these interactive views to the web to provide an easily accessible forum to visualize PPI. We have created a Web-based Protein Interaction Network Visualizer: PINV, an open source, native web application that facilitates the visualization of protein interactions (http://biosual.cbio.uct.ac.za/pinv.html). We developed PINV as a set of components that follow the protocol defined in BioJS and use the D3 library to create the graphic layouts. We demonstrate the use of PINV with multi-organism interaction networks for a predicted target from Mycobacterium tuberculosis, its interacting partners and its orthologs. CONCLUSIONS: The resultant tool provides an attractive view of complex, fully interactive networks with components that allow the querying, filtering and manipulation of the visible subset. Moreover, as a web resource, PINV simplifies sharing and publishing, activities which are vital in today's research collaborative environments. The source code is freely available for download at https://github.com/4ndr01d3/biosual.
Gustavo A. Salazar, Ayton Meintjes, Gaston K. Mazandu, Holifidy A. Rapanoël, Richard O. Akinola, Nicola J. Mulder
BMC Bioinform.6
2013 Identification of All Exact and Approximate Inverted Repeats in Regular and Weighted Sequences
Carl Barton, Costas S. Iliopoulos, Nicola J. Mulder, Bruce W. Watson
EANN (2)3
2013 Best practices in bioinformatics training for life scientists
abstract
The mountains of data thrusting from the new landscape of modern high-throughput biology are irrevocably changing biomedical research and creating a near-insatiable demand for training in data management and manipulation and data mining and analysis. Among life scientists, from clinicians to environmental researchers, a common theme is the need not just to use, and gain familiarity with, bioinformatics tools and resources but also to understand their underlying fundamental theoretical and practical concepts. Providing bioinformatics training to empower life scientists to handle and analyse their data efficiently, and progress their research, is a challenge across the globe. Delivering good training goes beyond traditional lectures and resource-centric demos, using interactivity, problem-solving exercises and cooperative learning to substantially enhance training quality and learning outcomes. In this context, this article discusses various pragmatic criteria for identifying training needs and learning objectives, for selecting suitable trainees and trainers, for developing and maintaining training skills and evaluating training quality. Adherence to these criteria may help not only to guide course organizers and trainers on the path towards bioinformatics training excellence but, importantly, also to improve the training experience for life scientists.
Allegra Via, Thomas Blicher, Erik Bongcam-Rudloff, Michelle D. Brazas, Catherine Brooksbank, Aidan Budd, Javier De Las Rivas, Jacqueline Dreyer, Pedro L. Fernandes, Celia W. G. van Gelder, Joachim Jacob, Rafael C. Jiménez, Jane E. Loveland, Federico Morán, Nicola J. Mulder, Tommi H. Nyrönen, Kristian Rother, Maria Victoria Schneider, Terri K. Attwood
Briefings Bioinform.15
2013 iAnn: an event sharing platform for the life sciences
abstract
SUMMARY: We present iAnn, an open source community-driven platform for dissemination of life science events, such as courses, conferences and workshops. iAnn allows automatic visualisation and integration of customised event reports. A central repository lies at the core of the platform: curators add submitted events, and these are subsequently accessed via web services. Thus, once an iAnn widget is incorporated into a website, it permanently shows timely relevant information as if it were native to the remote site. At the same time, announcements submitted to the repository are automatically disseminated to all portals that query the system. To facilitate the visualization of announcements, iAnn provides powerful filtering options and views, integrated in Google Maps and Google Calendar. All iAnn widgets are freely available. AVAILABILITY: http://iann.pro/iannviewer CONTACT: [email protected].
Rafael C. Jiménez, Juan P. Albar, Jong Bhak, Marie-Claude Blatter, Thomas Blicher, Michelle D. Brazas, Catherine Brooksbank, Aidan Budd, Javier De Las Rivas, Jacqueline Dreyer, Marc A. van Driel, Michael J. Dunn, Pedro L. Fernandes, Celia W. G. van Gelder, Henning Hermjakob, Vassilios Ioannidis, David Phillip Judge, Pascal Kahlem, Eija Korpelainen, Hans-Joachim Kraus, Jane E. Loveland, Christine Mayer, Jennifer McDowall, Federico Morán, Nicola J. Mulder, Tommi H. Nyrönen, Kristian Rother, Gustavo A. Salazar, Reinhard Schneider 0002, Allegra Via, Jose M. Villaveces, Maria Victoria Schneider, Terri K. Attwood, Manuel Corpas
Bioinform.25
2013 DaGO-Fun: tool for Gene Ontology-based functional analysis using term information content measures
abstract
BACKGROUND: The use of Gene Ontology (GO) data in protein analyses have largely contributed to the improved outcomes of these analyses. Several GO semantic similarity measures have been proposed in recent years and provide tools that allow the integration of biological knowledge embedded in the GO structure into different biological analyses. There is a need for a unified tool that provides the scientific community with the opportunity to explore these different GO similarity measure approaches and their biological applications. RESULTS: We have developed DaGO-Fun, an online tool available at http://web.cbio.uct.ac.za/ITGOM, which incorporates many different GO similarity measures for exploring, analyzing and comparing GO terms and proteins within the context of GO. It uses GO data and UniProt proteins with their GO annotations as provided by the Gene Ontology Annotation (GOA) project to precompute GO term information content (IC), enabling rapid response to user queries. CONCLUSIONS: The DaGO-Fun online tool presents the advantage of integrating all the relevant IC-based GO similarity measures, including topology- and annotation-based approaches to facilitate effective exploration of these measures, thus enabling users to choose the most relevant approach for their application. Furthermore, this tool includes several biological applications related to GO semantic similarity scores, including the retrieval of genes based on their GO annotations, the clustering of functionally related genes within a set, and term enrichment analysis.
Gaston K. Mazandu, Nicola J. Mulder
BMC Bioinform.2
2011 From sets to graphs: towards a realistic enrichment analysis of transcriptomic systems
abstract
MOTIVATION: Current gene set enrichment approaches do not take interactions and associations between set members into account. Mutual activation and inhibition causing positive and negative correlation among set members are thus neglected. As a consequence, inconsistent regulations and contextless expression changes are reported and, thus, the biological interpretation of the result is impeded. RESULTS: We analyzed established gene set enrichment methods and their result sets in a large-scale investigation of 1000 expression datasets. The reported statistically significant gene sets exhibit only average consistency between the observed patterns of differential expression and known regulatory interactions. We present Gene Graph Enrichment Analysis (GGEA) to detect consistently and coherently enriched gene sets, based on prior knowledge derived from directed gene regulatory networks. Firstly, GGEA improves the concordance of pairwise regulation with individual expression changes in respective pairs of regulating and regulated genes, compared with set enrichment methods. Secondly, GGEA yields result sets where a large fraction of relevant expression changes can be explained by nearby regulators, such as transcription factors, again improving on set-based methods. Thirdly, we demonstrate in additional case studies that GGEA can be applied to human regulatory pathways, where it sensitively detects very specific regulation processes, which are altered in tumors of the central nervous system. GGEA significantly increases the detection of gene sets where measured positively or negatively correlated expression patterns coincide with directed inducing or repressing relationships, thus facilitating further interpretation of gene expression data. AVAILABILITY: The method and accompanying visualization capabilities have been bundled into an R package and tied to a grahical user interface, the Galaxy workflow environment, that is running as a web server. CONTACT: [email protected]; [email protected].
Ludwig Geistlinger, Gergely Csaba, Robert Küffner, Nicola J. Mulder, Ralf Zimmer
Bioinform.4
2011 Dasty3, a WEB framework for DAS
abstract
MOTIVATION: Dasty3 is a highly interactive and extensible Web-based framework. It provides a rich Application Programming Interface upon which it is possible to develop specialized clients capable of retrieving information from DAS sources as well as from data providers not using the DAS protocol. Dasty3 provides significant improvements on previous Web-based frameworks and is implemented using the 1.6 DAS specification. AVAILABILITY: Dasty3 is an open-source tool freely available at http://www.ebi.ac.uk/dasty/ under the terms of the GNU General public license. Source and documentation can be found at http://code.google.com/p/dasty/. CONTACT: [email protected].
Jose M. Villaveces, Rafael C. Jiménez, Leyla Jael Castro, Gustavo A. Salazar, Bernat Gel, Nicola J. Mulder, Maria Jesus Martin, Alexander García Castro, Henning Hermjakob
Bioinform.6
2011 Investigating the effect of paralogs on microarray gene-set analysis
abstract
BACKGROUND: In order to interpret the results obtained from a microarray experiment, researchers often shift focus from analysis of individual differentially expressed genes to analyses of sets of genes. These gene-set analysis (GSA) methods use previously accumulated biological knowledge to group genes into sets and then aim to rank these gene sets in a way that reflects their relative importance in the experimental situation in question. We suspect that the presence of paralogs affects the ability of GSA methods to accurately identify the most important sets of genes for subsequent research. RESULTS: We show that paralogs, which typically have high sequence identity and similar molecular functions, also exhibit high correlation in their expression patterns. We investigate this correlation as a potential confounding factor common to current GSA methods using Indygene http://www.cbio.uct.ac.za/indygene, a web tool that reduces a supplied list of genes so that it includes no pairwise paralogy relationships above a specified sequence similarity threshold. We use the tool to reanalyse previously published microarray datasets and determine the potential utility of accounting for the presence of paralogs. CONCLUSIONS: The Indygene tool efficiently removes paralogy relationships from a given dataset and we found that such a reduction, performed prior to GSA, has the ability to generate significantly different results that often represent novel and plausible biological hypotheses. This was demonstrated for three different GSA approaches when applied to the reanalysis of previously published microarray datasets and suggests that the redundancy and non-independence of paralogs is an important consideration when dealing with GSA methodologies.
Andre J. Faure, Cathal Seoighe, Nicola J. Mulder
BMC Bioinform.3
2011 DAS Writeback: A Collaborative Annotation System
abstract
BACKGROUND: Centralised resources such as GenBank and UniProt are perfect examples of the major international efforts that have been made to integrate and share biological information. However, additional data that adds value to these resources needs a simple and rapid route to public access. The Distributed Annotation System (DAS) provides an adequate environment to integrate genomic and proteomic information from multiple sources, making this information accessible to the community. DAS offers a way to distribute and access information but it does not provide domain experts with the mechanisms to participate in the curation process of the available biological entities and their annotations. RESULTS: We designed and developed a Collaborative Annotation System for proteins called DAS Writeback. DAS writeback is a protocol extension of DAS to provide the functionalities of adding, editing and deleting annotations. We implemented this new specification as extensions of both a DAS server and a DAS client. The architecture was designed with the involvement of the DAS community and it was improved after performing usability experiments emulating a real annotation task. CONCLUSIONS: We demonstrate that DAS Writeback is effective, usable and will provide the appropriate environment for the creation and evolution of community protein annotation.
Gustavo A. Salazar, Rafael C. Jiménez, Alexander García Castro, Henning Hermjakob, Nicola J. Mulder, Edwin H. Blake
BMC Bioinform.5
2002 Applications of InterPro in Protein Annotation and Genome Analysis
abstract
The applications of InterPro span a range of biologically important areas that includes automatic annotation of protein sequences and genome analysis. In automatic annotation of protein sequences InterPro has been utilised to provide reliable characterisation of sequences, identifying them as candidates for functional annotation. Rules based on the InterPro characterisation are stored and operated through a database called RuleBase. RuleBase is used as the main tool in the sequence database group at the EBI to apply automatic annotation to unknown sequences. The annotated sequences are stored and distributed in the TrEMBL protein sequence database. InterPro also provides a means to carry out statistical and comparative analyses of whole genomes. In the Proteome Analysis Database, InterPro analyses have been combined with other analyses based on CluSTr, the Gene Ontology (GO) and structural information on the proteins.
Margaret Biswas, Joseph F. O'Rourke, Evelyn Camon, Gillian Fraser, Alexander Kanapin, Youla Karavidopoulou, Paul J. Kersey, Evgenia V. Kriventseva, Virginie Mittard, Nicola J. Mulder, Isabelle Phan, Florence Servant, Rolf Apweiler
Briefings Bioinform.10
2002 InterPro: An Integrated Documentation Resource for Protein Families, Domains and Functional Sites
abstract
The exponential increase in the submission of nucleotide sequences to the nucleotide sequence database by genome sequencing centres has resulted in a need for rapid, automatic methods for classification of the resulting protein sequences. There are several signature and sequence cluster-based methods for protein classification, each resource having distinct areas of optimum application owing to the differences in the underlying analysis methods. In recognition of this, InterPro was developed as an integrated documentation resource for protein families, domains and functional sites, to rationalise the complementary efforts of the individual protein signature database projects. The member databases - PRINTS, PROSITE, Pfam, ProDom, SMART and TIGRFAMs - form the InterPro core. Related signatures from each member database are unified into single InterPro entries. Each InterPro entry includes a unique accession number, functional descriptions and literature references, and links are made back to the relevant member database(s). Release 4.0 of InterPro (November 2001) contains 4,691 entries, representing 3,532 families, 1,068 domains, 74 repeats and 15 sites of post-translational modification (PTMs) encoded by different regular expressions, profiles, fingerprints and hidden Markov models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (2,141,621 InterPro hits from 586,124 SWISS-PROT and TrEMBL protein sequences). The database is freely accessible for text- and sequence-based searches.
Nicola J. Mulder, Rolf Apweiler, Terri K. Attwood, Amos Bairoch, Alex Bateman, David Binns, Margaret Biswas, Paul Bradley, Peer Bork, Philipp Bucher, Richard R. Copley, Emmanuel Courcelle, Richard Durbin, Laurent Falquet, Wolfgang Fleischmann, Jérôme Gouzy, Sam Griffiths-Jones, Daniel H. Haft, Henning Hermjakob, Nicolas Hulo, Daniel Kahn, Alexander Kanapin, Maria Krestyaninova, Rodrigo Lopez, Ivica Letunic, Sue Orchard, Marco Pagni, David Peyruc, Chris P. Ponting, Florence Servant, Christian J. A. Sigrist
Briefings Bioinform.1
2002 Interactive InterPro-based comparisons of proteins in whole genomes
abstract
Abstract Motivation: The SWISS-PROT group at the EBI has developed the Proteome Analysis Database utilizing existing resources and providing comprehensive and integrated comparative analysis of the predicted protein coding sequencesof the complete genomes of bacteria, archaea and eukaryotes. The Proteome Analysis Database is accompanied by a program that has been designed to carry out interactive InterPro proteome comparisons for any one proteome against any other one or more of the proteomes in the database. Availability: http://www.ebi.ac.uk/proteome/comparisons.html Contact: [email protected]; [email protected] * To whom all correspondence should be addressed.
Alexander Kanapin, Rolf Apweiler, Margaret Biswas, Wolfgang Fleischmann, Youla Karavidopoulou, Paul J. Kersey, Evgenia V. Kriventseva, Virginie Mittard, Nicola J. Mulder, Thomas M. Oinn, Isabelle Phan, Florence Servant, Evgeny M. Zdobnov
Bioinform.9
2000 InterPro-an integrated documentation resource for protein families, domains and functional sites
abstract
MOTIVATION: InterPro is a new integrated documentation resource for protein families, domains and functional sites, developed initially as a means of rationalising the complementary efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. RESULTS: Merged annotations from PRINTS, PROSITE and Pfam form the InterPro core. Each combined InterPro entry includes functional descriptions and literature references, and links are made back to the relevant parent database(s), allowing users to see at a glance whether a particular family or domain has associated patterns, profiles, fingerprints, etc. Merged and individual entries (i.e. those that have no counterpart in the companion resources) are assigned unique accession numbers. Release 1.2 of InterPro (June 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification (PTMs) encoded by 6581 different regular expressions, profiles, fingerprints and Hidden Markov Models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1000000 hits from 264333 different proteins out of 384572 in SWISS-PROT and TrEMBL).
Rolf Apweiler, Terri K. Attwood, Amos Bairoch, Alex Bateman, Ewan Birney, Margaret Biswas, Philipp Bucher, Lorenzo Cerutti, Florence Corpet, Michael D. R. Croning, Richard Durbin, Laurent Falquet, Wolfgang Fleischmann, Jérôme Gouzy, Henning Hermjakob, Nicolas Hulo, Inge Jonassen, Daniel Kahn, Alexander Kanapin, Youla Karavidopoulou, Rodrigo Lopez, Beate Marx, Nicola J. Mulder, Thomas M. Oinn, Marco Pagni, Florence Servant, Christian J. A. Sigrist, Evgeny M. Zdobnov
Bioinform.23