VLDB 2026 Research / reviewers in the wild / expert
Emile R. Chimusa
dblp:173/3655
· DBLP profile ↗
17ranked-venue papers
4as first author
5since 2021 · last 2021
0000-0001-8846-2047ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Simulation of African and non-African low and high coverage whole genome sequence data to assess variant calling approachesabstractCurrent variant calling (VC) approaches have been designed to leverage populations of long-range haplotypes and were benchmarked using populations of European descent, whereas most genetic diversity is found in non-European such as Africa populations. Working with these genetically diverse populations, VC tools may produce false positive and false negative results, which may produce misleading conclusions in prioritization of mutations, clinical relevancy and actionability of genes. The most prominent question is which tool or pipeline has a high rate of sensitivity and precision when analysing African data with either low or high sequence coverage, given the high genetic diversity and heterogeneity of this data. Here, a total of 100 synthetic Whole Genome Sequencing (WGS) samples, mimicking the genetics profile of African and European subjects for different specific coverage levels (high/low), have been generated to assess the performance of nine different VC tools on these contrasting datasets. The performances of these tools were assessed in false positive and false negative call rates by comparing the simulated golden variants to the variants identified by each VC tool. Combining our results on sensitivity and positive predictive value (PPV), VarDict [PPV = 0.999 and Matthews correlation coefficient (MCC) = 0.832] and BCFtools (PPV = 0.999 and MCC = 0.813) perform best when using African population data on high and low coverage data. Overall, current VC tools produce high false positive and false negative rates when analysing African compared with European data. This highlights the need for development of VC approaches with high sensitivity and precision tailored for populations characterized by high genetic variations and low linkage disequilibrium. Shatha Alosaimi, Noëlle van Biljon, Denis Awany, Prisca K. Thami, Joel Defo, Jacquiline W. Mugo, Christian D. Bope, Gaston K. Mazandu, Nicola J. Mulder, Emile R. Chimusa |
Briefings Bioinform. | 10 |
| 2021 | Heritability jointly explained by host genotype and microbiome: will improve traits prediction?abstractAs we observe the $70$th anniversary of the publication by Robertson that formalized the notion of 'heritability', geneticists remain puzzled by the problem of missing/hidden heritability, where heritability estimates from genome-wide association studies (GWASs) fall short of that from twin-based studies. Many possible explanations have been offered for this discrepancy, including existence of genetic variants poorly captured by existing arrays, dominance, epistasis and unaccounted-for environmental factors; albeit these remain controversial. We believe a substantial part of this problem could be solved or better understood by incorporating the host's microbiota information in the GWAS model for heritability estimation and may also increase human traits prediction for clinical utility. This is because, despite empirical observations such as (i) the intimate role of the microbiome in many complex human phenotypes, (ii) the overlap between genetic variants associated with both microbiome attributes and complex diseases and (iii) the existence of heritable bacterial taxa, current GWAS models for heritability estimate do not take into account the contributory role of the microbiome. Furthermore, heritability estimate from twin-based studies does not discern microbiome component of the observed total phenotypic variance. Here, we summarize the concept of heritability in GWAS and microbiome-wide association studies, focusing on its estimation, from a statistical genetics perspective. We then discuss a possible statistical method to incorporate the microbiome in the estimation of heritability in host GWAS. Denis Awany, Emile R. Chimusa |
Briefings Bioinform. | 2 |
| 2021 | Drug response in association with pharmacogenomics and pharmacomicrobiomics: towards a better personalized medicineabstractResearchers have long been presented with the challenge imposed by the role of genetic heterogeneity in drug response. For many years, Pharmacogenomics and pharmacomicrobiomics has been investigating the influence of an individual's genetic background to drug response and disposition. More recently, the human gut microbiome has proven to play a crucial role in the way patients respond to different therapeutic drugs and it has been shown that by understanding the composition of the human microbiome, we can improve the drug efficacy and effectively identify drug targets. However, our knowledge on the effect of host genetics on specific gut microbes related to variation in drug metabolizing enzymes, the drug remains limited and therefore limits the application of joint host-microbiome genome-wide association studies. In this paper, we provide a historical overview of the complex interactions between the host, human microbiome and drugs. While discussing applications, challenges and opportunities of these studies, we draw attention to the critical need for inclusion of diverse populations and the development of an innovative and combined pharmacogenomics and pharmacomicrobiomics approach, that may provide an important basis in personalized medicine. Radia Hassan, Imane Allali, Francis E. Agamah, Samar S. M. Elsheikh, Nicholas E. Thomford, Collet Dandara, Emile R. Chimusa |
Briefings Bioinform. | 7 |
| 2021 | Reviewing and assessing existing meta-analysis models and toolsabstractOver the past few years, meta-analysis has become popular among biomedical researchers for detecting biomarkers across multiple cohort studies with increased predictive power. Combining datasets from different sources increases sample size, thus overcoming the issue related to limited sample size from each individual study and boosting the predictive power. This leads to an increased likelihood of more accurately predicting differentially expressed genes/proteins or significant biomarkers underlying the biological condition of interest. Currently, several meta-analysis methods and tools exist, each having its own strengths and limitations. In this paper, we survey existing meta-analysis methods, and assess the performance of different methods based on results from different datasets as well as assessment from prior knowledge of each method. This provides a reference summary of meta-analysis models and tools, which helps to guide end-users on the choice of appropriate models or tools for given types of datasets and enables developers to consider current advances when planning the development of new meta-analysis models and more practical integrative tools. Funmilayo L. Makinde, Milaine S. S. Tchamga, James Jafali, Segun A. Fatumo, Emile R. Chimusa, Nicola J. Mulder, Gaston K. Mazandu |
Briefings Bioinform. | 5 |
| 2021 | IHP-PING - generating integrated human protein-protein interaction networks on-the-flyabstractAdvances in high-throughput sequencing technologies have resulted in an exponential growth of publicly accessible biological datasets. In the 'big data' driven 'post-genomic' context, much work is being done to explore human protein-protein interactions (PPIs) for a systems level based analysis to uncover useful signals and gain more insights to advance current knowledge and answer specific biological and health questions. These PPIs are experimentally or computationally predicted, stored in different online databases and some of PPI resources are updated regularly. As with many biological datasets, such regular updates continuously render older PPI datasets potentially outdated. Moreover, while many of these interactions are shared between these online resources, each resource includes its own identified PPIs and none of these databases exhaustively contains all existing human PPI maps. In this context, it is essential to enable the integration of or combining interaction datasets from different resources, to generate a PPI map with increased coverage and confidence. To allow researchers to produce an integrated human PPI datasets in real-time, we introduce the integrated human protein-protein interaction network generator (IHP-PING) tool. IHP-PING is a flexible python package which generates a human PPI network from freely available online resources. This tool extracts and integrates heterogeneous PPI datasets to generate a unified PPI network, which is stored locally for further applications. Gaston K. Mazandu, Christopher Hooper, Kenneth Opap, Funmilayo L. Makinde, Victoria Nembaware, Nicholas E. Thomford, Emile R. Chimusa, Ambroise Wonkam, Nicola J. Mulder |
Briefings Bioinform. | 7 |
| 2020 | Computational/in silico methods in drug target and lead predictionabstractDrug-like compounds are most of the time denied approval and use owing to the unexpected clinical side effects and cross-reactivity observed during clinical trials. These unexpected outcomes resulting in significant increase in attrition rate centralizes on the selected drug targets. These targets may be disease candidate proteins or genes, biological pathways, disease-associated microRNAs, disease-related biomarkers, abnormal molecular phenotypes, crucial nodes of biological network or molecular functions. This is generally linked to several factors, including incomplete knowledge on the drug targets and unpredicted pharmacokinetic expressions upon target interaction or off-target effects. A method used to identify targets, especially for polygenic diseases, is essential and constitutes a major bottleneck in drug development with the fundamental stage being the identification and validation of drug targets of interest for further downstream processes. Thus, various computational methods have been developed to complement experimental approaches in drug discovery. Here, we present an overview of various computational methods and tools applied in predicting or validating drug targets and drug-like molecules. We provide an overview on their advantages and compare these methods to identify effective methods which likely lead to optimal results. We also explore major sources of drug failure considering the challenges and opportunities involved. This review might guide researchers on selecting the most efficient approach or technique during the computational drug discovery process. Francis E. Agamah, Gaston K. Mazandu, Radia Hassan, Christian D. Bope, Nicholas E. Thomford, Anita Ghansah, Emile R. Chimusa |
Briefings Bioinform. | 7 |
| 2020 | Dating admixture events is unsolved problem in multi-way admixed populationsabstractAdvances in human sequencing technologies, coupled with statistical and computational tools, have fostered the development of methods for dating admixture events. These methods have merits and drawbacks in estimating admixture events in multi-way admixed populations. Here, we first provide a comprehensive review and comparison of current methods pertinent to dating admixture events. Second, we assess various admixture dating tools. We do so by performing various simulations. Third, we apply the top two assessed methods to real data of a uniquely admixed population from South Africa. Results reveal that current dating admixture models are not sufficiently equipped to estimate ancient admixtures events and to identify multi-faceted admixture events in complex multi-way admixed populations. We conclude with a discussion of research areas where further work on dating admixture-based methods is needed. Emile R. Chimusa, Joel Defo, Prisca K. Thami, Denis Awany, Delesa D. Mulisa, Imane Allali, Hassan Ghazal, Ahmed Moussa, Gaston K. Mazandu |
Briefings Bioinform. | 1 |
| 2020 | FRANC: a unified framework for multi-way local ancestry deconvolution with high density SNP dataabstractAbstract Several thousand genomes have been completed with millions of variants identified in the human deoxyribonucleic acid sequences. These genomic variations, especially those introduced by admixture, significantly contribute to a remarkable phenotypic variability with medical and/or evolutionary implications. Elucidating local ancestry estimates is necessary for a better understanding of genomic variation patterns throughout modern human evolution and adaptive processes, and consequences in human heredity and health. However, existing local ancestry deconvolution tools are accessible as individual scripts, each requiring input and producing output in its own complex format. This limits the user’s ability to retrieve local ancestry estimates. We introduce a unified framework for multi-way local ancestry inference, FRANC, integrating eight existing state-of-the-art local ancestry deconvolution tools. FRANC is an adaptable, expandable and portable tool that manipulates tool-specific inputs, deconvolutes ancestry and standardizes tool-specific results. To facilitate both medical and population genetics studies, FRANC requires convenient and easy to manipulate input files and allows users to choose output formats to ease their use in further potential local ancestry deconvolution applications. Ephifania Geza, Nicola J. Mulder, Emile R. Chimusa, Gaston K. Mazandu |
Briefings Bioinform. | 3 |
| 2019 | Post genome-wide association analysis: dissecting computational pathway/network-based approachesabstractOver thousands of genetic associations to diseases have been identified by genome-wide association studies (GWASs), which conceptually is a single-marker-based approach. There are potentially many uses of these identified variants, including a better understanding of the pathogenesis of diseases, new leads for studying underlying risk prediction and clinical prediction of treatment. However, because of inadequate power, GWAS might miss disease genes and/or pathways with weak genetic or strong epistatic effects. Driven by the need to extract useful information from GWAS summary statistics, post-GWAS approaches (PGAs) were introduced. Here, we dissect and discuss advances made in pathway/network-based PGAs, with a particular focus on protein-protein interaction networks that leverage GWAS summary statistics by combining effects of multiple loci, subnetworks or pathways to detect genetic signals associated with complex diseases. We conclude with a discussion of research areas where further work on summary statistic-based methods is needed. Emile R. Chimusa, Shareefa Dalvie, Collet Dandara, Ambroise Wonkam, Gaston K. Mazandu |
Briefings Bioinform. | 1 |
| 2019 | A comprehensive survey of models for dissecting local ancestry deconvolution in human genomeabstractOver the past decade, studies of admixed populations have increasingly gained interest in both medical and population genetics. These studies have so far shed light on the patterns of genetic variation throughout modern human evolution and have improved our understanding of the demographics and adaptive processes of human populations. To date, there exist about 20 methods or tools to deconvolve local ancestry. These methods have merits and drawbacks in estimating local ancestry in multiway admixed populations. In this article, we survey existing ancestry deconvolution methods, with special emphasis on multiway admixture, and compare these methods based on simulation results reported by different studies, computational approaches used, including mathematical and statistical models, and biological challenges related to each method. This should orient users on the choice of an appropriate method or tool for given population admixture characteristics and update researchers on current advances, challenges and opportunities behind existing ancestry deconvolution methods. Ephifania Geza, Jacquiline W. Mugo, Nicola J. Mulder, Ambroise Wonkam, Emile R. Chimusa, Gaston K. Mazandu |
Briefings Bioinform. | 5 |
| 2018 | Large-scale data-driven integrative framework for extracting essential targets and processes from disease-associated gene data setsabstractPopulations worldwide currently face several public health challenges, including growing prevalence of infections and the emergence of new pathogenic organisms. The cost and risk associated with drug development make the development of new drugs for several diseases, especially orphan or rare diseases, unappealing to the pharmaceutical industry. Proof of drug safety and efficacy is required before market approval, and rigorous testing makes the drug development process slow, expensive and frequently result in failure. This failure is often because of the use of irrelevant targets identified in the early steps of the drug discovery process, suggesting that target identification and validation are cornerstones for the success of drug discovery and development. Here, we present a large-scale data-driven integrative computational framework to extract essential targets and processes from an existing disease-associated data set and enhance target selection by leveraging drug-target-disease association at the systems level. We applied this framework to tuberculosis and Ebola virus diseases combining heterogeneous data from multiple sources, including protein-protein functional interaction, functional annotation and pharmaceutical data sets. Results obtained demonstrate the effectiveness of the pipeline, leading to the extraction of essential drug targets and to the rational use of existing approved drugs. This provides an opportunity to move toward optimal target-based strategies for screening available drugs and for drug discovery. There is potential for this model to bridge the gap in the production of orphan disease therapies, offering a systematic approach to predict new uses for existing drugs, thereby harnessing their full therapeutic potential. Gaston K. Mazandu, Emile R. Chimusa, Kayleigh Rutherford, Elsa-Gayle Zekeng, Zoe Z. Gebremariam, Maryam Y. Onifade, Nicola J. Mulder |
Briefings Bioinform. | 2 |
| 2017 | Gene Ontology semantic similarity tools: survey on features and challenges for biological knowledge discoveryabstractGene Ontology (GO) semantic similarity tools enable retrieval of semantic similarity scores, which incorporate biological knowledge embedded in the GO structure for comparing or classifying different proteins or list of proteins based on their GO annotations. This facilitates a better understanding of biological phenomena underlying the corresponding experiment and enables the identification of processes pertinent to different biological conditions. Currently, about 14 tools are available, which may play an important role in improving protein analyses at the functional level using different GO semantic similarity measures. Here we survey these tools to provide a comprehensive view of the challenges and advances made in this area to avoid redundant effort in developing features that already exist, or implementing ideas already proven to be obsolete in the context of GO. This helps researchers, tool developers, as well as end users, understand the underlying semantic similarity measures implemented through knowledge of pertinent features of, and issues related to, a particular tool. This should empower users to make appropriate choices for their biological applications and ensure effective knowledge discovery based on GO annotations. Gaston K. Mazandu, Emile R. Chimusa, Nicola J. Mulder |
Briefings Bioinform. | 2 |
| 2017 | A multi-scenario genome-wide medical population genetics simulation frameworkabstractMOTIVATION: Recent technological advances in high-throughput sequencing and genotyping have facilitated an improved understanding of genomic structure and disease-associated genetic factors. In this context, simulation models can play a critical role in revealing various evolutionary and demographic effects on genomic variation, enabling researchers to assess existing and design novel analytical approaches. Although various simulation frameworks have been suggested, they do not account for natural selection in admixture processes. Most are tailored to a single chromosome or a genomic region, very few capture large-scale genomic data, and most are not accessible for genomic communities. RESULTS: Here we develop a multi-scenario genome-wide medical population genetics simulation framework called 'FractalSIM'. FractalSIM has the capability to accurately mimic and generate genome-wide data under various genetic models on genetic diversity, genomic variation affecting diseases and DNA sequence patterns of admixed and/or homogeneous populations. Moreover, the framework accounts for natural selection in both homogeneous and admixture processes. The outputs of FractalSIM have been assessed using popular tools, and the results demonstrated its capability to accurately mimic real scenarios. They can be used to evaluate the performance of a range of genomic tools from ancestry inference to genome-wide association studies. AVAILABILITY AND IMPLEMENTATION: The FractalSIM package is available at http://www.cbio.uct.ac.za/FractalSIM. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jacquiline W. Mugo, Ephifania Geza, Joel Defo, Samar S. M. Elsheikh, Gaston K. Mazandu, Nicola J. Mulder, Emile R. Chimusa |
Bioinform. | 7 |
| 2017 | Assessing computational genomics skills: Our experience in the H3ABioNet African bioinformatics networkabstractThe H3ABioNet pan-African bioinformatics network, which is funded to support the Human Heredity and Health in Africa (H3Africa) program, has developed node-assessment exercises to gauge the ability of its participating research and service groups to analyze typical genome-wide datasets being generated by H3Africa research groups. We describe a framework for the assessment of computational genomics analysis skills, which includes standard operating procedures, training and test datasets, and a process for administering the exercise. We present the experiences of 3 research groups that have taken the exercise and the impact on their ability to manage complex projects. Finally, we discuss the reasons why many H3ABioNet nodes have declined so far to participate and potential strategies to encourage them to do so. C. Victor Jongeneel, Ovokeraye Achinike-Oduaran, Ezekiel F. Adebiyi, Marion O. Adebiyi, Seun Adeyemi, Bola Akanle, Shaun Aron, Efejiro Ashano, Hocine Bendou, Gerrit Botha, Emile R. Chimusa, Ananyo Choudhury, Ravikiran Donthu, Jenny Drnevich, Oluwadamilare Falola, Christopher J. Fields, Scott Hazelhurst, Liesl Hendry, Itunuoluwa Isewon, Radhika S. Khetani, Judit Kumuthini, Magambo Phillip Kimuda, Lerato E. Magosi, Liudmila S. Mainzer, Suresh Maslamoney, Mamana Mbiyavanga, Ayton Meintjes, Danny Mugutso, Phelelani T. Mpangase, Richard Munthali, Victoria Nembaware, Andrew Ndhlovu, Trust Odia, Adaobi Okafor, Olaleye Oladipo, Sumir Panji, Venesa Pillay, Gloria Rendon, Dhriti Sengupta, Nicola J. Mulder |
PLoS Comput. Biol. | 11 |
| 2016 | ancGWAS: a post genome-wide association study method for interaction, pathway and ancestry analysis in homogeneous and admixed populationsabstractMOTIVATION: Despite numerous successful Genome-wide Association Studies (GWAS), detecting variants that have low disease risk still poses a challenge. GWAS may miss disease genes with weak genetic effects or strong epistatic effects due to the single-marker testing approach commonly used. GWAS may thus generate false negative or inconclusive results, suggesting the need for novel methods to combine effects of single nucleotide polymorphisms within a gene to increase the likelihood of fully characterizing the susceptibility gene. RESULTS: We developed ancGWAS, an algebraic graph-based centrality measure that accounts for linkage disequilibrium in identifying significant disease sub-networks by integrating the association signal from GWAS data sets into the human protein-protein interaction (PPI) network. We validated ancGWAS using an association study result from a breast cancer data set and the simulation of interactive disease loci in the simulation of a complex admixed population, as well as pathway-based GWAS simulation. This new approach holds promise for deconvoluting the interactions between genes underlying the pathogenesis of complex diseases. Results obtained yield a novel central breast cancer sub-network of the human interactome implicated in the proteoglycan syndecan-mediated signaling events pathway which is known to play a major role in mesenchymal tumor cell proliferation, thus providing further insights into breast cancer pathogenesis. AVAILABILITY AND IMPLEMENTATION: The ancGWAS package and documents are available at http://www.cbio.uct.ac.za/~emile/software.html. Emile R. Chimusa, Mamana Mbiyavanga, Gaston K. Mazandu, Nicola J. Mulder |
Bioinform. | 1 |
| 2016 | A-DaGO-Fun: an adaptable Gene Ontology semantic similarity-based functional analysis toolabstractSUMMARY: Gene Ontology (GO) semantic similarity measures are being used for biological knowledge discovery based on GO annotations by integrating biological information contained in the GO structure into data analyses. To empower users to quickly compute, manipulate and explore these measures, we introduce A-DaGO-Fun (ADaptable Gene Ontology semantic similarity-based Functional analysis). It is a portable software package integrating all known GO information content-based semantic similarity measures and relevant biological applications associated with these measures. A-DaGO-Fun has the advantage not only of handling datasets from the current high-throughput genome-wide applications, but also allowing users to choose the most relevant semantic similarity approach for their biological applications and to adapt a given module to their needs. AVAILABILITY AND IMPLEMENTATION: A-DaGO-Fun is freely available to the research community at http://web.cbio.uct.ac.za/ITGOM/adagofun. It is implemented in Linux using Python under free software (GNU General Public Licence). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gaston K. Mazandu, Emile R. Chimusa, Mamana Mbiyavanga, Nicola J. Mulder |
Bioinform. | 2 |
| 2015 | "Broadband" Bioinformatics Skills Transfer with the Knowledge Transfer Programme (KTP): Educational Model for Upliftment and Sustainable DevelopmentabstractA shortage of practical skills and relevant expertise is possibly the primary obstacle to social upliftment and sustainable development in Africa. The "omics" fields, especially genomics, are increasingly dependent on the effective interpretation of large and complex sets of data. Despite abundant natural resources and population sizes comparable with many first-world countries from which talent could be drawn, countries in Africa still lag far behind the rest of the world in terms of specialized skills development. Moreover, there are serious concerns about disparities between countries within the continent. The multidisciplinary nature of the bioinformatics field, coupled with rare and depleting expertise, is a critical problem for the advancement of bioinformatics in Africa. We propose a formalized matchmaking system, which is aimed at reversing this trend, by introducing the Knowledge Transfer Programme (KTP). Instead of individual researchers travelling to other labs to learn, researchers with desirable skills are invited to join African research groups for six weeks to six months. Visiting researchers or trainers will pass on their expertise to multiple people simultaneously in their local environments, thus increasing the efficiency of knowledge transference. In return, visiting researchers have the opportunity to develop professional contacts, gain industry work experience, work with novel datasets, and strengthen and support their ongoing research. The KTP develops a network with a centralized hub through which groups and individuals are put into contact with one another and exchanges are facilitated by connecting both parties with potential funding sources. This is part of the PLOS Computational Biology Education collection. Emile R. Chimusa, Mamana Mbiyavanga, Velaphi Masilela, Judit Kumuthini |
PLoS Comput. Biol. | 1 |