VLDB 2026 Research / reviewers in the wild / expert
Peter N. Robinson
dblp:34/3156
· DBLP profile ↗
39ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0002-0736-9199ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 36 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tabular implicit deep neural networks ensembles for the prediction of pathogenic genetic variants in Mendelian diseasesabstractMendelian genetic diseases comprise approximately described disorders, yet the genetic basis is known only for about half of them, and a molecular diagnosis often remains difficult or unresolved. In this context, machine learning methods play a key role. However, identifying pathogenic variants in non-coding regions of the human genome is particularly challenging, as they are vastly outnumbered by neutral variants. This extreme imbalance causes standard machine learning approaches to exhibit a strong predictive bias towards the majority class, significantly limiting their sensitivity. Building on recent advances in imbalance-aware and ensemble learning methods, we propose two novel deep learning models for predicting pathogenic non-coding variants in Mendelian diseases. The first model, T-ResNet (Tabular Residual Neural Network), adopts a modular architecture with residual connections, along with a mini-batch balancing strategy to address class imbalance. This design simplifies hyperparameter optimization while mitigating vanishing-gradient effects. The second model, TIDE-Var (Tabular Implicit Deep neural network Ensembles for Variant prediction), leverages the TabM and BatchEnsemble models to build an implicit ensemble of deep neural networks trained jointly by minimizing a common objective function, and partially sharing learning parameters. Learner-specific adapter parameters promote diversity among the implicit base learners while keeping the ensemble computationally efficient. Genome-wide experiments show that T-ResNet achieves an average Area Under the Precision-Recall Curve (AUPRC) of , comparable to that of HyperSMURF , a state-of-the-art method. In contrast, TIDE-Var yields significantly better results than HyperSMURF (AUPRC ) and slightly better than XGBoost , one of the top-methods for the classification of tabular data. Ablation studies confirm that regularized joint ensemble learning with partially shared and base-learner specific adapter learning parameters are key factors in achieving high predictive performance, essential to improve the diagnostic yield for patients with rare genetic diseases. Federico Stacchietti, Marco Nicolini, Leonardo Chimirri, Peter N. Robinson, Elena Casiraghi, Giorgio Valentini |
Neurocomputing | 4 |
| 2025 | Biasing second-order random walk sampling for heterogeneous graph embedding *abstractWe present heterogeneous-node2vec, a novel method that leverages the well-known node2vec algorithm to enable the generation of random-walk samples in a heterogeneous context. Specifically, we propose a strategy to bias the random walk, enabling type-aware transitions between different node and edge types. We evaluate the proposed technique on node-label prediction tasks, applied to various real-world, complex networks. A comparison with state-of-the-art techniques for heterogeneous graph embedding demonstrates that our strategy achieves competitive results for node-label prediction. This evidences that graph representation methods based on heterogeneous random-walk sampling can attain strong performance on standard supervised tasks when the sampling procedure incorporates the semantic information defined by the type heterogeneity of entities within the graph. This approach provides an effective and scalable solution for representing and learning from complex heterogeneous graphs. Mauricio Soto Gomez, Carlos Cano, Justin T. Reese, Peter N. Robinson, Giorgio Valentini, Elena Casiraghi |
IJCNN | 4 |
| 2025 | Intrinsic-dimension analysis for guiding dimensionality reduction and data fusion in multi-omics data processingabstractMulti-omics data have revolutionized biomedical research by providing a comprehensive understanding of biological systems and the molecular mechanisms of disease development. However, analyzing multi-omics data is challenging due to high dimensionality and limited sample sizes, necessitating proper data-reduction pipelines to ensure reliable analyses. Additionally, its multimodal nature requires effective data-integration pipelines. While several dimensionality reduction and data fusion algorithms have been proposed, crucial aspects are often overlooked. Specifically, the choice of projection space dimension is typically heuristic and uniformly applied across all omics, neglecting the unique high dimension small sample size challenges faced by individual omics. This paper introduces a novel multi-modal dimensionality reduction pipeline tailored to individual views. By leveraging intrinsic dimensionality estimators, we assess the curse-of-dimensionality impact on each view and propose a two-step reduction strategy for significantly affected views, combining feature selection with feature extraction. Compared to traditional uniform reduction pipelines in a crucial and supervised multi-omics analysis setting, our approach shows significant improvement. Additionally, we explore three effective unsupervised multi-omics data fusion methods rooted in the main data fusion strategies to gain insights into their performance under crucial, yet overlooked, settings. Jessica Gliozzo, Mauricio Soto Gomez, Valentina Guarino, Arturo Bonometti, Alberto Cabri, Emanuele Cavalleri, Justin T. Reese, Peter N. Robinson, Marco Mesiti, Giorgio Valentini, Elena Casiraghi |
Artif. Intell. Medicine | 8 |
| 2025 | Oncopacket: integration of cancer research data using GA4GH phenopacketsabstractSUMMARY: Lack of data integration remains a significant impediment to cancer research, and many analyses still require customized software to transform and prepare cancer data. We describe a software package to harmonize genetic and clinical cancer data into the GA4GH Phenopacket schema, an ISO standard for representing clinical case data. We integrated demographic, mutation, morphology, diagnosis, intervention, and survival data using case data from the National Cancer Institute for 12 cancer types. The Phenopacket standard provides a foundation for downstream use, including sophisticated statistical and AI/ML analyses. We demonstrate fitness for purpose by using the integrated data to recapitulate a known association between mutations in the gene encoding isocitrate dehydrogenase 1 and survival time in brain cancer patients. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at: https://github.com/monarch-initiative/oncopacket (archived at 10.5281/zenodo.15353125). Michael L. Sierk, Daniel Danis, Sujay Patil, Nobal Kishor, Rajdeep Mondal, Qingrong Chen, Chunhua Yan, Monica C. Munoz-Torres, Daoud M. Meerzaman, Peter N. Robinson, Justin T. Reese |
Bioinform. | 11 |
| 2025 | Towards a standard benchmark for phenotype-driven variant and gene prioritisation algorithms: PhEval - Phenotypic inference Evaluation frameworkabstractBACKGROUND: Computational approaches to support rare disease diagnosis are challenging to build, requiring the integration of complex data types such as ontologies, gene-to-phenotype associations, and cross-species data into variant and gene prioritisation algorithms (VGPAs). However, the performance of VGPAs has been difficult to measure and is impacted by many factors, for example, ontology structure, annotation completeness or changes to the underlying algorithm. Assertions of the capabilities of VGPAs are often not reproducible, in part because there is no standardised, empirical framework and openly available patient data to assess the efficacy of VGPAs-ultimately hindering the development of effective prioritisation tools. RESULTS: In this paper, we present our benchmarking tool, PhEval, which aims to provide a standardised and empirical framework to evaluate phenotype-driven VGPAs. The inclusion of standardised test corpora and test corpus generation tools in the PhEval suite of tools allows open benchmarking and comparison of methods on standardised data sets. CONCLUSIONS: PhEval and the standardised test corpora solve the issues of patient data availability and experimental tooling configuration when benchmarking and comparing rare disease VGPAs. By providing standardised data on patient cohorts from real-world case-reports and controlling the configuration of evaluated VGPAs, PhEval enables transparent, portable, comparable and reproducible benchmarking of VGPAs. As these tools are often a key component of many rare disease diagnostic pipelines, a thorough and standardised method of assessment is essential for improving patient diagnosis and care. Yasemin Bridges, Vinicius de Souza, Katherina G. Cortes, Melissa A. Haendel, Nomi L. Harris, Daniel R. Korn, Nikolaos M. Marinakis, Nicolas Matentzoglu, James Alastair McLaughlin, Chris Mungall, Aaron Odell, David Osumi-Sutherland, Peter N. Robinson, Damian Smedley, Julius O. B. Jacobsen |
BMC Bioinform. | 13 |
| 2025 | National COVID Cohort Collaborative data enhancements: a path for expanding common data modelsabstractOBJECTIVE: To support long COVID research in National COVID Cohort Collaborative (N3C), the N3C Phenotype and Data Acquisition team created data designs to aid contributing sites in enhancing their data. Enhancements include long COVID specialty clinic indicator; Admission, Discharge, and Transfer transactions; patient-level social determinants of health; and in-hospital use of oxygen supplementation. MATERIALS AND METHODS: For each enhancement, we defined the scope and wrote guidance on how to prepare and populate the data in a standardized way. RESULTS: As of June 2024, 29 sites have added at least one data enhancement to their N3C pipeline. DISCUSSION: The use of common data models is critical to the success of N3C; however, these data models cannot account for all needs. Project-driven data enhancement is required. This should be done in a standardized way in alignment with common data model specifications. Our approach offers a useful pathway for enhancing data to improve fit for purpose. CONCLUSION: In this initiative, we rapidly produced project-specific data modeling guidance and documentation in support of long COVID research while maintaining a commitment to terminology standards and harmonized data. Kellie M. Walters, Marshall Clark, Sofia Dard, Stephanie S. Hong, Elizabeth Kelly, Kristin Kostka, Adam M. Lee, Robert T. Miller, Michele Morris, Matvey Palchuk, Emily R. Pfaff, Adam B. Wilcox, Alexis Graves, Alfred Anzalone, Amin Manna, Amit Saha, Amy Olex, Andrea Zhou, Andrew E. Williams, Andrew Southerland, Andrew T. Girvin, Anita Walden, Anjali A Sharathkumar, Benjamin R. C. Amor, Benjamin Bates, Brian Hendricks, Caleb Alexander, Carolyn T. Bramante, Cavin Ward-Caviness, Charisse R. Madlock-Brown, Christine Suver, Christopher G. Chute, Christopher Dillon, Chunlei Wu, Clare Schmitt, Cliff Takemoto, Dan Housman, Davera Gabriel, David Eichmann, Diego Mazzotti, Don Brown, Eilis A. Boudreau, Elaine L. Hill, Elizabeth Zampino, Emily Carlson Marti, Evan French, Farrukh M. Koraishy, Federico Mariona, Fred W. Prior, George Sokos, Greg Martin, Harold P. Lehmann, Heidi Spratt, Hemalkumar Mehta, Hythem Sidky, J. W. Awori Hayanga, Jami Pincavitch, Jaylyn Clark, Jeremy Richard Harper, Jessica Islam, Jin Ge, Joel Gagnier, Joel H. Saltz, Johanna Loomba, John Buse, Jomol P. Mathew, Joni L. Rutter, Julie A. McMurry, Justin Guinney, Justin Starren, Karen Crowley, Katie Rebecca Bradwell, Ken Wilkins, Kenneth R. Gersing, Kenrick Dwain Cato, Kimberly Murray, Lavance Northington, Lee Allan Pyles, Leonie Misquitta, Lesley Cottrell, Lili M. Portilla, Mariam Deacy, Mark M. Bissell, Mary Emmett, Mary Morrison Saltz, Melissa A. Haendel, Meredith C. B. Adams, Meredith Temple-O'Connor, Michael G. Kurilla, Nabeel Qureshi, Nasia Safdar, Nicole Garbarini, Noha Sharafeldin, Ofer Sadan, Patricia A. Francis, Penny Wung Burgoon, Peter N. Robinson, Philip R. O. Payne, Rafael Fuentes, Randeep Jawa, Rebecca Erwin-Cohen, Rena Patel, Richard A. Moffitt, Richard L. Zhu, Rishi Kamaleswaran, Robert Hurley, Saiju Pyarajan, Samuel G. Michael, Samuel Bozzette, Sandeep Mallipattu, Satyanarayana Vedula, Scott Chapman, Shawn T. O'Neil, Soko Setoguchi, Tellen D. Bennett, Tiffany Callahan, Umit Topaloglu, Usman Sheikh, Valery Gordon, Vignesh Subbian, Warren A. Kibbe, Wenndy Hernandez, Will Beasley, Will Cooper, William Hillegass, Xiaohan Tanner Zhang |
J. Am. Medical Informatics Assoc. | 99 |
| 2024 | Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learningabstractMOTIVATION: Creating knowledge bases and ontologies is a time consuming task that relies on manual curation. AI/NLP approaches can assist expert curators in populating these knowledge bases, but current approaches rely on extensive training data, and are not able to populate arbitrarily complex nested knowledge schemas. RESULTS: Here we present Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES), a Knowledge Extraction approach that relies on the ability of Large Language Models (LLMs) to perform zero-shot learning and general-purpose query answering from flexible prompts and return information conforming to a specified schema. Given a detailed, user-defined knowledge schema and an input text, SPIRES recursively performs prompt interrogation against an LLM to obtain a set of responses matching the provided schema. SPIRES uses existing ontologies and vocabularies to provide identifiers for matched elements. We present examples of applying SPIRES in different domains, including extraction of food recipes, multi-species cellular signaling pathways, disease treatments, multi-step drug mechanisms, and chemical to disease relationships. Current SPIRES accuracy is comparable to the mid-range of existing Relation Extraction methods, but greatly surpasses an LLM's native capability of grounding entities with unique identifiers. SPIRES has the advantage of easy customization, flexibility, and, crucially, the ability to perform new tasks in the absence of any new training data. This method supports a general strategy of leveraging the language interpreting capabilities of LLMs to assemble knowledge bases, assisting manual knowledge curation and acquisition while supporting validation with publicly-available databases and ontologies external to the LLM. AVAILABILITY AND IMPLEMENTATION: SPIRES is available as part of the open source OntoGPT package: https://github.com/monarch-initiative/ontogpt. J. Harry Caufield, Harshad Hegde, Vincent Emonet, Nomi L. Harris, Marcin P. Joachimiak, Nicolas Matentzoglu, HyeongSik Kim 0001, Sierra A. T. Moxon, Justin T. Reese, Melissa A. Haendel, Peter N. Robinson, Chris Mungall |
Bioinform. | 11 |
| 2024 | FastHPOCR: pragmatic, fast, and accurate concept recognition using the human phenotype ontologyabstractMOTIVATION: Human Phenotype Ontology (HPO)-based phenotype concept recognition (CR) underpins a faster and more effective mechanism to create patient phenotype profiles or to document novel phenotype-centred knowledge statements. While the increasing adoption of large language models (LLMs) for natural language understanding has led to several LLM-based solutions, we argue that their intrinsic resource-intensive nature is not suitable for realistic management of the phenotype CR lifecycle. Consequently, we propose to go back to the basics and adopt a dictionary-based approach that enables both an immediate refresh of the ontological concepts as well as efficient re-analysis of past data. RESULTS: We developed a dictionary-based approach using a pre-built large collection of clusters of morphologically equivalent tokens-to address lexical variability and a more effective CR step by reducing the entity boundary detection strictly to candidates consisting of tokens belonging to ontology concepts. Our method achieves state-of-the-art results (0.76 F1 on the GSC+ corpus) and a processing efficiency of 10 000 publication abstracts in 5 s. AVAILABILITY AND IMPLEMENTATION: FastHPOCR is available as a Python package installable via pip. The source code is available at https://github.com/tudorgroza/fast_hpo_cr. A Java implementation of FastHPOCR will be made available as part of the Fenominal Java library available at https://github.com/monarch-initiative/fenominal. The up-to-date GCS-2024 corpus is available at https://github.com/tudorgroza/code-for-papers/tree/main/gsc-2024. Tudor Groza, Dylan Gration, Gareth Baynam, Peter N. Robinson |
Bioinform. | 4 |
| 2023 | KG-Hub - building and exchanging biological knowledge graphsabstractMOTIVATION: Knowledge graphs (KGs) are a powerful approach for integrating heterogeneous data and making inferences in biology and many other domains, but a coherent solution for constructing, exchanging, and facilitating the downstream use of KGs is lacking. RESULTS: Here we present KG-Hub, a platform that enables standardized construction, exchange, and reuse of KGs. Features include a simple, modular extract-transform-load pattern for producing graphs compliant with Biolink Model (a high-level data model for standardizing biological data), easy integration of any OBO (Open Biological and Biomedical Ontologies) ontology, cached downloads of upstream data sources, versioned and automatically updated builds with stable URLs, web-browsable storage of KG artifacts on cloud infrastructure, and easy reuse of transformed subgraphs across projects. Current KG-Hub projects span use cases including COVID-19 research, drug repurposing, microbial-environmental interactions, and rare disease research. KG-Hub is equipped with tooling to easily analyze and manipulate KGs. KG-Hub is also tightly integrated with graph machine learning (ML) tools which allow automated graph ML, including node embeddings and training of models for link prediction and node classification. AVAILABILITY AND IMPLEMENTATION: https://kghub.org. J. Harry Caufield, Tim E. Putman, Kevin Schaper, Deepak R. Unni, Harshad Hegde, Tiffany Callahan, Luca Cappelletti, Sierra A. T. Moxon, Vida Ravanmehr, Seth Carbon, Lauren E. Chan, Katherina G. Cortes, Kent A. Shefchek, Glass Elsarboukh, James P. Balhoff, Tommaso Fontana, Nicolas Matentzoglu, Richard M. Bruskiewich, Anne E. Thessen, Nomi L. Harris, Monica C. Munoz-Torres, Melissa A. Haendel, Peter N. Robinson, Marcin P. Joachimiak, Chris Mungall, Justin T. Reese |
Bioinform. | 23 |
| 2023 | Term-BLAST-like alignment tool for concept recognition in noisy clinical textsabstractMOTIVATION: Methods for concept recognition (CR) in clinical texts have largely been tested on abstracts or articles from the medical literature. However, texts from electronic health records (EHRs) frequently contain spelling errors, abbreviations, and other nonstandard ways of representing clinical concepts. RESULTS: Here, we present a method inspired by the BLAST algorithm for biosequence alignment that screens texts for potential matches on the basis of matching k-mer counts and scores candidates based on conformance to typical patterns of spelling errors derived from 2.9 million clinical notes. Our method, the Term-BLAST-like alignment tool (TBLAT) leverages a gold standard corpus for typographical errors to implement a sequence alignment-inspired method for efficient entity linkage. We present a comprehensive experimental comparison of TBLAT with five widely used tools. Experimental results show an increase of 10% in recall on scientific publications and 20% increase in recall on EHR records (when compared against the next best method), hence supporting a significant enhancement of the entity linking task. The method can be used stand-alone or as a complement to existing approaches. AVAILABILITY AND IMPLEMENTATION: Fenominal is a Java library that implements TBLAT for named CR of Human Phenotype Ontology terms and is available at https://github.com/monarch-initiative/fenominal under the GNU General Public License v3.0. Tudor Groza, Honghan Wu, Marcel E. Dinger, Daniel Danis, Coleman Hilton, Anita Bagley, Jon R. Davids, Ling Luo 0001, Zhiyong Lu, Peter N. Robinson |
Bioinform. | 10 |
| 2023 | An expectation-maximization framework for comprehensive prediction of isoform-specific functionsabstractMOTIVATION: Advances in RNA sequencing technologies have achieved an unprecedented accuracy in the quantification of mRNA isoforms, but our knowledge of isoform-specific functions has lagged behind. There is a need to understand the functional consequences of differential splicing, which could be supported by the generation of accurate and comprehensive isoform-specific gene ontology annotations. RESULTS: We present isoform interpretation, a method that uses expectation-maximization to infer isoform-specific functions based on the relationship between sequence and functional isoform similarity. We predicted isoform-specific functional annotations for 85 617 isoforms of 17 900 protein-coding human genes spanning a range of 17 430 distinct gene ontology terms. Comparison with a gold-standard corpus of manually annotated human isoform functions showed that isoform interpretation significantly outperforms state-of-the-art competing methods. We provide experimental evidence that functionally related isoforms predicted by isoform interpretation show a higher degree of domain sharing and expression correlation than functionally related genes. We also show that isoform sequence similarity correlates better with inferred isoform function than with gene-level function. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, and resource files are freely available under a GNU3 license at https://github.com/TheJacksonLaboratory/isopretEM and https://zenodo.org/record/7594321. Guy Karlebach, Leigh Carmody, Jagadish Chandrabose Sundaramurthi, Elena Casiraghi, Justin T. Reese, Chris Mungall, Giorgio Valentini, Peter N. Robinson |
Bioinform. | 9 |
| 2023 | A method for comparing multiple imputation techniques: A case study on the U.S. national COVID cohort collaborativeabstractHealthcare datasets obtained from Electronic Health Records have proven to be extremely useful for assessing associations between patients' predictors and outcomes of interest. However, these datasets often suffer from missing values in a high proportion of cases, whose removal may introduce severe bias. Several multiple imputation algorithms have been proposed to attempt to recover the missing information under an assumed missingness mechanism. Each algorithm presents strengths and weaknesses, and there is currently no consensus on which multiple imputation algorithm works best in a given scenario. Furthermore, the selection of each algorithm's parameters and data-related modeling choices are also both crucial and challenging. In this paper we propose a novel framework to numerically evaluate strategies for handling missing data in the context of statistical analysis, with a particular focus on multiple imputation techniques. We demonstrate the feasibility of our approach on a large cohort of type-2 diabetes patients provided by the National COVID Cohort Collaborative (N3C) Enclave, where we explored the influence of various patient characteristics on outcomes related to COVID-19. Our analysis included classic multiple imputation techniques as well as simple complete-case Inverse Probability Weighted models. Extensive experiments show that our approach can effectively highlight the most promising and performant missing-data handling strategy for our case study. Moreover, our methodology allowed a better understanding of the behavior of the different models and of how it changed as we modified their parameters. Our method is general and can be applied to different research fields and on datasets containing heterogeneous types. Elena Casiraghi, Rachel Wong, Margaret Hall, Ben D. Coleman, Marco Notaro, Michael D. Evans, Jena S. Tronieri, Hannah Blau, Bryan Laraway, Tiffany Callahan, Lauren E. Chan, Carolyn T. Bramante, John B. Buse, Richard A. Moffitt, Til Sturmer, Steven G. Johnson, Yu Raymond Shao, Justin T. Reese, Peter N. Robinson, Alberto Paccanaro, Giorgio Valentini, Jared D. Huling, Kenneth Wilkins |
J. Biomed. Informatics | 19 |
| 2022 | Computing Minimal Boolean Models of Gene Regulatory NetworksabstractModels of Gene Regulatory Networks (GRNs) capture the dynamics of the regulatory processes that occur within the cell as a means to understand the variability observed in gene expression between different conditions. Arguably the simplest mathematical construct used for modeling is the Boolean network, which dictates a set of logical rules for transition between states described as Boolean vectors. Due to the complexity of gene regulation and the limitations of experimental technologies, in most cases knowledge about regulatory interactions and Boolean states is partial. In addition, the logical rules themselves are not known a-priori. Our goal in this work is to create an algorithm that finds the network that fits the data optimally, and identify the network states that correspond to the noise-free data. We present a novel methodology for integrating experimental data and performing a search for the optimal consistent structure via optimization of a linear objective function under a set of linear constraints. In addition, we extend our methodology into a heuristic that alleviates the computational complexity of the problem for datasets that are generated by single-cell RNA-Sequencing(scRNA-Seq). We demonstrate the effectiveness of these tools using a public scRNA-Seq dataset and the GRN that is associated with it. Our methodology wm enable researchers to obtain a better understanding of the dynamics of gene regulatory networks and their biological role. Guy Karlebach, Peter N. Robinson |
BIBM | 2 |
| 2022 | Evaluation of phenotype-driven gene prioritization methods for Mendelian diseasesabstractYuan et al. recently described an independent evaluation of several phenotype-driven gene prioritization methods for Mendelian disease on two separate, clinical datasets. Although they attempted to use default settings for each tool, we describe three key differences from those we currently recommend for our Exomiser and PhenIX tools. These influence how variant frequency, quality and predicted pathogenicity are used for filtering and prioritization. We propose that these differences account for much of the discrepancy in performance between that reported by them (15-26% diagnoses ranked top by Exomiser) and previously published reports by us and others (72-77%). On a set of 161 singleton samples, we show using these settings increases performance from 34% to 72% and suggest a reassessment of Exomiser and PhenIX on their datasets using these would show a similar uplift. Julius O. B. Jacobsen, Catherine Kelly, Valentina Cipriani, Peter N. Robinson, Damian Smedley |
Briefings Bioinform. | 4 |
| 2022 | Demonstrating an approach for evaluating synthetic geospatial and temporal epidemiologic data utility: results from analyzing >1.8 million SARS-CoV-2 tests in the United States National COVID Cohort Collaborative (N3C)abstractOBJECTIVE: This study sought to evaluate whether synthetic data derived from a national coronavirus disease 2019 (COVID-19) dataset could be used for geospatial and temporal epidemic analyses. MATERIALS AND METHODS: Using an original dataset (n = 1 854 968 severe acute respiratory syndrome coronavirus 2 tests) and its synthetic derivative, we compared key indicators of COVID-19 community spread through analysis of aggregate and zip code-level epidemic curves, patient characteristics and outcomes, distribution of tests by zip code, and indicator counts stratified by month and zip code. Similarity between the data was statistically and qualitatively evaluated. RESULTS: In general, synthetic data closely matched original data for epidemic curves, patient characteristics, and outcomes. Synthetic data suppressed labels of zip codes with few total tests (mean = 2.9 ± 2.4; max = 16 tests; 66% reduction of unique zip codes). Epidemic curves and monthly indicator counts were similar between synthetic and original data in a random sample of the most tested (top 1%; n = 171) and for all unsuppressed zip codes (n = 5819), respectively. In small sample sizes, synthetic data utility was notably decreased. DISCUSSION: Analyses on the population-level and of densely tested zip codes (which contained most of the data) were similar between original and synthetically derived datasets. Analyses of sparsely tested populations were less similar and had more data suppression. CONCLUSION: In general, synthetic data were successfully used to analyze geospatial and temporal trends. Analyses using small sample sizes or populations were limited, in part due to purposeful data label suppression-an attribute disclosure countermeasure. Users should consider data fitness for use in these cases. Jason A. Thomas, Randi E. Foraker, Noa Zamstein, Jon D. Morrow, Philip R. O. Payne, Adam B. Wilcox, Melissa A. Haendel, Christopher G. Chute, Kenneth R. Gersing, Anita Walden, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Justin Starren, Christine Suver, Chunlei Wu, Davera Gabriel, Stephanie S. Hong, Kristin Kostka, Harold P. Lehmann, Richard A. Moffitt, Michele Morris, Matvey Palchuk, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Mark M. Bissell, Marshall Clark, Andrew T. Girvin, Adam M. Lee, Robert T. Miller, Kellie M. Walters, Yooree Chae, Connor Cook, Alexandra Dest, Racquel R. Dietz, Thomas Dillon, Patricia A. Francis, Rafael Fuentes, Alexis Graves, Andrew J. Neumann, Shawn T. O'Neil, Usman Sheikh, Andréa M. Volz, Elizabeth Zampino, Christopher P. Austin, Samuel Bozzette, Mariam Deacy, Nicole Garbarini, Michael G. Kurilla, Samuel G. Michael, Joni L. Rutter, Meredith Temple-O'Connor, Katie Rebecca Bradwell, Amin Manna, Nabeel Qureshi, Mary Morrison Saltz, Julie A. McMurry, Carolyn T. Bramante, Jeremy Richard Harper, Wenndy Hernandez, Farrukh M. Koraishy, Federico Mariona, Saidulu Mattapally, Amit Saha, Satyanarayana Vedula, Yujuan Fu, Nisha Mathews, Ofer Mendelevitch |
J. Am. Medical Informatics Assoc. | 17 |
| 2022 | PhenoRerank: A re-ranking model for phenotypic concept recognition pre-trained on human phenotype ontologyabstractThe study aims at developing a neural network model to improve the performance of Human Phenotype Ontology (HPO) concept recognition tools. We used the terms, definitions, and comments about the phenotypic concepts in the HPO database to train our model. The document to be analyzed is first split into sentences and annotated with a base method to generate candidate concepts. The sentences, along with the candidate concepts, are then fed into the pre-trained model for re-ranking. Our model comprises the pre-trained BlueBERT and a feature selection module, followed by a contrastive loss. We re-ranked the results generated by three robust HPO annotation tools and compared the performance against most of the existing approaches. The experimental results show that our model can improve the performance of the existing methods. Significantly, it boosted 3.0% and 5.6% in F1 score on the two evaluated datasets compared with the base methods. It removed more than 80% of the false positives predicted by the base methods, resulting in up to 18% improvement in precision. Our model utilizes the descriptive data in the ontology and the contextual information in the sentences for re-ranking. The results indicate that the additional information and the re-ranking model can significantly enhance the precision of HPO concept recognition compared with the base method. Shankai Yan, Ling Luo 0001, Po-Ting Lai, Daniel Veltri, Andrew J. Oler, Sandhya Xirasagar, Rajarshi Ghosh, Morgan Similuk, Peter N. Robinson, Zhiyong Lu |
J. Biomed. Informatics | 9 |
| 2021 | COVID-19 Mortality Prediction among Patients with Cancer Using a Large National Cohort
Noha Sharafeldin, Vithal Madhira, Katie R. Bradwell, Qianqian Song 0002, Benjamin Bates, Yu R. Shao, Jing Su 0003, Alfred Anzalone, Timothym Bergquist, Sarah Cutrona, Ben S. Gerber, Peter N. Robinson, Justin Guinney, Umit Topaloglu |
AMIA | 13 |
| 2021 | PhenoTagger: a hybrid method for phenotype concept recognition using human phenotype ontologyabstractMOTIVATION: Automatic phenotype concept recognition from unstructured text remains a challenging task in biomedical text mining research. Previous works that address the task typically use dictionary-based matching methods, which can achieve high precision but suffer from lower recall. Recently, machine learning-based methods have been proposed to identify biomedical concepts, which can recognize more unseen concept synonyms by automatic feature learning. However, most methods require large corpora of manually annotated data for model training, which is difficult to obtain due to the high cost of human annotation. RESULTS: In this article, we propose PhenoTagger, a hybrid method that combines both dictionary and machine learning-based methods to recognize Human Phenotype Ontology (HPO) concepts in unstructured biomedical text. We first use all concepts and synonyms in HPO to construct a dictionary, which is then used to automatically build a distantly supervised training dataset for machine learning. Next, a cutting-edge deep learning model is trained to classify each candidate phrase (n-gram from input sentence) into a corresponding concept label. Finally, the dictionary and machine learning-based prediction results are combined for improved performance. Our method is validated with two HPO corpora, and the results show that PhenoTagger compares favorably to previous methods. In addition, to demonstrate the generalizability of our method, we retrained PhenoTagger using the disease ontology MEDIC for disease concept recognition to investigate the effect of training on different ontologies. Experimental results on the NCBI disease corpus show that PhenoTagger without requiring manually annotated training data achieves competitive performance as compared with state-of-the-art supervised methods. AVAILABILITYAND IMPLEMENTATION: The source code, API information and data for PhenoTagger are freely available at https://github.com/ncbi-nlp/PhenoTagger. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ling Luo 0001, Shankai Yan, Po-Ting Lai, Daniel Veltri, Andrew J. Oler, Sandhya Xirasagar, Rajarshi Ghosh, Morgan Similuk, Peter N. Robinson, Zhiyong Lu |
Bioinform. | 9 |
| 2021 | HEMDAG: a family of modular and scalable hierarchical ensemble methods to improve Gene Ontology term predictionabstractMOTIVATION: Automated protein function prediction is a complex multi-class, multi-label, structured classification problem in which protein functions are organized in a controlled vocabulary, according to the Gene Ontology (GO). 'Hierarchy-unaware' classifiers, also known as 'flat' methods, predict GO terms without exploiting the inherent structure of the ontology, potentially violating the True-Path-Rule (TPR) that governs the GO, while 'hierarchy-aware' approaches, even if they obey the TPR, do not always show clear improvements with respect to flat methods, or do not scale well when applied to the full GO. RESULTS: To overcome these limitations, we propose Hierarchical Ensemble Methods for Directed Acyclic Graphs (HEMDAG), a family of highly modular hierarchical ensembles of classifiers, able to build upon any flat method and to provide 'TPR-safe' predictions, by leveraging a combination of isotonic regression and TPR learning strategies. Extensive experiments on synthetic and real data across several organisms firstly show that HEMDAG can be used as a general tool to improve the predictions of flat classifiers, and secondly that HEMDAG is competitive versus state-of-the-art hierarchy-aware learning methods proposed in the last CAFA international challenges. AVAILABILITY AND IMPLEMENTATION: Fully tested R code freely available at https://anaconda.org/bioconda/r-hemdag. Tutorial and documentation at https://hemdag.readthedocs.io. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marco Notaro, Marco Frasca 0001, Alessandro Petrini, Jessica Gliozzo, Elena Casiraghi, Peter N. Robinson, Giorgio Valentini |
Bioinform. | 6 |
| 2021 | The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deploymentabstractOBJECTIVE: Coronavirus disease 2019 (COVID-19) poses societal challenges that require expeditious data and knowledge sharing. Though organizational clinical data are abundant, these are largely inaccessible to outside researchers. Statistical, machine learning, and causal analyses are most successful with large-scale data beyond what is available in any given organization. Here, we introduce the National COVID Cohort Collaborative (N3C), an open science community focused on analyzing patient-level data from many centers. MATERIALS AND METHODS: The Clinical and Translational Science Award Program and scientific community created N3C to overcome technical, regulatory, policy, and governance barriers to sharing and harmonizing individual-level clinical data. We developed solutions to extract, aggregate, and harmonize data across organizations and data models, and created a secure data enclave to enable efficient, transparent, and reproducible collaborative analytics. RESULTS: Organized in inclusive workstreams, we created legal agreements and governance for organizations and researchers; data extraction scripts to identify and ingest positive, negative, and possible COVID-19 cases; a data quality assurance and harmonization pipeline to create a single harmonized dataset; population of the secure data enclave with data, machine learning, and statistical analytics tools; dissemination mechanisms; and a synthetic data pilot to democratize data access. CONCLUSIONS: The N3C has demonstrated that a multisite collaborative learning health network can overcome barriers to rapidly build a scalable infrastructure incorporating multiorganizational clinical data for COVID-19 analytics. We expect this effort to save lives by enabling rapid collaboration among clinicians, researchers, and data scientists to identify treatments and specialized care and thereby reduce the immediate and long-term impacts of COVID-19. Melissa A. Haendel, Christopher G. Chute, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Philip R. O. Payne, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Christine Suver, John Wilbanks, Adam B. Wilcox, Andrew E. Williams, Chunlei Wu, Clair Blacketer, Robert L. Bradford, James J. Cimino, Marshall Clark, Evan W. Colmenares, Patricia A. Francis, Davera Gabriel, Alexis Graves, Raju Hemadri, Stephanie S. Hong, George Hripcsak, Dazhi Jiao, Jeffrey G. Klann, Kristin Kostka, Adam M. Lee, Harold P. Lehmann, Lora Lingrey, Robert T. Miller, Michele Morris, Shawn N. Murphy, Karthik Natarajan, Matvey Palchuk, Usman Sheikh, Harold R. Solbrig, Shyam Visweswaran, Anita Walden, Kellie M. Walters, Griffin M. Weber, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Andrew T. Girvin, Amin Manna, Nabeel Qureshi, Michael G. Kurilla, Samuel G. Michael, Lili M. Portilla, Joni L. Rutter, Christopher P. Austin, Kenneth R. Gersing |
J. Am. Medical Informatics Assoc. | 9 |
| 2017 | A likelihood ratio-based method to predict exact pedigrees for complex families from next-generation sequencing dataabstractMOTIVATION: Next generation sequencing technology considerably changed the way we screen for pathogenic mutations in rare Mendelian disorders. However, the identification of the disease-causing mutation amongst thousands of variants of partly unknown relevance is still challenging and efficient techniques that reduce the genomic search space play a decisive role. Often segregation- or linkage analysis are used to prioritize candidates, however, these approaches require correct information about the degree of relationship among the sequenced samples. For quality assurance an automated control of pedigree structures and sample assignment is therefore highly desirable in order to detect label mix-ups that might otherwise corrupt downstream analysis. RESULTS: We developed an algorithm based on likelihood ratios that discriminates between different classes of relationship for an arbitrary number of genotyped samples. By identifying the most likely class we are able to reconstruct entire pedigrees iteratively, even for highly consanguineous families. We tested our approach on exome data of different sequencing studies and achieved high precision for all pedigree predictions. By analyzing the precision for varying degrees of relatedness or inbreeding we could show that a prediction is robust down to magnitudes of a few hundred loci. AVAILABILITY AND IMPLEMENTATION: A java standalone application that computes the relationships between multiple samples as well as a Rscript that visualizes the pedigree information is available for download as well as a web service at www.gene-talk.de CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Verena Heinrich, Tom Kamphans, Stefan Mundlos, Peter N. Robinson, Peter M. Krawitz |
Bioinform. | 4 |
| 2017 | Prediction of Human Phenotype Ontology terms by means of hierarchical ensemble methodsabstractBACKGROUND: The prediction of human gene-abnormal phenotype associations is a fundamental step toward the discovery of novel genes associated with human disorders, especially when no genes are known to be associated with a specific disease. In this context the Human Phenotype Ontology (HPO) provides a standard categorization of the abnormalities associated with human diseases. While the problem of the prediction of gene-disease associations has been widely investigated, the related problem of gene-phenotypic feature (i.e., HPO term) associations has been largely overlooked, even if for most human genes no HPO term associations are known and despite the increasing application of the HPO to relevant medical problems. Moreover most of the methods proposed in literature are not able to capture the hierarchical relationships between HPO terms, thus resulting in inconsistent and relatively inaccurate predictions. RESULTS: We present two hierarchical ensemble methods that we formally prove to provide biologically consistent predictions according to the hierarchical structure of the HPO. The modular structure of the proposed methods, that consists in a "flat" learning first step and a hierarchical combination of the predictions in the second step, allows the predictions of virtually any flat learning method to be enhanced. The experimental results show that hierarchical ensemble methods are able to predict novel associations between genes and abnormal phenotypes with results that are competitive with state-of-the-art algorithms and with a significant reduction of the computational complexity. CONCLUSIONS: Hierarchical ensembles are efficient computational methods that guarantee biologically meaningful predictions that obey the true path rule, and can be used as a tool to improve and make consistent the HPO terms predictions starting from virtually any flat learning method. The implementation of the proposed methods is available as an R package from the CRAN repository. Marco Notaro, Max Schubach, Peter N. Robinson, Giorgio Valentini |
BMC Bioinform. | 3 |
| 2016 | The digital revolution in phenotypingabstractPhenotypes have gained increased notoriety in the clinical and biological domain owing to their application in numerous areas such as the discovery of disease genes and drug targets, phylogenetics and pharmacogenomics. Phenotypes, defined as observable characteristics of organisms, can be seen as one of the bridges that lead to a translation of experimental findings into clinical applications and thereby support 'bench to bedside' efforts. However, to build this translational bridge, a common and universal understanding of phenotypes is required that goes beyond domain-specific definitions. To achieve this ambitious goal, a digital revolution is ongoing that enables the encoding of data in computer-readable formats and the data storage in specialized repositories, ready for integration, enabling translational research. While phenome research is an ongoing endeavor, the true potential hidden in the currently available data still needs to be unlocked, offering exciting opportunities for the forthcoming years. Here, we provide insights into the state-of-the-art in digital phenotyping, by means of representing, acquiring and analyzing phenotype data. In addition, we provide visions of this field for future research work that could enable better applications of phenotype data. Anika Oellrich, Nigel Collier, Tudor Groza, Dietrich Rebholz-Schuhmann, Nigam H. Shah, Olivier Bodenreider, Mary Regina Boland, Ivo I. Georgiev, Kevin M. Livingston, Augustin Luna, Ann-Marie Mallon, Prashanti Manda, Peter N. Robinson, Gabriella Rustici, Michelle Simon, Rainer Winnenburg, Michel Dumontier |
Briefings Bioinform. | 14 |
| 2015 | IMSEQ - a fast and error aware approach to immunogenetic sequence analysisabstractMOTIVATION: Recombined T- and B-cell receptor repertoires are increasingly being studied using next generation sequencing (NGS) in order to interrogate the repertoire composition as well as changes in the distribution of receptor clones under different physiological and disease states. This type of analysis requires efficient and unambiguous clonotype assignment to a large number of NGS read sequences, including the identification of the incorporated V and J gene segments and the CDR3 sequence. Current tools have deficits with respect to performance, accuracy and documentation of their underlying algorithms and usage. RESULTS: We present IMSEQ, a method to derive clonotype repertoires from NGS data with sophisticated routines for handling errors stemming from PCR and sequencing artefacts. The application can handle different kinds of input data originating from single- or paired-end sequencing in different configurations and is generic regarding the species and gene of interest. We have carefully evaluated our method with simulated and real world data and show that IMSEQ is superior to other tools with respect to its clonotyping as well as standalone error correction and runtime performance. AVAILABILITY AND IMPLEMENTATION: IMSEQ was implemented in C++ using the SeqAn library for efficient sequence analysis. It is freely available under the GPLv2 open source license and can be downloaded at www.imtools.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected] or [email protected]. Léon Kuchenbecker, Mikalai Nienen, Jochen Hecht, Avidan U. Neumann, Nina Babel, Knut Reinert, Peter N. Robinson |
Bioinform. | 7 |
| 2015 | Strategies to improve the performance of rare variant association studies by optimizing the selection of controlsabstractMOTIVATION: When analyzing a case group of patients with ultra-rare disorders the ethnicities are often diverse and the data quality might vary. The population substructure in the case group as well as the heterogeneous data quality can cause substantial inflation of test statistics and result in spurious associations in case-control studies if not properly adjusted for. Existing techniques to correct for confounding effects were especially developed for common variants and are not applicable to rare variants. RESULTS: We analyzed strategies to select suitable controls for cases that are based on similarity metrics that vary in their weighting schemes. We simulated different disease entities on real exome data and show that a similarity-based selection scheme can help to reduce false positive associations and to optimize the performance of the statistical tests. Especially when data quality as well as ethnicities vary a lot in the case group, a matching approach that puts more weight on rare variants shows the best performance. We reanalyzed collections of unrelated patients with Kabuki make-up syndrome, Hyperphosphatasia with Mental Retardation syndrome and Catel-Manzke syndrome for which the disease genes were recently described. We show that rare variant association tests are more sensitive and specific in identifying the disease gene than intersection filters and should thus be considered as a favorable approach in analyzing even small patient cohorts. AVAILABILITY AND IMPLEMENTATION: Datasets used in our analysis are available at ftp://ftp.1000genomes.ebi.ac.uk./vol1/ftp/ CONTACT: : [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Verena Heinrich, Thorsten Dickhaus, Jochen Hecht, Peter N. Robinson, Stefan Mundlos, Tom Kamphans, Peter M. Krawitz |
Bioinform. | 5 |
| 2014 | Walking the interactome for candidate prioritization in exome sequencing studies of Mendelian diseasesabstractMOTIVATION: Whole-exome sequencing (WES) has opened up previously unheard of possibilities for identifying novel disease genes in Mendelian disorders, only about half of which have been elucidated to date. However, interpretation of WES data remains challenging. RESULTS: Here, we analyze protein-protein association (PPA) networks to identify candidate genes in the vicinity of genes previously implicated in a disease. The analysis, using a random-walk with restart (RWR) method, is adapted to the setting of WES by developing a composite variant-gene relevance score based on the rarity, location and predicted pathogenicity of variants and the RWR evaluation of genes harboring the variants. Benchmarking using known disease variants from 88 disease-gene families reveals that the correct gene is ranked among the top 10 candidates in ≥50% of cases, a figure which we confirmed using a prospective study of disease genes identified in 2012 and PPA data produced before that date. We implement our method in a freely available Web server, ExomeWalker, that displays a ranked list of candidates together with information on PPAs, frequency and predicted pathogenicity of the variants to allow quick and effective searches for candidates that are likely to reward closer investigation. AVAILABILITY AND IMPLEMENTATION: http://compbio.charite.de/ExomeWalker CONTACT: : [email protected]. Damian Smedley, Sebastian Köhler 0001, Johanna Christina Czeschik, Joanna S. Amberger, Carol A. Bocchini, Ada Hamosh, Julian Veldboer, Tomasz Zemojtel, Peter N. Robinson |
Bioinform. | 9 |
| 2014 | Clinical phenotype-based gene prioritization: an initial study using semantic similarity and the human phenotype ontologyabstractBACKGROUND: Exome sequencing is a promising method for diagnosing patients with a complex phenotype. However, variant interpretation relative to patient phenotype can be challenging in some scenarios, particularly clinical assessment of rare complex phenotypes. Each patient's sequence reveals many possibly damaging variants that must be individually assessed to establish clear association with patient phenotype. To assist interpretation, we implemented an algorithm that ranks a given set of genes relative to patient phenotype. The algorithm orders genes by the semantic similarity computed between phenotypic descriptors associated with each gene and those describing the patient. Phenotypic descriptor terms are taken from the Human Phenotype Ontology (HPO) and semantic similarity is derived from each term's information content. RESULTS: Model validation was performed via simulation and with clinical data. We simulated 33 Mendelian diseases with 100 patients per disease. We modeled clinical conditions by adding noise and imprecision, i.e. phenotypic terms unrelated to the disease and terms less specific than the actual disease terms. We ranked the causative gene against all 2488 HPO annotated genes. The median causative gene rank was 1 for the optimal and noise cases, 12 for the imprecision case, and 60 for the imprecision with noise case. Additionally, we examined a clinical cohort of subjects with hearing impairment. The disease gene median rank was 22. However, when also considering the patient's exome data and filtering non-exomic and common variants, the median rank improved to 3. CONCLUSIONS: Semantic similarity can rank a causative gene highly within a gene list relative to patient phenotype characteristics, provided that imprecision is mitigated. The clinical case results suggest that phenotype rank combined with variant analysis provides significant improvement over the individual approaches. We expect that this combined prioritization approach may increase accuracy and decrease effort for clinical genetic diagnosis. Aaron J. Masino, Elizabeth T. Dechene, Matthew C. Dulik, Alisha Wilkens, Nancy B. Spinner, Ian D. Krantz, Jeffrey W. Pennington, Peter N. Robinson, Peter S. White |
BMC Bioinform. | 8 |
| 2012 | Bayesian ontology querying for accurate and noise-tolerant semantic searchesabstractMOTIVATION: Ontologies provide a structured representation of the concepts of a domain of knowledge as well as the relations between them. Attribute ontologies are used to describe the characteristics of the items of a domain, such as the functions of proteins or the signs and symptoms of disease, which opens the possibility of searching a database of items for the best match to a list of observed or desired attributes. However, naive search methods do not perform well on realistic data because of noise in the data, imprecision in typical queries and because individual items may not display all attributes of the category they belong to. RESULTS: We present a method for combining ontological analysis with Bayesian networks to deal with noise, imprecision and attribute frequencies and demonstrate an application of our method as a differential diagnostic support system for human genetics. AVAILABILITY: We provide an implementation for the algorithm and the benchmark at http://compbio.charite.de/boqa/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary Material for this article is available at Bioinformatics online. Sebastian Bauer 0002, Sebastian Köhler 0001, Marcel H. Schulz, Peter N. Robinson |
Bioinform. | 4 |
| 2011 | Model-based gene set analysis for BioconductorabstractUNLABELLED: Gene Ontology and other forms of gene-category analysis play a major role in the evaluation of high-throughput experiments in molecular biology. Single-category enrichment analysis procedures such as Fisher's exact test tend to flag large numbers of redundant categories as significant, which can complicate interpretation. We have recently developed an approach called model-based gene set analysis (MGSA), that substantially reduces the number of redundant categories returned by the gene-category analysis. In this work, we present the Bioconductor package mgsa, which makes the MGSA algorithm available to users of the R language. Our package provides a simple and flexible application programming interface for applying the approach. AVAILABILITY: The mgsa package has been made available as part of Bioconductor 2.8. It is released under the conditions of the Artistic license 2.0. CONTACT: [email protected]; [email protected]. Sebastian Bauer 0002, Peter N. Robinson, Julien Gagneur |
Bioinform. | 2 |
| 2011 | Identity-by-descent filtering of exome sequence data for disease-gene identification in autosomal recessive disordersabstractMOTIVATION: Next-generation sequencing and exome-capture technologies are currently revolutionizing the way geneticists screen for disease-causing mutations in rare Mendelian disorders. However, the identification of causal mutations is challenging due to the sheer number of variants that are identified in individual exomes. Although databases such as dbSNP or HapMap can be used to reduce the plethora of candidate genes by filtering out common variants, the remaining set of genes still remains on the order of dozens. RESULTS: Our algorithm uses a non-homogeneous hidden Markov model that employs local recombination rates to identify chromosomal regions that are identical by descent (IBD = 2) in children of consanguineous or non-consanguineous parents solely based on genotype data of siblings derived from high-throughput sequencing platforms. Using simulated and real exome sequence data, we show that our algorithm is able to reduce the search space for the causative disease gene to a fifth or a tenth of the entire exome. AVAILABILITY: An R script and an accompanying tutorial are available at http://compbio.charite.de/index.php/ibd2.html. Christian Rödelsperger, Peter M. Krawitz, Sebastian Bauer 0002, Jochen Hecht, Abigail W. Bigham, Michael Bamshad, Birgit Jonske de Condor, Michal R. Schweiger, Peter N. Robinson |
Bioinform. | 9 |
| 2011 | Improving ontologies by automatic reasoning and evaluation of logical definitionsabstractBACKGROUND: Ontologies are widely used to represent knowledge in biomedicine. Systematic approaches for detecting errors and disagreements are needed for large ontologies with hundreds or thousands of terms and semantic relationships. A recent approach of defining terms using logical definitions is now increasingly being adopted as a method for quality control as well as for facilitating interoperability and data integration. RESULTS: We show how automated reasoning over logical definitions of ontology terms can be used to improve ontology structure. We provide the Java software package GULO (Getting an Understanding of LOgical definitions), which allows fast and easy evaluation for any kind of logically decomposed ontology by generating a composite OWL ontology from appropriate subsets of the referenced ontologies and comparing the inferred relationships with the relationships asserted in the target ontology. As a case study we show how to use GULO to evaluate the logical definitions that have been developed for the Mammalian Phenotype Ontology (MPO). CONCLUSIONS: Logical definitions of terms from biomedical ontologies represent an important resource for error and disagreement detection. GULO gives ontology curators a fast and simple tool for validation of their work. Sebastian Köhler 0001, Sebastian Bauer 0002, Chris Mungall, Gabriele Carletti, Cynthia L. Smith, Paul N. Schofield, Georgios V. Gkoutos, Peter N. Robinson |
BMC Bioinform. | 8 |
| 2011 | Exact score distribution computation for ontological similarity searchesabstractBACKGROUND: Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., finding functionally related proteins with the Gene Ontology or phenotypically similar diseases with the Human Phenotype Ontology (HPO). We have recently shown that the performance of semantic similarity searches can be improved by ranking results according to the probability of obtaining a given score at random rather than by the scores themselves. However, to date, there are no algorithms for computing the exact distribution of semantic similarity scores, which is necessary for computing the exact P-value of a given score. RESULTS: In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a simple null hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the HPO. It is shown that exact P-value calculation improves clinical diagnosis using the HPO compared to approaches based on sampling. CONCLUSIONS: The new algorithm enables for the first time exact P-value calculation via exact score distribution computation for ontology similarity searches. The approach is applicable to any ontology for which the annotation-propagation rule holds and can improve any bioinformatic method that makes only use of the raw similarity scores. The algorithm was implemented in Java, supports any ontology in OBO format, and is available for non-commercial and academic usage under: https://compbio.charite.de/svn/hpo/trunk/src/tools/significance/ Marcel H. Schulz, Sebastian Köhler 0001, Sebastian Bauer 0002, Peter N. Robinson |
BMC Bioinform. | 4 |
| 2010 | Microindel detection in short-read sequence dataabstractMOTIVATION: Several recent studies have demonstrated the effectiveness of resequencing and single nucleotide variant (SNV) detection by deep short-read sequencing platforms. While several reliable algorithms are available for automated SNV detection, the automated detection of microindels in deep short-read data presents a new bioinformatics challenge. RESULTS: We systematically analyzed how the short-read mapping tools MAQ, Bowtie, Burrows-Wheeler alignment tool (BWA), Novoalign and RazerS perform on simulated datasets that contain indels and evaluated how indels affect error rates in SNV detection. We implemented a simple algorithm to compute the equivalent indel region eir, which can be used to process the alignments produced by the mapping tools in order to perform indel calling. Using simulated data that contains indels, we demonstrate that indel detection works well on short-read data: the detection rate for microindels (<4 bp) is >90%. Our study provides insights into systematic errors in SNV detection that is based on ungapped short sequence read alignments. Gapped alignments of short sequence reads can be used to reduce this error and to detect microindels in simulated short-read data. A comparison with microindels automatically identified on the ABI Sanger and Roche 454 platform indicates that microindel detection from short sequence reads identifies both overlapping and distinct indels. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peter M. Krawitz, Christian Rödelsperger, Marten Jäger, Luke Jostins, Sebastian Bauer 0002, Peter N. Robinson |
Bioinform. | 6 |
| 2009 | Exact Score Distribution Computation for Similarity Searches in Ontologies
Marcel H. Schulz, Sebastian Köhler 0001, Sebastian Bauer 0002, Martin Vingron, Peter N. Robinson |
WABI | 5 |
| 2008 | Ontologizer 2.0 - a multifunctional tool for GO term enrichment analysis and data explorationabstractUNLABELLED: The Ontologizer is a Java application that can be used to perform statistical analysis for overrepresentation of Gene Ontology (GO) terms in sets of genes or proteins derived from an experiment. The Ontologizer implements the standard approach to statistical analysis based on the one-sided Fisher's exact test, the novel parent-child method, as well as topology-based algorithms. A number of multiple-testing correction procedures are provided. The Ontologizer allows users to visualize data as a graph including all significantly overrepresented GO terms and to explore the data by linking GO terms to all genes/proteins annotated to the term and by linking individual terms to child terms. AVAILABILITY: The Ontologizer application is available under the terms of the GNU GPL. It can be started as a WebStart application from the project homepage, where source code is also provided: http://compbio.charite.de/ontologizer. REQUIREMENTS: Ontologizer requires a Java SE 5.0 compliant Java runtime engine and GraphViz for the optional graph visualization tool. Sebastian Bauer 0002, Steffen Grossmann, Martin Vingron, Peter N. Robinson |
Bioinform. | 4 |
| 2007 | Improved detection of overrepresentation of Gene-Ontology annotations with parent-child analysisabstractMOTIVATION: High-throughput experiments such as microarray hybridizations often yield long lists of genes found to share a certain characteristic such as differential expression. Exploring Gene Ontology (GO) annotations for such lists of genes has become a widespread practice to get first insights into the potential biological meaning of the experiment. The standard statistical approach to measuring overrepresentation of GO terms cannot cope with the dependencies resulting from the structure of GO because they analyze each term in isolation. Especially the fact that annotations are inherited from more specific descendant terms can result in certain types of false-positive results with potentially misleading biological interpretation, a phenomenon which we term the inheritance problem. RESULTS: We present here a novel approach to analysis of GO term overrepresentation that determines overrepresentation of terms in the context of annotations to the term's parents. This approach reduces the dependencies between the individual term's measurements, and thereby avoids producing false-positive results owing to the inheritance problem. ROC analysis using study sets with overrepresented GO terms showed a clear advantage for our approach over the standard algorithm with respect to the inheritance problem. Although there can be no gold standard for exploratory methods such as analysis of GO term overrepresentation, analysis of biological datasets suggests that our algorithm tends to identify the core GO terms that are most characteristic of the dataset being analyzed. Steffen Grossmann, Sebastian Bauer 0002, Peter N. Robinson, Martin Vingron |
Bioinform. | 3 |
| 2006 | An Improved Statistic for Detecting Over-Represented Gene Ontology Annotations in Gene Sets
Steffen Grossmann, Sebastian Bauer 0002, Peter N. Robinson, Martin Vingron |
RECOMB | 3 |
| 2006 | HotSwap for bioinformatics: A STRAP tutorialabstractBACKGROUND: Bioinformatics applications are now routinely used to analyze large amounts of data. Application development often requires many cycles of optimization, compiling, and testing. Repeatedly loading large datasets can significantly slow down the development process. We have incorporated HotSwap functionality into the protein workbench STRAP, allowing developers to create plugins using the Java HotSwap technique. RESULTS: Users can load multiple protein sequences or structures into the main STRAP user interface, and simultaneously develop plugins using an editor of their choice such as Emacs. Saving changes to the Java file causes STRAP to recompile the plugin and automatically update its user interface without requiring recompilation of STRAP or reloading of protein data. This article presents a tutorial on how to develop HotSwap plugins. STRAP is available at http://strapjava.de and http://www.charite.de/bioinf/strap. CONCLUSION: HotSwap is a useful and time-saving technique for bioinformatics developers. HotSwap can be used to efficiently develop bioinformatics applications that require loading large amounts of data into memory. Christoph Gille, Peter N. Robinson |
BMC Bioinform. | 2 |
| 2004 | Ontologizing gene-expression microarray data: characterizing clusters with Gene OntologyabstractAn XML-based Java application is described that provides a function-oriented overview of the results of cluster analysis of gene-expression microarray data based on Gene Ontology terms and associations. The application generates one HTML page with listings of the frequencies of explicit and implicit Gene Ontology annotations for each cluster, and separate, linked pages with listings of explicit annotations for each gene in a cluster. Peter N. Robinson, Andreas Wollstein, Ulrike Böhme, Bradley J. Beattie |
Bioinform. | 1 |