EDBT 2026 Demo / reviewers in the wild / expert
Joel Dudley
dblp:47/2610 · also Joel T. Dudley
· DBLP profile ↗
26ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0002-7036-6492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 2 first-authorDatabases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
8 papers |
Medical and health informatics · 51% Bioinformatics and computational biology · 49% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Medical and health informatics
electronic health records |
0.6 | 2 | 2019 | PatientExploreR: an extensible application for dynamic visualization of patient clinical history from electronic health records in the OMOP common data model · Bioinform. 2019 Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networks · Bioinform. 2016 |
Medical and health informatics
computer-aided diagnosis |
0.4 | 1 | 2019 | CANDI: an R package and Shiny app for annotating radiographs and evaluating computer-aided diagnosis · Bioinform. 2019 |
Medical and health informatics
medical visualization |
0.4 | 1 | 2019 | PatientExploreR: an extensible application for dynamic visualization of patient clinical history from electronic health records in the OMOP common data model · Bioinform. 2019 |
Medical and health informatics › clinical text processing
patient history analysis |
0.4 | 1 | 2019 | PatientExploreR: an extensible application for dynamic visualization of patient clinical history from electronic health records in the OMOP common data model · Bioinform. 2019 |
Bioinformatics and computational biology › bioimage informatics › cellular image analysis
automated cell classification |
0.3 | 1 | 2017 | Automated cell type discovery and classification through knowledge transfer · Bioinform. 2017 |
Bioinformatics and computational biology › single-cell analysis
cell type annotation |
0.3 | 1 | 2017 | Automated cell type discovery and classification through knowledge transfer · Bioinform. 2017 |
Bioinformatics and computational biology
single-cell analysis |
0.3 | 1 | 2017 | Automated cell type discovery and classification through knowledge transfer · Bioinform. 2017 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
disease network analysis |
0.2 | 1 | 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networks · Bioinform. 2016 |
Bioinformatics and computational biology
phenomics |
0.2 | 1 | 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networks · Bioinform. 2016 |
Recommender systems › collaborative filtering
matrix factorization |
0.2 | 1 | 2016 | Healthcare Data Mining with Matrix Models · KDD 2016 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.1 | 1 | 2011 | ProfileChaser: searching microarray repositories based on genome-wide patterns of differential expression · Bioinform. 2011 |
Bioinformatics and computational biology › drug discovery
drug repositioning |
0.1 | 1 | 2016 | Healthcare Data Mining with Matrix Models · KDD 2016 |
Medical and health informatics › clinical data analysis
multimorbidity analysis |
0.1 | 1 | 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networks · Bioinform. 2016 |
Bioinformatics and computational biology › phylogenetics › evolutionary history reconstruction
divergence time estimation |
0.1 | 1 | 2006 | TimeTree: a public knowledge-base of divergence times among organisms · Bioinform. 2006 |
Bioinformatics and computational biology
phylogenetics |
0.1 | 1 | 2006 | TimeTree: a public knowledge-base of divergence times among organisms · Bioinform. 2006 |
Bioinformatics and computational biology › genomics › genome analysis
genome-wide analysis |
0.0 | 1 | 2007 | Bioinformatics software for biologists in the genomics era · Bioinform. 2007 |
Bioinformatics and computational biology › sequence analysis
sequence comparison |
0.0 | 1 | 2007 | Bioinformatics software for biologists in the genomics era · Bioinform. 2007 |
Bioinformatics and computational biology
knowledge base |
0.0 | 1 | 2006 | TimeTree: a public knowledge-base of divergence times among organisms · Bioinform. 2006 |
Methods — techniques the papers use, named apart from their topics
tensor factorization · 0.5non-negative matrix factorization · 0.5joint matrix factorization · 0.5segmentation · 0.4r shiny · 0.4image captioning · 0.4deep learning · 0.4classification · 0.4machine learning · 0.3knowledge transfer · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Sepsis in the era of data-driven medicine: personalizing risks, diagnoses, treatments and prognosesabstractSepsis is a series of clinical syndromes caused by the immunological response to infection. The clinical evidence for sepsis could typically attribute to bacterial infection or bacterial endotoxins, but infections due to viruses, fungi or parasites could also lead to sepsis. Regardless of the etiology, rapid clinical deterioration, prolonged stay in intensive care units and high risk for mortality correlate with the incidence of sepsis. Despite its prevalence and morbidity, improvement in sepsis outcomes has remained limited. In this comprehensive review, we summarize the current landscape of risk estimation, diagnosis, treatment and prognosis strategies in the setting of sepsis and discuss future challenges. We argue that the advent of modern technologies such as in-depth molecular profiling, biomedical big data and machine intelligence methods will augment the treatment and prevention of sepsis. The volume, variety, veracity and velocity of heterogeneous data generated as part of healthcare delivery and recent advances in biotechnology-driven therapeutics and companion diagnostics may provide a new wave of approaches to identify the most at-risk sepsis patients and reduce the symptom burden in patients within shorter turnaround times. Developing novel therapies by leveraging modern drug discovery strategies including computational drug repositioning, cell and gene-therapy, clustered regularly interspaced short palindromic repeats -based genetic editing systems, immunotherapy, microbiome restoration, nanomaterial-based therapy and phage therapy may help to develop treatments to target sepsis. We also provide empirical evidence for potential new sepsis targets including FER and STARD3NL. Implementing data-driven methods that use real-time collection and analysis of clinical variables to trace, track and treat sepsis-related adverse outcomes will be key. Understanding the root and route of sepsis and its comorbid conditions that complicate treatment outcomes and lead to organ dysfunction may help to facilitate identification of most at-risk patients and prevent further deterioration. To conclude, leveraging the advances in precision medicine, biomedical data science and translational bioinformatics approaches may help to develop better strategies to diagnose and treat sepsis in the next decade. Andrew C. Liu, Krishna Patel, Ramya Dhatri Vunikili, Kipp W. Johnson, Fahad Jibrin Abdu, Shivani Kamath Belman, Benjamin S. Glicksberg, Pratyush Tandale, Roberto Fontanez, Oommen K. Mathew, Andrew Kasarskis, Priyabrata Mukherjee, Lakshminarayanan Subramanian, Joel Dudley, Khader Shameer |
Briefings Bioinform. | 14 |
| 2019 | CANDI: an R package and Shiny app for annotating radiographs and evaluating computer-aided diagnosisabstractMOTIVATION: Radiologists have used algorithms for Computer-Aided Diagnosis (CAD) for decades. These algorithms use machine learning with engineered features, and there have been mixed findings on whether they improve radiologists' interpretations. Deep learning offers superior performance but requires more training data and has not been evaluated in joint algorithm-radiologist decision systems. RESULTS: We developed the Computer-Aided Note and Diagnosis Interface (CANDI) for collaboratively annotating radiographs and evaluating how algorithms alter human interpretation. The annotation app collects classification, segmentation, and image captioning training data, and the evaluation app randomizes the availability of CAD tools to facilitate clinical trials on radiologist enhancement. AVAILABILITY AND IMPLEMENTATION: Demonstrations and source code are hosted at (https://candi.nextgenhealthcare.org), and (https://github.com/mbadge/candi), respectively, under GPL-3 license. SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online. Marcus A. Badgeley, Manway Liu, Benjamin S. Glicksberg, Mark M. Shervey, John R. Zech, Khader Shameer, Joseph Lehar, Eric K. Oermann, Michael V. McConnell, Thomas M. Snyder, Joel Dudley |
Bioinform. | 11 |
| 2019 | PatientExploreR: an extensible application for dynamic visualization of patient clinical history from electronic health records in the OMOP common data modelabstractMOTIVATION: Electronic health records (EHRs) are quickly becoming omnipresent in healthcare, but interoperability issues and technical demands limit their use for biomedical and clinical research. Interactive and flexible software that interfaces directly with EHR data structured around a common data model (CDM) could accelerate more EHR-based research by making the data more accessible to researchers who lack computational expertise and/or domain knowledge. RESULTS: We present PatientExploreR, an extensible application built on the R/Shiny framework that interfaces with a relational database of EHR data in the Observational Medical Outcomes Partnership CDM format. PatientExploreR produces patient-level interactive and dynamic reports and facilitates visualization of clinical data without any programming required. It allows researchers to easily construct and export patient cohorts from the EHR for analysis with other software. This application could enable easier exploration of patient-level data for physicians and researchers. PatientExploreR can incorporate EHR data from any institution that employs the CDM for users with approved access. The software code is free and open source under the MIT license, enabling institutions to install and users to expand and modify the application for their own purposes. AVAILABILITY AND IMPLEMENTATION: PatientExploreR can be freely obtained from GitHub: https://github.com/BenGlicksberg/PatientExploreR. We provide instructions for how researchers with approved access to their institutional EHR can use this package. We also release an open sandbox server of synthesized patient data for users without EHR access to explore: http://patientexplorer.ucsf.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin S. Glicksberg, Boris Oskotsky, Phyllis Thangaraj, Nicholas Giangreco, Marcus A. Badgeley, Kipp W. Johnson, Debajyoti Datta, Vivek A. Rudrapatna, Nadav Rappoport, Mark M. Shervey, Riccardo Miotto, Theodore C. Goldstein, Eugenia Rutenberg, Remi Frazier, Sharat Israni, Rick Larsen, Bethany Percha, Li Li 0062, Joel Dudley, Nicholas P. Tatonetti, Atul J. Butte |
Bioinform. | 20 |
| 2018 | Deep learning for healthcare: review, opportunities and challengesabstractGaining knowledge and actionable insights from complex, high-dimensional and heterogeneous biomedical data remains a key challenge in transforming health care. Various types of data have been emerging in modern biomedical research, including electronic health records, imaging, -omics, sensor data and text, which are complex, heterogeneous, poorly annotated and generally unstructured. Traditional data mining and statistical learning approaches typically need to first perform feature engineering to obtain effective and more robust features from those data, and then build prediction or clustering models on top of them. There are lots of challenges on both steps in a scenario of complicated data and lacking of sufficient domain knowledge. The latest advances in deep learning technologies provide new effective paradigms to obtain end-to-end learning models from complex data. In this article, we review the recent literature on applying deep learning technologies to advance the health care domain. Based on the analyzed work, we suggest that deep learning approaches could be the vehicle for translating big biomedical data into improved human health. However, we also note limitations and needs for improved methods development and applications, especially in terms of ease-of-understanding for domain experts and citizen scientists. We discuss such challenges and suggest developing holistic and meaningful interpretable architectures to bridge deep learning models and human interpretability. Riccardo Miotto, Fei Wang 0001, Shuang Wang 0002, Xiaoqian Jiang, Joel Dudley |
Briefings Bioinform. | 5 |
| 2018 | Systematic analyses of drugs and disease indications in RepurposeDB reveal pharmacological, biological and epidemiological factors influencing drug repositioningabstractIncrease in global population and growing disease burden due to the emergence of infectious diseases (Zika virus), multidrug-resistant pathogens, drug-resistant cancers (cisplatin-resistant ovarian cancer) and chronic diseases (arterial hypertension) necessitate effective therapies to improve health outcomes. However, the rapid increase in drug development cost demands innovative and sustainable drug discovery approaches. Drug repositioning, the discovery of new or improved therapies by reevaluation of approved or investigational compounds, solves a significant gap in the public health setting and improves the productivity of drug development. As the number of drug repurposing investigations increases, a new opportunity has emerged to understand factors driving drug repositioning through systematic analyses of drugs, drug targets and associated disease indications. However, such analyses have so far been hampered by the lack of a centralized knowledgebase, benchmarking data sets and reporting standards. To address these knowledge and clinical needs, here, we present RepurposeDB, a collection of repurposed drugs, drug targets and diseases, which was assembled, indexed and annotated from public data. RepurposeDB combines information on 253 drugs [small molecules (74.30%) and protein drugs (25.29%)] and 1125 diseases. Using RepurposeDB data, we identified pharmacological (chemical descriptors, physicochemical features and absorption, distribution, metabolism, excretion and toxicity properties), biological (protein domains, functional process, molecular mechanisms and pathway cross talks) and epidemiological (shared genetic architectures, disease comorbidities and clinical phenotype similarities) factors mediating drug repositioning. Collectively, RepurposeDB is developed as the reference database for drug repositioning investigations. The pharmacological, biological and epidemiological principles of drug repositioning identified from the meta-analyses could augment therapeutic development. Khader Shameer, Benjamin S. Glicksberg, Rachel Hodos, Kipp W. Johnson, Marcus A. Badgeley, Ben Readhead, Max S. Tomlinson, Timothy O'Connor, Riccardo Miotto, Brian A. Kidd, Rong Chen 0006, Avi Ma'ayan, Joel Dudley |
Briefings Bioinform. | 13 |
| 2018 | Uncovering exposures responsible for birth season - disease effects: a global studyabstractOBJECTIVE: Birth month and climate impact lifetime disease risk, while the underlying exposures remain largely elusive. We seek to uncover distal risk factors underlying these relationships by probing the relationship between global exposure variance and disease risk variance by birth season. MATERIAL AND METHODS: This study utilizes electronic health record data from 6 sites representing 10.5 million individuals in 3 countries (United States, South Korea, and Taiwan). We obtained birth month-disease risk curves from each site in a case-control manner. Next, we correlated each birth month-disease risk curve with each exposure. A meta-analysis was then performed of correlations across sites. This allowed us to identify the most significant birth month-exposure relationships supported by all 6 sites while adjusting for multiplicity. We also successfully distinguish relative age effects (a cultural effect) from environmental exposures. RESULTS: Attention deficit hyperactivity disorder was the only identified relative age association. Our methods identified several culprit exposures that correspond well with the literature in the field. These include a link between first-trimester exposure to carbon monoxide and increased risk of depressive disorder (R = 0.725, confidence interval [95% CI], 0.529-0.847), first-trimester exposure to fine air particulates and increased risk of atrial fibrillation (R = 0.564, 95% CI, 0.363-0.715), and decreased exposure to sunlight during the third trimester and increased risk of type 2 diabetes mellitus (R = -0.816, 95% CI, -0.5767, -0.929). CONCLUSION: A global study of birth month-disease relationships reveals distal risk factors involved in causal biological pathways that underlie them. Mary Regina Boland, Pradipta Parhi, Li Li 0062, Riccardo Miotto, Robert J. Carroll, Usman Iqbal, Phung Anh Nguyen, Martijn J. Schuemie, Seng Chan You, Donahue Smith, Sean D. Mooney, Patrick B. Ryan, Yu-Chuan Li, Rae Woong Park, Joshua C. Denny, Joel Dudley, George Hripcsak, Pierre Gentine, Nicholas P. Tatonetti |
J. Am. Medical Informatics Assoc. | 16 |
| 2017 | Translational bioinformatics in the era of real-time biomedical, health care and wellness data streamsabstractMonitoring and modeling biomedical, health care and wellness data from individuals and converging data on a population scale have tremendous potential to improve understanding of the transition to the healthy state of human physiology to disease setting. Wellness monitoring devices and companion software applications capable of generating alerts and sharing data with health care providers or social networks are now available. The accessibility and clinical utility of such data for disease or wellness research are currently limited. Designing methods for streaming data capture, real-time data aggregation, machine learning, predictive analytics and visualization solutions to integrate wellness or health monitoring data elements with the electronic medical records (EMRs) maintained by health care providers permits better utilization. Integration of population-scale biomedical, health care and wellness data would help to stratify patients for active health management and to understand clinically asymptomatic patients and underlying illness trajectories. In this article, we discuss various health-monitoring devices, their ability to capture the unique state of health represented in a patient and their application in individualized diagnostics, prognosis, clinical or wellness intervention. We also discuss examples of translational bioinformatics approaches to integrating patient-generated data with existing EMRs, personal health records, patient portals and clinical data repositories. Briefly, translational bioinformatics methods, tools and resources are at the center of these advances in implementing real-time biomedical and health care analytics in the clinical setting. Furthermore, these advances are poised to play a significant role in clinical decision-making and implementation of data-driven medicine and wellness care. Khader Shameer, Marcus A. Badgeley, Riccardo Miotto, Benjamin S. Glicksberg, Joseph W. Morgan, Joel Dudley |
Briefings Bioinform. | 6 |
| 2017 | Automated cell type discovery and classification through knowledge transferabstractMOTIVATION: Recent advances in mass cytometry allow simultaneous measurements of up to 50 markers at single-cell resolution. However, the high dimensionality of mass cytometry data introduces computational challenges for automated data analysis and hinders translation of new biological understanding into clinical applications. Previous studies have applied machine learning to facilitate processing of mass cytometry data. However, manual inspection is still inevitable and becoming the barrier to reliable large-scale analysis. RESULTS: We present a new algorithm called utomated ell-type iscovery and lassification (ACDC) that fully automates the classification of canonical cell populations and highlights novel cell types in mass cytometry data. Evaluations on real-world data show ACDC provides accurate and reliable estimations compared to manual gating results. Additionally, ACDC automatically classifies previously ambiguous cell types to facilitate discovery. Our findings suggest that ACDC substantially improves both reliability and interpretability of results obtained from high-dimensional mass cytometry profiling data. AVAILABILITY AND IMPLEMENTATION: A Python package (Python 3) and analysis scripts for reproducing the results are availability on https://bitbucket.org/dudleylab/acdc . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao-Chih Lee, Roman Kosoy, Christine E. Becker, Joel Dudley, Brian A. Kidd |
Bioinform. | 4 |
| 2017 | Predicting age by mining electronic medical records with deep learning characterizes differences between chronological and physiological age
Zichen Wang 0002, Li Li 0062, Benjamin S. Glicksberg, Ariel Israel, Joel Dudley, Avi Ma'ayan |
J. Biomed. Informatics | 5 |
| 2016 | Deep Learning to Predict Patient Future Diseases from the Electronic Health Records
Riccardo Miotto, Li Li 0062, Joel Dudley |
ECIR | 3 |
| 2016 | Healthcare Data Mining with Matrix ModelsabstractIn the last decade, advances in high-throughput technologies, growth of clinical data warehouses, and rapid accumulation of biomedical knowledge provided unprecedented opportunities and challenges to researchers in biomedical informatics. One distinct solution, to efficiently conduct big data analytics for biomedical problems, is the application of matrix computation and factorization methods such as non-negative matrix factorization, joint matrix factorization, tensor factorization. Compared to probabilistic and information theoretic approaches, matrix-based methods are fast, easy to understand and implement. In this tutorial, we provide a review of recent advances in algorithms and methods using matrix and their potential applications in biomedical informatics. We survey various related articles from data mining venues as well as from biomedical informatics venues to share with the audience key problems and trends in matrix computation research, with different novel applications such as drug repositioning, personalized medicine, and electronic phenotyping. Fei Wang 0001, Ping Zhang 0016, Joel Dudley |
KDD | 3 |
| 2016 | Interpreting functional effects of coding variants: challenges in proteome-scale prediction, annotation and assessmentabstractAccurate assessment of genetic variation in human DNA sequencing studies remains a nontrivial challenge in clinical genomics and genome informatics. Ascribing functional roles and/or clinical significances to single nucleotide variants identified from a next-generation sequencing study is an important step in genome interpretation. Experimental characterization of all the observed functional variants is yet impractical; thus, the prediction of functional and/or regulatory impacts of the various mutations using in silico approaches is an important step toward the identification of functionally significant or clinically actionable variants. The relationships between genotypes and the expressed phenotypes are multilayered and biologically complex; such relationships present numerous challenges and at the same time offer various opportunities for the design of in silico variant assessment strategies. Over the past decade, many bioinformatics algorithms have been developed to predict functional consequences of single nucleotide variants in the protein coding regions. In this review, we provide an overview of the bioinformatics resources for the prediction, annotation and visualization of coding single nucleotide variants. We discuss the currently available approaches and major challenges from the perspective of protein sequence, structure, function and interactions that require consideration when interpreting the impact of putatively functional variants. We also discuss the relevance of incorporating integrated workflows for predicting the biomedical impact of the functionally important variations encoded in a genome, exome or transcriptome. Finally, we propose a framework to classify variant assessment approaches and strategies for incorporation of variant assessment within electronic health records. Khader Shameer, Lokesh P. Tripathi, Krishna R. Kalari, Joel Dudley, Ramanathan Sowdhamini |
Briefings Bioinform. | 4 |
| 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networksabstractMOTIVATION: Underrepresentation of racial groups represents an important challenge and major gap in phenomics research. Most of the current human phenomics research is based primarily on European populations; hence it is an important challenge to expand it to consider other population groups. One approach is to utilize data from EMR databases that contain patient data from diverse demographics and ancestries. The implications of this racial underrepresentation of data can be profound regarding effects on the healthcare delivery and actionability. To the best of our knowledge, our work is the first attempt to perform comparative, population-scale analyses of disease networks across three different populations, namely Caucasian (EA), African American (AA) and Hispanic/Latino (HL). RESULTS: We compared susceptibility profiles and temporal connectivity patterns for 1988 diseases and 37 282 disease pairs represented in a clinical population of 1 025 573 patients. Accordingly, we revealed appreciable differences in disease susceptibility, temporal patterns, network structure and underlying disease connections between EA, AA and HL populations. We found 2158 significantly comorbid diseases for the EA cohort, 3265 for AA and 672 for HL. We further outlined key disease pair associations unique to each population as well as categorical enrichments of these pairs. Finally, we identified 51 key 'hub' diseases that are the focal points in the race-centric networks and of particular clinical importance. Incorporating race-specific disease comorbidity patterns will produce a more accurate and complete picture of the disease landscape overall and could support more precise understanding of disease relationships and patient management towards improved clinical outcomes. CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin S. Glicksberg, Li Li 0062, Marcus A. Badgeley, Khader Shameer, Roman Kosoy, Noam D. Beckmann, Nam Pho, Jörg Hakenberg, Kristin L. Ayers, Gabriel E. Hoffman, Shuyu Dan Li, Eric E. Schadt, Chirag J. Patel, Rong Chen 0006, Joel Dudley |
Bioinform. | 16 |
| 2012 | Integrative Approach to Pain Genetics Identifies Pain Sensitivity Loci across DiseasesabstractIdentifying human genes relevant for the processing of pain requires difficult-to-conduct and expensive large-scale clinical trials. Here, we examine a novel integrative paradigm for data-driven discovery of pain gene candidates, taking advantage of the vast amount of existing disease-related clinical literature and gene expression microarray data stored in large international repositories. First, thousands of diseases were ranked according to a disease-specific pain index (DSPI), derived from Medical Subject Heading (MESH) annotations in MEDLINE. Second, gene expression profiles of 121 of these human diseases were obtained from public sources. Third, genes with expression variation significantly correlated with DSPI across diseases were selected as candidate pain genes. Finally, selected candidate pain genes were genotyped in an independent human cohort and prospectively evaluated for significant association between variants and measures of pain sensitivity. The strongest signal was with rs4512126 (5q32, ABLIM3, P = 1.3×10⁻¹⁰) for the sensitivity to cold pressor pain in males, but not in females. Significant associations were also observed with rs12548828, rs7826700 and rs1075791 on 8q22.2 within NCALD (P = 1.7×10⁻⁴, 1.8×10⁻⁴, and 2.2×10⁻⁴ respectively). Our results demonstrate the utility of a novel paradigm that integrates publicly available disease-specific gene expression data with clinical data curated from MEDLINE to facilitate the discovery of pain-relevant genes. This data-derived list of pain gene candidates enables additional focused and efficient biological studies validating additional candidates. David Ruau, Joel Dudley, Rong Chen 0006, Nicholas G. Phillips, Gary E. Swan, Laura Lazzeroni, J. David Clark, Atul J. Butte, Martin S. Angst |
PLoS Comput. Biol. | 2 |
| 2011 | Exploiting drug-disease relationships for computational drug repositioningabstractFinding new uses for existing drugs, or drug repositioning, has been used as a strategy for decades to get drugs to more patients. As the ability to measure molecules in high-throughput ways has improved over the past decade, it is logical that such data might be useful for enabling drug repositioning through computational methods. Many computational predictions for new indications have been borne out in cellular model systems, though extensive animal model and clinical trial-based validation are still pending. In this review, we show that computational methods for drug repositioning can be classified in two axes: drug based, where discovery initiates from the chemical perspective, or disease based, where discovery initiates from the clinical perspective of disease or its pathology. Newer algorithms for computational drug repositioning will likely span these two axes, will take advantage of newer types of molecular measurements, and will certainly play a role in reducing the global burden of disease. Joel Dudley, Tarangini Deshpande, Atul J. Butte |
Briefings Bioinform. | 1 |
| 2011 | ProfileChaser: searching microarray repositories based on genome-wide patterns of differential expressionabstractSUMMARY: We introduce ProfileChaser, a web server that allows for querying the Gene Expression Omnibus based on genome-wide patterns of differential expression. Using a novel, content-based approach, ProfileChaser retrieves expression profiles that match the differentially regulated transcriptional programs in a user-supplied experiment. This analysis identifies statistical links to similar expression experiments from the vast array of publicly available data on diseases, drugs, phenotypes and other experimental conditions. AVAILABILITY: http://profilechaser.stanford.edu CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jesse M. Engreitz, Rong Chen 0006, Alexander A. Morgan, Joel Dudley, Rohan Mallelwar, Atul J. Butte |
Bioinform. | 4 |
| 2011 | Comparison of automated and human assignment of MeSH terms on publicly-available molecular datasetsabstractPublicly available molecular datasets can be used for independent verification or investigative repurposing, but depends on the presence, consistency and quality of descriptive annotations. Annotation and indexing of molecular datasets using well-defined controlled vocabularies or ontologies enables accurate and systematic data discovery, yet the majority of molecular datasets available through public data repositories lack such annotations. A number of automated annotation methods have been developed; however few systematic evaluations of the quality of annotations supplied by application of these methods have been performed using annotations from standing public data repositories. Here, we compared manually-assigned Medical Subject Heading (MeSH) annotations associated with experiments by data submitters in the PRoteomics IDEntification (PRIDE) proteomics data repository to automated MeSH annotations derived through the National Center for Biomedical Ontology Annotator and National Library of Medicine MetaMap programs. These programs were applied to free-text annotations for experiments in PRIDE. As many submitted datasets were referenced in publications, we used the manually curated MeSH annotations of those linked publications in MEDLINE as "gold standard". Annotator and MetaMap exhibited recall performance 3-fold greater than that of the manual annotations. We connected PRIDE experiments in a network topology according to shared MeSH annotations and found 373 distinct clusters, many of which were found to be biologically coherent by network analysis. The results of this study suggest that both Annotator and MetaMap are capable of annotating public molecular datasets with a quality comparable, and often exceeding, that of the actual data submitters, highlighting a continuous need to improve and apply automated methods to molecular datasets in public data repositories to maximize their value and utility. David Ruau, Michael Mbagwu, Joel Dudley, Vijay Krishnan, Atul J. Butte |
J. Biomed. Informatics | 3 |
| 2010 | Latent physiological factors of complex human diseases revealed by independent component analysis of clinarraysabstractBACKGROUND: Diagnosis and treatment of patients in the clinical setting is often driven by known symptomatic factors that distinguish one particular condition from another. Treatment based on noticeable symptoms, however, is limited to the types of clinical biomarkers collected, and is prone to overlooking dysfunctions in physiological factors not easily evident to medical practitioners. We used a vector-based representation of patient clinical biomarkers, or clinarrays, to search for latent physiological factors that underlie human diseases directly from clinical laboratory data. Knowledge of these factors could be used to improve assessment of disease severity and help to refine strategies for diagnosis and monitoring disease progression. RESULTS: Applying Independent Component Analysis on clinarrays built from patient laboratory measurements revealed both known and novel concomitant physiological factors for asthma, types 1 and 2 diabetes, cystic fibrosis, and Duchenne muscular dystrophy. Serum sodium was found to be the most significant factor for both type 1 and type 2 diabetes, and was also significant in asthma. TSH3, a measure of thyroid function, and blood urea nitrogen, indicative of kidney function, were factors unique to type 1 diabetes respective to type 2 diabetes. Platelet count was significant across all the diseases analyzed. CONCLUSIONS: The results demonstrate that large-scale analyses of clinical biomarkers using unsupervised methods can offer novel insights into the pathophysiological basis of human disease, and suggest novel clinical utility of established laboratory measurements. David P. Chen, Joel Dudley, Atul J. Butte |
BMC Bioinform. | 2 |
| 2010 | Content-based microarray search using differential expression profilesabstractBACKGROUND: With the expansion of public repositories such as the Gene Expression Omnibus (GEO), we are rapidly cataloging cellular transcriptional responses to diverse experimental conditions. Methods that query these repositories based on gene expression content, rather than textual annotations, may enable more effective experiment retrieval as well as the discovery of novel associations between drugs, diseases, and other perturbations. RESULTS: We develop methods to retrieve gene expression experiments that differentially express the same transcriptional programs as a query experiment. Avoiding thresholds, we generate differential expression profiles that include a score for each gene measured in an experiment. We use existing and novel dimension reduction and correlation measures to rank relevant experiments in an entirely data-driven manner, allowing emergent features of the data to drive the results. A combination of matrix decomposition and p-weighted Pearson correlation proves the most suitable for comparing differential expression profiles. We apply this method to index all GEO DataSets, and demonstrate the utility of our approach by identifying pathways and conditions relevant to transcription factors Nanog and FoxO3. CONCLUSIONS: Content-based gene expression search generates relevant hypotheses for biological inquiry. Experiments across platforms, tissue types, and protocols inform the analysis of new datasets. Jesse M. Engreitz, Alexander A. Morgan, Joel Dudley, Rong Chen 0006, Rahul Thathoo, Russ B. Altman, Atul J. Butte |
BMC Bioinform. | 3 |
| 2010 | An integrative method for scoring candidate genes from association studies: application to warfarin dosingabstractBACKGROUND: A key challenge in pharmacogenomics is the identification of genes whose variants contribute to drug response phenotypes, which can include severe adverse effects. Pharmacogenomics GWAS attempt to elucidate genotypes predictive of drug response. However, the size of these studies has severely limited their power and potential application. We propose a novel knowledge integration and SNP aggregation approach for identifying genes impacting drug response. Our SNP aggregation method characterizes the degree to which uncommon alleles of a gene are associated with drug response. We first use pre-existing knowledge sources to rank pharmacogenes by their likelihood to affect drug response. We then define a summary score for each gene based on allele frequencies and train linear and logistic regression classifiers to predict drug response phenotypes. RESULTS: We applied our method to a published warfarin GWAS data set comprising 181 individuals. We find that our method can increase the power of the GWAS to identify both VKORC1 and CYP2C9 as warfarin pharmacogenes, where the original analysis had only identified VKORC1. Additionally, we find that our method can be used to discriminate between low-dose (AUROC=0.886) and high-dose (AUROC=0.764) responders. CONCLUSIONS: Our method offers a new route for candidate pharmacogene discovery from pharmacogenomics GWAS, and serves as a foundation for future work in methods for predictive pharmacogenomics. Nicholas P. Tatonetti, Joel Dudley, Hersh Sagreiya, Atul J. Butte, Russ B. Altman |
BMC Bioinform. | 2 |
| 2010 | Differentially Expressed RNA from Public Microarray Data Identifies Serum Protein Biomarkers for Cross-Organ Transplant Rejection and Other ConditionsabstractSerum proteins are routinely used to diagnose diseases, but are hard to find due to low sensitivity in screening the serum proteome. Public repositories of microarray data, such as the Gene Expression Omnibus (GEO), contain RNA expression profiles for more than 16,000 biological conditions, covering more than 30% of United States mortality. We hypothesized that genes coding for serum- and urine-detectable proteins, and showing differential expression of RNA in disease-damaged tissues would make ideal diagnostic protein biomarkers for those diseases. We showed that predicted protein biomarkers are significantly enriched for known diagnostic protein biomarkers in 22 diseases, with enrichment significantly higher in diseases for which at least three datasets are available. We then used this strategy to search for new biomarkers indicating acute rejection (AR) across different types of transplanted solid organs. We integrated three biopsy-based microarray studies of AR from pediatric renal, adult renal and adult cardiac transplantation and identified 45 genes upregulated in all three. From this set, we chose 10 proteins for serum ELISA assays in 39 renal transplant patients, and discovered three that were significantly higher in AR. Interestingly, all three proteins were also significantly higher during AR in the 63 cardiac transplant recipients studied. Our best marker, serum PECAM1, identified renal AR with 89% sensitivity and 75% specificity, and also showed increased expression in AR by immunohistochemistry in renal, hepatic and cardiac transplant biopsies. Our results demonstrate that integrating gene expression microarray measurements from disease samples and even publicly-available data sets can be a powerful, fast, and cost-effective strategy for the discovery of new diagnostic serum protein biomarkers. Rong Chen 0006, Tara K. Sigdel, Li Li 0062, Neeraja Kambham, Joel Dudley, Szu-chuan Hsieh, R. Bryan Klassen, Amery Chen, Tuyen Caohuu, Alexander A. Morgan, Hannah A. Valantine, Kiran K. Khush, Minnie M. Sarwal, Atul J. Butte |
PLoS Comput. Biol. | 5 |
| 2010 | Network-Based Elucidation of Human Disease Similarities Reveals Common Functional Modules Enriched for Pluripotent Drug TargetsabstractCurrent work in elucidating relationships between diseases has largely been based on pre-existing knowledge of disease genes. Consequently, these studies are limited in their discovery of new and unknown disease relationships. We present the first quantitative framework to compare and contrast diseases by an integrated analysis of disease-related mRNA expression data and the human protein interaction network. We identified 4,620 functional modules in the human protein network and provided a quantitative metric to record their responses in 54 diseases leading to 138 significant similarities between diseases. Fourteen of the significant disease correlations also shared common drugs, supporting the hypothesis that similar diseases can be treated by the same drugs, allowing us to make predictions for new uses of existing drugs. Finally, we also identified 59 modules that were dysregulated in at least half of the diseases, representing a common disease-state "signature". These modules were significantly enriched for genes that are known to be drug targets. Interestingly, drugs known to target these genes/proteins are already known to treat significantly more diseases than drugs targeting other genes/proteins, highlighting the importance of these core modules as prime therapeutic opportunities. Silpa Suthram, Joel Dudley, Annie P. Chiang, Rong Chen 0006, Trevor J. Hastie, Atul J. Butte |
PLoS Comput. Biol. | 2 |
| 2009 | A Quick Guide for Developing Effective Bioinformatics Programming SkillsabstractBioinformatics programming skills are becoming a necessity across many facets of biology and medicine, owed in part to the continuing explosion of biological data Joel Dudley, Atul J. Butte |
PLoS Comput. Biol. | 1 |
| 2008 | MEGA: A biologist-centric software for evolutionary analysis of DNA and protein sequencesabstractThe Molecular Evolutionary Genetics Analysis (MEGA) software is a desktop application designed for comparative analysis of homologous gene sequences either from multigene families or from different species with a special emphasis on inferring evolutionary relationships and patterns of DNA and protein evolution. In addition to the tools for statistical analysis of data, MEGA provides many convenient facilities for the assembly of sequence data sets from files or web-based repositories, and it includes tools for visual presentation of the results obtained in the form of interactive phylogenetic trees and evolutionary distance matrices. Here we discuss the motivation, design principles and priorities that have shaped the development of MEGA. We also discuss how MEGA might evolve in the future to assist researchers in their growing need to analyze large data set using new computational methods. Sudhir Kumar 0001, Masatoshi Nei, Joel Dudley, Koichiro Tamura |
Briefings Bioinform. | 3 |
| 2007 | Bioinformatics software for biologists in the genomics eraabstractMOTIVATION: The genome sequencing revolution is approaching a landmark figure of 1000 completely sequenced genomes. Coupled with fast-declining, per-base sequencing costs, this influx of DNA sequence data has encouraged laboratory scientists to engage large datasets in comparative sequence analyses for making evolutionary, functional and translational inferences. However, the majority of the scientists at the forefront of experimental research are not bioinformaticians, so a gap exists between the user-friendly software needed and the scripting/programming infrastructure often employed for the analysis of large numbers of genes, long genomic segments and groups of sequences. We see an urgent need for the expansion of the fundamental paradigms under which biologist-friendly software tools are designed and developed to fulfill the needs of biologists to analyze large datasets by using sophisticated computational methods. We argue that the design principles need to be sensitive to the reality that comparatively small teams of biologists have historically developed some of the most popular biological software packages in molecular evolutionary analysis. Furthermore, biological intuitiveness and investigator empowerment need to take precedence over the current supposition that biologists should re-tool and become programmers when analyzing genome scale datasets. Sudhir Kumar 0001, Joel Dudley |
Bioinform. | 2 |
| 2006 | TimeTree: a public knowledge-base of divergence times among organismsabstractUNLABELLED: Biologists and other scientists routinely need to know times of divergence between species and to construct phylogenies calibrated to time (timetrees). Published studies reporting time estimates from molecular data have been increasing rapidly, but the data have been largely inaccessible to the greater community of scientists because of their complexity. TimeTree brings these data together in a consistent format and uses a hierarchical structure, corresponding to the tree of life, to maximize their utility. Results are presented and summarized, allowing users to quickly determine the range and robustness of time estimates and the degree of consensus from the published literature. AVAILABILITY: TimeTree is available at http://www.timetree.net S. Blair Hedges, Joel Dudley, Sudhir Kumar 0001 |
Bioinform. | 2 |