EDBT 2026 Demo / reviewers in the wild / expert
Chirag J. Patel
dblp:42/11538
· DBLP profile ↗
11ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0002-8756-8525ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 64% Medical and health informatics · 36% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Medical and health informatics › biomedical data science
biomedical database querying |
0.3 | 1 | 2018 | Rcupcake: an R package for querying and analyzing biomedical data through the BD2K PIC-SURE RESTful API · Bioinform. 2018 |
Bioinformatics and computational biology › genomics
genomic data analysis |
0.3 | 1 | 2018 | Aether: leveraging linear programming for optimal cloud computing in genomics · Bioinform. 2018 |
Medical and health informatics
precision medicine |
0.3 | 1 | 2018 | Rcupcake: an R package for querying and analyzing biomedical data through the BD2K PIC-SURE RESTful API · Bioinform. 2018 |
Cloud and datacenter computing › resource management
cloud resource management |
0.3 | 1 | 2018 | Aether: leveraging linear programming for optimal cloud computing in genomics · Bioinform. 2018 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
disease network analysis |
0.2 | 1 | 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networks · Bioinform. 2016 |
Medical and health informatics
electronic health records |
0.2 | 1 | 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networks · Bioinform. 2016 |
Bioinformatics and computational biology
phenomics |
0.2 | 1 | 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networks · Bioinform. 2016 |
Bioinformatics and computational biology › data integration
cross-platform data integration |
0.2 | 1 | 2015 | aRrayLasso: a network-based approach to microarray interconversion · Bioinform. 2015 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.2 | 1 | 2015 | aRrayLasso: a network-based approach to microarray interconversion · Bioinform. 2015 |
Bioinformatics and computational biology
transcriptomics |
0.2 | 1 | 2015 | aRrayLasso: a network-based approach to microarray interconversion · Bioinform. 2015 |
Bioinformatics and computational biology › statistical genetics
gene-environment interaction |
0.1 | 1 | 2012 | Data-driven integration of epidemiological and toxicological data to select candidate interacting genes and environmental factors in association with disease · Bioinform. 2012 |
Medical and health informatics › clinical data analysis
multimorbidity analysis |
0.1 | 1 | 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networks · Bioinform. 2016 |
Bioinformatics and computational biology › multi-omics data integration
gene expression data integration |
0.1 | 1 | 2015 | aRrayLasso: a network-based approach to microarray interconversion · Bioinform. 2015 |
Bioinformatics and computational biology
data integration |
0.0 | 1 | 2012 | Data-driven integration of epidemiological and toxicological data to select candidate interacting genes and environmental factors in association with disease · Bioinform. 2012 |
Methods — techniques the papers use, named apart from their topics
linear programming · 0.7data visualization · 0.3RESTful API · 0.3network analysis · 0.2lasso-penalized generalized linear model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A new method for estimating the probability of causal relationships from observational data: Application to the study of the short-term effects of air pollution on cardiovascular and respiratory disease
Bryan Andrews, Chirayu Wongchokprasitti, Shyam Visweswaran, Chirag M. Lakhani, Chirag J. Patel, Gregory F. Cooper |
Artif. Intell. Medicine | 5 |
| 2020 | Scalability and cost-effectiveness analysis of whole genome-wide association studies on Google Cloud Platform and Amazon Web ServicesabstractOBJECTIVE: Advancements in human genomics have generated a surge of available data, fueling the growth and accessibility of databases for more comprehensive, in-depth genetic studies. METHODS: We provide a straightforward and innovative methodology to optimize cloud configuration in order to conduct genome-wide association studies. We utilized Spark clusters on both Google Cloud Platform and Amazon Web Services, as well as Hail (http://doi.org/10.5281/zenodo.2646680) for analysis and exploration of genomic variants dataset. RESULTS: Comparative evaluation of numerous cloud-based cluster configurations demonstrate a successful and unprecedented compromise between speed and cost for performing genome-wide association studies on 4 distinct whole-genome sequencing datasets. Results are consistent across the 2 cloud providers and could be highly useful for accelerating research in genetics. CONCLUSIONS: We present a timely piece for one of the most frequently asked questions when moving to the cloud: what is the trade-off between speed and cost? Inès Krissaane, Carlos De Niz, Alba Gutiérrez-Sacristán, Gabor Korodi, Nneka Ede, Ranjay Kumar, Jessica Lyons, Arjun K. Manrai, Chirag J. Patel, Isaac S. Kohane, Paul Avillach |
J. Am. Medical Informatics Assoc. | 9 |
| 2020 | A systematic machine learning and data type comparison yields metagenomic predictors of infant age, sex, breastfeeding, antibiotic usage, country of origin, and delivery typeabstractThe microbiome is a new frontier for building predictors of human phenotypes. However, machine learning in the microbiome is fraught with issues of reproducibility, driven in large part by the wide range of analytic models and metagenomic data types available. We aimed to build robust metagenomic predictors of host phenotype by comparing prediction performances and biological interpretation across 8 machine learning methods and 4 different types of metagenomic data. Using 1,570 samples from 300 infants, we fit 7,865 models for 6 host phenotypes. We demonstrate the dependence of accuracy on algorithm choice and feature definition in microbiome data and propose a framework for building microbiome-derived indicators of host phenotype. We additionally identify biological features predictive of age, sex, breastfeeding status, historical antibiotic usage, country of origin, and delivery type. Our complete results can be viewed at http://apps.chiragjpgroup.org/ubiome_predictions/. Alan Le Goallec, Braden T. Tierney, Jacob M. Luber, Evan M. Cofer, Aleksandar D. Kostic, Chirag J. Patel |
PLoS Comput. Biol. | 6 |
| 2018 | A review of validation strategies for computational drug repositioningabstractRepositioning of previously approved drugs is a promising methodology because it reduces the cost and duration of the drug development pipeline and reduces the likelihood of unforeseen adverse events. Computational repositioning is especially appealing because of the ability to rapidly screen candidates in silico and to reduce the number of possible repositioning candidates. What is unclear, however, is how useful such methods are in producing clinically efficacious repositioning hypotheses. Furthermore, there is no agreement in the field over the proper way to perform validation of in silico predictions, and in fact no systematic review of repositioning validation methodologies. To address this unmet need, we review the computational repositioning literature and capture studies in which authors claimed to have validated their work. Our analysis reveals widespread variation in the types of strategies, predictions made and databases used as 'gold standards'. We highlight a key weakness of the most commonly used strategy and propose a path forward for the consistent analytic validation of repositioning techniques. Adam S. Brown, Chirag J. Patel |
Briefings Bioinform. | 2 |
| 2018 | Rcupcake: an R package for querying and analyzing biomedical data through the BD2K PIC-SURE RESTful APIabstractMotivation: In the era of big data and precision medicine, the number of databases containing clinical, environmental, self-reported and biochemical variables is increasing exponentially. Enabling the experts to focus on their research questions rather than on computational data management, access and analysis is one of the most significant challenges nowadays. Results: We present Rcupcake, an R package that contains a variety of functions for leveraging different databases through the BD2K PIC-SURE RESTful API and facilitating its query, analysis and interpretation. The package offers a variety of analysis and visualization tools, including the study of the phenotype co-occurrence and prevalence, according to multiple layers of data, such as phenome, exposome or genome. Availability and implementation: The package is implemented in R and is available under Mozilla v2 license from GitHub (https://github.com/hms-dbmi/Rcupcake). Two reproducible case studies are also available (https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/SSCcaseStudy_v01.ipynb, https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/NHANEScaseStudy_v01.ipynb). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Alba Gutiérrez-Sacristán, Romain Guedj, Gabor Korodi, Jason Stedman, Laura Inés Furlong, Chirag J. Patel, Isaac S. Kohane, Paul Avillach |
Bioinform. | 6 |
| 2018 | Aether: leveraging linear programming for optimal cloud computing in genomicsabstractMotivation: Across biology, we are seeing rapid developments in scale of data production without a corresponding increase in data analysis capabilities. Results: Here, we present Aether (http://aether.kosticlab.org), an intuitive, easy-to-use, cost-effective and scalable framework that uses linear programming to optimally bid on and deploy combinations of underutilized cloud computing resources. Our approach simultaneously minimizes the cost of data analysis and provides an easy transition from users' existing HPC pipelines. Availability and implementation: Data utilized are available at https://pubs.broadinstitute.org/diabimmune and with EBI SRA accession ERP005989. Source code is available at (https://github.com/kosticlab/aether). Examples, documentation and a tutorial are available at http://aether.kosticlab.org. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Jacob M. Luber, Braden T. Tierney, Evan M. Cofer, Chirag J. Patel, Aleksandar D. Kostic |
Bioinform. | 4 |
| 2017 | MeSHDD: Literature-based drug-drug similarity for drug repositioningabstractOBJECTIVE: Drug repositioning is a promising methodology for reducing the cost and duration of the drug discovery pipeline. We sought to develop a computational repositioning method leveraging annotations in the literature, such as Medical Subject Heading (MeSH) terms. METHODS: We developed software to determine significantly co-occurring drug-MeSH term pairs and a method to estimate pair-wise literature-derived distances between drugs. RESULTS: We found that literature-based drug-drug similarities predicted the number of shared indications across drug-drug pairs. Clustering drugs based on their similarity revealed both known and novel drug indications. We demonstrate the utility of our approach by generating repositioning hypotheses for the commonly used diabetes drug metformin. CONCLUSION: Our study demonstrates that literature-derived similarity is useful for identifying potential repositioning opportunities. We provided open-source code and deployed a free-to-use, interactive application to explore our database of similarity-based drug clusters (available at http://apps.chiragjpgroup.org/MeSHDD/ ). Adam S. Brown, Chirag J. Patel |
J. Am. Medical Informatics Assoc. | 2 |
| 2016 | Comparative analyses of population-scale phenomic data in electronic medical records reveal race-specific disease networksabstractMOTIVATION: Underrepresentation of racial groups represents an important challenge and major gap in phenomics research. Most of the current human phenomics research is based primarily on European populations; hence it is an important challenge to expand it to consider other population groups. One approach is to utilize data from EMR databases that contain patient data from diverse demographics and ancestries. The implications of this racial underrepresentation of data can be profound regarding effects on the healthcare delivery and actionability. To the best of our knowledge, our work is the first attempt to perform comparative, population-scale analyses of disease networks across three different populations, namely Caucasian (EA), African American (AA) and Hispanic/Latino (HL). RESULTS: We compared susceptibility profiles and temporal connectivity patterns for 1988 diseases and 37 282 disease pairs represented in a clinical population of 1 025 573 patients. Accordingly, we revealed appreciable differences in disease susceptibility, temporal patterns, network structure and underlying disease connections between EA, AA and HL populations. We found 2158 significantly comorbid diseases for the EA cohort, 3265 for AA and 672 for HL. We further outlined key disease pair associations unique to each population as well as categorical enrichments of these pairs. Finally, we identified 51 key 'hub' diseases that are the focal points in the race-centric networks and of particular clinical importance. Incorporating race-specific disease comorbidity patterns will produce a more accurate and complete picture of the disease landscape overall and could support more precise understanding of disease relationships and patient management towards improved clinical outcomes. CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin S. Glicksberg, Li Li 0062, Marcus A. Badgeley, Khader Shameer, Roman Kosoy, Noam D. Beckmann, Nam Pho, Jörg Hakenberg, Kristin L. Ayers, Gabriel E. Hoffman, Shuyu Dan Li, Eric E. Schadt, Chirag J. Patel, Rong Chen 0006, Joel Dudley |
Bioinform. | 14 |
| 2016 | ksRepo: a generalized platform for computational drug repositioningabstractBACKGROUND: Repositioning approved drug and small molecules in novel therapeutic areas is of key interest to the pharmaceutical industry. A number of promising computational techniques have been developed to aid in repositioning, however, the majority of available methodologies require highly specific data inputs that preclude the use of many datasets and databases. There is a clear unmet need for a generalized methodology that enables the integration of multiple types of both gene expression data and database schema. RESULTS: ksRepo eliminates the need for a single microarray platform as input and allows for the use of a variety of drug and chemical exposure databases. We tested ksRepo's performance on a set of five prostate cancer datasets using the Comparative Toxicogenomics Database (CTD) as our database of gene-compound interactions. ksRepo successfully predicted significance for five frontline prostate cancer therapies, representing a significant enrichment from over 7000 CTD compounds, and achieved specificity similar to other repositioning methods. CONCLUSIONS: We present ksRepo, which enables investigators to use any data inputs for computational drug repositioning. ksRepo is implemented in a series of four functions in the R statistical environment under a BSD3 license. Source code is freely available at http://github.com/adam-sam-brown/ksRepo. A vignette is provided to aid users in performing ksRepo analysis. Adam S. Brown, Sek Won Kong, Isaac S. Kohane, Chirag J. Patel |
BMC Bioinform. | 4 |
| 2015 | aRrayLasso: a network-based approach to microarray interconversionabstractUNLABELLED: Robust conversion between microarray platforms is needed to leverage the wide variety of microarray expression studies that have been conducted to date. Currently available conversion methods rely on manufacturer annotations, which are often incomplete, or on direct alignment of probes from different platforms, which often fail to yield acceptable genewise correlation. Here, we describe aRrayLasso, which uses the Lasso-penalized generalized linear model to model the relationships between individual probes in different probe sets. We have implemented aRrayLasso in a set of five open-source R functions that allow the user to acquire data from public sources such as Gene Expression Omnibus, train a set of Lasso models on that data and directly map one microarray platform to another. aRrayLasso significantly predicts expression levels with similar fidelity to technical replicates of the same RNA pool, demonstrating its utility in the integration of datasets from different platforms. AVAILABILITY AND IMPLEMENTATION: All functions are available, along with descriptions, at https://github.com/adam-sam-brown/aRrayLasso. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Adam S. Brown, Chirag J. Patel |
Bioinform. | 2 |
| 2012 | Data-driven integration of epidemiological and toxicological data to select candidate interacting genes and environmental factors in association with diseaseabstractMOTIVATION: Complex diseases, such as Type 2 Diabetes Mellitus (T2D), result from the interplay of both environmental and genetic factors. However, most studies investigate either the genetics or the environment and there are a few that study their possible interaction in context of disease. One key challenge in documenting interactions between genes and environment includes choosing which of each to test jointly. Here, we attempt to address this challenge through a data-driven integration of epidemiological and toxicological studies. Specifically, we derive lists of candidate interacting genetic and environmental factors by integrating findings from genome-wide and environment-wide association studies. Next, we search for evidence of toxicological relationships between these genetic and environmental factors that may have an etiological role in the disease. We illustrate our method by selecting candidate interacting factors for T2D. Chirag J. Patel, Rong Chen 0006, Atul J. Butte |
Bioinform. | 1 |