EDBT 2026 Demo / reviewers in the wild / expert
Michael Snyder 0001
dblp:50/5034 · also Michael P. Snyder
· DBLP profile ↗
39ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0003-0784-7987ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Longitudinal urine metabolic profiling and gestational age prediction in human pregnancyabstractPregnancy is a vital period affecting both maternal and fetal health, with impacts on maternal metabolism, fetal growth, and long-term development. While the maternal metabolome undergoes significant changes during pregnancy, longitudinal shifts in maternal urine have been largely unexplored. In this study, we applied liquid chromatography-mass spectrometry-based untargeted metabolomics to analyze 346 maternal urine samples collected throughout pregnancy from 36 women with diverse backgrounds and clinical profiles. Key metabolite changes included glucocorticoids, lipids, and amino acid derivatives, indicating systematic pathway alterations. We also developed a machine learning model to accurately predict gestational age using urine metabolites, offering a non-invasive pregnancy dating method. Additionally, we demonstrated the ability of the urine metabolome to predict time-to-delivery, providing a complementary tool for prenatal care and delivery planning. This study highlights the clinical potential of urine untargeted metabolomics in obstetric care. Xiaotao Shen, Songjie Chen, Monika Avina, Hanyah Zackriah, Laura Jelliffe-Pawlowski, Larry Rand, Michael Snyder 0001 |
Briefings Bioinform. | 8 |
| 2024 | PRS-Net: Interpretable Polygenic Risk Scores via Geometric Learning
Han Li 0018, Jianyang Zeng 0001, Michael Snyder 0001 |
RECOMB | 3 |
| 2022 | Deep learning-based pseudo-mass spectrometry imaging analysis for precision medicineabstractLiquid chromatography-mass spectrometry (LC-MS)-based untargeted metabolomics provides systematic profiling of metabolic. Yet, its applications in precision medicine (disease diagnosis) have been limited by several challenges, including metabolite identification, information loss and low reproducibility. Here, we present the deep-learning-based Pseudo-Mass Spectrometry Imaging (deepPseudoMSI) project (https://www.deeppseudomsi.org/), which converts LC-MS raw data to pseudo-MS images and then processes them by deep learning for precision medicine, such as disease diagnosis. Extensive tests based on real data demonstrated the superiority of deepPseudoMSI over traditional approaches and the capacity of our method to achieve an accurate individualized diagnosis. Our framework lays the foundation for future metabolic-based precision medicine. Xiaotao Shen, Wei Shao 0008, Chuchu Wang, Songjie Chen, Mirabela Rusu, Michael Snyder 0001 |
Briefings Bioinform. | 8 |
| 2022 | Robust identification of temporal biomarkers in longitudinal omics studiesabstractMOTIVATION: Longitudinal studies increasingly collect rich 'omics' data sampled frequently over time and across large cohorts to capture dynamic health fluctuations and disease transitions. However, the generation of longitudinal omics data has preceded the development of analysis tools that can efficiently extract insights from such data. In particular, there is a need for statistical frameworks that can identify not only which omics features are differentially regulated between groups but also over what time intervals. Additionally, longitudinal omics data may have inconsistencies, including non-uniform sampling intervals, missing data points, subject dropout and differing numbers of samples per subject. RESULTS: In this work, we developed OmicsLonDA, a statistical method that provides robust identification of time intervals of temporal omics biomarkers. OmicsLonDA is based on a semi-parametric approach, in which we use smoothing splines to model longitudinal data and infer significant time intervals of omics features based on an empirical distribution constructed through a permutation procedure. We benchmarked OmicsLonDA on five simulated datasets with diverse temporal patterns, and the method showed specificity greater than 0.99 and sensitivity greater than 0.87. Applying OmicsLonDA to the iPOP cohort revealed temporal patterns of genes, proteins, metabolites and microbes that are differentially regulated in male versus female subjects following a respiratory infection. In addition, we applied OmicsLonDA to a longitudinal multi-omics dataset of pregnant women with and without preeclampsia, and OmicsLonDA identified potential lipid markers that are temporally significantly different between the two groups. AVAILABILITY AND IMPLEMENTATION: We provide an open-source R package (https://bioconductor.org/packages/OmicsLonDA), to enable widespread use. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ahmed Metwally 0002, Tom Zhang, Ryan Kellogg, Wenyu Zhou, Kévin Contrepois, Hua Tang, Michael Snyder 0001 |
Bioinform. | 8 |
| 2022 | metID: an R package for automatable compound annotation for LC-MS-based dataabstractSUMMARY: Accurate and efficient compound annotation is a long-standing challenge for LC-MS-based data (e.g. untargeted metabolomics and exposomics). Substantial efforts have been devoted to overcoming this obstacle, whereas current tools are limited by the sources of spectral information used (in-house and public databases) and are not automated and streamlined. Therefore, we developed metID, an R package that combines information from all major databases for comprehensive and streamlined compound annotation. metID is a flexible, simple and powerful tool that can be installed on all platforms, allowing the compound annotation process to be fully automatic and reproducible. A detailed tutorial and a case study are provided in Supplementary Materials. AVAILABILITY AND IMPLEMENTATION: https://jaspershen.github.io/metID. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaotao Shen, Songjie Chen, Kévin Contrepois, Zheng-Jiang Zhu, Michael Snyder 0001 |
Bioinform. | 7 |
| 2022 | massDatabase: utilities for the operation of the public compound and pathway databaseabstractSUMMARY: One of the major challenges in liquid chromatography coupled to mass spectrometry data is converting many metabolic feature entries to biological function information, such as metabolite annotation and pathway enrichment, which are based on the compound and pathway databases. Multiple online databases have been developed. However, no tool has been developed for operating all these databases for biological analysis. Therefore, we developed massDatabase, an R package that operates the online public databases and combines with other tools for streamlined compound annotation and pathway enrichment. massDatabase is a flexible, simple and powerful tool that can be installed on all platforms, allowing the users to leverage all the online public databases for biological function mining. A detailed tutorial and a case study are provided in the Supplementary Material. AVAILABILITY AND IMPLEMENTATION: https://massdatabase.tidymass.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaotao Shen, Chuchu Wang, Michael Snyder 0001 |
Bioinform. | 3 |
| 2021 | Learning Personal Food Preferences via Food Logs EmbeddingabstractDiet management is key to managing chronic dis-eases such as diabetes. Automated food recommender systems may be able to assist by providing meal recommendations that conform to a user’s nutrition goals and food preferences. Current recommendation systems suffer from a lack of accuracy that is in part due to a lack of knowledge of food preferences. In this work, we propose a method for learning food preferences from food logs, a comprehensive but noisy source of information about users’ dietary habits. We also introduce accompanying metrics to evaluate personal learning food preferences. The method generates and compares word embeddings to identify the parent food category of each food entry and then calculates the most popular. Our proposed approach identifies 82% of a user’s ten most frequently eaten foods. Our method is publicly available on (https://github.com/aametwally/LearningFoodPreferences). Ahmed Metwally 0002, Ariel K. Leong, Aman Desai, Anvith Nagarjuna, Dalia Perelman, Michael Snyder 0001 |
BIBM | 6 |
| 2021 | Hummingbird: efficient performance prediction for executing genomic applications in the cloudabstractMOTIVATION: A major drawback of executing genomic applications on cloud computing facilities is the lack of tools to predict which instance type is the most appropriate, often resulting in an over- or under- matching of resources. Determining the right configuration before actually running the applications will save money and time. Here, we introduce Hummingbird, a tool for predicting performance of computing instances with varying memory and CPU on multiple cloud platforms. RESULTS: Our experiments on three major genomic data pipelines, including GATK HaplotypeCaller, GATK Mutect2 and ENCODE ATAC-seq, showed that Hummingbird was able to address applications in command line specified in JSON format or workflow description language (WDL) format, and accurately predicted the fastest, the cheapest and the most cost-efficient compute instances in an economic manner. AVAILABILITY AND IMPLEMENTATION: Hummingbird is available as an open source tool at: https://github.com/StanfordBioinformatics/Hummingbird. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Amir Bahmani, Ziye Xing, Vandhana Krishnan, Utsab Ray, Frank Mueller 0001, Amir Alavi, Philip S. Tsao, Michael Snyder 0001, Cuiping Pan |
Bioinform. | 8 |
| 2021 | RobNorm: model-based robust normalization method for labeled quantitative mass spectrometry proteomics dataabstractMOTIVATION: Data normalization is an important step in processing proteomics data generated in mass spectrometry experiments, which aims to reduce sample-level variation and facilitate comparisons of samples. Previously published methods for normalization primarily depend on the assumption that the distribution of protein expression is similar across all samples. However, this assumption fails when the protein expression data is generated from heterogenous samples, such as from various tissue types. This led us to develop a novel data-driven method for improved normalization to correct the systematic bias meanwhile maintaining underlying biological heterogeneity. RESULTS: To robustly correct the systematic bias, we used the density-power-weight method to down-weigh outliers and extended the one-dimensional robust fitting method described in the previous work to our structured data. We then constructed a robustness criterion and developed a new normalization algorithm, called RobNorm.In simulation studies and analysis of real data from the genotype-tissue expression project, we compared and evaluated the performance of RobNorm against other normalization methods. We found that the RobNorm approach exhibits the greatest reduction in systematic bias while maintaining across-tissue variation, especially for datasets from highly heterogeneous samples. AVAILABILITYAND IMPLEMENTATION: https://github.com/mwgrassgreen/RobNorm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Meng Wang 0007, Lihua Jiang, Ruiqi Jian, Joanne Y. Chan, Michael Snyder 0001, Hua Tang |
Bioinform. | 6 |
| 2021 | AdaTiSS: a novel data-Adaptive robust method for identifying Tissue Specificity ScoresabstractMOTIVATION: Accurately detecting tissue specificity (TS) in genes helps researchers understand tissue functions at the molecular level. The Genotype-Tissue Expression project is one of the publicly available data resources, providing large-scale gene expressions across multiple tissue types. Multiple tissue comparisons and heterogeneous tissue expression make it challenging to accurately identify tissue specific gene expression. How to distinguish the inlier expression from the outlier expression becomes important to build the population level information and further quantify the TS. There still lacks a robust and data-adaptive TS method taking into account heterogeneities of the data. RESULTS: We found that the key to identify tissue specific gene expression is to properly define a concept of expression population. In a linear regression problem, we developed a novel data-adaptive robust estimation approach (AdaReg) based on density-power-weight under unknown outlier distribution and non-vanishing outlier proportion. The Gaussian-population mixture model was considered in the setting of identifying TS. We took into account heterogeneities of gene expression and applied the robust data-adaptive procedure to estimate the population parameters. With the well-estimated population parameters, we constructed the AdaTiSS algorithm.Our AdaTiSS profiled TS for each gene and each tissue, which standardized the gene expression in terms of TS. We provided a new robust and powerful tool to the literature of defining TS. AVAILABILITY AND IMPLEMENTATION: https://github.com/mwgrassgreen/AdaTiSS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Meng Wang 0007, Lihua Jiang, Michael Snyder 0001 |
Bioinform. | 3 |
| 2021 | Benchmarking workflows to assess performance and suitability of germline variant calling pipelines in clinical diagnostic assaysabstractBACKGROUND: Benchmarking the performance of complex analytical pipelines is an essential part of developing Lab Developed Tests (LDT). Reference samples and benchmark calls published by Genome in a Bottle (GIAB) consortium have enabled the evaluation of analytical methods. The performance of such methods is not uniform across the different genomic regions of interest and variant types. Several benchmarking methods such as hap.py, vcfeval, and vcflib are available to assess the analytical performance characteristics of variant calling algorithms. However, assessing the performance characteristics of an overall LDT assay still requires stringing together several such methods and experienced bioinformaticians to interpret the results. In addition, these methods are dependent on the hardware, operating system and other software libraries, making it impossible to reliably repeat the analytical assessment, when any of the underlying dependencies change in the assay. Here we present a scalable and reproducible, cloud-based benchmarking workflow that is independent of the laboratory and the technician executing the workflow, or the underlying compute hardware used to rapidly and continually assess the performance of LDT assays, across their regions of interest and reportable range, using a broad set of benchmarking samples. RESULTS: The benchmarking workflow was used to evaluate the performance characteristics for secondary analysis pipelines commonly used by Clinical Genomics laboratories in their LDT assays such as the GATK HaplotypeCaller v3.7 and the SpeedSeq workflow based on FreeBayes v0.9.10. Five reference sample truth sets generated by Genome in a Bottle (GIAB) consortium, six samples from the Personal Genome Project (PGP) and several samples with validated clinically relevant variants from the Centers for Disease Control were used in this work. The performance characteristics were evaluated and compared for multiple reportable ranges, such as whole exome and the clinical exome. CONCLUSIONS: We have implemented a benchmarking workflow for clinical diagnostic laboratories that generates metrics such as specificity, precision and sensitivity for germline SNPs and InDels within a reportable range using whole exome or genome sequencing data. Combining these benchmarking results with validation using known variants of clinical significance in publicly available cell lines, we were able to establish the performance of variant calling pipelines in a clinical setting. Vandhana Krishnan, Sowmithri Utiramerur, Zena Ng, Somalee Datta, Michael Snyder 0001, Euan A. Ashley |
BMC Bioinform. | 5 |
| 2021 | Swarm: A federated cloud framework for large-scale variant analysisabstractGenomic data analysis across multiple cloud platforms is an ongoing challenge, especially when large amounts of data are involved. Here, we present Swarm, a framework for federated computation that promotes minimal data motion and facilitates crosstalk between genomic datasets stored on various cloud platforms. We demonstrate its utility via common inquiries of genomic variants across BigQuery in the Google Cloud Platform (GCP), Athena in the Amazon Web Services (AWS), Apache Presto and MySQL. Compared to single-cloud platforms, the Swarm framework significantly reduced computational costs, run-time delays and risks of security breach and privacy violation. Amir Bahmani, Kyle Ferriter, Vandhana Krishnan, Arash Alavi 0003, Amir Alavi, Philip S. Tsao, Michael Snyder 0001, Cuiping Pan |
PLoS Comput. Biol. | 7 |
| 2020 | Classifying non-small cell lung cancer types and transcriptomic subtypes using convolutional neural networksabstractOBJECTIVE: Non-small cell lung cancer is a leading cause of cancer death worldwide, and histopathological evaluation plays the primary role in its diagnosis. However, the morphological patterns associated with the molecular subtypes have not been systematically studied. To bridge this gap, we developed a quantitative histopathology analytic framework to identify the types and gene expression subtypes of non-small cell lung cancer objectively. MATERIALS AND METHODS: We processed whole-slide histopathology images of lung adenocarcinoma (n = 427) and lung squamous cell carcinoma patients (n = 457) in the Cancer Genome Atlas. We built convolutional neural networks to classify histopathology images, evaluated their performance by the areas under the receiver-operating characteristic curves (AUCs), and validated the results in an independent cohort (n = 125). RESULTS: To establish neural networks for quantitative image analyses, we first built convolutional neural network models to identify tumor regions from adjacent dense benign tissues (AUCs > 0.935) and recapitulated expert pathologists' diagnosis (AUCs > 0.877), with the results validated in an independent cohort (AUCs = 0.726-0.864). We further demonstrated that quantitative histopathology morphology features identified the major transcriptomic subtypes of both adenocarcinoma and squamous cell carcinoma (P < .01). DISCUSSION: Our study is the first to classify the transcriptomic subtypes of non-small cell lung cancer using fully automated machine learning methods. Our approach does not rely on prior pathology knowledge and can discover novel clinically relevant histopathology patterns objectively. The developed procedure is generalizable to other tumor types or diseases. Kun-Hsing Yu, Gerald J. Berry, Christopher Ré, Russ B. Altman, Michael Snyder 0001, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Classifying Non-Small Cell Lung Cancer Histopathology Types and Transcriptomic Subtypes using Convolutional Neural Networks
Kun-Hsing Yu, Gerald J. Berry, Christopher Ré, Russ B. Altman, Michael Snyder 0001, Isaac S. Kohane |
AMIA | 6 |
| 2019 | Multiomics modeling of the immunome, transcriptome, microbiome, proteome and metabolome adaptations during human pregnancyabstractMotivation: Multiple biological clocks govern a healthy pregnancy. These biological mechanisms produce immunologic, metabolomic, proteomic, genomic and microbiomic adaptations during the course of pregnancy. Modeling the chronology of these adaptations during full-term pregnancy provides the frameworks for future studies examining deviations implicated in pregnancy-related pathologies including preterm birth and preeclampsia. Results: We performed a multiomics analysis of 51 samples from 17 pregnant women, delivering at term. The datasets included measurements from the immunome, transcriptome, microbiome, proteome and metabolome of samples obtained simultaneously from the same patients. Multivariate predictive modeling using the Elastic Net (EN) algorithm was used to measure the ability of each dataset to predict gestational age. Using stacked generalization, these datasets were combined into a single model. This model not only significantly increased predictive power by combining all datasets, but also revealed novel interactions between different biological modalities. Future work includes expansion of the cohort to preterm-enriched populations and in vivo analysis of immune-modulating interventions based on the mechanisms identified. Availability and implementation: Datasets and scripts for reproduction of results are available through: https://nalab.stanford.edu/multiomics-pregnancy/. Supplementary information: Supplementary data are available at Bioinformatics online. Mohammad Sajjad Ghaemi, Daniel B. DiGiulio, Kévin Contrepois, Benjamin J. Callahan, Thuy T. M. Ngo, Brittany Lee-McMullen, Benoit Lehallier, Anna Robaczewska, David Mcilwain, Yael Rosenberg-Hasson, Ronald J. Wong, Cecele Quaintance, Anthony Culos, Natalie Stanley, Athena Tanada, Amy Tsai, Dyani Gaudilliere, Edward Ganio, Xiaoyuan Han, Kazuo Ando, Leslie McNeil, Martha Tingle, Paul H. Wise, Ivana Maric, Marina Sirota, Tony Wyss-Coray, Virginia D. Winn, Maurice L. Druzin, Ronald Gibbs, Gary L. Darmstadt, David B. Lewis, Vahid Partovi Nia, Bruno Agard, Robert Tibshirani, Garry P. Nolan, Michael Snyder 0001, David A. Relman, Stephen R. Quake, Gary M. Shaw, David K. Stevenson, Martin S. Angst, Brice Gaudilliere, Nima Aghaeepour |
Bioinform. | 36 |
| 2018 | Unraveling the Molecular Basis of Lung Adenocarcinoma Dedifferentiation and Prognosis by Integrating Omics and Histopathology
Kun-Hsing Yu, Gerald J. Berry, Daniel L. Rubin, Christopher Ré, Russ B. Altman, Michael Snyder 0001 |
AMIA | 6 |
| 2018 | Omics AnalySIs System for PRecision Oncology (OASISPRO): a web-based omics analysis tool for clinical phenotype predictionabstractSUMMARY: Precision oncology is an approach that accounts for individual differences to guide cancer management. Omics signatures have been shown to predict clinical traits for cancer patients. However, the vast amount of omics information poses an informatics challenge in systematically identifying patterns associated with health outcomes, and no general purpose data mining tool exists for physicians, medical researchers and citizen scientists without significant training in programming and bioinformatics. To bridge this gap, we built the Omics AnalySIs System for PRecision Oncology (OASISPRO), a web-based system to mine the quantitative omics information from The Cancer Genome Atlas (TCGA). This system effectively visualizes patients' clinical profiles, executes machine-learning algorithms of choice on the omics data and evaluates the prediction performance using held-out test sets. With this tool, we successfully identified genes strongly associated with tumor stage, and accurately predicted patients' survival outcomes in many cancer types, including adrenocortical carcinoma. By identifying the links between omics and clinical phenotypes, this system will facilitate omics studies on precision cancer medicine and contribute to establishing personalized cancer treatment plans. AVAILABILITY AND IMPLEMENTATION: This web-based tool is available at http://tinyurl.com/oasispro; source codes are available at http://tinyurl.com/oasisproSourceCode. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kun-Hsing Yu, Michael R. Fitzpatrick, Luke Pappas, Warren Chan, Jessica Kung, Michael Snyder 0001 |
Bioinform. | 6 |
| 2017 | Predicting Non-Small Cell Lung Cancer Diagnosis and Prognosis by Fully Automated Microscopic Pathology Image Features
Kun-Hsing Yu, Ce Zhang 0001, Gerald J. Berry, Russ B. Altman, Christopher Ré, Daniel L. Rubin, Michael Snyder 0001 |
AMIA | 7 |
| 2017 | GATTACA: Lightweight Metagenomic Binning Using Kmer Counting
Victoria Popic, Volodymyr Kuleshov, Michael Snyder 0001, Serafim Batzoglou |
RECOMB | 3 |
| 2017 | Cloud-based interactive analytics for terabytes of genomic variants dataabstractMOTIVATION: Large scale genomic sequencing is now widely used to decipher questions in diverse realms such as biological function, human diseases, evolution, ecosystems, and agriculture. With the quantity and diversity these data harbor, a robust and scalable data handling and analysis solution is desired. RESULTS: We present interactive analytics using a cloud-based columnar database built on Dremel to perform information compression, comprehensive quality controls, and biological information retrieval in large volumes of genomic data. We demonstrate such Big Data computing paradigms can provide orders of magnitude faster turnaround for common genomic analyses, transforming long-running batch jobs submitted via a Linux shell into questions that can be asked from a web browser in seconds. Using this method, we assessed a study population of 475 deeply sequenced human genomes for genomic call rate, genotype and allele frequency distribution, variant density across the genome, and pharmacogenomic information. AVAILABILITY AND IMPLEMENTATION: Our analysis framework is implemented in Google Cloud Platform and BigQuery. Codes are available at https://github.com/StanfordBioinformatics/mvp_aaa_codelabs. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cuiping Pan, Gregory McInnes, Nicole Deflaux, Michael Snyder 0001, Jonathan Bingham, Somalee Datta, Philip S. Tsao |
Bioinform. | 4 |
| 2016 | Genome assembly from synthetic long read cloudsabstractMOTIVATION: Despite rapid progress in sequencing technology, assembling de novo the genomes of new species as well as reconstructing complex metagenomes remains major technological challenges. New synthetic long read (SLR) technologies promise significant advances towards these goals; however, their applicability is limited by high sequencing requirements and the inability of current assembly paradigms to cope with combinations of short and long reads. RESULTS: Here, we introduce Architect, a new de novo scaffolder aimed at SLR technologies. Unlike previous assembly strategies, Architect does not require a costly subassembly step; instead it assembles genomes directly from the SLR's underlying short reads, which we refer to as read clouds This enables a 4- to 20-fold reduction in sequencing requirements and a 5-fold increase in assembly contiguity on both genomic and metagenomic datasets relative to state-of-the-art assembly strategies aimed directly at fully subassembled long reads. AVAILABILITY AND IMPLEMENTATION: Our source code is freely available at https://github.com/kuleshov/architect CONTACT: [email protected]. Volodymyr Kuleshov, Michael Snyder 0001, Serafim Batzoglou |
Bioinform. | 2 |
| 2016 | Systematic evaluation of the impact of ChIP-seq read designs on genome coverage, peak identification, and allele-specific binding detectionabstractBACKGROUND: Chromatin immunoprecipitation followed by sequencing (ChIP-seq) experiments revolutionized genome-wide profiling of transcription factors and histone modifications. Although maturing sequencing technologies allow these experiments to be carried out with short (36-50 bps), long (75-100 bps), single-end, or paired-end reads, the impact of these read parameters on the downstream data analysis are not well understood. In this paper, we evaluate the effects of different read parameters on genome sequence alignment, coverage of different classes of genomic features, peak identification, and allele-specific binding detection. RESULTS: We generated 101 bps paired-end ChIP-seq data for many transcription factors from human GM12878 and MCF7 cell lines. Systematic evaluations using in silico variations of these data as well as fully simulated data, revealed complex interplay between the sequencing parameters and analysis tools, and indicated clear advantages of paired-end designs in several aspects such as alignment accuracy, peak resolution, and most notably, allele-specific binding detection. CONCLUSIONS: Our work elucidates the effect of design on the downstream analysis and provides insights to investigators in deciding sequencing parameters in ChIP-seq experiments. We present the first systematic evaluation of the impact of ChIP-seq designs on allele-specific binding detection and highlights the power of pair-end designs in such studies. Samuel G. Younkin, Trupti Kawli, Michael Snyder 0001, Sündüz Keles |
BMC Bioinform. | 5 |
| 2015 | Mango: a bias-correcting ChIA-PET analysis pipelineabstractMOTIVATION: Chromatin Interaction Analysis by Paired-End Tag sequencing (ChIA-PET) is an established method for detecting genome-wide looping interactions at high resolution. Current ChIA-PET analysis software packages either fail to correct for non-specific interactions due to genomic proximity or only address a fraction of the steps required for data processing. We present Mango, a complete ChIA-PET data analysis pipeline that provides statistical confidence estimates for interactions and corrects for major sources of bias including differential peak enrichment and genomic proximity. RESULTS: Comparison to the existing software packages, ChIA-PET Tool and ChiaSig revealed that Mango interactions exhibit much better agreement with high-resolution Hi-C data. Importantly, Mango executes all steps required for processing ChIA-PET datasets, whereas ChiaSig only completes 20% of the required steps. Application of Mango to multiple available ChIA-PET datasets permitted the independent rediscovery of known trends in chromatin loops including enrichment of CTCF, RAD21, SMC3 and ZNF143 at the anchor regions of interactions and strong bias for convergent CTCF motifs. AVAILABILITY AND IMPLEMENTATION: Mango is open source and distributed through github at https://github.com/dphansti/mango. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Douglas H. Phanstiel, Alan P. Boyle, Nastaran Heidari, Michael Snyder 0001 |
Bioinform. | 4 |
| 2014 | Sushi.R: flexible, quantitative and integrative genomic visualizations for publication-quality multi-panel figuresabstractMOTIVATION: Interpretation and communication of genomic data require flexible and quantitative tools to analyze and visualize diverse data types, and yet, a comprehensive tool to display all common genomic data types in publication quality figures does not exist to date. To address this shortcoming, we present Sushi.R, an R/Bioconductor package that allows flexible integration of genomic visualizations into highly customizable, publication-ready, multi-panel figures from common genomic data formats including Browser Extensible Data (BED), bedGraph and Browser Extensible Data Paired-End (BEDPE). Sushi.R is open source and made publicly available through GitHub (https://github.com/dphansti/Sushi) and Bioconductor (http://bioconductor.org/packages/release/bioc/html/Sushi.html). Douglas H. Phanstiel, Alan P. Boyle, Carlos L. Araya, Michael Snyder 0001 |
Bioinform. | 4 |
| 2012 | VAT: a computational framework to functionally annotate variants in personal genomes within a cloud-computing environmentabstractUNLABELLED: The functional annotation of variants obtained through sequencing projects is generally assumed to be a simple intersection of genomic coordinates with genomic features. However, complexities arise for several reasons, including the differential effects of a variant on alternatively spliced transcripts, as well as the difficulty in assessing the impact of small insertions/deletions and large structural variants. Taking these factors into consideration, we developed the Variant Annotation Tool (VAT) to functionally annotate variants from multiple personal genomes at the transcript level as well as obtain summary statistics across genes and individuals. VAT also allows visualization of the effects of different variants, integrates allele frequencies and genotype data from the underlying individuals and facilitates comparative analysis between different groups of individuals. VAT can either be run through a command-line interface or as a web application. Finally, in order to enable on-demand access and to minimize unnecessary transfers of large data files, VAT can be run as a virtual machine in a cloud-computing environment. AVAILABILITY AND IMPLEMENTATION: VAT is implemented in C and PHP. The VAT web service, Amazon Machine Image, source code and detailed documentation are available at vat.gersteinlab.org. Lukas Habegger, Suganthi Balasubramanian, David Z. Chen, Ekta Khurana, Andrea Sboner, Arif Ozgun Harmanci, Joel S. Rozowsky, Declan Clarke, Michael Snyder 0001, Mark Gerstein |
Bioinform. | 9 |
| 2012 | Copy Number Variation detection from 1000 Genomes project exon capture sequencing dataabstractBACKGROUND: DNA capture technologies combined with high-throughput sequencing now enable cost-effective, deep-coverage, targeted sequencing of complete exomes. This is well suited for SNP discovery and genotyping. However there has been little attention devoted to Copy Number Variation (CNV) detection from exome capture datasets despite the potentially high impact of CNVs in exonic regions on protein function. RESULTS: As members of the 1000 Genomes Project analysis effort, we investigated 697 samples in which 931 genes were targeted and sampled with 454 or Illumina paired-end sequencing. We developed a rigorous Bayesian method to detect CNVs in the genes, based on read depth within target regions. Despite substantial variability in read coverage across samples and targeted exons, we were able to identify 107 heterozygous deletions in the dataset. The experimentally determined false discovery rate (FDR) of the cleanest dataset from the Wellcome Trust Sanger Institute is 12.5%. We were able to substantially improve the FDR in a subset of gene deletion candidates that were adjacent to another gene deletion call (17 calls). The estimated sensitivity of our call-set was 45%. CONCLUSIONS: This study demonstrates that exonic sequencing datasets, collected both in population based and medical sequencing projects, will be a useful substrate for detecting genic CNV events, particularly deletions. Based on the number of events we found and the sensitivity of the methods in the present dataset, we estimate on average 16 genic heterozygous deletions per individual genome. Our power analysis informs ongoing and future projects about sequencing depth and uniformity of read coverage required for efficient detection. Jiantao Wu, Krzysztof R. Grzeda, Chip Stewart, Fabian Grubert, Alexander E. Urban, Michael Snyder 0001, Gabor T. Marth |
BMC Bioinform. | 6 |
| 2011 | RSEQtools: a modular framework to analyze RNA-Seq data using compact, anonymized data summariesabstractSUMMARY: The advent of next-generation sequencing for functional genomics has given rise to quantities of sequence information that are often so large that they are difficult to handle. Moreover, sequence reads from a specific individual can contain sufficient information to potentially identify and genetically characterize that person, raising privacy concerns. In order to address these issues, we have developed the Mapped Read Format (MRF), a compact data summary format for both short and long read alignments that enables the anonymization of confidential sequence information, while allowing one to still carry out many functional genomics studies. We have developed a suite of tools (RSEQtools) that use this format for the analysis of RNA-Seq experiments. These tools consist of a set of modules that perform common tasks such as calculating gene expression values, generating signal tracks of mapped reads and segmenting that signal into actively transcribed regions. Moreover, the tools can readily be used to build customizable RNA-Seq workflows. In addition to the anonymization afforded by MRF, this format also facilitates the decoupling of the alignment of reads from downstream analyses. AVAILABILITY AND IMPLEMENTATION: RSEQtools is implemented in C and the source code is available at http://rseqtools.gersteinlab.org/. Lukas Habegger, Andrea Sboner, Tara A. Gianoulis, Joel S. Rozowsky, Ashish Agarwal, Michael Snyder 0001, Mark Gerstein |
Bioinform. | 6 |
| 2011 | Construction and Analysis of an Integrated Regulatory Network Derived from High-Throughput Sequencing DataabstractWe present a network framework for analyzing multi-level regulation in higher eukaryotes based on systematic integration of various high-throughput datasets. The network, namely the integrated regulatory network, consists of three major types of regulation: TF→gene, TF→miRNA and miRNA→gene. We identified the target genes and target miRNAs for a set of TFs based on the ChIP-Seq binding profiles, the predicted targets of miRNAs using annotated 3'UTR sequences and conservation information. Making use of the system-wide RNA-Seq profiles, we classified transcription factors into positive and negative regulators and assigned a sign for each regulatory interaction. Other types of edges such as protein-protein interactions and potential intra-regulations between miRNAs based on the embedding of miRNAs in their host genes were further incorporated. We examined the topological structures of the network, including its hierarchical organization and motif enrichment. We found that transcription factors downstream of the hierarchy distinguish themselves by expressing more uniformly at various tissues, have more interacting partners, and are more likely to be essential. We found an over-representation of notable network motifs, including a FFL in which a miRNA cost-effectively shuts down a transcription factor and its target. We used data of C. elegans from the modENCODE project as a primary model to illustrate our framework, but further verified the results using other two data sets. As more and more genome-wide ChIP-Seq and RNA-Seq data becomes available in the near future, our methods of data integration have various potential applications. Koon-Kiu Yan, Woochang Hwang, Nitin Bhardwaj, Joel S. Rozowsky, Zhi John Lu, Pedro Alves, Masaomi Kato, Michael Snyder 0001, Mark Gerstein |
PLoS Comput. Biol. | 11 |
| 2011 | Measuring the Evolutionary Rewiring of Biological NetworksabstractWe have accumulated a large amount of biological network data and expect even more to come. Soon, we anticipate being able to compare many different biological networks as we commonly do for molecular sequences. It has long been believed that many of these networks change, or "rewire", at different rates. It is therefore important to develop a framework to quantify the differences between networks in a unified fashion. We developed such a formalism based on analogy to simple models of sequence evolution, and used it to conduct a systematic study of network rewiring on all the currently available biological networks. We found that, similar to sequences, biological networks show a decreased rate of change at large time divergences, because of saturation in potential substitutions. However, different types of biological networks consistently rewire at different rates. Using comparative genomics and proteomics data, we found a consistent ordering of the rewiring rates: transcription regulatory, phosphorylation regulatory, genetic interaction, miRNA regulatory, protein interaction, and metabolic pathway network, from fast to slow. This ordering was found in all comparisons we did of matched networks between organisms. To gain further intuition on network rewiring, we compared our observed rewirings with those obtained from simulation. We also investigated how readily our formalism could be mapped to other network contexts; in particular, we showed how it could be applied to analyze changes in a range of "commonplace" networks such as family trees, co-authorships and linux-kernel function dependencies. Chong Shou, Nitin Bhardwaj, Hugo Y. K. Lam, Koon-Kiu Yan, Philip M. Kim, Michael Snyder 0001, Mark Gerstein |
PLoS Comput. Biol. | 6 |
| 2010 | MOTIPS: Automated Motif Analysis for Predicting Targets of Modular Protein DomainsabstractBACKGROUND: Many protein interactions, especially those involved in signaling, involve short linear motifs consisting of 5-10 amino acid residues that interact with modular protein domains such as the SH3 binding domains and the kinase catalytic domains. One straightforward way of identifying these interactions is by scanning for matches to the motif against all the sequences in a target proteome. However, predicting domain targets by motif sequence alone without considering other genomic and structural information has been shown to be lacking in accuracy. RESULTS: We developed an efficient search algorithm to scan the target proteome for potential domain targets and to increase the accuracy of each hit by integrating a variety of pre-computed features, such as conservation, surface propensity, and disorder. The integration is performed using naïve Bayes and a training set of validated experiments. CONCLUSIONS: By integrating a variety of biologically relevant features to predict domain targets, we demonstrated a notably improved prediction of modular protein domain targets. Combined with emerging high-resolution data of domain specificities, we believe that our approach can assist in the reconstruction of many signaling pathways. Hugo Y. K. Lam, Philip M. Kim, Janine Mok, Raffi Tonikian, Sachdev S. Sidhu, Benjamin E. Turk, Michael Snyder 0001, Mark Gerstein |
BMC Bioinform. | 7 |
| 2009 | Integrating Sequencing Technologies in Personal Genomics: Optimal Low Cost Reconstruction of Structural VariantsabstractThe goal of human genome re-sequencing is obtaining an accurate assembly of an individual's genome. Recently, there has been great excitement in the development of many technologies for this (e.g. medium and short read sequencing from companies such as 454 and SOLiD, and high-density oligo-arrays from Affymetrix and NimbelGen), with even more expected to appear. The costs and sensitivities of these technologies differ considerably from each other. As an important goal of personal genomics is to reduce the cost of re-sequencing to an affordable point, it is worthwhile to consider optimally integrating technologies. Here, we build a simulation toolbox that will help us optimally combine different technologies for genome re-sequencing, especially in reconstructing large structural variants (SVs). SV reconstruction is considered the most challenging step in human genome re-sequencing. (It is sometimes even harder than de novo assembly of small genomes because of the duplications and repetitive sequences in the human genome.) To this end, we formulate canonical problems that are representative of issues in reconstruction and are of small enough scale to be computationally tractable and simulatable. Using semi-realistic simulations, we show how we can combine different technologies to optimally solve the assembly at low cost. With mapability maps, our simulations efficiently handle the inhomogeneous repeat-containing structure of the human genome and the computational complexity of practical assembly algorithms. They quantitatively show how combining different read lengths is more cost-effective than using one length, how an optimal mixed sequencing strategy for reconstructing large novel SVs usually also gives accurate detection of SNPs/indels, how paired-end reads can improve reconstruction efficiency, and how adding in arrays is more efficient than just sequencing for disentangling some complex SVs. Our strategy should facilitate the sequencing of human genomes at maximum accuracy and low cost. Jiang Du 0003, Robert D. Bjornson, Zhengdong D. Zhang, Michael Snyder 0001, Mark Gerstein |
PLoS Comput. Biol. | 5 |
| 2008 | Modeling ChIP Sequencing In Silico with ApplicationsabstractChIP sequencing (ChIP-seq) is a new method for genomewide mapping of protein binding sites on DNA. It has generated much excitement in functional genomics. To score data and determine adequate sequencing depth, both the genomic background and the binding sites must be properly modeled. To develop a computational foundation to tackle these issues, we first performed a study to characterize the observed statistical nature of this new type of high-throughput data. By linking sequence tags into clusters, we show that there are two components to the distribution of tag counts observed in a number of recent experiments: an initial power-law distribution and a subsequent long right tail. Then we develop in silico ChIP-seq, a computational method to simulate the experimental outcome by placing tags onto the genome according to particular assumed distributions for the actual binding sites and for the background genomic sequence. In contrast to current assumptions, our results show that both the background and the binding sites need to have a markedly nonuniform distribution in order to correctly model the observed ChIP-seq data, with, for instance, the background tag counts modeled by a gamma distribution. On the basis of these results, we extend an existing scoring approach by using a more realistic genomic-background model. This enables us to identify transcription-factor binding sites in ChIP-seq data in a statistically rigorous fashion. Zhengdong D. Zhang, Joel S. Rozowsky, Michael Snyder 0001, Joseph T. Chang, Mark Gerstein |
PLoS Comput. Biol. | 3 |
| 2006 | A supervised hidden markov model framework for efficiently segmenting tiling array data in transcriptional and chIP-chip experiments: systematically incorporating validated biological knowledgeabstractMOTIVATION: Large-scale tiling array experiments are becoming increasingly common in genomics. In particular, the ENCODE project requires the consistent segmentation of many different tiling array datasets into 'active regions' (e.g. finding transfrags from transcriptional data and putative binding sites from ChIP-chip experiments). Previously, such segmentation was done in an unsupervised fashion mainly based on characteristics of the signal distribution in the tiling array data itself. Here we propose a supervised framework for doing this. It has the advantage of explicitly incorporating validated biological knowledge into the model and allowing for formal training and testing. METHODOLOGY: In particular, we use a hidden Markov model (HMM) framework, which is capable of explicitly modeling the dependency between neighboring probes and whose extended version (the generalized HMM) also allows explicit description of state duration density. We introduce a formal definition of the tiling-array analysis problem, and explain how we can use this to describe sampling small genomic regions for experimental validation to build up a gold-standard set for training and testing. We then describe various ideal and practical sampling strategies (e.g. maximizing signal entropy within a selected region versus using gene annotation or known promoters as positives for transcription or ChIP-chip data, respectively). RESULTS: For the practical sampling and training strategies, we show how the size and noise in the validated training data affects the performance of an HMM applied to the ENCODE transcriptional and ChIP-chip experiments. In particular, we show that the HMM framework is able to efficiently process tiling array data as well as or better than previous approaches. For the idealized sampling strategies, we show how we can assess their performance in a simulation framework and how a maximum entropy approach, which samples sub-regions with very different signal intensities, gives the maximally performing gold-standard. This latter result has strong implications for the optimum way medium-scale validation experiments should be carried out to verify the results of the genome-scale tiling array experiments. Jiang Du 0003, Joel S. Rozowsky, Jan O. Korbel, Zhengdong D. Zhang, Thomas E. Royce, Martin H. Schultz, Michael Snyder 0001, Mark Gerstein |
Bioinform. | 7 |
| 2002 | YMD: a microarray database for large-scale gene expression analysis
Kei-Hoi Cheung, Kevin P. White, Janet Hager, Mark Gerstein, Valerie Reinke, Kenneth Nelson, Peter Masiar, Ranjana Srivastava, Yuli Li, Hongyu Zhao 0003, David B. Allison, Michael Snyder 0001, Perry L. Miller, Kenneth R. Williams |
AMIA | 14 |
| 2002 | Fast Optimal Genome Tiling with Applications to Microarray Design and Homology Search
Piotr Berman, Paul Bertone, Bhaskar DasGupta, Mark Gerstein, Ming-Yang Kao, Michael Snyder 0001 |
WABI | 6 |
| 2002 | A dynamic approach to mapping coordinates between microplates and microarraysabstractThe retrieval of useful data from spotted microarray slides requires keeping track of which microplate wells and DNA sample corresponds to each spot on each array slide. Existing approaches are closely coupled with the type of arrayer in use and are computer operating-system-specific. To support the microarray researcher community at large who use different arrayers and computer platforms, increased flexibility, generality, and portability of these approaches are required. In this paper, we describe a general algorithm that correlates the well positions of DNA samples in each microplate to the positions of the spots on each array slide. Based on this algorithm, we have implemented a flexible and platform-independent program named MicroArray Convolutor (MAC) that provides a Web solution allowing the user to: (a) import a text file that identifies the DNA samples and their well locations, (b) select a transformation method that converts data in 96-well plate format into 384-well plate format, and (c) specify the output format of the array lists dependant on the configuration of the array platform as well as the downstream analysis software chosen for the array. MAC and its source code can be accessed via the following Web address: http://ymd.med.yale.edu/kei-cgi/kc_mac_dev8.pl. Kei-Hoi Cheung, Janet Hager, Kenneth Nelson, Kevin P. White, Yuli Li, Michael Snyder 0001, Kenneth R. Williams, Perry L. Miller |
J. Biomed. Informatics | 6 |
| 2001 | A metadata framework for interoperating heterogeneous genome data using XML
Kei-Hoi Cheung, Aniruddha M. Deshpande, Nick P. Tosches, S. Nath, Perry L. Miller, Michael Snyder 0001 |
AMIA | 8 |
| 2001 | An XML Application For Genomic Data InteroperationabstractAs the eXtensible Markup Language (XML) becomes a popular or standard language for exchanging data over the Internet/Web, there are a growing number of genome Web sites that make their data available in XML format. Publishing genomic data in XML format alone would not be that useful if there is a lack of development of software applications that could take advantage of the XML technology to process these XML-formatted data. This paper illustrates the usefulness of XML in representing and interoperating genomic data between two different data sources (Snyder's laboratory at Yale and SGD at Stanford). In particular, we compare the locations of transposon insertions in the yeast DNA sequences that have been identified by BLAST searches with the chromosomal locations of the yeast open reading frames (ORFs) stored in SGD. Such a comparison allows us to characterize the transposon insertions by indicating whether they fall into any ORFs (which may potentially encode proteins that possess essential biological functions). To implement this XML-based interoperation, we used NCBIs "blastall" (which gives an XML output option) and SGD's yeast nucleotide sequence dataset to establish a local blast server. Also, we converted the SGD's ORF location data file (which is available in tab-delimited formal) into an XML document based on the BIOML (BIOpolymer Markup Language) standard. Kei-Hoi Cheung, Yang Liu 0023, Michael Snyder 0001, Mark Gerstein, Perry L. Miller |
BIBE | 4 |
| 2000 | Graphically-enabled integration of bioinformatics tools allowing parallel execution
Kei-Hoi Cheung, Perry L. Miller, Andrew H. Sherman, Stephen B. Weston, Eric Stratmann, Martin H. Schultz, Michael Snyder 0001 |
AMIA | 7 |