EDBT 2026 Demo / reviewers in the wild / expert
Mikhail G. Dozmorov
dblp:92/7103
· DBLP profile ↗
28ranked-venue papers
11as first author
5since 2021 · last 2026
0000-0002-0086-8358ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 11 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond blacklists: a critical assessment of exclusion set generation strategies and alternative approachesabstractMOTIVATION: Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Alternatively, "sponge" or decoy sequences have been proposed to reduce alignment artifacts. RESULTS: We examined the widely used Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to sensitivity to input data, aligner choice, and read length. We further explored the use of "sponge" sequences-unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA-as an alternative approach. We additionally investigated the effect of the T2T-CHM13 genome assembly on improving biological signals. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets, and recommend the use of the T2T-CHM13 assembly or, for the hg38 genome assembly, "sponge" sequences as an alignment-guided strategy for reducing artifacts and improving functional genomics analyses. Brydon P. G. Wall, Jonathan D. Ogata, My Nguyen, Amy L. Olex, Konstantinos V. Floros, Anthony C. Faber, Joseph L. McClay, J. Chuck Harrell, Mikhail G. Dozmorov |
Bioinform. | 9 |
| 2023 | matchRanges: generating null hypothesis genomic ranges via covariate-matched samplingabstractMOTIVATION: Deriving biological insights from genomic data commonly requires comparing attributes of selected genomic loci to a null set of loci. The selection of this null set is non-trivial, as it requires careful consideration of potential covariates, a problem that is exacerbated by the non-uniform distribution of genomic features including genes, enhancers, and transcription factor binding sites. Propensity score-based covariate matching methods allow the selection of null sets from a pool of possible items while controlling for multiple covariates; however, existing packages do not operate on genomic data classes and can be slow for large data sets making them difficult to integrate into genomic workflows. RESULTS: To address this, we developed matchRanges, a propensity score-based covariate matching method for the efficient and convenient generation of matched null ranges from a set of background ranges within the Bioconductor framework. AVAILABILITY AND IMPLEMENTATION: Package: https://bioconductor.org/packages/nullranges, Code: https://github.com/nullranges, Documentation: https://nullranges.github.io/nullranges. Eric S. Davis, Wancen Mu, Stuart Lee, Mikhail G. Dozmorov, Michael I. Love, Douglas H. Phanstiel |
Bioinform. | 4 |
| 2023 | bootRanges: flexible generation of null sets of genomic ranges for hypothesis testingabstractMOTIVATION: Enrichment analysis is a widely utilized technique in genomic analysis that aims to determine if there is a statistically significant association between two sets of genomic features. To conduct this type of hypothesis testing, an appropriate null model is typically required. However, the null distribution that is commonly used can be overly simplistic and may result in inaccurate conclusions. RESULTS: bootRanges provides fast functions for generation of block bootstrapped genomic ranges representing the null hypothesis in enrichment analysis. As part of a modular workflow, bootRanges offers greater flexibility for computing various test statistics leveraging other Bioconductor packages. We show that shuffling or permutation schemes may result in overly narrow test statistic null distributions and over-estimation of statistical significance, while creating new range sets with a block bootstrap preserves local genomic correlation structure and generates more reliable null distributions. It can also be used in more complex analyses, such as accessing correlations between cis-regulatory elements (CREs) and genes across cell types or providing optimized thresholds, e.g. log fold change (logFC) from differential analysis. AVAILABILITY AND IMPLEMENTATION: bootRanges is freely available in the R/Bioconductor package nullranges hosted at https://bioconductor.org/packages/nullranges. Wancen Mu, Eric S. Davis, Stuart Lee, Mikhail G. Dozmorov, Douglas H. Phanstiel, Michael I. Love |
Bioinform. | 4 |
| 2023 | excluderanges: exclusion sets for T2T-CHM13, GRCm39, and other genome assembliesabstractSUMMARY: Exclusion regions are sections of reference genomes with abnormal pileups of short sequencing reads. Removing reads overlapping them improves biological signal, and these benefits are most pronounced in differential analysis settings. Several labs created exclusion region sets, available primarily through ENCODE and Github. However, the variety of exclusion sets creates uncertainty which sets to use. Furthermore, gap regions (e.g. centromeres, telomeres, short arms) create additional considerations in generating exclusion sets. We generated exclusion sets for the latest human T2T-CHM13 and mouse GRCm39 genomes and systematically assembled and annotated these and other sets in the excluderanges R/Bioconductor data package, also accessible via the BEDbase.org API. The package provides unified access to 82 GenomicRanges objects covering six organisms, multiple genome assemblies, and types of exclusion regions. For human hg38 genome assembly, we recommend hg38.Kundaje.GRCh38_unified_blacklist as the most well-curated and annotated, and sets generated by the Blacklist tool for other organisms. AVAILABILITY AND IMPLEMENTATION: https://bioconductor.org/packages/excluderanges/. Package website: https://dozmorovlab.github.io/excluderanges/. Jonathan D. Ogata, Wancen Mu, Eric S. Davis, Bingjie Xue, J. Chuck Harrell, Nathan C. Sheffield, Douglas H. Phanstiel, Michael I. Love, Mikhail G. Dozmorov |
Bioinform. | 9 |
| 2022 | preciseTAD: a transfer learning framework for 3D domain boundary prediction at base-pair resolutionabstractMOTIVATION: Chromosome conformation capture technologies (Hi-C) revealed extensive DNA folding into discrete 3D domains, such as Topologically Associating Domains and chromatin loops. The correct binding of CTCF and cohesin at domain boundaries is integral in maintaining the proper structure and function of these 3D domains. 3D domains have been mapped at the resolutions of 1 kilobase and above. However, it has not been possible to define their boundaries at the resolution of boundary-forming proteins. RESULTS: To predict domain boundaries at base-pair resolution, we developed preciseTAD, an optimized transfer learning framework trained on high-resolution genome annotation data. In contrast to current TAD/loop callers, preciseTAD-predicted boundaries are strongly supported by experimental evidence. Importantly, this approach can accurately delineate boundaries in cells without Hi-C data. preciseTAD provides a powerful framework to improve our understanding of how genomic regulators are shaping the 3D structure of the genome at base-pair resolution. AVAILABILITY AND IMPLEMENTATION: preciseTAD is an R/Bioconductor package available at https://bioconductor.org/packages/preciseTAD/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Spiro C. Stilianoudakis, Maggie A. Marshall, Mikhail G. Dozmorov |
Bioinform. | 3 |
| 2020 | SpectralTAD: an R package for defining a hierarchy of topologically associated domains using spectral clusteringabstractBACKGROUND: The three-dimensional (3D) structure of the genome plays a crucial role in gene expression regulation. Chromatin conformation capture technologies (Hi-C) have revealed that the genome is organized in a hierarchy of topologically associated domains (TADs), sub-TADs, and chromatin loops. Identifying such hierarchical structures is a critical step in understanding genome regulation. Existing tools for TAD calling are frequently sensitive to biases in Hi-C data, depend on tunable parameters, and are computationally inefficient. METHODS: To address these challenges, we developed a novel sliding window-based spectral clustering framework that uses gaps between consecutive eigenvectors for TAD boundary identification. RESULTS: Our method, implemented in an R package, SpectralTAD, detects hierarchical, biologically relevant TADs, has automatic parameter selection, is robust to sequencing depth, resolution, and sparsity of Hi-C data. SpectralTAD outperforms four state-of-the-art TAD callers in simulated and experimental settings. We demonstrate that TAD boundaries shared among multiple levels of the TAD hierarchy were more enriched in classical boundary marks and more conserved across cell lines and tissues. In contrast, boundaries of TADs that cannot be split into sub-TADs showed less enrichment and conservation, suggesting their more dynamic role in genome regulation. CONCLUSION: SpectralTAD is available on Bioconductor, http://bioconductor.org/packages/SpectralTAD/ . Kellen Garrison Cresswell, John C. Stansfield, Mikhail G. Dozmorov |
BMC Bioinform. | 3 |
| 2020 | Correction to: SpectralTAD: an R package for defining a hierarchy of topologically associated domains using spectral clusteringabstractAn amendment to this paper has been published and can be accessed via the original article. Kellen Garrison Cresswell, John C. Stansfield, Mikhail G. Dozmorov |
BMC Bioinform. | 3 |
| 2020 | A method for estimating coherence of molecular mechanisms in major human disease and traitsabstractBACKGROUND: Phenotypes such as height and intelligence, are thought to be a product of the collective effects of multiple phenotype-associated genes and interactions among their protein products. High/low degree of interactions is suggestive of coherent/random molecular mechanisms, respectively. Comparing the degree of interactions may help to better understand the coherence of phenotype-specific molecular mechanisms and the potential for therapeutic intervention. However, direct comparison of the degree of interactions is difficult due to different sizes and configurations of phenotype-associated gene networks. METHODS: We introduce a metric for measuring coherence of molecular-interaction networks as a slope of internal versus external distributions of the degree of interactions. The internal degree distribution is defined by interaction counts within a phenotype-specific gene network, while the external degree distribution counts interactions with other genes in the whole protein-protein interaction (PPI) network. We present a novel method for normalizing the coherence estimates, making them directly comparable. RESULTS: Using STRING and BioGrid PPI databases, we compared the coherence of 116 phenotype-associated gene sets from GWAScatalog against size-matched KEGG pathways (the reference for high coherence) and random networks (the lower limit of coherence). We observed a range of coherence estimates for each category of phenotypes. Metabolic traits and diseases were the most coherent, while psychiatric disorders and intelligence-related traits were the least coherent. We demonstrate that coherence and modularity measures capture distinct network properties. CONCLUSIONS: We present a general-purpose method for estimating and comparing the coherence of molecular-interaction gene networks that accounts for the network size and shape differences. Our results highlight gaps in our current knowledge of genetics and molecular mechanisms of complex phenotypes and suggest priorities for future GWASs. Mikhail G. Dozmorov, Kellen Garrison Cresswell, Silviu-Alin Bacanu, Carl F. Craver, Mark Reimers, Kenneth S. Kendler |
BMC Bioinform. | 1 |
| 2019 | Disease classification: from phenotypic similarity to integrative genomics and beyondabstractA fundamental challenge of modern biomedical research is understanding how diseases that are similar on the phenotypic level are similar on the molecular level. Integration of various genomic data sets with the traditionally used phenotypic disease similarity revealed novel genetic and molecular mechanisms and blurred the distinction between monogenic (Mendelian) and complex diseases. Network-based medicine has emerged as a complementary approach for identifying disease-causing genes, genetic mediators, disruptions in the underlying cellular functions and for drug repositioning. The recent development of machine and deep learning methods allow for leveraging real-life information about diseases to refine genetic and phenotypic disease relationships. This review describes the historical development and recent methodological advancements for studying disease classification (nosology). Mikhail G. Dozmorov |
Briefings Bioinform. | 1 |
| 2019 | multiHiCcompare: joint normalization and comparative analysis of complex Hi-C experimentsabstractMOTIVATION: With the development of chromatin conformation capture technology and its high-throughput derivative Hi-C sequencing, studies of the three-dimensional interactome of the genome that involve multiple Hi-C datasets are becoming available. To account for the technology-driven biases unique to each dataset, there is a distinct need for methods to jointly normalize multiple Hi-C datasets. Previous attempts at removing biases from Hi-C data have made use of techniques which normalize individual Hi-C datasets, or, at best, jointly normalize two datasets. RESULTS: Here, we present multiHiCcompare, a cyclic loess regression-based joint normalization technique for removing biases across multiple Hi-C datasets. In contrast to other normalization techniques, it properly handles the Hi-C-specific decay of chromatin interaction frequencies with the increasing distance between interacting regions. multiHiCcompare uses the general linear model framework for comparative analysis of multiple Hi-C datasets, adapted for the Hi-C-specific decay of chromatin interaction frequencies. multiHiCcompare outperforms other methods when detecting a priori known chromatin interaction differences from jointly normalized datasets. Applied to the analysis of auxin-treated versus untreated experiments, and CTCF depletion experiments, multiHiCcompare was able to recover the expected epigenetic and gene expression signatures of loss of chromatin interactions and reveal novel insights. AVAILABILITY AND IMPLEMENTATION: multiHiCcompare is freely available on GitHub and as a Bioconductor R package https://bioconductor.org/packages/multiHiCcompare. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John C. Stansfield, Kellen Garrison Cresswell, Mikhail G. Dozmorov |
Bioinform. | 3 |
| 2018 | HiCcompare: an R-package for joint normalization and comparison of HI-C datasetsabstractBACKGROUND: Changes in spatial chromatin interactions are now emerging as a unifying mechanism orchestrating the regulation of gene expression. Hi-C sequencing technology allows insight into chromatin interactions on a genome-wide scale. However, Hi-C data contains many DNA sequence- and technology-driven biases. These biases prevent effective comparison of chromatin interactions aimed at identifying genomic regions differentially interacting between, e.g., disease-normal states or different cell types. Several methods have been developed for normalizing individual Hi-C datasets. However, they fail to account for biases between two or more Hi-C datasets, hindering comparative analysis of chromatin interactions. RESULTS: We developed a simple and effective method, HiCcompare, for the joint normalization and differential analysis of multiple Hi-C datasets. The method introduces a distance-centric analysis and visualization of the differences between two Hi-C datasets on a single plot that allows for a data-driven normalization of biases using locally weighted linear regression (loess). HiCcompare outperforms methods for normalizing individual Hi-C datasets and methods for differential analysis (diffHiC, FIND) in detecting a priori known chromatin interaction differences while preserving the detection of genomic structures, such as A/B compartments. CONCLUSIONS: HiCcompare is able to remove between-dataset bias present in Hi-C matrices. It also provides a user-friendly tool to allow the scientific community to perform direct comparisons between the growing number of pre-processed Hi-C datasets available at online repositories. HiCcompare is freely available as a Bioconductor R package https://bioconductor.org/packages/HiCcompare/ . John C. Stansfield, Kellen Garrison Cresswell, Vladimir I. Vladimirov, Mikhail G. Dozmorov |
BMC Bioinform. | 4 |
| 2017 | Epigenomic annotation-based interpretation of genomic data: from enrichment analysis to machine learningabstractMOTIVATION: One of the goals of functional genomics is to understand the regulatory implications of experimentally obtained genomic regions of interest (ROIs). Most sequencing technologies now generate ROIs distributed across the whole genome. The interpretation of these genome-wide ROIs represents a challenge as the majority of them lie outside of functionally well-defined protein coding regions. Recent efforts by the members of the International Human Epigenome Consortium have generated volumes of functional/regulatory data (reference epigenomic datasets), effectively annotating the genome with epigenomic properties. Consequently, a wide variety of computational tools has been developed utilizing these epigenomic datasets for the interpretation of genomic data. RESULTS: The purpose of this review is to provide a structured overview of practical solutions for the interpretation of ROIs with the help of epigenomic data. Starting with epigenomic enrichment analysis, we discuss leading tools and machine learning methods utilizing epigenomic and 3D genome structure data. The hierarchy of tools and methods reviewed here presents a practical guide for the interpretation of genome-wide ROIs within an epigenomic context. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mikhail G. Dozmorov |
Bioinform. | 1 |
| 2017 | Proceedings of the 2017 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe XIVth Annual MidSouth Computational Biology and Bioinformatics Society was held in Little Rock, AR From March 23-25st 2017 and was co-hosted by University of Arkansas at Little Rock, University of Arkansas for Medical Sciences, Little Rock, AR and National Center for Toxicological Research, Jefferson, AR. The fourteenth annual conference entitled “ Make them Safer Make them Better: Bioinformatics and the Development of Therapeuticals ”. There were 220 conference registrants and 129 abstracts submitted, including 62 oral and 67 poster presentations. Jonathan D. Wren, Mikhail G. Dozmorov, Inimary T. Toby, Bindu Nanduri, Ramin Homayouni, Prashanti Manda, Shraddha Thakkar |
BMC Bioinform. | 2 |
| 2016 | GenomeRunner web server: regulatory similarity and differences define the functional impact of SNP setsabstractMOTIVATION: The growing amount of regulatory data from the ENCODE, Roadmap Epigenomics and other consortia provides a wealth of opportunities to investigate the functional impact of single nucleotide polymorphisms (SNPs). Yet, given the large number of regulatory datasets, researchers are posed with a challenge of how to efficiently utilize them to interpret the functional impact of SNP sets. RESULTS: We developed the GenomeRunner web server to automate systematic statistical analysis of SNP sets within a regulatory context. Besides defining the functional impact of SNP sets, GenomeRunner implements novel regulatory similarity/differential analyses, and cell type-specific regulatory enrichment analysis. Validated against literature- and disease ontology-based approaches, analysis of 39 disease/trait-associated SNP sets demonstrated that the functional impact of SNP sets corresponds to known disease relationships. We identified a group of autoimmune diseases with SNPs distinctly enriched in the enhancers of T helper cell subpopulations, and demonstrated relevant cell type-specificity of the functional impact of other SNP sets. In summary, we show how systematic analysis of genomic data within a regulatory context can help interpreting the functional impact of SNP sets. AVAILABILITY AND IMPLEMENTATION: GenomeRunner web server is freely available at http://www.integrativegenomics.org/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mikhail G. Dozmorov, Lukas R. Cara, Cory B. Giles, Jonathan D. Wren |
Bioinform. | 1 |
| 2016 | Improving sensitivity of linear regression-based cell type-specific differential expression deconvolution with per-gene vs. global significance thresholdabstractBACKGROUND: The goal of many human disease-oriented studies is to detect molecular mechanisms different between healthy controls and patients. Yet, commonly used gene expression measurements from blood samples suffer from variability of cell composition. This variability hinders the detection of differentially expressed genes and is often ignored. Combined with cell counts, heterogeneous gene expression may provide deeper insights into the gene expression differences on the cell type-specific level. Published computational methods use linear regression to estimate cell type-specific differential expression, and a global cutoff to judge significance, such as False Discovery Rate (FDR). Yet, they do not consider many artifacts hidden in high-dimensional gene expression data that may negatively affect linear regression. In this paper we quantify the parameter space affecting the performance of linear regression (sensitivity of cell type-specific differential expression detection) on a per-gene basis. RESULTS: We evaluated the effect of sample sizes, cell type-specific proportion variability, and mean squared error on sensitivity of cell type-specific differential expression detection using linear regression. Each parameter affected variability of cell type-specific expression estimates and, subsequently, the sensitivity of differential expression detection. We provide the R package, LRCDE, which performs linear regression-based cell type-specific differential expression (deconvolution) detection on a gene-by-gene basis. Accounting for variability around cell type-specific gene expression estimates, it computes per-gene t-statistics of differential detection, p-values, t-statistic-based sensitivity, group-specific mean squared error, and several gene-specific diagnostic metrics. CONCLUSIONS: The sensitivity of linear regression-based cell type-specific differential expression detection differed for each gene as a function of mean squared error, per group sample sizes, and variability of the proportions of target cell (cell type being analyzed). We demonstrate that LRCDE, which uses Welch's t-test to compare per-gene cell type-specific gene expression estimates, is more sensitive in detecting cell type-specific differential expression at α < 0.05 missed by the global false discovery rate threshold FDR < 0.3. Edmund R. Glass, Mikhail G. Dozmorov |
BMC Bioinform. | 2 |
| 2016 | Proceedings of the 2016 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe MidSouth Computational Biology and Bioinformatics Society (MCBIOS) held its thirteenth annual conference themed “Precision Medicine and Data Sciences” at the University of Memphis, FedEx Institute of Technology in Memphis, Tennessee on March 3–5, 2016. There were 156 conference registrants and 117 abstracts submitted, including 63 oral and 54 poster presentations. Jonathan D. Wren, Inimary T. Toby, Huxiao Hong, Bindu Nanduri, Rakesh Kaundal, Mikhail G. Dozmorov, Shraddha Thakkar |
BMC Bioinform. | 6 |
| 2015 | Detrimental effects of duplicate reads and low complexity regions on RNA- and ChIP-seq dataabstractBACKGROUND: Adapter trimming and removal of duplicate reads are common practices in next-generation sequencing pipelines. Sequencing reads ambiguously mapped to repetitive and low complexity regions can also be problematic for accurate assessment of the biological signal, yet their impact on sequencing data has not received much attention. We investigate how trimming the adapters, removing duplicates, and filtering out reads overlapping low complexity regions influence the significance of biological signal in RNA- and ChIP-seq experiments. METHODS: We assessed the effect of data processing steps on the alignment statistics and the functional enrichment analysis results of RNA- and ChIP-seq data. We compared differentially processed RNA-seq data with matching microarray data on the same patient samples to determine whether changes in pre-processing improved correlation between the two. We have developed a simple tool to remove low complexity regions, RepeatSoaker, available at https://github.com/mdozmorov/RepeatSoaker, and tested its effect on the alignment statistics and the results of the enrichment analyses. RESULTS: Both adapter trimming and duplicate removal moderately improved the strength of biological signals in RNA-seq and ChIP-seq data. Aggressive filtering of reads overlapping with low complexity regions, as defined by RepeatMasker, further improved the strength of biological signals, and the correlation between RNA-seq and microarray gene expression data. CONCLUSIONS: Adapter trimming and duplicates removal, coupled with filtering out reads overlapping low complexity regions, is shown to increase the quality and reliability of detecting biological signals in RNA-seq and ChIP-seq data. Mikhail G. Dozmorov, Indra Adrianto, Cory B. Giles, Edmund R. Glass, Stuart B. Glenn, Courtney Montgomery, Kathy L. Sivils, Lorin E. Olson, Tomoaki Iwayama, Willard M. Freeman, Christopher J. Lessard, Jonathan D. Wren |
BMC Bioinform. | 1 |
| 2014 | Proceedings of the 2014 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe MidSouth Computational Biology and Bioinformatics Society (MCBIOS 2014) held its eleventh annual conference at the Wes Watkins Center at Oklahoma State University, Stillwater on March 7-8, 2014. The theme was " From Genome to Phenome: Connecting the Dots ". Conference Chair this year was Rakesh Kaundal, who is also one of the MCBIOS board members, and conference committee members were Ulrich K. Melcher and Doris Kupfer. The current president is Andy Perkins and Cesar Compadre was elected as President-Elect for 2015-16. There were 154 registrants and a total of 125 abstracts submitted (50 oral and 75 poster presentations). Jonathan D. Wren, Mikhail G. Dozmorov, Dennis Burian, Andy D. Perkins, Peter Hoyt, Rakesh Kaundal |
BMC Bioinform. | 2 |
| 2013 | Systematic classification of non-coding RNAs by epigenomic similarityabstractEven though only 1.5% of the human genome is translated into proteins, recent reports indicate that most of it is transcribed into non-coding RNAs (ncRNAs), which are becoming the subject of increased scientific interest. We hypothesized that examining how different classes of ncRNAs co-localized with annotated epigenomic elements could help understand the functions, regulatory mechanisms, and relationships among ncRNA families. We examined 15 different ncRNA classes for statistically significant genomic co-localizations with cell type-specific chromatin segmentation states, transcription factor binding sites (TFBSs), and histone modification marks using GenomeRunner ( http://www.genomerunner.org ). P-values were obtained using a Chi-square test and corrected for multiple testing using the Benjamini-Hochberg procedure. We clustered and visualized the ncRNA classes by the strength of their statistical enrichments and depletions. We found piwi-interacting RNAs (piRNAs) to be depleted in regions containing activating histone modification marks, such as H3K4 mono-, di- and trimethylation, H3K27 acetylation, as well as certain TFBSs. piRNAs were further depleted in active promoters, weak transcription, and transcription elongation regions, and enriched in repressed and heterochromatic regions. Conversely, transfer RNAs (tRNAs) were depleted in heterochromatin regions and strongly enriched in regions containing activating H3K4 di- and trimethylation marks, H2az histone variant, and a variety of TFBSs. Interestingly, regions containing CTCF insulator protein binding sites were associated with tRNAs. tRNAs were also enriched in the active, weak and poised promoters and, surprisingly, in regions with repetitive/copy number variations. Searching for statistically significant associations between ncRNA classes and epigenomic elements permits detection of potential functional and/or regulatory relationships among ncRNA classes, and suggests cell type-specific biological roles of ncRNAs. Mikhail G. Dozmorov, Cory B. Giles, Kristi A. Koelsch, Jonathan D. Wren |
BMC Bioinform. | 1 |
| 2013 | mirCoX: a database of miRNA-mRNA expression correlations derived from RNA-seq meta-analysisabstractBACKGROUND: Experimentally validated co-expression correlations between miRNAs and genes are a valuable resource to corroborate observations about miRNA/mRNA changes after experimental perturbations, as well as compare miRNA target predictions with empirical observations. For example, when a given miRNA is transcribed, true targets of that miRNA should tend to have lower expression levels relative to when the miRNA is not expressed. METHODS: We processed publicly available human RNA-seq experiments obtained from NCBI's Sequence Read Archive (SRA) to identify miRNA-mRNA co-expression trends and summarized them in terms of their Pearson's Correlation Coefficient (PCC) and significance. RESULTS: We found that sequence-derived parameters from TargetScan and miRanda were predictive of co-expression, and that TargetScan- and miRanda-derived gene-miRNA pairs tend to have anti-correlated expression patterns in RNA-seq data compared to controls. We provide this data for download and as a web application available at http://wrenlab.org/mirCoX/. CONCLUSION: This database of empirically established miRNA-mRNA transcriptional correlations will help to corroborate experimental observations and could be used to help refine and validate miRNA target predictions. Cory B. Giles, Reshmi Girijadevi, Mikhail G. Dozmorov, Jonathan D. Wren |
BMC Bioinform. | 3 |
| 2013 | Proceedings of the 2013 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe tenth annual conference of the MidSouth Computational Biology and Bioinformatics Society (MCBIOS 2013), "The 10th Anniversary in a Decade of Change: Discovery in a Sea of Data", took place at the Stoney Creek Inn & Conference Center in Columbia, Missouri on April 5-6, 2013. This year's Conference Chairs were Gordon Springer and Chi-Ren Shyu from the University of Missouri and Edward Perkins from the US Army Corps of Engineers Engineering Research and Development Center, who is also the current MCBIOS President (2012-3). There were 151 registrants and a total of 111 abstracts (51 oral presentations and 60 poster session abstracts). Jonathan D. Wren, Mikhail G. Dozmorov, Dennis Burian, Rakesh Kaundal, Andy D. Perkins, Edward J. Perkins, Doris M. Kupfer, Gordon K. Springer |
BMC Bioinform. | 2 |
| 2012 | GenomeRunner: automating genome explorationabstractMOTIVATION: One of the challenges in interpreting high-throughput genomic studies such as a genome-wide associations, microarray or ChIP-seq is their open-ended nature-once a set of experimentally identified regions is identified as statistically significant, at least two questions arise: (i) besides P-value, do any of these significant regions stand out in terms of biological implications? (ii) Does the set of significant regions, as a whole, have anything in common genome wide? These issues are difficult to address because of the growing number of annotated genomic features (e.g. single nucleotide polymorphisms, transcription factor binding sites, methylation peaks, etc.), and it is difficult to know a priori which features would be most fruitful to analyze. Our goal is to provide partial automation of this process to begin examining associations between experimental features and annotated genomic regions in a hypothesis-free, data-driven manner. RESULTS: We created GenomeRunner-a tool for automating annotation and enrichment of genomic features of interest (FOI) with annotated genomic features (GFs), in different organisms. Besides simple association of FOIs with known GFs GenomeRunner tests whether the enriched FOIs, as a group, are statistically associated with a large and growing set of genomic features. AVAILABILITY: GenomeRunner setup files and source code are freely available at http://sourceforge.net/projects/genomerunner. CONTACT: [email protected]; [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mikhail G. Dozmorov, Lukas R. Cara, Cory B. Giles, Jonathan D. Wren |
Bioinform. | 1 |
| 2012 | Proceedings of the 2012 MidSouth computational biology and bioinformatics society (MCBIOS) conferenceabstractThe ninth annual conference of the MidSouth Computational Biology and Bioinformatics Society (MCBIOS 2012), "Making Sense of the Omics Data Deluge", took place in Oxford, Mississippi February 17-8 2012. This year's Conference Chairs were Dr. Dawn Wilkins, of the University of Mississippi and Dr. Doris Kupfer, also the current MCBIOS President (2011-2), from the Federal Aviation Administration. There were 170 registrants and a total of 106 abstracts (34 oral presentations and 72 poster session abstracts). Jonathan D. Wren, Mikhail G. Dozmorov, Dennis Burian, Rakesh Kaundal, Susan M. Bridges, Doris M. Kupfer |
BMC Bioinform. | 2 |
| 2011 | Predicting gene ontology from a global meta-analysis of 1-color microarray experimentsabstractBACKGROUND: Global meta-analysis (GMA) of microarray data to identify genes with highly similar co-expression profiles is emerging as an accurate method to predict gene function and phenotype, even in the absence of published data on the gene(s) being analyzed. With a third of human genes still uncharacterized, this approach is a promising way to direct experiments and rapidly understand the biological roles of genes. To predict function for genes of interest, GMA relies on a guilt-by-association approach to identify sets of genes with known functions that are consistently co-expressed with it across different experimental conditions, suggesting coordinated regulation for a specific biological purpose. Our goal here is to define how sample, dataset size and ranking parameters affect prediction performance. RESULTS: 13,000 human 1-color microarrays were downloaded from GEO for GMA analysis. Prediction performance was benchmarked by calculating the distance within the Gene Ontology (GO) tree between predicted function and annotated function for sets of 100 randomly selected genes. We find the number of new predicted functions rises as more datasets are added, but begins to saturate at a sample size of approximately 2,000 experiments. For the gene set used to predict function, we find precision to be higher with smaller set sizes, yet with correspondingly poor recall and, as set size is increased, recall and F-measure also tend to increase but at the cost of precision. CONCLUSIONS: Of the 20,813 genes expressed in 50 or more experiments, at least one predicted GO category was found for 72.5% of them. Of the 5,720 genes without GO annotation, 4,189 had at least one predicted ontology using top 40 co-expressed genes for prediction analysis. For the remaining 1,531 genes without GO predictions or annotations, ~17% (257 genes) had sufficient co-expression data yet no statistically significantly overrepresented ontologies, suggesting their regulation may be more complex. Mikhail G. Dozmorov, Cory B. Giles, Jonathan D. Wren |
BMC Bioinform. | 1 |
| 2011 | High-throughput processing and normalization of one-color microarrays for transcriptional meta-analysesabstractBACKGROUND: Microarray experiments are becoming increasingly common in biomedical research, as is their deposition in publicly accessible repositories, such as Gene Expression Omnibus (GEO). As such, there has been a surge in interest to use this microarray data for meta-analytic approaches, whether to increase sample size for a more powerful analysis of a specific disease (e.g. lung cancer) or to re-examine experiments for reasons different than those examined in the initial, publishing study that generated them. For the average biomedical researcher, there are a number of practical barriers to conducting such meta-analyses such as manually aggregating, filtering and formatting the data. Methods to automatically process large repositories of microarray data into a standardized, directly comparable format will enable easier and more reliable access to microarray data to conduct meta-analyses. METHODS: We present a straightforward, simple but robust against potential outliers method for automatic quality control and pre-processing of tens of thousands of single-channel microarray data files. GEO GDS files are quality checked by comparing parametric distributions and quantile normalized to enable direct comparison of expression level for subsequent meta-analyses. RESULTS: 13,000 human 1-color experiments were processed to create a single gene expression matrix that subsets can be extracted from to conduct meta-analyses. Interestingly, we found that when conducting a global meta-analysis of gene-gene co-expression patterns across all 13,000 experiments to predict gene function, normalization had minimal improvement over using the raw data. CONCLUSIONS: Normalization of microarray data appears to be of minimal importance on analyses based on co-expression patterns when the sample size is on the order of thousands microarray datasets. Smaller subsets, however, are more prone to aberrations and artefacts, and effective means of automating normalization procedures not only empowers meta-analytic approaches, but aids in reproducibility by providing a standard way of approaching the problem.Data availability: matrix containing normalized expression of 20,813 genes across 13,000 experiments is available for download at . Source code for GDS files pre-processing is available from the authors upon request. Mikhail G. Dozmorov, Jonathan D. Wren |
BMC Bioinform. | 1 |
| 2011 | Proceedings of the 2011 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe Eighth Annual Conference of the MidSouth Computational Biology and Bioinformatics Society (MCBIOS’2011) was held in College Station, Texas on April 1-2, 2011. The Conference General Chair was Ulisses Braga-Neto, the MCBIOS President for the 2010-2011 term, from Texas A&M University. There were nearly 200 registrants and 140 abstracts were submitted, divided into 48 oral presentation abstracts and 92 poster session abstracts. Jonathan D. Wren, Doris M. Kupfer, Edward J. Perkins, Susan M. Bridges, Stephen Winters-Hilt, Mikhail G. Dozmorov, Ulisses Braga-Neto |
BMC Bioinform. | 6 |
| 2008 | From microarray to biology: an integrated experimental, statistical and in silico analysis of how the extracellular matrix modulates the phenotype of cancer cellsabstractA statistically robust and biologically-based approach for analysis of microarray data is described that integrates independent biological knowledge and data with a global F-test for finding genes of interest that minimizes the need for replicates when used for hypothesis generation. First, each microarray is normalized to its noise level around zero. The microarray dataset is then globally adjusted by robust linear regression. Second, genes of interest that capture significant responses to experimental conditions are selected by finding those that express significantly higher variance than those expressing only technical variability. Clustering expression data and identifying expression-independent properties of genes of interest including upstream transcriptional regulatory elements (TREs), ontologies and networks or pathways organizes the data into a biologically meaningful system. We demonstrate that when the number of genes of interest is inconveniently large, identifying a subset of "beacon genes" representing the largest changes will identify pathways or networks altered by biological manipulation. The entire dataset is then used to complete the picture outlined by the "beacon genes." This allow construction of a structured model of a system that can generate biologically testable hypotheses. We illustrate this approach by comparing cells cultured on plastic or an extracellular matrix which organizes a dataset of over 2,000 genes of interest from a genome wide scan of transcription. The resulting model was confirmed by comparing the predicted pattern of TREs with experimental determination of active transcription factors. Mikhail G. Dozmorov, Kimberly D. Kyker, Paul J. Hauser, Ricardo Saban, David D. Buethe, Igor Dozmorov, Michael Centola, Daniel J. Culkin, Robert E. Hurst |
BMC Bioinform. | 1 |
| 2007 | Systems biology approach for mapping the response of human urothelial cells to infection by Enterococcus faecalisabstractBACKGROUND: To better understand the response of urinary epithelial (urothelial) cells to Enterococcus faecalis, a uropathogen that exhibits resistance to multiple antibiotics, a genome-wide scan of gene expression was obtained as a time series from urothelial cells growing as a layered 3-dimensional culture similar to normal urothelium. We herein describe a novel means of analysis that is based on deconvolution of gene variability into technical and biological components. RESULTS: Analysis of the expression of 21,521 genes from 30 minutes to 10 hours post infection, showed 9553 genes were expressed 3 standard deviations (SD) above the system zero-point noise in at least 1 time point. The asymmetric distribution of relative variances of the expressed genes was deconvoluted into technical variation (with a 6.5% relative SD) and biological variation components (>3 SD above the mode technical variability). These 1409 hypervariable (HV) genes encapsulated the effect of infection on gene expression. Pathway analysis of the HV genes revealed an orchestrated response to infection in which early events included initiation of immune response, cytoskeletal rearrangement and cell signaling followed at the end by apoptosis and shutting down cell metabolism. The number of poorly annotated genes in the earliest time points suggests heretofore unknown processes likely also are involved. CONCLUSION: Enterococcus infection produced an orchestrated response by the host cells involving several pathways and transcription factors that potentially drive these pathways. The early time points potentially identify novel targets for enhancing the host response. These approaches combine rigorous statistical principles with a biological context and are readily applied by biologists. Mikhail G. Dozmorov, Kimberly D. Kyker, Ricardo Saban, Nathan Shankar, Arto S. Baghdayan, Michael Centola, Robert E. Hurst |
BMC Bioinform. | 1 |