EDBT 2026 Demo / reviewers in the wild / expert
Shyr Yu
dblp:03/2320 · also Yu Shyr
· DBLP profile ↗
33ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0003-2086-9670ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 8 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | stImage: a versatile framework for optimizing spatial transcriptomic analysis through customizable deep histology and location informed integrationabstractSpatial transcriptomics (ST) integrates gene expression data with the spatial organization of cells and their associated histology, offering unprecedented insights into tissue biology. While existing methods incorporate either location-based or histology-informed information, none fully synergize gene expression, histological features, and precise spatial coordinates within a unified framework. Moreover, these methods often exhibit inconsistent performance across diverse datasets and conditions. Here, we introduce stImage, an open-source R package that provides a comprehensive and flexible solution for ST analysis. By generating deep learning-derived histology features and offering 54 integrative strategies, stImage seamlessly combines transcriptional profiles, histology images, and spatial information. We demonstrate stImage's effectiveness across multiple datasets, underscoring its ability to guide users toward the most suitable integration strategy using diagnostic graph. Our results highlight how stImage can optimize ST, consistently improving biological insights and advancing our understanding of tissue architecture. stImage is freely available at https://github.com/YuWang-VUMC/stImage. Haichun Yang, Ruining Deng, Yuankai Huo, Qi Liu 0024, Shyr Yu, Shilin Zhao |
Briefings Bioinform. | 6 |
| 2024 | A distribution-free and analytic method for power and sample size calculation in single-cell differential expressionabstractMOTIVATION: Differential expression analysis in single-cell transcriptomics unveils cell type-specific responses to various treatments or biological conditions. To ensure the robustness and reliability of the analysis, it is essential to have a solid experimental design with ample statistical power and sample size. However, existing methods for power and sample size calculation often assume a specific distribution for single-cell transcriptomics data, potentially deviating from the true data distribution. Moreover, they commonly overlook cell-cell correlations within individual samples, posing challenges in accurately representing biological phenomena. Additionally, due to the complexity of deriving an analytic formula, most methods employ time-consuming simulation-based strategies. RESULTS: We propose an analytic-based method named scPS for calculating power and sample sizes based on generalized estimating equations. scPS stands out by making no assumptions about the data distribution and considering cell-cell correlations within individual samples. scPS is a rapid and powerful approach for designing experiments in single-cell differential expression analysis. AVAILABILITY AND IMPLEMENTATION: scPS is freely available at https://github.com/cyhsuTN/scPS and Zenodo https://zenodo.org/records/13375996. Chih-Yuan Hsu, Qi Liu 0024, Shyr Yu |
Bioinform. | 3 |
| 2024 | scKWARN: Kernel-weighted-average robust normalization for single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA-seq normalization is an essential step to correct unwanted biases caused by sequencing depth, capture efficiency, dropout, and other technical factors. Existing normalization methods primarily reduce biases arising from sequencing depth by modeling count-depth relationship and/or assuming a specific distribution for read counts. However, these methods may lead to over or under-correction due to presence of technical biases beyond sequencing depth and the restrictive assumption on models and distributions. RESULTS: We present scKWARN, a Kernel Weighted Average Robust Normalization designed to correct known or hidden technical confounders without assuming specific data distributions or count-depth relationships. scKWARN generates a pseudo expression profile for each cell by borrowing information from its fuzzy technical neighbors through a kernel smoother. It then compares this profile against the reference derived from cells with the same bimodality patterns to determine the normalization factor. As demonstrated in both simulated and real datasets, scKWARN outperforms existing methods in removing a variety of technical biases while preserving true biological heterogeneity. AVAILABILITY AND IMPLEMENTATION: scKWARN is freely available at https://github.com/cyhsuTN/scKWARN. Chih-Yuan Hsu, Qi Liu 0024, Shyr Yu |
Bioinform. | 4 |
| 2024 | FindAdapt: A python package for fast and accurate adapter detection in small RNA sequencingabstractAdapter trimming is an essential step for analyzing small RNA sequencing data, where reads are generally longer than target RNAs ranging from 18 to 30 bp. Most adapter trimming tools require adapter information as input. However, adapter information is hard to access, specified incorrectly, or not provided with publicly available datasets, hampering their reproducibility and reusability. Manual identification of adapter patterns from raw reads is labor-intensive and error-prone. Moreover, the use of randomized adapters to reduce ligation biases during library preparation makes adapter detection even more challenging. Here, we present FindAdapt, a Python package for fast and accurate detection of adapter patterns without relying on prior information. We demonstrated that FindAdapt was far superior to existing approaches. It identified adapters successfully in 180 simulation datasets with diverse read structures and 3,184 real datasets covering a variety of commercial and customized small RNA library preparation kits. FindAdapt is stand-alone software that can be easily integrated into small RNA sequencing analysis pipelines. Hua-Chang Chen, Jing Wang 0026, Shyr Yu, Qi Liu 0024 |
PLoS Comput. Biol. | 3 |
| 2022 | Comprehensive evaluation of noise reduction methods for single-cell RNA sequencing dataabstractNormalization and batch correction are critical steps in processing single-cell RNA sequencing (scRNA-seq) data, which remove technical effects and systematic biases to unmask biological signals of interest. Although a number of computational methods have been developed, there is no guidance for choosing appropriate procedures in different scenarios. In this study, we assessed the performance of 28 scRNA-seq noise reduction procedures in 55 scenarios using simulated and real datasets. The scenarios accounted for multiple biological and technical factors that greatly affect the denoising performance, including relative magnitude of batch effects, the extent of cell population imbalance, the complexity of cell group structures, the proportion and the similarity of nonoverlapping cell populations, dropout rates and variable library sizes. We used multiple quantitative metrics and visualization of low-dimensional cell embeddings to evaluate the performance on batch mixing while preserving the original cell group and gene structures. Based on our results, we specified technical or biological factors affecting the performance of each method and recommended proper methods in different scenarios. In addition, we highlighted one challenging scenario where most methods failed and resulted in overcorrection. Our studies not only provided a comprehensive guideline for selecting suitable noise reduction procedures but also pointed out unsolved issues in the field, especially the urgent need of developing metrics for assessing batch correction on imperceptible cell-type mixing. Shih-Kai Chu, Shilin Zhao, Shyr Yu, Qi Liu 0024 |
Briefings Bioinform. | 3 |
| 2022 | Dysregulated ligand-receptor interactions from single-cell transcriptomicsabstractMOTIVATION: Intracellular communication is crucial to many biological processes, such as differentiation, development, homeostasis and inflammation. Single-cell transcriptomics provides an unprecedented opportunity for studying cell-cell communications mediated by ligand-receptor interactions. Although computational methods have been developed to infer cell type-specific ligand-receptor interactions from one single-cell transcriptomics profile, there is lack of approaches considering ligand and receptor simultaneously to identifying dysregulated interactions across conditions from multiple single-cell profiles. RESULTS: We developed scLR, a statistical method for examining dysregulated ligand-receptor interactions between two conditions. scLR models the distribution of the product of ligands and receptors expressions and accounts for inter-sample variances and small sample sizes. scLR achieved high sensitivity and specificity in simulation studies. scLR revealed important cytokine signaling between macrophages and proliferating T cells during severe acute COVID-19 infection, and activated TGF-β signaling from alveolar type II cells in the pathogenesis of pulmonary fibrosis. AVAILABILITY AND IMPLEMENTATION: scLR is freely available at https://github.com/cyhsuTN/scLR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qi Liu 0024, Chih-Yuan Hsu, Jia Li 0027, Shyr Yu |
Bioinform. | 4 |
| 2021 | PheWAS-ME: a web-app for interactive exploration of multimorbidity patterns in PheWASabstractSUMMARY: Electronic health records (EHRs) linked with a DNA biobank provide unprecedented opportunities for biomedical research in precision medicine. The Phenome-wide association study (PheWAS) is a widely used technique for the evaluation of relationships between genetic variants and a large collection of clinical phenotypes recorded in EHRs. PheWAS analyses are typically presented as static tables and charts of summary statistics obtained from statistical tests of association between a genetic variant and individual phenotypes. Comorbidities are common and typically lead to complex, multivariate gene-disease association signals that are challenging to interpret. Discovering and interrogating multimorbidity patterns and their influence in PheWAS is difficult and time-consuming. We present PheWAS-ME: an interactive dashboard to visualize individual-level genotype and phenotype data side-by-side with PheWAS analysis results, allowing researchers to explore multimorbidity patterns and their associations with a genetic variant of interest. We expect this application to enrich PheWAS analyses by illuminating clinical multimorbidity patterns present in the data. AVAILABILITY AND IMPLEMENTATION: A demo PheWAS-ME application is publicly available at https://prod.tbilab.org/phewas_me/. Sample datasets are provided for exploration with the option to upload custom PheWAS results and corresponding individual-level data. Online versions of the appendices are available at https://prod.tbilab.org/phewas_me_info/. The source code is available as an R package on GitHub (https://github.com/tbilab/multimorbidity_explorer). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nick Strayer, Jana Shirey-Rice, Shyr Yu, Joshua C. Denny, Jill M. Pulley, Yaomin Xu |
Bioinform. | 3 |
| 2021 | MetaGSCA: A tool for meta-analysis of gene set differential coexpressionabstractAnalyses of gene set differential coexpression may shed light on molecular mechanisms underlying phenotypes and diseases. However, differential coexpression analyses of conceptually similar individual studies are often inconsistent and underpowered to provide definitive results. Researchers can greatly benefit from an open-source application facilitating the aggregation of evidence of differential coexpression across studies and the estimation of more robust common effects. We developed Meta Gene Set Coexpression Analysis (MetaGSCA), an analytical tool to systematically assess differential coexpression of an a priori defined gene set by aggregating evidence across studies to provide a definitive result. In the kernel, a nonparametric approach that accounts for the gene-gene correlation structure is used to test whether the gene set is differentially coexpressed between two comparative conditions, from which a permutation test p-statistic is computed for each individual study. A meta-analysis is then performed to combine individual study results with one of two options: a random-intercept logistic regression model or the inverse variance method. We demonstrated MetaGSCA in case studies investigating two human diseases and identified pathways highly relevant to each disease across studies. We further applied MetaGSCA in a pan-cancer analysis with hundreds of major cellular pathways in 11 cancer types. The results indicated that a majority of the pathways identified were dysregulated in the pan-cancer scenario, many of which have been previously reported in the cancer literature. Our analysis with randomly generated gene sets showed excellent specificity, indicating that the significant pathways/gene sets identified by MetaGSCA are unlikely false positives. MetaGSCA is a user-friendly tool implemented in both forms of a Web-based application and an R package "MetaGSCA". It enables comprehensive meta-analyses of gene set differential coexpression data, with an optional module of post hoc pathway crosstalk network analysis to identify and visualize pathways having similar coexpression profiles. Haocan Song, Jiapeng He, Olufunmilola Oyebamiji, Huining Kang, Jie Ping, Scott Ness, Shyr Yu |
PLoS Comput. Biol. | 9 |
| 2020 | Bayesian Inference of Lymph Node Ratio Estimation and Survival Prognosis for Breast Cancer PatientsabstractOBJECTIVE: We evaluated the prognostic value of lymph node ratio (LNR) for the survival of breast cancer patients using Bayesian inference. METHODS: Data on 5,279 women with infiltrating duct and lobular carcinoma breast cancer, diagnosed from 2006-2010, was obtained from the NCI SEER Cancer Registry. A prognostic modeling framework was proposed using Bayesian inference to estimate the impact of LNR in breast cancer survival. Based on the proposed model, we then developed a web application for estimating LNR and predicting overall survival. RESULTS: The final survival model with LNR outperformed the other models considered (C-statistic 0.71). Compared to directly measured LNR, estimated LNR slightly increased the accuracy of the prognostic model. Model diagnostics and predictive performance confirmed the effectiveness of Bayesian modeling and the prognostic value of the LNR in predicting breast cancer survival. CONCLUSION: The estimated LNR was found to have a significant predictive value for the overall survival of breast cancer patients. SIGNIFICANCE: We used Bayesian inference to estimate LNR which was then used to predict overall survival. The models were developed from a large population-based cancer registry. We also built a user-friendly web application for individual patient survival prognosis. The diagnostic value of the LNR and the effectiveness of the proposed model were evaluated by comparisons with existing prediction models. Jing Teng, Assem Abdygametova, Bian Ma, Shyr Yu |
IEEE J. Biomed. Health Informatics | 6 |
| 2019 | scRNABatchQC: multi-samples quality control for single cell RNA-seq dataabstractSUMMARY: Single cell RNA sequencing is a revolutionary technique to characterize inter-cellular transcriptomics heterogeneity. However, the data are noise-prone because gene expression is often driven by both technical artifacts and genuine biological variations. Proper disentanglement of these two effects is critical to prevent spurious results. While several tools exist to detect and remove low-quality cells in one single cell RNA-seq dataset, there is lack of approach to examining consistency between sample sets and detecting systematic biases, batch effects and outliers. We present scRNABatchQC, an R package to compare multiple sample sets simultaneously over numerous technical and biological features, which gives valuable hints to distinguish technical artifact from biological variations. scRNABatchQC helps identify and systematically characterize sources of variability in single cell transcriptome data. The examination of consistency across datasets allows visual detection of biases and outliers. AVAILABILITY AND IMPLEMENTATION: scRNABatchQC is freely available at https://github.com/liuqivandy/scRNABatchQC as an R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qi Liu 0024, Quanhu Sheng, Jie Ping, Marisol Adelina Ramirez, Ken S. Lau, Robert J. Coffey, Shyr Yu |
Bioinform. | 7 |
| 2018 | Power and sample size calculations for high-throughput sequencing-based experimentsabstractPower/sample size (power) analysis estimates the likelihood of successfully finding the statistical significance in a data set. There has been a growing recognition of the importance of power analysis in the proper design of experiments. Power analysis is complex, yet necessary for the success of large studies. It is important to design a study that produces statistically accurate and reliable results. Power computation methods have been well established for both microarray-based gene expression studies and genotyping microarray-based genome-wide association studies. High-throughput sequencing (HTS) has greatly enhanced our ability to conduct biomedical studies at the highest possible resolution (per nucleotide). However, the complexity of power computations is much greater for sequencing data than for the simpler genotyping array data. Research on methods of power computations for HTS-based studies has been recently conducted but is not yet well known or widely used. In this article, we describe the power computation methods that are currently available for a range of HTS-based studies, including DNA sequencing, RNA-sequencing, microbiome sequencing and chromatin immunoprecipitation sequencing. Most importantly, we review the methods of power analysis for several types of sequencing data and guide the reader to the relevant methods for each data type. Chung-I Li, David C. Samuels, Ying-Yong Zhao, Shyr Yu |
Briefings Bioinform. | 4 |
| 2018 | Strategies for processing and quality control of Illumina genotyping arraysabstractIllumina genotyping arrays have powered thousands of large-scale genome-wide association studies over the past decade. Yet, because of the tremendous volume and complicated genetic assumptions of Illumina genotyping data, processing and quality control (QC) of these data remain a challenge. Thorough QC ensures the accurate identification of single-nucleotide polymorphisms and is required for the correct interpretation of genetic association results. By processing genotyping data on > 100 000 subjects from >10 major Illumina genotyping arrays, we have accumulated extensive experience in handling some of the most peculiar scenarios related to the processing and QC of Illumina genotyping data. Here, we describe strategies for processing Illumina genotyping data from the raw data to an analysis ready format, and we elaborate on the necessary QC procedures required at each processing step. High-quality Illumina genotyping data sets can be obtained by following our detailed QC strategies. Shilin Zhao, Jing Wang 0026, David C. Samuels, Quanghu Sheng, Shyr Yu |
Briefings Bioinform. | 5 |
| 2018 | RnaSeqSampleSize: real data based sample size estimation for RNA sequencingabstractBACKGROUND: One of the most important and often neglected components of a successful RNA sequencing (RNA-Seq) experiment is sample size estimation. A few negative binomial model-based methods have been developed to estimate sample size based on the parameters of a single gene. However, thousands of genes are quantified and tested for differential expression simultaneously in RNA-Seq experiments. Thus, additional issues should be carefully addressed, including the false discovery rate for multiple statistic tests, widely distributed read counts and dispersions for different genes. RESULTS: To solve these issues, we developed a sample size and power estimation method named RnaSeqSampleSize, based on the distributions of gene average read counts and dispersions estimated from real RNA-seq data. Datasets from previous, similar experiments such as the Cancer Genome Atlas (TCGA) can be used as a point of reference. Read counts and their dispersions were estimated from the reference's distribution; using that information, we estimated and summarized the power and sample size. RnaSeqSampleSize is implemented in R language and can be installed from Bioconductor website. A user friendly web graphic interface is provided at http://cqs.mc.vanderbilt.edu/shiny/RnaSeqSampleSize/ . CONCLUSIONS: RnaSeqSampleSize provides a convenient and powerful way for power and sample size estimation for an RNAseq experiment. It is also equipped with several unique features, including estimation for interested genes or pathway, power curve visualization, and parameter optimization. Shilin Zhao, Chung-I Li, Quanhu Sheng, Shyr Yu |
BMC Bioinform. | 5 |
| 2017 | StrandScript: evaluation of Illumina genotyping array design and strand correctionabstractSUMMARY: After the introduction of high-throughput sequencing, genotyping arrays continue to be a viable source for conducting large-scale genetic studies. Currently, Illumina is one of the largest genotyping array manufacturers. One technical issue that has always plagued the post-processing of Illumina genotyping array data is the strand definition. Against convention, Illumina uses their own definition of strand, which is inconsistent with the standard reference forward and reverse definition. This issue has been a major obstacle in the consistency of reporting, meta-analysis and correct interpretation of phenotype association results. To date, the strand issue has not been adequately addressed, prompting us to develop StrandScript, a tool that can convert all genotyping data generated from Illumina genotyping arrays to the reference forward strand. StrandScript works independently of the Illumina array version and is future proof for newer Illumina array designs. Furthermore, StrandScript can examine an Illumina genotyping array manifest file and can detect all problematic SNPs, including SNPs with wrong RS ID and SNPs with mismatched probe sequences. Here, we introduce StrandScript's design and development, and demonstrate its effectiveness using real genotyping data. AVAILABILITY AND IMPLEMENTATION: https://github.com/seasky002002/Strandscript. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jing Wang 0026, David C. Samuels, Shyr Yu |
Bioinform. | 3 |
| 2016 | Mitochondria sequence mapping strategies and practicability of mitochondria variant detection from exome and RNA sequencing dataabstractThe rapid progress in high-throughput sequencing has significantly enriched our capacity for studying the mitochondrial DNA (mtDNA). In addition to performing specific mitochondrial targeted sequencing, an increasingly popular alternative approach is using the off-target reads from exome sequencing to infer mtDNA variants, including single nucleotide polymorphisms (SNPs) and heteroplasmy. However, the effectiveness and practicality of this approach have not been tested. Recently, RNAseq data have also been suggested as a good source for alternative data mining, but whether mitochondrial variants can be detected from RNAseq data has not been validated. We designed a study to evaluate the practicability of mtDNA variant detection using exome and RNA sequencing data. Five breast cancer cell lines were sequenced through mitochondrial targeted, exome, and RNA sequencing. Mitochondrial targeted sequencing was used as the gold standard to compute the validation and false discovery rates of SNP and heteroplasmy detection in exome and RNAseq data. We found that exome and RNA sequencing can accurately detect mitochondrial SNPs. However, the lower false discovery rate makes exome sequencing a better choice for heteroplasmy detection than RNAseq. Furthermore, we examined three alignment strategies and found that aligning reads directly to the mitochondrial reference genome or aligning reads to the nuclear and mitochondrial references genomes simultaneously produced the best results, and that aligning to the nuclear genome first and afterwards to the mitochondrial genome performed poorly. In conclusion, our study provides important guidelines for future studies that intend to use either exome sequencing or RNAseq data to infer mitochondrial SNPs and heteroplasmy. David C. Samuels, Brian D. Lehmann, Thomas Stricker, Jennifer A. Pietenpol, Shyr Yu |
Briefings Bioinform. | 6 |
| 2016 | Proceedings of the 15th Annual UT-KBRIN Bioinformatics Summit 2016: Cadiz, KY, USA. 8-10 April 2016abstractI1 Proceedings of the Fifteenth Annual UT- KBRIN Bioinformatics Summit 2016 Eric C. Rouchka, Julia H. Chariker, Benjamin J. Harrison, Juw Won Park P1 CC-PROMISE: Projection onto the Most Interesting Statistical Evidence (PROMISE) with Canonical Correlation to integrate gene expression and methylation data with multiple pharmacologic and clinical endpoints Xueyuan Cao, Stanley Pounds, Susana Raimondi, James Downing, Raul Ribeiro, Jeffery Rubnitz, Jatinder Lamba P2 Integration of microRNA-mRNA interaction networks with gene expression data to increase experimental power Bernie J Daigle, Jr. P3 Designing and writing software for in silico subtractive hybridization of large eukaryotic genomes Deborah Burgess, Stephanie Gehrlich, John C Carmen P4 Tracking the molecular evolution of Pax gene Nicholas Johnson; Chandrakanth Emani P5 Identifying genetic differences in thermally dimorphic and state specific fungi using in silico genomic comparison Stephanie Gehrlich, Deborah Burgess, John C Carmen P6 Identification of conserved genomic regions and variation therein amongst Cetartiodactyla species using next generation sequencing Kalpani De Silva, Michael P Heaton, Theodore S Kalbfleisch P7 Mining physiological data to identify patients with similar medical events and phenotypes Teeradache Viangteeravat, Rahul Mudunuri, Oluwaseun Ajayi, Fatih Şen, Eunice Y Huang P8 Smart brief for home health monitoring Mohammad Mohebbi, Luaire Florian, Douglas J Jackson, John F Naber P9 Side-effect term matching for computational adverse drug reaction predictions AKM Sabbir, Sally R Ellingson P10 Enrichment vs robustness: A comparison of transcriptomic data clustering metrics Yuping Lu, Charles A Phillips, Michael A Langston P11 Deep neural networks for transcriptome-based cancer classification Rahul K Sevakula, Raghuveer Thirukovalluru, Nishchal K. Verma, Yan Cui P12 Motif discovery using K-means clustering Mohammed Sayed, Juw Won Park P13 Large scale discovery of active enhancers from nascent RNA sequencing Jing Wang, Qi Liu, Yu Shyr P14 Computationally characterizing genomic pipelines and benchmarking results using GATK best practices on the high performance computing cluster at the University of Kentucky Xiaofei Zhang, Sally R Ellingson P15 Development of approaches enabling the identification of abnormal gene expression from RNA-Seq in personalized oncology Naresh Prodduturi, Gavin R Oliver, Diane Grill, Jie Na, Jeanette Eckel-Passow, Eric W Klee P16 Processing RNA-Seq data of plants infected with coffee ringspot virus Michael M Goodin, Mark Farman, Harrison Inocencio, Chanyong Jang, Jerzy W Jaromczyk, Neil Moore, Kelly Sovacool P17 Comparative transcriptomics of three Acinetobacter baumanii clinical isolates with different antibiotic resistance patterns Leon Dent, Mike Izban, Sammed Mandape, Shruti Sakhare, Siddharth Pratap, Dana Marshall P18 Metagenomic assessment of possible microbial contamination in the equine reference genome assembly M Scotty DePriest, James N MacLeod, Theodore S Kalbfleisch P19 Molecular evolution of cancer driver genes Chandrakanth Emani, Hanady Adam, Ethan Blandford, Joel Campbell, Joshua Castlen, Brittany Dixon, Ginger Gilbert, Aaron Hall, Philip Kreisle, Jessica Lasher, Bethany Oakes, Allison Speer, Maximilian Valentine P20 Biorepository Laboratory Information Management System Naga Satya V Rao Nagisetty, Rony Jose, Teeradache Viangteeravat, Robert Rooney, David Hains Eric C. Rouchka, Julia H. Chariker, Benjamin J. Harrison, Juw Won Park, Xueyuan Cao, Stan Pounds, Susana C. Raimondi, James R. Downing, Raul C. Ribeiro, Jeffrey Rubnitz, Jatinder Lamba, Bernie J. Daigle Jr., Deborah Burgess, Stephanie Gehrlich, John C. Carmen, Chandrakanth Emani, Kalpani De Silva, Michael P. Heaton, Ted Kalbfleisch, Teeradache Viangteeravat, Rahul Mudunuri, Oluwaseun Ajayi, Fatih Sen, Eunice Y. Huang, Mohammad Mohebbi, Luaire Florian, Douglas J. Jackson, John F. Naber, Akm Sabbir, Sally R. Ellingson, Yuping Lu, Charles A. Phillips, Michael A. Langston, Rahul Kumar Sevakula, Raghuveer Thirukovalluru, Nishchal K. Verma, Yan Cui 0001, Mohammed Sayed, Jing Wang 0026, Qi Liu 0024, Shyr Yu, Naresh Prodduturi, Gavin R. Oliver, Diane Grill, Jie Na, Jeanette Eckel-Passow, Eric W. Klee, Michael M. Goodin, Mark L. Farman, Harrison Inocencio, Chanyong Jang, Jerzy W. Jaromczyk, Neil Moore, Kelly L. Sovacool, Leon Dent, Mike Izban, Sammed N. Mandape, Shruti S. Sakhare, Siddharth Pratap, Dana Marshall, M. Scotty Depriest, James N. MacLeod, Hanady Adam, Ethan Blandford, Joel Campbell, Joshua Castlen, Brittany Dixon, Ginger Gilbert, Aaron Hall, Philip Kreisle, Jessica Lasher, Bethany Oakes, Allison Speer, Maximilian Valentine, Naga Satya Venkateswara Ra Nagisetty, Rony Jose, Robert W. Rooney, David Hains |
BMC Bioinform. | 42 |
| 2016 | Evidence Combination From an Evolutionary Game Theory PerspectiveabstractDempster-Shafer evidence theory is a primary methodology for multisource information fusion because it is good at dealing with uncertain information. This theory provides a Dempster's rule of combination to synthesize multiple evidences from various information sources. However, in some cases, counter-intuitive results may be obtained based on that combination rule. Numerous new or improved methods have been proposed to suppress these counter-intuitive results based on perspectives, such as minimizing the information loss or deviation. Inspired by evolutionary game theory, this paper considers a biological and evolutionary perspective to study the combination of evidences. An evolutionary combination rule (ECR) is proposed to help find the most biologically supported proposition in a multievidence system. Within the proposed ECR, we develop a Jaccard matrix game to formalize the interaction between propositions in evidences, and utilize the replicator dynamics to mimick the evolution of propositions. Experimental results show that the proposed ECR can effectively suppress the counter-intuitive behaviors appeared in typical paradoxes of evidence theory, compared with many existing methods. Properties of the ECR, such as solution's stability and convergence, have been mathematically proved as well. Xinyang Deng, Deqiang Han, Jean Dezert, Yong Deng 0001, Shyr Yu |
IEEE Trans. Cybern. | 5 |
| 2015 | PheWAS Network Analysis and Visualization
Yaomin Xu, Todd L. Edwards, Lisa Bastarache, Rebecca N. Jerome, Shilin Zho, Eric Torstenson, Wei-Qi Wei, Jana Shirey-Rice, Erica A. Bowton, Shyr Yu, Jill M. Pulley, Joshua C. Denny |
AMIA | 10 |
| 2015 | Genome measures used for quality control are dependent on gene function and ancestryabstractMOTIVATION: The transition/transversion (Ti/Tv) ratio and heterozygous/nonreference-homozygous (het/nonref-hom) ratio have been commonly computed in genetic studies as a quality control (QC) measurement. Additionally, these two ratios are helpful in our understanding of the patterns of DNA sequence evolution. RESULTS: To thoroughly understand these two genomic measures, we performed a study using 1000 Genomes Project (1000G) released genotype data (N=1092). An additional two datasets (N=581 and N=6) were used to validate our findings from the 1000G dataset. We compared the two ratios among continental ancestry, genome regions and gene functionality. We found that the Ti/Tv ratio can be used as a quality indicator for single nucleotide polymorphisms inferred from high-throughput sequencing data. The Ti/Tv ratio varies greatly by genome region and functionality, but not by ancestry. The het/nonref-hom ratio varies greatly by ancestry, but not by genome regions and functionality. Furthermore, extreme guanine + cytosine content (either high or low) is negatively associated with the Ti/Tv ratio magnitude. Thus, when performing QC assessment using these two measures, care must be taken to apply the correct thresholds based on ancestry and genome region. Failure to take these considerations into account at the QC stage will bias any following analysis. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jing Wang 0026, Leon Raskin, David C. Samuels, Shyr Yu |
Bioinform. | 4 |
| 2015 | Population structure analysis on 2504 individuals across 26 ancestries using bioinformatics approachesabstractBackground Characterizing genetic diversity is crucial for reconstructing human evolution and for understanding the genetic basis of complex diseases; however, human population genetics are very complicated. Previously, we proved that based on the Hardy-Weinberg equilibrium, the heterozygous vs. non-reference homozygous single nucleotide polymorphism (SNP) ratio (het/nonref-hom) is two [1]. Later, we found that this ratio is race dependent, with African being the most genetically diverse race and Asian being the most homozygous [2]. This observation prompted us to conduct further study to understand the reasoning behind this diversity. Jing Wang 0026, David C. Samuels, Shyr Yu |
BMC Bioinform. | 3 |
| 2015 | Practicality of identifying mitochondria variants from exome and RNAseq dataabstractBackground The rapid progress in high throughput sequencing technology has significantly enriched our capability to study mitochondria genomes. Other than performing mitochondria targeted sequencing, an increasingly popular alternative approach is to utilize the off-target reads from exome sequencing to infer mitochondria genomic variants including SNP and heteroplasmy[1-9]. However, the effectiveness and practicality of such an approach has not been tested. Recently, RNAseq data has also been suggested as good source for alternative data mining[10,11], but whether mitochondria variants are minable has not been studied. David C. Samuels, Brian D. Lehmann, Thomas Stricker, Jennifer A. Pietenpol, Shyr Yu |
BMC Bioinform. | 6 |
| 2014 | Detection of internal exon deletion with exon DelabstractBACKGROUND: Exome sequencing allows researchers to study the human genome in unprecedented detail. Among the many types of variants detectable through exome sequencing, one of the most over looked types of mutation is internal deletion of exons. Internal exon deletions are the absence of consecutive exons in a gene. Such deletions have potentially significant biological meaning, and they are often too short to be considered copy number variation. Therefore, to the need for efficient detection of such deletions using exome sequencing data exists. RESULTS: We present ExonDel, a tool specially designed to detect homozygous exon deletions efficiently. We tested ExonDel on exome sequencing data generated from 16 breast cancer cell lines and identified both novel and known IEDs. Subsequently, we verified our findings using RNAseq and PCR technologies. Further comparisons with multiple sequencing-based CNV tools showed that ExonDel is capable of detecting unique IEDs not found by other CNV tools. CONCLUSIONS: ExonDel is an efficient way to screen for novel and known IEDs using exome sequencing data. ExonDel and its source code can be downloaded freely at https://github.com/slzhao/ExonDel. Shilin Zhao, Brian D. Lehmann, Quanhu Sheng, Timothy M. Shaver, Thomas Stricker, Jennifer A. Pietenpol, Shyr Yu |
BMC Bioinform. | 8 |
| 2014 | DupChecker: a bioconductor package for checking high-throughput genomic data redundancy in meta-analysisabstractBACKGROUND: Meta-analysis has become a popular approach for high-throughput genomic data analysis because it often can significantly increase power to detect biological signals or patterns in datasets. However, when using public-available databases for meta-analysis, duplication of samples is an often encountered problem, especially for gene expression data. Not removing duplicates could lead false positive finding, misleading clustering pattern or model over-fitting issue, etc in the subsequent data analysis. RESULTS: We developed a Bioconductor package Dupchecker that efficiently identifies duplicated samples by generating MD5 fingerprints for raw data. A real data example was demonstrated to show the usage and output of the package. CONCLUSIONS: Researchers may not pay enough attention to checking and removing duplicated samples, and then data contamination could make the results or conclusions from meta-analysis questionable. We suggest applying DupChecker to examine all gene expression data sets before any data analysis step. Quanhu Sheng, Shyr Yu |
BMC Bioinform. | 2 |
| 2014 | Heatmap3: an improved heatmap package with more powerful and convenient featuresabstractBackgroundHeat map and clustering are used frequently in expression analysis studies for data visualization.Simple clustering and heat map can be produced from the "heatmap" function in R language.However, it has some limitations in producing advanced graphics and is not highly customizable.Thus, we developed an R package heatmap 3 which significantly improves the original heatmap by adding more powerful and convenient features and providing a highly customizable interface. Shilin Zhao, Quanhu Sheng, Shyr Yu |
BMC Bioinform. | 4 |
| 2014 | CLIP-EZ: a computational tool for HITS-CLIP data analysisabstractBackground Mapping the binding regions of mRNA-binding proteins is critical to the understanding of their regulatory roles in cellular processes. Recent development in experimental technologies combines high throughput sequencing with crosslink immunoprecipitation (HITS-CLIP), which has the merit of detecting RNA-protein interaction sites at a high resolution to single nucleotide level. Analysis of such data typically involves many steps, and teasing out true signals from noise requires crosslink induced mutations (CIMS) analysis, peak identification, integration of the two signal types, and some downstream analysis such as motif finding and conservation evaluation. To our knowledge, there is a lack of a single computational tool that can perform all the tasks as mentioned in an easily accessible manner. Despite the fact that there are several tools available, each performs an individual task. Xue Zhong, Qi Liu 0024, Shyr Yu |
BMC Bioinform. | 3 |
| 2013 | MitoSeek: extracting mitochondria information and performing high-throughput mitochondria sequencing analysisabstractMOTIVATION: Exome capture kits have capture efficiencies that range from 40 to 60%. A significant amount of off-target reads are from the mitochondrial genome. These unintentionally sequenced mitochondrial reads provide unique opportunities to study the mitochondria genome. RESULTS: MitoSeek is an open-source software tool that can reliably and easily extract mitochondrial genome information from exome and whole genome sequencing data. MitoSeek evaluates mitochondrial genome alignment quality, estimates relative mitochondrial copy numbers and detects heteroplasmy, somatic mutation and structural variants of the mitochondrial genome. MitoSeek can be set up to run in parallel or serial on large exome sequencing datasets. AVAILABILITY: https://github.com/riverlee/MitoSeek Chung-I Li, Shyr Yu, David C. Samuels |
Bioinform. | 4 |
| 2013 | Sample size calculation based on exact test for assessing differential expression analysis in RNA-seq dataabstractBACKGROUND: Sample size calculation is an important issue in the experimental design of biomedical research. For RNA-seq experiments, the sample size calculation method based on the Poisson model has been proposed; however, when there are biological replicates, RNA-seq data could exhibit variation significantly greater than the mean (i.e. over-dispersion). The Poisson model cannot appropriately model the over-dispersion, and in such cases, the negative binomial model has been used as a natural extension of the Poisson model. Because the field currently lacks a sample size calculation method based on the negative binomial model for assessing differential expression analysis of RNA-seq data, we propose a method to calculate the sample size. RESULTS: We propose a sample size calculation method based on the exact test for assessing differential expression analysis of RNA-seq data. CONCLUSIONS: The proposed sample size calculation method is straightforward and not computationally intensive. Simulation studies to evaluate the performance of the proposed sample size method are presented; the results indicate our method works well, with achievement of desired power. Chung-I Li, Pei-Fang Su, Shyr Yu |
BMC Bioinform. | 3 |
| 2011 | Wave-spec: a preprocessing package for mass spectrometry dataabstractAbstract Summary:Wave-spec is a pre-processing package for mass spectrometry (MS) data. The package includes several novel algorithms that overcome conventional difficulties with the pre-processing of such data. In this application note, we demonstrate step-by-step use of this package on a real-world MALDI dataset. Availability: The package can be downloaded at http://www.vicc.org/biostatistics/supp.php. A shared mailbox ([email protected]) also is available for questions regarding application of the package. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Ming Li 0022, Joan Zhang, Heidi Chen, Shyr Yu |
Bioinform. | 5 |
| 2009 | A novel comprehensive wave-form MS data processing methodabstractMOTIVATION: Mass spectrometry (MS) can generate high-throughput protein profiles for biomedical research to discover biologically related protein patterns/biomarkers. The noisy functional MS data collected by current technologies, however, require consistent, sensitive and robust data-processing techniques for successful biomedical application. Therefore, it is important to detect features precisely for each spectrum, quantify them well and assign a unique label to features from the same protein/peptide across spectra. RESULTS: In this article, we propose a new comprehensive MS data preprocessing package, Wave-spec, which includes several novel algorithms. It can overcome several conventional difficulties. Wave-spec can be applied to multiple types of MS data generated with different MS technologies. Results from this new package were evaluated and compared to several existing approaches based on a MALDI-TOF MS dataset. AVAILABILITY: An example of MATLAB scripts used to implement the methods described in this article, along with Supplementary Figures, can be found at http://www.vicc.org/biostatistics/supp.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ming Li 0022, Don Hong, Dean Billheimer, Huiming Li, Baogang J. Xu, Shyr Yu |
Bioinform. | 7 |
| 2009 | Web-based interrogation of gene expression signatures using EXALTabstractBACKGROUND: Widespread use of high-throughput techniques such as microarrays to monitor gene expression levels has resulted in an explosive growth of data sets in public domains. Integration and exploration of these complex and heterogeneous data have become a major challenge. RESULTS: The EXALT (EXpression signature AnaLysis Tool) online program enables meta-analysis of gene expression profiles derived from publically accessible sources. Searches can be executed online against two large databases currently containing more than 28,000 gene expression signatures derived from GEO (Gene Expression Omnibus) and published expression profiles of human cancer. Comparisons among gene expression signatures can be performed with homology analysis and co-expression analysis. Results can be visualized instantly in a plot or a heat map. Three typical use cases are illustrated. CONCLUSIONS: The EXALT online program is uniquely suited for discovering relationships among transcriptional profiles and searching gene expression patterns derived from diverse physiological and pathological settings. The EXALT online program is freely available for non-commercial users from http://seq.mc.vanderbilt.edu/exalt/. Qingchao Qiu, Lu Xie, Joseph Fullerton, Shyr Yu, Alfred L. George Jr., Yajun Yi |
BMC Bioinform. | 6 |
| 2008 | Use of normalization methods for analysis of microarrays containing a high degree of gene effectsabstractBACKGROUND: High-throughput microarrays are widely used to study gene expression across tissues and developmental stages. Analysis of gene expression data is challenging in these experiments due to the presence of significant percentages of differentially expressed genes (DEG) observed between tissues and developmental stages. Data normalization methods that are widely used today are not designed for data with a large proportion of tissue or gene effects. RESULTS: In our current study, we describe a novel two-dimensional nonparametric normalization method for analyzing microarray data which functions well in the absence or presence of large numbers of gene effects. Rather than relying on an assumption of low variability among most genes, the method implements a unique peak selection strategy to distinguish DEG from genes that are invariant in expression, prior to nonlinear curve fitting. We compared the method under simulated and experimental conditions with five alternative nonlinear normalization approaches: quantile, lowess, robust lowess, invariant set, and cross-correlation (Xcorr). Simulations included various percentages of simulated DEG and the experimental data used is from publicly available datasets known to be difficult to analyze due to the presence of approximately 34% DEG. CONCLUSION: We have demonstrated that the new method provides considerable improvement in the accuracy of data normalization when large proportions of gene effects are present. The performance improvement is mostly attributed to its variable selection component, which is designed to separate expression invariant genes from DEG. Adding this key component of the new method to alternative normalization approaches rescues the most of the sensitivity of these methods to gene effects. The results indicate that our method may be used without prior knowledge of or assumptions about housekeeping genes to normalize microarrays that are quite different. Terri T. Ni, William J. Lemon, Shyr Yu, Tao P. Zhong |
BMC Bioinform. | 3 |
| 2000 | Registration of Physical Space to Laparoscopic Image Space for use in Minimally Invasive Hepatic SurgeryabstractWhile laparoscopes are used for numerous minimally invasive (MI) procedures, MI liver resection and ablative surgery is infrequently performed. The paucity of cases is due to the restriction of the field of view by the laparoscope and the difficulty in determining tumor location and margins under video guidance. By merging MI surgery with interactive, image-guided surgery (IIGS), we hope to overcome localization difficulties present in laparoscopic liver procedures. One key component of any IIGS system is the development of accurate registration techniques to map image space to physical or patient space. This manuscript focuses on the accuracy and analysis of the direct linear transformation (DLT) method to register physical space with laparoscopic image space on both distorted and distortion-corrected video images. Experiments were conducted on a liver-sized plastic phantom affixed with 20 markers at various depths. After localizing the points in both physical and laparoscopic image space, registration accuracy was assessed for different combinations and numbers of control points (n) to determine the quantity necessary to develop a robust registration matrix. For n = 11, average target registration error (TRE) was 0.70 +/- 0.20 mm. We also studied the effects of distortion correction on registration accuracy. For the particular distortion correction method and laparoscope used in our experiments, there was no statistical significance between physical to image registration error for distorted and corrected images. In cases where a minimum number of control points (n = 6) are acquired, the DLT is often not stable and the mathematical process can lead to high TRE values. Mathematical filters developed through the analysis of the DLT were used to prospectively eliminate outlier cases where the TRE was high. For n = 6, prefilter average TRE was 17.4 +/- 153 mm for all trials; when the filters were applied, average TRE decreased to 1.64 +/- 1.10 mm for the remaining trials. James D. Stefansic, Alan J. Herline, Shyr Yu, William C. Chapman, J. Michael Fitzpatrick, Robert L. Galloway |
IEEE Trans. Medical Imaging | 3 |
| 1998 | Visual Assessment of the Accuracy of Retrospective Registration of MR and CT Images of the BrainabstractIn a previous study we demonstrated that automatic retrospective registration algorithms can frequently register magnetic resonance (MR) and computed tomography (CT) images of the brain with an accuracy of better than 2 mm, but in that same study we found that such algorithms sometimes fail, leading to errors of 6 mm or more. Before these algorithms can be used routinely in the clinic, methods must be provided for distinguishing between registration solutions that are clinically satisfactory and those that are not. One approach is to rely on a human observer to inspect the registration results and reject images that have been registered with insufficient accuracy. In this paper, we present a methodology for evaluating the efficacy of the visual assessment of registration accuracy. Since the clinical requirements for level of registration accuracy are likely to be application dependent, we have evaluated the accuracy of the observer's estimate relative to six thresholds: 1-6 mm. The performance of the observers was evaluated relative to the registration solution obtained using external fiducial markers that are screwed into the patient's skull and that are visible in both MR and CT images. This fiducial marker system provides the gold standard for our study. Its accuracy is shown to be approximately 0.5 mm. Two experienced, blinded observers viewed five pairs of clinical MR and CT brain images, each of which had each been misregistered with respect to the gold standard solution. Fourteen misregistrations were assessed for each image pair with misregistration errors distributed between 0 and 10 mm with approximate uniformity. For each misregistered image pair each observer estimated the registration error (in millimeters) at each of five locations distributed around the head using each of three assessment methods. These estimated errors were compared with the errors as measured by the gold standard to determine agreement relative to each of the six thresholds, where agreement means that the two errors lie on the same side of the threshold. The effect of error in the gold standard itself is taken into account in the analysis of the assessment methods. The results were analyzed by means of the Kappa statistic, the agreement rate, and the area of receiver-operating-characteristic (ROC) curves. No assessment performed well at 1 mm, but all methods performed well at 2 mm and higher. For these five thresholds, two methods agreed with the standard at least 80% of the time and exhibited mean ROC areas greater than 0.84. One of these same methods exhibited Kappa statistics that indicated good agreement relative to chance (Kappa > 0.6) between the pooled observers and the standard for these same five thresholds. Further analysis demonstrates that the results depend strongly on the choice of the distribution of misregistration errors presented to the observers. J. Michael Fitzpatrick, Derek L. G. Hill, Shyr Yu, Jay B. West, Colin Studholme, Calvin R. Maurer Jr. |
IEEE Trans. Medical Imaging | 3 |