EDBT 2026 Demo / reviewers in the wild / expert
Guanghua Xiao
dblp:59/751
· DBLP profile ↗
23ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advances in predicting omics profiles from imaging dataabstractWhile traditional imaging techniques, such as histopathology, are often part of clinical workflows, molecular profiling remains more difficult to conduct and is less cost-effective. Thus, the prediction of molecular 'omics' data directly from imaging has emerged as an appealing alternative. While existing reviews have mentioned image-based prediction of biomarkers within specific disease contexts, this review provides a comprehensive overview of current methods that leverage imaging to predict (i) DNA-based aberrations, (ii) bulk transcriptomic profiles, (iii) single-cell transcriptomics, and (iv) spatial transcriptomics across disease contexts and imaging modalities. To address the complexity of these predictive tasks, we find that many studies employ cutting-edge deep learning strategies for image processing, feature extraction, feature aggregation, and downstream molecular prediction. In this review, we highlight the diverse applications of both deep learning-based and modern statistical frameworks designed for image-based omics prediction. The insights gleaned from these inferred molecular data have broad clinical relevance and will continue to improve our understanding of the relationships between molecular and visual features, paving the way for new diagnostic and therapeutic applications. Alexa H. Beachum, Yuansheng Zhou, Qiwei Li 0001, Guanghua Xiao, Lin Xu 0005 |
Briefings Bioinform. | 5 |
| 2026 | GNN-EGG: Graph neural network explanations via graph generation
Art Taychameekiatchai, Liwei Jia, Zhikai Chi, Yueshuang Xu, Guanghua Xiao, Xiaowei Zhan |
Neurocomputing | 6 |
| 2025 | ClinBench: A Standardized Multi-Domain Framework for Evaluating Large Language Models in Clinical Information ExtractionabstractLarge Language Models (LLMs) offer substantial promise for clinical natural language processing (NLP); however, a lack of standardized benchmarking methodologies limits their objective evaluation and practical translation. To address this gap, we introduce ClinBench, an open-source, multi-model, multi-domain benchmarking framework. ClinBench is designed for the rigorous evaluation of LLMs on important structured information extraction tasks (e.g., tumor staging, histologic diagnoses, atrial fibrillation, and social determinants of health) from unstructured clinical notes. The framework standardizes the evaluation pipeline by: (i) operating on consistently structured input datasets; (ii) employing dynamic, YAML-based prompting for uniform task definition; and (iii) enforcing output validation via JSON schemas, supporting robust comparison across diverse LLM architectures. We demonstrate ClinBench through a large-scale study of 11 prominent LLMs (e.g., GPT-4o series, LLaMA3 variants, Mixtral) across three clinical domains using configurations of public datasets (TCGA for lung cancer, MIMIC-IV-ECG for atrial fibrillation, and MIMIC notes for SDOH). Our results reveal significant performance-efficiency trade-offs. For example, when averaged across the four benchmarked clinical extraction tasks, GPT-3.5-turbo achieved a mean F1 score of 0.83 with a mean runtime of 16.8 minutes. In comparison, LLaMA3.1-70b obtained a similar mean F1 of 0.82 but required a substantially longer mean runtime of 42.7 minutes. GPT-4o-mini also presented a favorable balance with a mean F1 of 0.81 and a mean runtime of 13.4 minutes. ClinBench provides a unified, extensible framework and empirical insights for reproducible, fair LLM benchmarking in clinical NLP. By enabling transparent and standardized evaluation, this work advances data-centric AI research, informs model selection based on performance, cost, and clinical priorities, and supports the effective integration of LLMs into healthcare. The framework and evaluation code are publicly available at https://github.com/ismaelvillanuevamiranda/ClinBench/. Ismael Villanueva-Miranda, Zifan Gu, Donghan M. Yang, Kuroush Nezafati, Jingwei Huang 0002, Peifeng Ruan, Xiaowei Zhan, Guanghua Xiao |
NeurIPS | 8 |
| 2025 | SBDH-Reader: a large language model-powered method for extracting social and behavioral determinants of health from clinical notesabstractOBJECTIVE: Social and behavioral determinants of health (SBDH) are increasingly recognized as essential for prognostication and informing targeted interventions. Clinical notes often contain details about SBDH in unstructured format. Conventional extraction methods for these data tend to be labor intensive, inaccurate, and/or unscalable. In this study, we aim to develop and validate a large language model (LLM)-powered method to extract structured SBDH data from clinical notes through prompt engineering. MATERIALS AND METHODS: We developed SBDH-Reader to extract 6 categories of granular SBDH data by prompting GPT-4o, including employment, housing, marital status, and substance use including alcohol, tobacco, and drug use. SBDH-Reader was developed using 7225 notes from 6382 patients in the MIMIC-III database (2001-2012) and externally validated using 971 notes from 437 patients at The University of Texas Southwestern Medical Center (UTSW; 2022-2023). We evaluated SBDH-Reader's performance against human-annotated ground truths based on precision, recall, F1, and confusion matrix. RESULTS: When tested on the UTSW validation set, SBDH-Reader achieved a macro-average F1 ranging from 0.94 to 0.98 across 6 SBDH categories. For clinically relevant adverse attributes, F1 ranged from 0.96 (employment; housing) to 0.99 (tobacco use). When extracting any adverse attributes across all SBDH categories, SBDH-Reader achieved an F1 of 0.97, recall of 0.97, and precision of 0.98 in the independent validation set. DISCUSSION: SBDH-Reader demonstrated strong performance in extracting structured SBDH data through effective prompt engineering of a general-purpose LLM, without the need for task-specific fine-tuning. Its modular design and adaptability to diverse datasets and documentation patterns support its applicability in real-world clinical settings. CONCLUSION: SBDH-Reader has the potential to serve as a scalable and effective method for collecting real-time, patient-level SBDH data to support clinical research and care. Zifan Gu, Lesi He, Awais Naeem, Pui Man Chan, Asim Mohamed, Hafsa Khalil, Yujia Guo, Jingwei Huang 0002, Ismael Villanueva-Miranda, Ying Ding 0001, Wenqi Shi 0002, Matthew E. Dupre, Guanghua Xiao, Eric D. Peterson, Ann Marie Navar, Donghan M. Yang |
J. Am. Medical Informatics Assoc. | 13 |
| 2024 | SegPath-YOLO: A High-Speed, High-Accuracy Pathology Image Segmentation and Tumor Microenvironment Feature Extraction ToolabstractIn the field of digital pathology, accurate and rapid segmentation of pathology images are critical for cancer diagnosis and prognosis. This paper presents SegPath-YOLO, a novel tool in digital pathology for nuclei segmentation and tumor microenvironment feature extraction tool, leveraging the robustness of YOLO series with significant enhancements for accuracy and speed. SegPath-YOLO addresses critical challenges in pathology image analysis, particularly in handling overlapping and high-density cellular structures and ensuring rapid processing without sacrificing precision. The novelty of SegPath-YOLO lies in its Segmentation and Overlapping-Aware Loss, which utilizes a binary overlap mask to identify and enhance the loss in overlapping regions. In conjunction with PathNuclei attention mechanisms, SegPath-YOLO not only refines segmentation results but also contributes to a deeper characterization and quantification of the tumor microenvironment, significantly aiding in survival outcome predictions. The comprehensive evaluations demonstrate that SegPath-YOLO achieves superior performance compared to existing tools, including the YOLOv8 by effectively balancing computational efficiency with high accuracy. We validate SegPath-YOLO’s across two different tissue types, it demonstrates its potential as a high generalizability tool in digital pathology. Overall, SegPath-YOLO is a state of the art tool for pathologist and researcher, enabling the extraction of reliable and clinically relevant data from complex pathology image. The SegPath-YOLO can be acessed at https://github.com/yaober/SegPath-YOLO. Ruichen Rong, Guanghua Xiao |
BIBM | 4 |
| 2024 | BayeSMART: Bayesian clustering of multi-sample spatially resolved transcriptomics dataabstractThe field of spatially resolved transcriptomics (SRT) has greatly advanced our understanding of cellular microenvironments by integrating spatial information with molecular data collected from multiple tissue sections or individuals. However, methods for multi-sample spatial clustering are lacking, and existing methods primarily rely on molecular information alone. This paper introduces BayeSMART, a Bayesian statistical method designed to identify spatial domains across multiple samples. BayeSMART leverages artificial intelligence (AI)-reconstructed single-cell level information from the paired histology images of multi-sample SRT datasets while simultaneously considering the spatial context of gene expression. The AI integration enables BayeSMART to effectively interpret the spatial domains. We conducted case studies using four datasets from various tissue types and SRT platforms, and compared BayeSMART with alternative multi-sample spatial clustering approaches and a number of state-of-the-art methods for single-sample SRT analysis, demonstrating that it surpasses existing methods in terms of clustering accuracy, interpretability, and computational efficiency. BayeSMART offers new insights into the spatial organization of cells in multi-sample SRT data. Yanghong Guo, Bencong Zhu, Ruichen Rong, Guanghua Xiao, Lin Xu 0005, Qiwei Li 0001 |
Briefings Bioinform. | 6 |
| 2024 | MetaNorm: incorporating meta-analytic priors into normalization of NanoString nCounter dataabstractMOTIVATION: Non-informative or diffuse prior distributions are widely employed in Bayesian data analysis to maintain objectivity. However, when meaningful prior information exists and can be identified, using an informative prior distribution to accurately reflect current knowledge may lead to superior outcomes and great efficiency. RESULTS: We propose MetaNorm, a Bayesian algorithm for normalizing NanoString nCounter gene expression data. MetaNorm is based on RCRnorm, a powerful method designed under an integrated series of hierarchical models that allow various sources of error to be explained by different types of probes in the nCounter system. However, a lack of accurate prior information, weak computational efficiency, and instability of estimates that sometimes occur weakens the approach despite its impressive performance. MetaNorm employs priors carefully constructed from a rigorous meta-analysis to leverage information from large public data. Combined with additional algorithmic enhancements, MetaNorm improves RCRnorm by yielding more stable estimation of normalized values, better convergence diagnostics and superior computational efficiency. AVAILABILITY AND IMPLEMENTATION: R Code for replicating the meta-analysis and the normalization function can be found at github.com/jbarth216/MetaNorm. Jackson Barth, Yuqiu Yang, Guanghua Xiao, Xinlei Wang 0001 |
Bioinform. | 3 |
| 2024 | Navigating electronic health record accuracy by examination of sex incongruent conditionsabstractOBJECTIVE: The increasing reliance on electronic health records (EHRs) for research and clinical care necessitates robust methods for assessing data quality and identifying inconsistencies. To address this need, we develop and apply the incongruence rate (IR) using sex-specific medical conditions. We also characterized participants with incongruent records to better understand the scope and nature of data discrepancies. MATERIALS AND METHODS: In this cross-sectional study, we used the All of Us Research Program's latest version 7 (v7) EHR data to identify prevalent sex-specific conditions and evaluated the occurrence of incongruent cases, quantified as IR. RESULTS: Among the 92 597 males and 152 551 females with condition occurrence data available from All of Us and sex-conformed gender, we identified 167 prevalent sex-specific conditions. Among the 37 537 biological males and 95 499 biological females with these sex-specific conditions, we detected an overall IR of 0.86%. Attempt to include non-cisgender participants result in inflated overall IR. Additionally, a significant proportion of participants with incongruent conditions also presented with conditions congruent to their biological sex, indicating a mix of accurate and erroneous records. These incongruences were not geographically or temporally isolated, suggesting systematic issues in EHR data integrity. DISCUSSION: Our findings call attention to the existence of systemic data incongruences in sex-specific conditions and the need for robust validation checks. Extending IR evaluation to non-cisgender participants or non-sex-based conditions remain a challenge. CONCLUSION: The sex condition-specific IR, when applied to adult populations, provides a valuable metric for data quality assessment in EHRs. Ralph J. Deberardinis, Xiaowei Zhan, Guanghua Xiao |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Contrastive Learning with Dynamic Weighting and Jigsaw Augmentation for Brain Tumor Classification in MRI
Guanghua Xiao, Jie Shen 0004, Zhe Chen 0004, Zhen Zhang 0019, Xiaomin Ge |
Neural Process. Lett. | 1 |
| 2022 | An Adaptive Hierarchical Concatenated Network With A Robust Loss Function For Image Denoising
Guanghua Xiao, Jie Shen 0004, Zhe Chen 0004, Zhen Zhang 0019 |
J. Grid Comput. | 1 |
| 2021 | Spatial molecular profiling: platforms, applications and analysis toolsabstractMolecular profiling technologies, such as genome sequencing and proteomics, have transformed biomedical research, but most such technologies require tissue dissociation, which leads to loss of tissue morphology and spatial information. Recent developments in spatial molecular profiling technologies have enabled the comprehensive molecular characterization of cells while keeping their spatial and morphological contexts intact. Molecular profiling data generate deep characterizations of the genetic, transcriptional and proteomic events of cells, while tissue images capture the spatial locations, organizations and interactions of the cells together with their morphology features. These data, together with cell and tissue imaging data, provide unprecedented opportunities to study tissue heterogeneity and cell spatial organization. This review aims to provide an overview of these recent developments in spatial molecular profiling technologies and the corresponding computational methods developed for analyzing such data. Minzhe Zhang, Thomas Sheffield, Xiaowei Zhan, Qiwei Li 0001, Donghan M. Yang, Yunguan Wang, Shidan Wang, Guanghua Xiao |
Briefings Bioinform. | 10 |
| 2021 | Assessing consistency across functional screening datasets in cancer cellsabstractMOTIVATION: Many high-throughput screening studies have been carried out in cancer cell lines to identify therapeutic agents and targets. Existing consistency assessment studies only examined two datasets at a time, with conclusions based on a subset of carefully selected features rather than considering global consistency of all the data. However, poor concordance can still be observed for a large part of the data even when selected features are highly consistent. RESULTS: In this study, we assembled nine compound screening datasets and three functional genomics datasets. We derived direct measures of consistency as well as indirect measures of consistency based on association between functional data and copy number-adjusted gene expression data. These results have been integrated into a web application-the Functional Data Consistency Explorer (FDCE), to allow users to make queries and generate interactive visualizations so that functional data consistency can be assessed for individual features of interest. AVAILABILITY AND IMPLEMENTATION: The FDCE web tool and we have developed and the functional data consistency measures we have generated are available at https://lccl.shinyapps.io/FDCE/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John D. Minna, Ralph J. Deberardinis, Guanghua Xiao |
Bioinform. | 5 |
| 2021 | Bayesian modeling of spatial molecular profiling data via Gaussian processabstractMOTIVATION: The location, timing and abundance of gene expression (both mRNA and proteins) within a tissue define the molecular mechanisms of cell functions. Recent technology breakthroughs in spatial molecular profiling, including imaging-based technologies and sequencing-based technologies, have enabled the comprehensive molecular characterization of single cells while preserving their spatial and morphological contexts. This new bioinformatics scenario calls for effective and robust computational methods to identify genes with spatial patterns. RESULTS: We represent a novel Bayesian hierarchical model to analyze spatial transcriptomics data, with several unique characteristics. It models the zero-inflated and over-dispersed counts by deploying a zero-inflated negative binomial model that greatly increases model stability and robustness. Besides, the Bayesian inference framework allows us to borrow strength in parameter estimation in a de novo fashion. As a result, the proposed model shows competitive performances in accuracy and robustness over existing methods in both simulation studies and two real data applications. AVAILABILITY AND IMPLEMENTATION: The related R/C++ source code is available at https://github.com/Minzhe/BOOST-GP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qiwei Li 0001, Minzhe Zhang, Guanghua Xiao |
Bioinform. | 4 |
| 2020 | VAMPr: VAriant Mapping and Prediction of antibiotic resistance via explainable features and machine learningabstractAntimicrobial resistance (AMR) is an increasing threat to public health. Current methods of determining AMR rely on inefficient phenotypic approaches, and there remains incomplete understanding of AMR mechanisms for many pathogen-antimicrobial combinations. Given the rapid, ongoing increase in availability of high-density genomic data for a diverse array of bacteria, development of algorithms that could utilize genomic information to predict phenotype could both be useful clinically and assist with discovery of heretofore unrecognized AMR pathways. To facilitate understanding of the connections between DNA variation and phenotypic AMR, we developed a new bioinformatics tool, variant mapping and prediction of antibiotic resistance (VAMPr), to (1) derive gene ortholog-based sequence features for protein variants; (2) interrogate these explainable gene-level variants for their known or novel associations with AMR; and (3) build accurate models to predict AMR based on whole genome sequencing data. We curated the publicly available sequencing data for 3,393 bacterial isolates from 9 species that contained AMR phenotypes for 29 antibiotics. We detected 14,615 variant genotypes and built 93 association and prediction models. The association models confirmed known genetic antibiotic resistance mechanisms, such as blaKPC and carbapenem resistance consistent with the accurate nature of our approach. The prediction models achieved high accuracies (mean accuracy of 91.1% for all antibiotic-pathogen combinations) internally through nested cross validation and were also validated using external clinical datasets. The VAMPr variant detection method, association and prediction models will be valuable tools for AMR research for basic scientists with potential for clinical applicability. Jiwoong Kim, David E. Greenberg, Reed Pifer, Shuang Jiang, Guanghua Xiao, Samuel A. Shelburne, Andrew Y. Koh, Xiaowei Zhan |
PLoS Comput. Biol. | 5 |
| 2019 | A comparative study of rank aggregation methods for partial and top ranked lists in genomic applicationsabstractRank aggregation (RA), the process of combining multiple ranked lists into a single ranking, has played an important role in integrating information from individual genomic studies that address the same biological question. In previous research, attention has been focused on aggregating full lists. However, partial and/or top ranked lists are prevalent because of the great heterogeneity of genomic studies and limited resources for follow-up investigation. To be able to handle such lists, some ad hoc adjustments have been suggested in the past, but how RA methods perform on them (after the adjustments) has never been fully evaluated. In this article, a systematic framework is proposed to define different situations that may occur based on the nature of individually ranked lists. A comprehensive simulation study is conducted to examine the performance characteristics of a collection of existing RA methods that are suitable for genomic applications under various settings simulated to mimic practical situations. A non-small cell lung cancer data example is provided for further comparison. Based on our numerical results, general guidelines about which methods perform the best/worst, and under what conditions, are provided. Also, we discuss key factors that substantially affect the performance of the different methods. Xinlei Wang 0001, Guanghua Xiao |
Briefings Bioinform. | 3 |
| 2019 | DIGREM: an integrated web-based platform for detecting effective multi-drug combinationsabstractMOTIVATION: Synergistic drug combinations are a promising approach to achieve a desirable therapeutic effect in complex diseases through the multi-target mechanism. However, in vivo screening of all possible multi-drug combinations remains cost-prohibitive. An effective and robust computational model to predict drug synergy in silico will greatly facilitate this process. RESULTS: We developed DIGREM (Drug-Induced Genomic Response models for identification of Effective Multi-drug combinations), an online tool kit that can effectively predict drug synergy. DIGREM integrates DIGRE, IUPUI_CCBB, gene set-based and correlation-based models for users to predict synergistic drug combinations with dose-response information and drug-treated gene expression profiles. AVAILABILITY AND IMPLEMENTATION: http://lce.biohpc.swmed.edu/drugcombination. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Minzhe Zhang, Sangin Lee, Guanghua Xiao |
Bioinform. | 4 |
| 2019 | GeNeCK: a web server for gene network construction and visualizationabstractBACKGROUND: Reverse engineering approaches to infer gene regulatory networks using computational methods are of great importance to annotate gene functionality and identify hub genes. Although various statistical algorithms have been proposed, development of computational tools to integrate results from different methods and user-friendly online tools is still lagging. RESULTS: We developed a web server that efficiently constructs gene networks from expression data. It allows the user to use ten different network construction methods (such as partial correlation-, likelihood-, Bayesian- and mutual information-based methods) and integrates the resulting networks from multiple methods. Hub gene information, if available, can be incorporated to enhance performance. CONCLUSIONS: GeNeCK is an efficient and easy-to-use web application for gene regulatory network construction. It can be accessed at http://lce.biohpc.swmed.edu/geneck . Minzhe Zhang, Qiwei Li 0001, Donghyeon Yu, Guanghua Xiao |
BMC Bioinform. | 7 |
| 2018 | Microvessel prediction in H&E Stained Pathology Images using fully convolutional neural networksabstractBACKGROUND: Pathological angiogenesis has been identified in many malignancies as a potential prognostic factor and target for therapy. In most cases, angiogenic analysis is based on the measurement of microvessel density (MVD) detected by immunostaining of CD31 or CD34. However, most retrievable public data is generally composed of Hematoxylin and Eosin (H&E)-stained pathology images, for which is difficult to get the corresponding immunohistochemistry images. The role of microvessels in H&E stained images has not been widely studied due to their complexity and heterogeneity. Furthermore, identifying microvessels manually for study is a labor-intensive task for pathologists, with high inter- and intra-observer variation. Therefore, it is important to develop automated microvessel-detection algorithms in H&E stained pathology images for clinical association analysis. RESULTS: In this paper, we propose a microvessel prediction method using fully convolutional neural networks. The feasibility of our proposed algorithm is demonstrated through experimental results on H&E stained images. Furthermore, the identified microvessel features were significantly associated with the patient clinical outcomes. CONCLUSIONS: This is the first study to develop an algorithm for automated microvessel detection in H&E stained pathology images. Faliu Yi, Lin Yang 0025, Shidan Wang, Guanghua Xiao |
BMC Bioinform. | 7 |
| 2017 | Enhanced construction of gene regulatory networks using hub gene informationabstractBACKGROUND: Gene regulatory networks reveal how genes work together to carry out their biological functions. Reconstructions of gene networks from gene expression data greatly facilitate our understanding of underlying biological mechanisms and provide new opportunities for biomarker and drug discoveries. In gene networks, a gene that has many interactions with other genes is called a hub gene, which usually plays an essential role in gene regulation and biological processes. In this study, we developed a method for reconstructing gene networks using a partial correlation-based approach that incorporates prior information about hub genes. Through simulation studies and two real-data examples, we compare the performance in estimating the network structures between the existing methods and the proposed method. RESULTS: In simulation studies, we show that the proposed strategy reduces errors in estimating network structures compared to the existing methods. When applied to Escherichia coli, the regulation network constructed by our proposed ESPACE method is more consistent with current biological knowledge than the SPACE method. Furthermore, application of the proposed method in lung cancer has identified hub genes whose mRNA expression predicts cancer progress and patient response to treatment. CONCLUSIONS: We have demonstrated that incorporating hub gene information in estimating network structures can improve the performance of the existing methods. Donghyeon Yu, Johan Lim, Xinlei Wang 0001, Faming Liang, Guanghua Xiao |
BMC Bioinform. | 5 |
| 2016 | Imaging-genetic data mapping for clinical outcome prediction via supervised conditional Gaussian graphical modelabstractImaging-genetic data mapping is important for clinical outcome prediction like survival analysis. In this paper, we propose a supervised conditional Gaussian graphical model (SuperCGGM) to uncover survival associated mapping between pathological images and genetic data. The proposed method integrates heterogeneous modal data into the survival model by weighted projection within the data. To obtain a sparse solution, we employ l-1 regularization to the partial log likelihood loss function and propose a cyclic coordinate ascent algorithm to solve it. It also gives a way to bridge the gap between the supervised model with conditional Gaussian graphical model (CGGM). Compared to nine state-of-the-art methods like SuperPCA, CGGM, etc., our method is superior due to its ability of integrating diverse information from heterogeneous modal data in a supervised way. The extensive experiments also show the strong power of SuperCGGM in mapping survival associated image and gene expression signatures. Xinliang Zhu, Jiawen Yao, Guanghua Xiao, Jaime Rodriguez-Canales, Edwin R. Parra, Carmen Behrens, Ignacio I. Wistuba, Junzhou Huang |
BIBM | 3 |
| 2013 | A powerful Bayesian meta-analysis method to integrate multiple gene set enrichment studiesabstractMOTIVATION: Much research effort has been devoted to the identification of enriched gene sets for microarray experiments. However, identified gene sets are often found to be inconsistent among independent studies. This is probably owing to the noisy data of microarray experiments coupled with small sample sizes of individual studies. Therefore, combining information from multiple studies is likely to improve the detection of truly enriched gene classes. As more and more data become available, it calls for statistical methods to integrate information from multiple studies, also known as meta-analysis, to improve the power of identifying enriched gene sets. RESULTS: We propose a Bayesian model that provides a coherent framework for joint modeling of both gene set information and gene expression data from multiple studies, to improve the detection of enriched gene sets by leveraging information from different sources available. One distinct feature of our method is that it directly models the gene expression data, instead of using summary statistics, when synthesizing studies. Besides, the proposed model is flexible and offers an appropriate treatment of between-study heterogeneities that frequently arise in the meta-analysis of microarray experiments. We show that under our Bayesian model, the full posterior conditionals all have known distributions, which greatly facilitates the MCMC computation. Simulation results show that the proposed method can improve the power of gene set enrichment meta-analysis, as opposed to existing methods developed by Shen and Tseng (2010, Bioinformatics, 26, 1316-1323), and it is not sensitive to mild or moderate deviations from the distributional assumption for gene expression data. We illustrate the proposed method through an application of combining eight lung cancer datasets for gene set enrichment analysis, which demonstrates the usefulness of the method. AVAILABILITY: http://qbrc.swmed.edu/software/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Min Chen 0014, Miao Zang, Xinlei Wang 0001, Guanghua Xiao |
Bioinform. | 4 |
| 2013 | SbacHTS: Spatial background noise correction for High-Throughput RNAi ScreeningabstractMOTIVATION: High-throughput cell-based phenotypic screening has become an increasingly important technology for discovering new drug targets and assigning gene functions. Such experiments use hundreds of 96-well or 384-well plates, to cover whole-genome RNAi collections and/or chemical compound files, and often collect measurements that are sensitive to spatial background noise whose patterns can vary across individual plates. Correcting these position effects can substantially improve measurement accuracy and screening success. RESULT: We developed SbacHTS (Spatial background noise correction for High-Throughput RNAi Screening) software for visualization, estimation and correction of spatial background noise in high-throughput RNAi screens. SbacHTS is supported on the Galaxy open-source framework with a user-friendly open access web interface. We find that SbacHTS software can effectively detect and correct spatial background noise, increase signal to noise ratio and enhance statistical detection power in high-throughput RNAi screening experiments. AVAILABILITY: http://www.galaxy.qbrc.org/ Min Soo Kim 0006, Michael A. White, Guanghua Xiao |
Bioinform. | 5 |
| 2012 | Probe mapping across multiple microarray platformsabstractAccess to gene expression data has become increasingly common in recent years; however, analysis has become more difficult as it is often desirable to integrate data from different platforms. Probe mapping across microarray platforms is the first and most crucial step for data integration. In this article, we systematically review and compare different approaches to map probes across seven platforms from different vendors: U95A, U133A and U133 Plus 2.0 from Affymetrix, Inc.; HT-12 v1, HT-12v2 and HT-12v3 from Illumina, Inc.; and 4112A from Agilent, Inc. We use a unique data set, which contains 56 lung cancer cell line samples-each of which has been measured by two different microarray platforms-to evaluate the consistency of expression measurement across platforms using different approaches. Based on the evaluation from the empirical data set, the BLAST alignment of the probe sequences to a recent revision of the Transcriptome generated better results than using annotations provided by Vendors or from Bioconductor's Annotate package. However, a combination of all three methods (deemed the 'Consensus Annotation') yielded the most consistent expression measurement across platforms. To facilitate data integration across microarray platforms for the research community, we develop a user-friendly web-based tool, an API and an R package to map data across different microarray platforms from Affymetrix, Illumina and Agilent. Information on all three can be found at http://qbrc.swmed.edu/software/probemapper/. Jeffrey D. Allen, Siling Wang, Min Chen 0014, Luc Girard, John D. Minna, Guanghua Xiao |
Briefings Bioinform. | 7 |