EDBT 2026 Demo / reviewers in the wild / expert
Pingzhao Hu
dblp:34/1943
· DBLP profile ↗
22ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-9546-2245ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CL-MFAP: A Contrastive Learning-Based Multimodal Foundation Model for Molecular Property Prediction and Antibiotic ScreeningabstractDue to the rise in antimicrobial resistance, identifying novel compounds with antibiotic potential is crucial for combatting this global health issue. However, traditional drug development methods are costly and inefficient. Recognizing the pressing need for more effective solutions, researchers have turned to machine learning techniques to streamline the prediction and development of novel antibiotic compounds. While foundation models have shown promise in antibiotic discovery, current mainstream efforts still fall short of fully leveraging the potential of multimodal molecular data. Recent studies suggest that contrastive learning frameworks utilizing multimodal data exhibit excellent performance in representation learning across various domains. Building upon this, we introduce CL-MFAP, an unsupervised contrastive learning (CL)-based multimodal foundation (MF) model specifically tailored for discovering small molecules with potential antibiotic properties (AP) using three types of molecular data. This model employs 1.6 million bioactive molecules with drug-like properties from the ChEMBL dataset to jointly pretrain three encoders: (1) a transformer-based encoder with rotary position embedding for processing SMILES strings; (2) another transformer-based encoder, incorporating a novel bi-level routing attention mechanism to handle molecular graph representations; and (3) a Morgan fingerprint encoder using a multilayer perceptron, to achieve the contrastive learning purpose. The CL-MFAP outperforms baseline models in antibiotic property prediction by effectively utilizing different molecular modalities and demonstrates superior domain-specific performance when fine-tuned for antibiotic-related property prediction tasks. Gen Zhou, Sugitha Janarthanan, Yutong Lu, Pingzhao Hu |
ICLR | 4 |
| 2025 | Uncertainty-Aware Multi-Objective Reinforcement Learning-Guided Diffusion Models for 3D De Novo Molecular DesignabstractDesigning de novo 3D molecules with desirable properties remains a fundamental challenge in drug discovery and molecular engineering. While diffusion models have demonstrated remarkable capabilities in generating high-quality 3D molecular structures, they often struggle to effectively control complex multi-objective constraints critical for real-world applications. In this study, we propose an uncertainty-aware Reinforcement Learning (RL) framework to guide the optimization of 3D molecular diffusion models toward multiple property objectives while enhancing the overall quality of the generated molecules. Our method leverages surrogate models with predictive uncertainty estimation to dynamically shape reward functions, facilitating balance across multiple optimization objectives. We comprehensively evaluate our framework across three benchmark datasets and multiple diffusion model architectures, consistently outperforming baselines for molecular quality and property optimization. Additionally, Molecular Dynamics (MD) simulations and ADMET profiling of top generated candidates indicate promising drug-like behavior and binding stability, comparable to known Epidermal Growth Factor Receptor (EGFR) inhibitors. Our results demonstrate the strong potential of RL-guided generative diffusion models for advancing automated molecular design. Lianghong Chen, Dongkyu Eugene Kim, Michael Domaratzki, Pingzhao Hu |
NeurIPS | 4 |
| 2025 | Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation LearningabstractMultimodal molecular models often suffer from 3D conformer unreliability and modality collapse, limiting their robustness and generalization. We propose MuMo, a structured multimodal fusion framework that addresses these challenges in molecular representation through two key strategies. To reduce the instability of conformer-dependent fusion, we design a Structured Fusion Pipeline (SFP) that combines 2D topology and 3D geometry into a unified and stable structural prior. To mitigate modality collapse caused by naive fusion, we introduce a Progressive Injection (PI) mechanism that asymmetrically integrates this prior into the sequence stream, preserving modality-specific modeling while enabling cross-modal enrichment. Built on a state space backbone, MuMo supports long-range dependency modeling and robust information propagation. Across 29 benchmark tasks from Therapeutics Data Commons (TDC) and MoleculeNet, MuMo achieves an average improvement of 2.7% over the best-performing baseline on each task, ranking first on 22 of them, including a 27% improvement on the LD50 task. These results validate its robustness to 3D conformer noise and the effectiveness of multimodal fusion in molecular representation. The code is available at: github.com/selmiss/MuMo. Zihao Jing, Yan Yi Li, Sugitha Janarthanan, Alana Deng, Pingzhao Hu |
NeurIPS | 6 |
| 2025 | Out of distribution learning in bioinformatics: advancements and challengesabstractIn the dynamic and complex field of bioinformatics, the development of machine learning models capable of accurately predicting and interpreting genomic data underpins many critical applications, from disease diagnosis to drug discovery. Traditional machine learning models, however, often fail when facing with out-of-distribution (OOD) samples that deviate from their training data, leading to significant performance degradation. This review paper delves into the realm of OOD learning within bioinformatics, highlighting its crucial role in enhancing model generalization and reliability across varied genomic datasets. We provide a comprehensive overview of recent advancements in OOD learning applications, detection techniques, and the integration of foundation models. The discussion extends to various bioinformatics sub-disciplines, including drug discovery, single cell genomics, and polygenic risk score analysis, underscoring how OOD learning has facilitated notable breakthroughs in these areas. Through detailed examination of different model architectures and methods designed to address distribution shifts, we explore the potential of OOD learning to overcome the inherent limitations of standard machine learning models in bioinformatics. This review paper can be served as a valuable resource for bioinformatics researchers, offering a detailed exploration of OOD learning's transformative impact on understanding complex genomic data and its implications for human health. Pingzhao Hu |
Briefings Bioinform. | 3 |
| 2024 | iNGNN-DTI: prediction of drug-target interaction with interpretable nested graph neural network and pretrained molecule modelsabstractMOTIVATION: Drug-target interaction (DTI) prediction aims to identify interactions between drugs and protein targets. Deep learning can automatically learn discriminative features from drug and protein target representations for DTI prediction, but challenges remain, making it an open question. Existing approaches encode drugs and targets into features using deep learning models, but they often lack explanations for underlying interactions. Moreover, limited labeled DTIs in the chemical space can hinder model generalization. RESULTS: We propose an interpretable nested graph neural network for DTI prediction (iNGNN-DTI) using pre-trained molecule and protein models. The analysis is conducted on graph data representing drugs and targets by using a specific type of nested graph neural network, in which the target graphs are created based on 3D structures using Alphafold2. This architecture is highly expressive in capturing substructures of the graph data. We use a cross-attention module to capture interaction information between the substructures of drugs and targets. To improve feature representations, we integrate features learned by models that are pre-trained on large unlabeled small molecule and protein datasets, respectively. We evaluate our model on three benchmark datasets, and it shows a consistent improvement on all baseline models in all datasets. We also run an experiment with previously unseen drugs or targets in the test set, and our model outperforms all of the baselines. Furthermore, the iNGNN-DTI can provide more insights into the interaction by visualizing the weights learned by the cross-attention module. AVAILABILITY AND IMPLEMENTATION: The source code of the algorithm is available at https://github.com/syan1992/iNGNN-DTI. Yan Yi Li, Carson K. Leung, Pingzhao Hu |
Bioinform. | 4 |
| 2024 | Computational frameworks integrating deep learning and statistical models in mining multimodal omics dataabstractBACKGROUND: In health research, multimodal omics data analysis is widely used to address important clinical and biological questions. Traditional statistical methods rely on the strong assumptions of distribution. Statistical methods such as testing and differential expression are commonly used in omics analysis. Deep learning, on the other hand, is an advanced computer science technique that is powerful in mining high-dimensional omics data for prediction tasks. Recently, integrative frameworks or methods have been developed for omics studies that combine statistical models and deep learning algorithms. METHODS AND RESULTS: The aim of these integrative frameworks is to combine the strengths of both statistical methods and deep learning algorithms to improve prediction accuracy while also providing interpretability and explainability. This review report discusses the current state-of-the-art integrative frameworks, their limitations, and potential future directions in survival and time-to-event longitudinal analysis, dimension reduction and clustering, regression and classification, feature selection, and causal and transfer learning. Leann Lac, Carson K. Leung, Pingzhao Hu |
J. Biomed. Informatics | 3 |
| 2024 | Conditional probabilistic diffusion model driven synthetic radiogenomic applications in breast cancerabstractThis study addresses the heterogeneity of Breast Cancer (BC) by employing a Conditional Probabilistic Diffusion Model (CPDM) to synthesize Magnetic Resonance Images (MRIs) based on multi-omic data, including gene expression, copy number variation, and DNA methylation. The lack of paired medical images and genomics data in previous studies presented a challenge, which the CPDM aims to overcome. The well-trained CPDM successfully generated synthetic MRIs for 726 TCGA-BRCA patients, who lacked actual MRIs, using their multi-omic profiles. Evaluation metrics such as Frechet's Inception Distance (FID), Mean Square Error (MSE), and Structural Similarity Index Measure (SSIM) demonstrated the CPDM's effectiveness, with an FID of 2.02, an MSE of 0.02, and an SSIM of 0.59 based on the 15-fold cross-validation. The synthetic MRIs were used to predict clinical attributes, achieving an Area Under the Receiver-Operating-Characteristic curve (AUROC) of 0.82 and an Area Under the Precision-Recall Curve (AUPRC) of 0.84 for predicting ER+/HER2+ subtypes. Additionally, the MRIs served to accurately predicted BC patient survival with a Concordance-index (C-index) score of 0.88, outperforming other baseline models. This research demonstrates the potential of CPDMs in generating MRIs based on BC patients' genomic profiles, offering valuable insights for radiogenomic research and advancements in precision medicine. The study provides a novel approach to understanding BC heterogeneity for early detection and personalized treatment. Lianghong Chen, Zi Huai Huang, Michael Domaratzki, Qian Liu 0015, Pingzhao Hu |
PLoS Comput. Biol. | 6 |
| 2024 | ST-CellSeg: Cell segmentation for imaging-based spatial transcriptomics using multi-scale manifold learningabstractSpatial transcriptomics has gained popularity over the past decade due to its ability to evaluate transcriptome data while preserving spatial information. Cell segmentation is a crucial step in spatial transcriptomic analysis, as it enables the avoidance of unpredictable tissue disentanglement steps. Although high-quality cell segmentation algorithms can aid in the extraction of valuable data, traditional methods are frequently non-spatial, do not account for spatial information efficiently, and perform poorly when confronted with the problem of spatial transcriptome cell segmentation with varying shapes. In this study, we propose ST-CellSeg, an image-based machine learning method for spatial transcriptomics that uses manifold for cell segmentation and is novel in its consideration of multi-scale information. We first construct a fully connected graph which acts as a spatial transcriptomic manifold. Using multi-scale data, we then determine the low-dimensional spatial probability distribution representation for cell segmentation. Using the adjusted Rand index (ARI), normalized mutual information (NMI), and Silhouette coefficient (SC) as model performance measures, the proposed algorithm significantly outperforms baseline models in selected datasets and is efficient in computational complexity. Youcheng Li, Leann Lac, Qian Liu 0015, Pingzhao Hu |
PLoS Comput. Biol. | 4 |
| 2022 | Molecular Property Prediction based on Bimodal Supervised Contrastive LearningabstractThe simplified molecular-input line-entry system (SMILES) and the molecular graph are commonly used in chem-informatics to represent a molecule. Transformers are widely used for encoding SMILES to learn the relationship between elements that are far away from each other, while Graph Convolutional Networks (GCNs) are popular in graph representation learning and mostly focus on local structures. Since different information can be extracted from the SMILES string and the molecular graph, their integration might benefit the molecular property prediction task. In this work, we propose a bimodal supervised contrastive learning (BSCL) framework to integrate the SMILES string and the molecular graph in a unified network. Furthermore, the vanilla supervised contrastive loss (SCL) is not suitable for regression tasks, hence we design a weighted SCL to solve the problem. Six publicly available molecular property datasets are used to evaluate the proposed BSCL method, and our results show that the proposed bimodal method is superior to using the SMILES string or the molecular graph alone. Our code is released at https://github.com syanl992/BSCL. Md. Mohaiminul Islam, Ehsan Zahedi, Mélaine Kuenemann, Hassan Chouaib, Pingzhao Hu |
BIBM | 6 |
| 2022 | Tightly integrated multiomics-based deep tensor survival model for time-to-event predictionabstractMOTIVATION: Multiomics cancer profiles provide essential signals for predicting cancer survival. It is challenging to reveal the complex patterns from multiple types of data and link them to survival outcomes. We aim to develop a new deep learning-based algorithm to integrate three types of high-dimensional omics data measured on the same individuals to improve cancer survival outcome prediction. RESULTS: We built a three-dimension tensor to integrate multi-omics cancer data and factorized it into two-dimension matrices of latent factors, which were fed into neural networks-based survival networks. The new algorithm and other multi-omics-based algorithms, as well as individual genomic-based survival analysis algorithms, were applied to the breast cancer data colon and rectal cancer data from The Cancer Genome Atlas (TCGA) program. We evaluated the goodness-of-fit using the concordance index (C-index) and Integrated Brier Score (IBS). We demonstrated that the proposed tight integration framework has better survival prediction performance than the models using individual genomic data and other conventional data integration methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/jasperzyzhang/DeepTensorSurvival. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jasper Zhongyuan Zhang, Pingzhao Hu |
Bioinform. | 3 |
| 2022 | Deep clustering of small molecules at large-scale via variational autoencoder embedding and K-meansabstractBACKGROUND: Converting molecules into computer-interpretable features with rich molecular information is a core problem of data-driven machine learning applications in chemical and drug-related tasks. Generally speaking, there are global and local features to represent a given molecule. As most algorithms have been developed based on one type of feature, a remaining bottleneck is to combine both feature sets for advanced molecule-based machine learning analysis. Here, we explored a novel analytical framework to make embeddings of the molecular features and apply them in the clustering of a large number of small molecules. RESULTS: In this novel framework, we first introduced a principal component analysis method encoding the molecule-specific atom and bond information. We then used a variational autoencoder (AE)-based method to make embeddings of the global chemical properties and the local atom and bond features. Next, using the embeddings from the encoded local and global features, we implemented and compared several unsupervised clustering algorithms to group the molecule-specific embeddings. The number of clusters was treated as a hyper-parameter and determined by the Silhouette method. Finally, we evaluated the corresponding results using three internal indices. Applying the analysis framework to a large chemical library of more than 47,000 molecules, we successfully identified 50 molecular clusters using the K-means method with 32 embeddings based on the AE method. We visualized the clustering result via t-SNE for the overall distribution of molecules and the similarity maps for the structural analysis of randomly selected cluster-specific molecules. CONCLUSIONS: This study developed a novel analytical framework that comprises a feature engineering scheme for molecule-specific atomic and bonding features and a deep learning-based embedding strategy for different molecular features. By applying the identified embeddings, we show their usefulness for clustering a large molecule dataset. Our novel analytic algorithms can be applied to any virtual library of chemical compounds with diverse molecular structures. Hence, these tools have the potential of optimizing drug discovery, as they can decrease the number of compounds to be screened in any drug screening campaign. Hamid Hadipour, Chengyou Liu, Rebecca L. Davis, Silvia T. Cardona, Pingzhao Hu |
BMC Bioinform. | 5 |
| 2022 | Semi-supervised COVID-19 CT image segmentation using deep generative modelsabstractBACKGROUND: A recurring problem in image segmentation is a lack of labelled data. This problem is especially acute in the segmentation of lung computed tomography (CT) of patients with Coronavirus Disease 2019 (COVID-19). The reason for this is simple: the disease has not been prevalent long enough to generate a great number of labels. Semi-supervised learning promises a way to learn from data that is unlabelled and has seen tremendous advancements in recent years. However, due to the complexity of its label space, those advancements cannot be applied to image segmentation. That being said, it is this same complexity that makes it extremely expensive to obtain pixel-level labels, making semi-supervised learning all the more appealing. This study seeks to bridge this gap by proposing a novel model that utilizes the image segmentation abilities of deep convolution networks and the semi-supervised learning abilities of generative models for chest CT images of patients with the COVID-19. RESULTS: We propose a novel generative model called the shared variational autoencoder (SVAE). The SVAE utilizes a five-layer deep hierarchy of latent variables and deep convolutional mappings between them, resulting in a generative model that is well suited for lung CT images. Then, we add a novel component to the final layer of the SVAE which forces the model to reconstruct the input image using a segmentation that must match the ground truth segmentation whenever it is present. We name this final model StitchNet. CONCLUSION: We compare StitchNet to other image segmentation models on a high-quality dataset of CT images from COVID-19 patients. We show that our model has comparable performance to the other segmentation models. We also explore the potential limitations and advantages in our proposed algorithm and propose some potential future research directions for this challenging issue. Judah Zammit, Daryl L. X. Fung, Qian Liu 0015, Carson K. Leung, Pingzhao Hu |
BMC Bioinform. | 5 |
| 2022 | Bayesian tensor factorization-drive breast cancer subtyping by integrating multi-omics data
Qian Liu 0015, Bowen Cheng, Yongwon Jin, Pingzhao Hu |
J. Biomed. Informatics | 4 |
| 2022 | A machine learning model trained on a high-throughput antibacterial screen increases the hit rate of drug discoveryabstractScreening for novel antibacterial compounds in small molecule libraries has a low success rate. We applied machine learning (ML)-based virtual screening for antibacterial activity and evaluated its predictive power by experimental validation. We first binarized 29,537 compounds according to their growth inhibitory activity (hit rate 0.87%) against the antibiotic-resistant bacterium Burkholderia cenocepacia and described their molecular features with a directed-message passing neural network (D-MPNN). Then, we used the data to train an ML model that achieved a receiver operating characteristic (ROC) score of 0.823 on the test set. Finally, we predicted antibacterial activity in virtual libraries corresponding to 1,614 compounds from the Food and Drug Administration (FDA)-approved list and 224,205 natural products. Hit rates of 26% and 12%, respectively, were obtained when we tested the top-ranked predicted compounds for growth inhibitory activity against B. cenocepacia, which represents at least a 14-fold increase from the previous hit rate. In addition, more than 51% of the predicted antibacterial natural compounds inhibited ESKAPE pathogens showing that predictions expand beyond the organism-specific dataset to a broad range of bacteria. Overall, the developed ML approach can be used for compound prioritization before screening, increasing the typical hit rate of drug discovery. A. S. M. Zisanur Rahman, Chengyou Liu, Hunter Sturm, Andrew M. Hogan, Rebecca L. Davis, Pingzhao Hu, Silvia T. Cardona |
PLoS Comput. Biol. | 6 |
| 2020 | DTF: Deep Tensor Factorization for predicting anticancer drug synergyabstractMOTIVATION: Combination therapies have been widely used to treat cancers. However, it is cost and time consuming to experimentally screen synergistic drug pairs due to the enormous number of possible drug combinations. Thus, computational methods have become an important way to predict and prioritize synergistic drug pairs. RESULTS: We proposed a Deep Tensor Factorization (DTF) model, which integrated a tensor factorization method and a deep neural network (DNN), to predict drug synergy. The former extracts latent features from drug synergy information while the latter constructs a binary classifier to predict the drug synergy status. Compared to the tensor-based method, the DTF model performed better in predicting drug synergy. The area under precision-recall curve (PR AUC) was 0.58 for DTF and 0.24 for the tensor method. We also compared the DTF model with DeepSynergy and logistic regression models, and found that the DTF outperformed the logistic regression model and achieved similar performance as DeepSynergy using several performance metrics for classification task. Applying the DTF model to predict missing entries in our drug-cell-line tensor, we identified novel synergistic drug combinations for 10 cell lines from the 5 cancer types. A literature survey showed that some of these predicted drug synergies have been identified in vivo or in vitro. Thus, the DTF model could be a valuable in silico tool for prioritizing novel synergistic drug combinations. AVAILABILITY AND IMPLEMENTATION: Source code and data are available at https://github.com/ZexuanSun/DTF-Drug-Synergy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zexuan Sun, Shujun Huang, Peiran Jiang, Pingzhao Hu, Zhiyong Lu |
Bioinform. | 4 |
| 2018 | An integrative network-based approach to identify novel disease genes and pathways: a case study in the context of inflammatory bowel diseaseabstractBACKGROUND: There are different and complicated associations between genes and diseases. Finding the causal associations between genes and specific diseases is still challenging. In this work we present a method to predict novel associations of genes and pathways with inflammatory bowel disease (IBD) by integrating information of differential gene expression, protein-protein interaction and known disease genes related to IBD. RESULTS: We downloaded IBD gene expression data from NCBI's Gene Expression Omnibus, performed statistical analysis to determine differentially expressed genes, collected known IBD genes from DisGeNet database, which were used to construct a IBD related PPI network with HIPPIE database. We adapted our graph-based clustering algorithm DPClusO to cluster the disease PPI network. We evaluated the statistical significance of the identified clusters in the context of determining the richness of IBD genes using Fisher's exact test and predicted novel genes related to IBD. We showed 93.8% of our predictions are correct in the context of other databases and published literatures related to IBD. CONCLUSIONS: Finding disease-causing genes is necessary for developing drugs with synergistic effect targeting many genes simultaneously. Here we present an approach to identify novel disease genes and pathways and discuss our approach in the context of IBD. The approach can be generalized to find disease-associated genes for other diseases. Ryohei Eguchi, Mohammad Bozlul Karim, Pingzhao Hu, Tetsuo Sato, Naoaki Ono, Shigehiko Kanaya, Md. Altaf-Ul-Amin |
BMC Bioinform. | 3 |
| 2015 | Discriminative learning of generative models: large margin multinomial mixture models for document classification
Hui Jiang 0001, Zhen-Yu Pan, Pingzhao Hu |
Pattern Anal. Appl. | 3 |
| 2012 | Gene network modular-based classification of microarray samplesabstractBACKGROUND: Molecular predictor is a new tool for disease diagnosis, which uses gene expression to classify diagnostic category of a patient. The statistical challenge for constructing such a predictor is that there are thousands of genes to predict for the disease categories, but only a small number of samples are available. RESULTS: We proposed a gene network modular-based linear discriminant analysis approach by integrating 'essential' correlation structure among genes into the predictor in order that the modules or cluster structures of genes, which are related to the diagnostic classes we look for, can have potential biological interpretation. We evaluated performance of the new method with other established classification methods using three real data sets. CONCLUSIONS: Our results show that the new approach has the advantage of computational simplicity and efficiency with relatively lower classification error rates than the compared methods in many cases. The modular-based linear discriminant analysis approach induced in the study has the potential to increase the power of discriminant analysis for which sample sizes are small and there are large number of genes in the microarray studies. Pingzhao Hu, Shelley B. Bull, Hui Jiang 0001 |
BMC Bioinform. | 1 |
| 2011 | Gene Network Modules-Based Liner Discriminant Analysis of Microarray Gene Expression Data
Pingzhao Hu, Shelley B. Bull, Hui Jiang 0001 |
ISBRA | 1 |
| 2010 | Predicting protein functions by relaxation labelling protein interaction networkabstractBACKGROUND: One of key issues in the post-genomic era is to assign functions to uncharacterized proteins. Since proteins seldom act alone; rather, they must interact with other biomolecular units to execute their functions. Thus, the functions of unknown proteins may be discovered through studying their interactions with proteins having known functions. Although many approaches have been developed for this purpose, one of main limitations in most of these methods is that the dependence among functional terms has not been taken into account. RESULTS: We developed a new network-based protein function prediction method which combines the likelihood scores of local classifiers with a relaxation labelling technique. The framework can incorporate the inter-relationship among functional labels into the function prediction procedure and allow us to efficiently discover relevant non-local dependence. We evaluated the performance of the new method with one other representative network-based function prediction method using E. coli protein functional association networks. CONCLUSION: Our results showed that the new method has better prediction performance than the previous method. The better predictive power of our method gives new insights about the importance of the dependence between functional terms in protein functional prediction. Pingzhao Hu, Hui Jiang 0001, Andrew Emili |
BMC Bioinform. | 1 |
| 2006 | Integrating Affymetrix microarray data sets using probe-level test statistic for predicting prostate cancerabstractMicroarray technology has previously been used to identify differentially expressed genes between tumor and normal prostate samples in a single study as well as in a synthesis involving multiple studies. When integrating results from several Affymetrix microarray datasets, previous studies have used probeset-level data which may lead to a loss of information contained at the probe-level. Here, we propose a new approach for combining results across studies, based on a probe-level test statistic. Each probe-level test statistic is transformed into an effect size measure for each probeset and a random-effects model (REM) is used to integrate effect sizes across studies. We compared statistical and biological significance of the prognostic gene expression signatures identified in the probe-level model (PLM) with those in the probeset-level model (PSLM). Support vector machines (SVMs)-based predictive models were built using these two sets of signatures and their performances were evaluated using independent test datasets. Our analyses show that the prognostic gene expression signatures identified through the probe-level test statistics are more strongly differentially expressed and have better prediction accuracy than signatures derived from a probeset-level model. Pingzhao Hu, Celia M. T. Greenwood, Joseph Beyene |
CIBCB | 1 |
| 2005 | Integrative analysis of multiple gene expression profiles with quality-adjusted effect size modelsabstractBACKGROUND: With the explosion of microarray studies, an enormous amount of data is being produced. Systematic integration of gene expression data from different sources increases statistical power of detecting differentially expressed genes and allows assessment of heterogeneity. The challenge, however, is in designing and implementing efficient analytic methodologies for combination of data generated by different research groups. RESULTS: We extended traditional effect size models to combine information from different microarray datasets by incorporating a quality measure for each gene in each study into the effect size estimation. We illustrated our method by integrating two datasets generated using different Affymetrix oligonucleotide types. Our results indicate that the proposed quality-adjusted weighting strategy for modelling inter-study variation of gene expression profiles not only increases consistency and decreases heterogeneous results between these two datasets, but also identifies many more differentially expressed genes than methods proposed previously. CONCLUSION: Data integration and synthesis is becoming increasingly important. We live in a high-throughput era where technologies constantly change leaving behind a trail of data with different forms, shapes and sizes. Statistical and computational methodologies are therefore critical for extracting the most out of these related but not identical sources of data. Pingzhao Hu, Celia M. T. Greenwood, Joseph Beyene |
BMC Bioinform. | 1 |