Shizheng Qiu

dblp:311/2098 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-0047-4199ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2025 MetImputBERT: a pretrained BERT framework for missing value imputation in NMR metabolomics data
abstract
Missing values in nuclear magnetic resonance metabolomics data compromise downstream clinical interpretation. Here, we present MetImputBERT, an imputation method based on a pretrained BERT framework. MetImputBERT uses the masks in the masked language model to simulate missing values and leverages predictions and reconstructions to these positions to simulate the imputation process. The learning of MetImputBERT is driven by minimizing the reconstruction error. MetImputBERT was pretrained on the largest metabolomics dataset to date, comprising data from over 230 000 individuals in the UK Biobank. When new datasets with missing values were encountered, MetImputBERT loaded the pretrained parameters and directly imputed the missing values by inferring their reconstructed estimates. MetImputBERT outperformed commonly used methods-K-nearest neighbors, multiple imputation by chained equations, and singular value decomposition-in imputation performance on two independent test sets. We provide an open-source Python tool that allows users to quickly impute missing values in their own NMR metabolomics data without any additional training.
Shizheng Qiu, Yang Hu 0008, Guiyou Liu, Yadong Wang 0001
Briefings Bioinform.1
2025 Proformer: a multimodal proteomics transformer model for multidisease early risk assessment
abstract
Early identification of individuals at high risk for chronic diseases is crucial for prevention and intervention, yet current risk assessment tools are disease-specific, require extensive clinical data collection, and cannot provide multidisease risk profiles from a single measurement. Several protein large language models have been developed for tasks such as protein structure prediction, function prediction, and sequence design. However, none of these models can be directly applied in clinical settings to predict an individual's future disease risk. Here, we present a multimodal proteomics Transformer (Proformer) model that integrates protein expression, sequence, and function information for multidisease risk assessment. We trained Proformer using real proteomics data from 47 124 individuals from the UK Biobank to evaluate its performance in discriminating the risk of 20 common chronic diseases. Proformer achieved state-of-the-art (SOTA) performance in all 20 diseases compared with five common machine learning and deep learning models. Compared to three common clinical predictors, Proformer's 10-year discriminative performance outperforms Age + Sex model for 19 diseases, outperforms the ASCVD risk score for 16 diseases, and outperforms the panel composed of 35 clinical variables for 11 diseases. These results were replicated in the Scotland and Wales cohort from UK Biobank. In conclusion, Proformer enabled users to directly obtain a 10-year risk report for common chronic diseases by inputting their individual proteomics data.
Shizheng Qiu, Yang Hu 0008, Yadong Wang 0001
Briefings Bioinform.1
2024 Gene ontology and conjunctional false discovery rate statistical framework revealed shared genetic mechanisms underlying glaucoma and myopia
abstract
Observational studies have demonstrated that refractive error (RE) is associated with an increased risk of primary open-angle glaucoma (POAG). However, the underlying mechanism through which RE influences the development of POAG remains unknown. In this study, we analyzed genome-wide association study (GWAS) data for RE (95,505 participants) and POAG (15,229 cases and 177,473 controls) to characterize polygenic architecture and identify genetic loci shared of these conditions. We integrated the conjunctional false discovery rate (FDR) statistical framework with Gene ontology to investigate the potential biological mechanisms underlying shared genetic loci. Our study revealed a substantial and previously underexplored shared genetic basis between these disorders. Genetic correlation analysis not only confirms overall genetic correlation between RE and POAG (r for genetic = -0.15, 95%CI: -0.21 to -0.09, P = 7.70E07) but also uncovers localized genetic associations, pointing to 12 genome regions of interest. MR analysis established a bidirectional causal relationship between RE and POAG. We identified a robust polygenic overlap between RE and POAG beyond what conventional genetic correlation analysis can reveal. Moreover, we found substantial polygenic overlap and identified 16 shared loci of RE and POAG at conjFDR < 0.01, of which four were novel. Functional annotation offered insights into potential biological mechanisms through which these genetic loci exert their influence, particularly within retinal tissues. Together, this study significantly advances our understanding of the genetic underpinnings that link RE and POAG, elucidating the complex genetic landscape that contributes to their co-occurrence.
Shizheng Qiu, Xuehui Zhang, Zhishuai Zhang
BIBM1
2023 Identification of risk genes and biological pathways influencing myopia via transcriptome association study and biomedical ontology methods
abstract
Refractive error remains one of the most common eye diseases in the world, which causes great burden on public health and economy every year. Although genome-wide association studies (GWAS) have identified susceptibility loci for refractive error, the credible causal genes associated with gene expression in brain are still unclear. In addition, the influence of RNA splicing events on refractive error in brain remains unknown. Here, we identified candidate refractive error gene targets and biological pathways. In stage 1, we carried out gene-based association analysis to identify refractive error risk genes. In stage 2, by using transcriptome-wide association studies (TWAS) and colocalization analysis, we identified cis-regulated genes in brain tissues that were positively related to refractive error from the whole gene region and the exon region after RNA splicing, respectively. In stage 3, we reveal biological pathways influencing refractive error using biological ontology and knowledge base methods. We identified 367 risk genes passed gene-based association test. A total of 43 genes in brain could be considered as causal genes using TWAS and colocalization analysis, and 27 were novel. After RNA splicing, RCBTB1 significantly affected refractive error. Risk genes for refractive error significantly enriched in biological processes related to synaptic transmission and acetylcholine. In conclusion, our analysis provides a better understanding of potential therapeutic targets and biological pathways for refractive error.
Shizheng Qiu, Yadong Wang 0001, Yang Hu 0008
BIBM1
2023 HNetGO: protein function prediction via heterogeneous network transformer
abstract
Protein function annotation is one of the most important research topics for revealing the essence of life at molecular level in the post-genome era. Current research shows that integrating multisource data can effectively improve the performance of protein function prediction models. However, the heavy reliance on complex feature engineering and model integration methods limits the development of existing methods. Besides, models based on deep learning only use labeled data in a certain dataset to extract sequence features, thus ignoring a large amount of existing unlabeled sequence data. Here, we propose an end-to-end protein function annotation model named HNetGO, which innovatively uses heterogeneous network to integrate protein sequence similarity and protein-protein interaction network information and combines the pretraining model to extract the semantic features of the protein sequence. In addition, we design an attention-based graph neural network model, which can effectively extract node-level features from heterogeneous networks and predict protein function by measuring the similarity between protein nodes and gene ontology term nodes. Comparative experiments on the human dataset show that HNetGO achieves state-of-the-art performance on cellular component and molecular function branches.
Xiaoshuai Zhang, Huannan Guo, Xuan Wang 0002, Kaitao Wu, Shizheng Qiu, Bo Liu 0023, Yadong Wang 0001, Yang Hu 0008, Junyi Li 0004
Briefings Bioinform.6
2022 Identification of the expression, prognostic value and cancer immunity of Gasdermin E based on multi-omics data, machine learning and gene ontology
abstract
Gasdermin E (GSDME)-mediated pyroptosis participates in the recruitment of tumor-infiltrating lymphocytes into the tumor microenvironment in primary breast cancers, melanoma and colorectal cancer. However, it is still unknown whether GSDME suppresses tumors in kidney cancer. Here, we comprehensively analyzed the mRNA expression, genetic alteration, prognosis value and cancer immunity of GSDME in three types of kidney cancer, including kidney renal clear cell carcinoma (KIRC), kidney renal papillary cell carcinoma (KIRP) and kidney chromophobe (KICH), based on multi-omics data, machine learning and gene ontology (GO). Compared with normal samples, GSDME expression was significantly down-regulated in KICH and up regulated in KIRP. The genetic alteration of GSDME changed significantly in the process of tumor development, and the results of the univariate cox regression model and log-rank test suggested GSDME as a potential prognostic marker for all the three kidney cancers. Immune infiltration and correlation analysis supported that GSDME-mediated pyroptosis played a critical role in kidney cancer and was highly associated with killer cells and macrophages in the tumor microenvironment. Most GSDME-related genes were enriched in membrane proteins and lysosomal, and enzyme activity-related pathways by GO enrichment, which was consistent with the pyroptosis process. Together, our work provided new insights and theoretical basis for GSDME as a drug target for the treatment of kidney cancer.
Shizheng Qiu, Yang Hu 0008
BIBM1
2022 MHCRoBERTa: pan-specific peptide-MHC class I binding prediction through transfer learning with label-agnostic protein sequences
abstract
Predicting the binding of peptide and major histocompatibility complex (MHC) plays a vital role in immunotherapy for cancer. The success of Alphafold of applying natural language processing (NLP) algorithms in protein secondary struction prediction has inspired us to explore the possibility of NLP methods in predicting peptide-MHC class I binding. Based on the above motivations, we propose the MHCRoBERTa method, RoBERTa pre-training approach, for predicting the binding affinity between type I MHC and peptides. Analysis of the results on benchmark dataset demonstrates that MHCRoBERTa can outperform other state-of-art prediction methods with an increase of the Spearman rank correlation coefficient (SRCC) value. Notably, our model gave a significant improvement on IC50 value. Our method has achieved SRCC value and AUC value as 0.785 and 0.817, respectively. Our SRCC value is 14.3% higher than NetMHCpan3.0 (the second highest SRCC value on pan-specific) and is 3% higher than MHCflurry (the second highest SRCC value on all methods). The AUC value is also better than any other pan-specific methods. Moreover, we visualize the multi-head self-attention for the token representation across the layers and heads by this method. Through the analysis of the representation of each layer and head, we can show whether the model has learned the syntax and semantics necessary to perform the prediction task well. All these results demonstrate that our model can accurately predict the peptide-MHC class I binding affinity and that MHCRoBERTa is a powerful tool for screening potential neoantigens for cancer immunotherapy. MHCRoBERTa is available as an open source software at github (https://github.com/FuxuWang/MHCRoBERTa).
Fuxu Wang, Haoyan Wang, Lizhuang Wang, Haoyu Lu, Shizheng Qiu, Tianyi Zang, Xinjun Zhang, Yang Hu 0008
Briefings Bioinform.5
2021 Differentially Expressed Mutant Genes Reveal Potential Prognostic Markers For Lung Adenocarcinoma
abstract
Lung adenocarcinoma is a serious lung cancer, belonging to the category of non-small cell lung cancer, accounting for 30 to 35 percent of the total, and lung adenocarcinoma is more common in women and non-smoking patients. In this study, LUAD samples of mutation genes and RNA-seq expression dataset were retrieved from TCGA database. WGCNA, Cox regression analysis were used to classify melanoma prognosis. It revealed that eleven mutant and differentially expressed prognosis biomarkers were significantly associated with the prognosis of patients. Therefore, detecting these gene mutations and exploring their corresponding expression could be valuable in predicting the prognosis of patients. The results of the high-throughput data mining provide important fundamental bioinformatics information and a relevant theoretical basis for further exploring the molecular pathogenesis of LUAD and assessing the prognosis of patients.
Yue Liu 0034, Shizheng Qiu, Yang Hu 0008, Yadong Wang 0001
BIBM2