Xiaobo Zhou 0005

dblp:13/6395-5 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0001-7191-6495ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 AllergenAI: a deep learning model predicting allergenicity based on protein sequence
abstract
BACKGROUND: Innovations in protein engineering offer promising solutions for redesigning allergenic proteins to minimize adverse reactions in sensitive individuals. Earlier models for predicting allergenicity have relied on the knowledge of physicochemical properties and sequence homology to assess the potential risk. However, to better understand the allergenic proteins’ sequence features, we need a novel sequence-based deep learning model for predicting allergenicity. RESULTS: We present a novel AI-based tool, AllergenAI, to quantify the allergenic potential of a protein’s sequence without using any other known features. Our study utilized allergenic protein sequence data archived in the three well-established databases, SDAP 2.0, COMPARE, and AlgPred 2, to train a convolutional neural network and assessed its prediction performance by cross-validation. We then used AllergenAI to find novel potential proteins of the cupin family in date palm, spinach, maize, and red clover plants with a high allergenicity score that might have an adverse allergenic effect on sensitive individuals. By analyzing the feature importance scores (FIS) of vicilins, we identified a proline-alanine-rich (P-A) motif in the top 50% of FIS regions that overlapped with known IgE epitope regions of vicilin allergens. We then used the approximately 1600 allergen structures in our SDAP database, in a pilot study to show the potential to incorporate 3D information in a CNN model. The prediction quality was slightly increased. CONCLUSION: Our allergenicity prediction study through the development of AllergenAI provides a foundation for identifying the critical features that distinguish allergenic proteins.
Surendra S. Negi, Chengyuan Yang, Xiaobo Zhou 0005, Catherine H. Schein, Werner Braun, Pora Kim
BMC Bioinform.4
2025 Cross-dataset EEG emotion recognition based on pre-trained Vision Transformer considering emotional sensitivity diversity
Fang Wang 0027, Yu-Chu Tian, Xiaobo Zhou 0005
Expert Syst. Appl.3
2024 Revealing chronic disease progression patterns using Gaussian process for stage inference
abstract
OBJECTIVE: The early stages of chronic disease typically progress slowly, so symptoms are usually only noticed until the disease is advanced. Slow progression and heterogeneous manifestations make it challenging to model the transition from normal to disease status. As patient conditions are only observed at discrete timestamps with varying intervals, an incomplete understanding of disease progression and heterogeneity affects clinical practice and drug development. MATERIALS AND METHODS: We developed the Gaussian Process for Stage Inference (GPSI) approach to uncover chronic disease progression patterns and assess the dynamic contribution of clinical features. We tested the ability of the GPSI to reliably stratify synthetic and real-world data for osteoarthritis (OA) in the Osteoarthritis Initiative (OAI), bipolar disorder (BP) in the Adolescent Brain Cognitive Development Study (ABCD), and hepatocellular carcinoma (HCC) in the UTHealth and The Cancer Genome Atlas (TCGA). RESULTS: First, GPSI identified two subgroups of OA based on image features, where these subgroups corresponded to different genotypes, indicating the bone-remodeling and overweight-related pathways. Second, GPSI differentiated BP into two distinct developmental patterns and defined the contribution of specific brain region atrophy from early to advanced disease stages, demonstrating the ability of the GPSI to identify diagnostic subgroups. Third, HCC progression patterns were well reproduced in the two independent UTHealth and TCGA datasets. CONCLUSION: Our study demonstrated that an unsupervised approach can disentangle temporal and phenotypic heterogeneity and identify population subgroups with common patterns of disease progression. Based on the differences in these features across stages, physicians can better tailor treatment plans and medications to individual patients.
Weiling Zhao, Angela Ross, Lei You 0005, Xiaobo Zhou 0005
J. Am. Medical Informatics Assoc.6
2023 Integrated mRNA sequence optimization using deep learning
abstract
The coronavirus disease of 2019 pandemic has catalyzed the rapid development of mRNA vaccines, whereas, how to optimize the mRNA sequence of exogenous gene such as severe acute respiratory syndrome coronavirus 2 spike to fit human cells remains a critical challenge. A new algorithm, iDRO (integrated deep-learning-based mRNA optimization), is developed to optimize multiple components of mRNA sequences based on given amino acid sequences of target protein. Considering the biological constraints, we divided iDRO into two steps: open reading frame (ORF) optimization and 5' untranslated region (UTR) and 3'UTR generation. In ORF optimization, BiLSTM-CRF (bidirectional long-short-term memory with conditional random field) is employed to determine the codon for each amino acid. In UTR generation, RNA-Bart (bidirectional auto-regressive transformer) is proposed to output the corresponding UTR. The results show that the optimized sequences of exogenous genes acquired the pattern of human endogenous gene sequence. In experimental validation, the mRNA sequence optimized by our method, compared with conventional method, shows higher protein expression. To the best of our knowledge, this is the first study by introducing deep-learning methods to integrated mRNA sequence optimization, and these results may contribute to the development of mRNA therapeutics.
Haoran Gong 0002, Jianguo Wen, Ruihan Luo, Yuzhou Feng, Hongguang Fu, Xiaobo Zhou 0005
Briefings Bioinform.7
2023 Genetic control of RNA editing in neurodegenerative disease
abstract
A-to-I RNA editing diversifies human transcriptome to confer its functional effects on the downstream genes or regulations, potentially involving in neurodegenerative pathogenesis. Its variabilities are attributed to multiple regulators, including the key factor of genetic variants. To comprehensively investigate the potentials of neurodegenerative disease-susceptibility variants from the view of A-to-I RNA editing, we analyzed matched genetic and transcriptomic data of 1596 samples across nine brain tissues and whole blood from two large consortiums, Accelerating Medicines Partnership-Alzheimer's Disease and Parkinson's Progression Markers Initiative. The large-scale and genome-wide identification of 95 198 RNA editing quantitative trait loci revealed the preferred genetic effects on adjacent editing events. Furthermore, to explore the underlying mechanisms of the genetic controls of A-to-I RNA editing, several top RNA-binding proteins were pointed out, such as EIF4A3, U2AF2, NOP58, FBL, NOP56 and DHX9, since their regulations on multiple RNA-editing events were probably interfered by these genetic variants. Moreover, these variants may also contribute to the variability of other molecular phenotypes associated with RNA editing, including the functions of 3 proteins, expressions of 277 genes and splicing of 449 events. All the analyses results shown in NeuroEdQTL (https://relab.xidian.edu.cn/NeuroEdQTL/) constituted a unique resource for the understanding of neurodegenerative pathogenesis from genotypes to phenotypes related to A-to-I RNA editing.
Sijia Wu, Qiuping Xue, Mengyuan Yang 0001, Pora Kim, Xiaobo Zhou 0005
Briefings Bioinform.6
2023 Real-time prediction of organ failures in patients with acute pancreatitis using longitudinal irregular data
Jiawei Luo 0002, Lan Lan 0003, Shixin Huang, Xiaoxi Zeng, Qu Xiang, Mengjiao Li, Weiling Zhao, Xiaobo Zhou 0005
J. Biomed. Informatics9
2023 Inferring evolutionary trajectories from cross-sectional transcriptomic data to mirror lung adenocarcinoma progression
abstract
Lung adenocarcinoma (LUAD) is a deadly tumor with dynamic evolutionary process. Although much endeavors have been made in identifying the temporal patterns of cancer progression, it remains challenging to infer and interpret the molecular alterations associated with cancer development and progression. To this end, we developed a computational approach to infer the progression trajectory based on cross-sectional transcriptomic data. Analysis of the LUAD data using our approach revealed a linear trajectory with three different branches for malignant progression, and the results showed consistency in three independent cohorts. We used the progression model to elucidate the potential molecular events in LUAD progression. Further analysis showed that overexpression of BUB1B, BUB1 and BUB3 promoted tumor cell proliferation and metastases by disturbing the spindle assembly checkpoint (SAC) in the mitosis. Aberrant mitotic spindle checkpoint signaling appeared to be one of the key factors promoting LUAD progression. We found the inferred cancer trajectory allows to identify LUAD susceptibility genetic variations using genome-wide association analysis. This result shows the opportunity for combining analysis of candidate genetic factors with disease progression. Furthermore, the trajectory showed clear evident mutation accumulation and clonal expansion along with the LUAD progression. Understanding how tumors evolve and identifying mutated genes will help guide cancer management. We investigated the clonal architectures and identified distinct clones and subclones in different LUAD branches. Validation of the model in multiple independent data sets and correlation analysis with clinical results demonstrate that our method is effective and unbiased.
Haoran Gong 0002, Zhengzheng Qiao, Tiangang Wang, Weiling Zhao, Xiaobo Zhou 0005
PLoS Comput. Biol.8
2022 A novel sagittal craniosynostosis classification system based on multi-view learning algorithm
Lei You 0005, Griffin Patrick Bins, Christopher Michael Runyan, Lisa David, Xiaobo Zhou 0005
Neural Comput. Appl.8
2021 miRactDB characterizes miRNA-gene relation switch between normal and cancer tissues across pan-cancer
abstract
It has been increasingly accepted that microRNA (miRNA) can both activate and suppress gene expression, directly or indirectly, under particular circumstances. Yet, a systematic study on the switch in their interaction pattern between activation and suppression and between normal and cancer conditions based on multi-omics evidences is not available. We built miRactDB, a database for miRNA-gene interaction, at https://ccsm.uth.edu/miRactDB, to provide a versatile resource and platform for annotation and interpretation of miRNA-gene relations. We conducted a comprehensive investigation on miRNA-gene interactions and their biological implications across tissue types in both tumour and normal conditions, based on TCGA, CCLE and GTEx databases. We particularly explored the genetic and epigenetic mechanisms potentially contributing to the positive correlation, including identification of miRNA binding sites in the gene coding sequence (CDS) and promoter regions of partner genes. Integrative analysis based on this resource revealed that top-ranked genes derived from TCGA tumour and adjacent normal samples share an overwhelming part of biological processes, which are quite different than those from CCLE and GTEx. The most active miRNAs predicted to target CDS and promoter regions are largely overlapped. These findings corroborate that adjacent normal tissues might have undergone significant molecular transformations towards oncogenesis before phenotypic and histological change; and there probably exists a small yet critical set of miRNAs that profoundly influence various cancer hallmark processes. miRactDB provides a unique resource for the cancer and genomics communities to screen, prioritize and rationalize their candidates of miRNA-gene interactions, in both normal and cancer scenarios.
Hua Tan, Pora Kim, Peiqing Sun, Xiaobo Zhou 0005
Briefings Bioinform.4
2021 ADeditome provides the genomic landscape of A-to-I RNA editing in Alzheimer's disease
abstract
A-to-I RNA editing, contributing to nearly 90% of all editing events in human, has been reported to involve in the pathogenesis of Alzheimer's disease (AD) due to its roles in brain development and immune regulation, such as the deficient editing of GluA2 Q/R related to cell death and memory loss. Currently, there are urgent needs for the systematic annotations of A-to-I RNA editing events in AD. Here, we built ADeditome, the annotation database of A-to-I RNA editing in AD available at https://ccsm.uth.edu/ADeditome, aiming to provide a resource and reference for functional annotation of A-to-I RNA editing in AD to identify therapeutically targetable genes in an individual. We detected 1676 363 editing sites in 1524 samples across nine brain regions from ROSMAP, MayoRNAseq and MSBB. For these editing events, we performed multiple functional annotations including identification of specific and disease stage associated editing events and the influence of editing events on gene expression, protein recoding, alternative splicing and miRNA regulation for all the genes, especially for AD-related genes in order to explore the pathology of AD. Combing all the analysis results, we found 108 010 and 26 168 editing events which may promote or inhibit AD progression, respectively. We also found 5582 brain region-specific editing events with potentially dual roles in AD across different brain regions. ADeditome will be a unique resource for AD and drug research communities to identify therapeutically targetable editing events. Significance: ADeditome is the first comprehensive resource of the functional genomics of individual A-to-I RNA editing events in AD, which will be useful for many researchers in the fields of AD pathology, precision medicine, and therapeutic researches.
Sijia Wu, Mengyuan Yang 0001, Pora Kim, Xiaobo Zhou 0005
Briefings Bioinform.4
2021 ExonSkipAD provides the functional genomic landscape of exon skipping events in Alzheimer's disease
abstract
Exon skipping (ES), the most common alternative splicing event, has been reported to contribute to diverse human diseases due to the loss of functional domains/sites or frameshifting of the open reading frame (ORF) and noticed as therapeutic targets. Accumulating transcriptomic studies of aging brains show the splicing disruption is a widespread hallmark of neurodegenerative diseases such as Alzheimer's disease (AD). Here, we built ExonSkipAD, the ES annotation database aiming to provide a resource/reference for functional annotation of ES events in AD and identify therapeutic targets in exon units. We identified 16 414 genes that have ~156 K, ~ 69 K, ~ 231 K ES events from the three representative AD cohorts of ROSMAP, MSBB and Mayo, respectively. For these ES events, we performed multiple functional annotations relating to ES mechanisms or downstream. Specifically, through the functional feature retention studies followed by the open reading frames (ORFs), we identified 275 important cellular regulators that might lose their cellular regulator roles due to exon skipping in AD. ExonSkipAD provides twelve categories of annotations: gene summary, gene structures and expression levels, exon skipping events with PSIs, ORF annotation, exon skipping events in the canonical protein sequence, 3'-UTR located exon skipping events lost miRNA-binding sites, SNversus in the skipped exons with a depth of coverage, AD stage-associated exon skipping events, splicing quantitative trait loci (sQTLs) in the skipped exons, correlation with RNA-binding proteins, and related drugs & diseases. ExonSkipAD will be a unique resource of transcriptomic diversity research for understanding the mechanisms of neurodegenerative disease development and identifying potential therapeutic targets in AD. Significance AS the first comprehensive resource of the functional genomics of the alternative splicing events in AD, ExonSkipAD will be useful for many researchers in the fields of pathology, AD genomics and precision medicine, and pharmaceutical and therapeutic researches.
Mengyuan Yang 0001, Ke Yiya, Pora Kim, Xiaobo Zhou 0005
Briefings Bioinform.4
2020 PredictFP2: A New Computational Model to Predict Fusion Peptide Domain in All Retroviruses
abstract
Fusion peptide (FP) is a pivotal domain for the entry of retrovirus into host cells to continue self-replication. The crucial role indicates that FP is a promising drug target for therapeutic intervention. A FP model proposed in our previous work is relatively not efficient to predict FP in retroviruses. Thus in this work, we come up with a new computational model to predict FP domains in all the retroviruses. It basically predicts FP domains through recognizing their start and end sites separately with SVM method combing the hydrophobicity knowledge of the subdomain around furin cleavage site. The classification accuracy rates are 91.91, 91.20 and 89.13 percent respectively corresponding to jack-knife, 10-fold cross-validation and 5-fold cross-validation test. Second, this model discovered 69,753 and 493 putative FPs after scanning amino acid sequences and HERV DNA sequences both without FP annotations. Subsequently, a statistical analysis was performed on the 69,753 putative FP sequences, which confirms that FP is a hydrophobic domain. Lastly, we depicted the distribution of the 493 putative FP sequences on each human chromosome and each HERV family, which shows that FP of HERV probably has chromosome and family preference.
Sijia Wu, Jie Tian 0001, Xiaobo Zhou 0005
IEEE ACM Trans. Comput. Biol. Bioinform.4
2019 Systematically understanding the immunity leading to CRPC progression
abstract
Prostate cancer (PCa) is the most commonly diagnosed malignancy and the second leading cause of cancer-related death in American men. Androgen deprivation therapy (ADT) has become a standard treatment strategy for advanced PCa. Although a majority of patients initially respond to ADT well, most of them will eventually develop castration-resistant PCa (CRPC). Previous studies suggest that ADT-induced changes in the immune microenvironment (mE) in PCa might be responsible for the failures of various therapies. However, the role of the immune system in CRPC development remains unclear. To systematically understand the immunity leading to CRPC progression and predict the optimal treatment strategy in silico, we developed a 3D Hybrid Multi-scale Model (HMSM), consisting of an ODE system and an agent-based model (ABM), to manipulate the tumor growth in a defined immune system. Based on our analysis, we revealed that the key factors (e.g. WNT5A, TRAIL, CSF1, etc.) mediated the activation of PC-Treg and PC-TAM interaction pathways, which induced the immunosuppression during CRPC progression. Our HMSM model also provided an optimal therapeutic strategy for improving the outcomes of PCa treatment.
Zhiwei Ji, Weiling Zhao, Hui-Kuan Lin, Xiaobo Zhou 0005
PLoS Comput. Biol.4