Jian Zhao 0034

dblp:70/2932-34 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-7857-4766ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 THOR: a TMB heterogeneity-adaptive optimization model predicts immunotherapy response using clonal genomic features in group-structured data
abstract
With the increasing number of indications for immune checkpoint inhibitors in early and advanced cancers, the prospect of a tumor-agnostic biomarker to prioritize patients is compelling. Tumor mutation burden (TMB) is a widely endorsed biomarker that quantifies nonsynonymous mutations within tumor DNA, essential for neoantigen production, which, in turn, correlates with the immune response and guides decision-making. However, the general clinical application of TMB-relying on simple mutational counts targeted at a single endpoint-does not adequately capture the complex clonal structure of tumors nor the multifaceted nature of prognostic indicators. This recognition has spurred the exploration of sophisticated high-dimensional regression techniques. Unfortunately, the limited cohort sizes in immunotherapy trials have hindered the full potential of these advanced methods. Our approach considers patient subgroups as related yet distinct entities, enabling precise tailoring and refinement to address subgroup-specific dynamics. Given the deficiencies and the constraints, we introduce a TMB heterogeneity-optimized regression (THOR). This innovative model enhances the predictive capabilities of TMB by integrating tumor clonality and a diverse spectrum of clinical endpoints, further augmented by fusion techniques across subgroups to facilitate robust data sharing and interpretation. Our simulations validate THOR's superiority in parameter estimation for statistical inference. Clinically, we assess the utility of THOR in a structured cohort of 238 cancer patients undergoing immunotherapy, supplemented by 2212 patients across 19 subgroups from public datasets. The forecast of the responses and comparison of survival hazards demonstrate that THOR significantly enhances patient stratification and prognostic predictions by incorporating complex immunogenetic biology and subgroup-specific dynamics.
Yanfang Guan, Xin Lai 0003, Yuqian Liu, Zhili Chang, Quan Wang 0004, Jian Zhao 0034, Shuanying Yang, Jiayin Wang 0002
Briefings Bioinform.9
2024 DeepIRES: a hybrid deep learning model for accurate identification of internal ribosome entry sites in cellular and viral mRNAs
abstract
The internal ribosome entry site (IRES) is a cis-regulatory element that can initiate translation in a cap-independent manner. It is often related to cellular processes and many diseases. Thus, identifying the IRES is important for understanding its mechanism and finding potential therapeutic strategies for relevant diseases since identifying IRES elements by experimental method is time-consuming and laborious. Many bioinformatics tools have been developed to predict IRES, but all these tools are based on structure similarity or machine learning algorithms. Here, we introduced a deep learning model named DeepIRES for precisely identifying IRES elements in messenger RNA (mRNA) sequences. DeepIRES is a hybrid model incorporating dilated 1D convolutional neural network blocks, bidirectional gated recurrent units, and self-attention module. Tenfold cross-validation results suggest that DeepIRES can capture deeper relationships between sequence features and prediction results than other baseline models. Further comparison on independent test sets illustrates that DeepIRES has superior and robust prediction capability than other existing methods. Moreover, DeepIRES achieves high accuracy in predicting experimental validated IRESs that are collected in recent studies. With the application of a deep learning interpretable analysis, we discover some potential consensus motifs that are related to IRES activities. In summary, DeepIRES is a reliable tool for IRES prediction and gives insights into the mechanism of IRES elements.
Jian Zhao 0034, Zhewei Chen, Meng Zhang 0046, Lingxiao Zou, Quan Wang 0004
Briefings Bioinform.1
2023 A Statistical Explainable Learning Model Optimizing Co-localization of Multidimensional Positivity Thresholds in Immunotherapy Decision-Supporting
abstract
Tumor Mutation Burden (TMB) serves as a recognized stratified biomarker for immunotherapy. However, its one-dimensional representation of non-synonymous genetic alterations has been contentious. Specifically, the uniform quantification of mutations by TMB, coupled with measurement inaccuracies, complicates the accurate determination of a positive threshold for classifying patients. Parallel to this, assessing immunotherapy benefits requires the joint analysis of multiscale endpoints, namely discrete tumor response and sequential time-to-event, presenting a pressing challenge for clinical computation. Recognizing the intertwined nature of these challenges, we address the inter-sample bias inherent in multidimensional mutation biomarkers within the framework of multiscale endpoint fusion analysis, aiming for a more robust and comprehensive patient stratification. By combining the concept of corrected-score with a soft-threshold strategy, and utilizing the attention mechanism alongside the multiple instance learning, we propose a statistically explainable learning model optimizing co-localization of multidimensional positivity thresholds for immunotherapy categorical decision-supporting.
Jian Zhao 0034, Jiayin Wang 0002, Quan Wang 0004
BIBM3
2022 csORF-finder: an effective ensemble learning framework for accurate identification of multi-species coding short open reading frames
abstract
Short open reading frames (sORFs) refer to the small nucleic fragments no longer than 303 nt in length that probably encode small peptides. To date, translatable sORFs have been found in both untranslated regions of messenger ribonucleic acids (RNAs; mRNAs) and long non-coding RNAs (lncRNAs), playing vital roles in a myriad of biological processes. As not all sORFs are translated or essentially translatable, it is important to develop a highly accurate computational tool for characterizing the coding potential of sORFs, thereby facilitating discovery of novel functional peptides. In light of this, we designed a series of ensemble models by integrating Efficient-CapsNet and LightGBM, collectively termed csORF-finder, to differentiate the coding sORFs (csORFs) from non-coding sORFs in Homo sapiens, Mus musculus and Drosophila melanogaster, respectively. To improve the performance of csORF-finder, we introduced a novel feature encoding scheme named trinucleotide deviation from expected mean (TDE) and computed all types of in-frame sequence-based features, such as i-framed-3mer, i-framed-CKSNAP and i-framed-TDE. Benchmarking results showed that these features could significantly boost the performance compared to the original 3-mer, CKSNAP and TDE features. Our performance comparisons showed that csORF-finder achieved a superior performance than the state-of-the-art methods for csORF prediction on multi-species and non-ATG initiation independent test datasets. Furthermore, we applied csORF-finder to screen the lncRNA datasets for identifying potential csORFs. The resulting data serve as an important computational repository for further experimental validation. We hope that csORF-finder can be exploited as a powerful platform for high-throughput identification of csORFs and functional characterization of these csORFs encoded peptides.
Meng Zhang 0046, Jian Zhao 0034, Chen Li 0021, Fang Ge, Bin Jiang 0001, Jiangning Song
Briefings Bioinform.2
2022 pHisPred: a tool for the identification of histidine phosphorylation sites by integrating amino acid patterns and properties
abstract
BACKGROUND: Protein histidine phosphorylation (pHis) plays critical roles in prokaryotic signal transduction pathways and various eukaryotic cellular processes. It is estimated to account for 6-10% of the phosphoproteome, however only hundreds of pHis sites have been discovered to date. Due to the inherent disadvantages of experimental methods, it is an urgent task for developing efficient computational approaches to identify pHis sites. RESULTS: Here, we present a novel tool, pHisPred, for accurately identifying pHis sites from protein sequences. We manually collected the largest number of experimental validated pHis sites to build benchmark datasets. Using randomized tenfold CV, the weighted SVM-RBF model shows the best performance than other four commonly used classification models (LR, KNN, RF, and MLP). From ten thousands of features, 140 and 150 most informative features were individually selected out for eukaryotic and prokaryotic models. The average AUC and F1-score values of pHisPred were (0.81, 0.40) and (0.78, 0.46) for tenfold CV on the eukaryotic and prokaryotic training datasets, respectively. In addition, pHisPred significantly outperforms other tools on testing datasets, in particular on the eukaryotic one. CONCLUSION: We implemented a python program of pHisPred, which is freely available for non-commercial use at https://github.com/xiaofengsong/pHisPred . Moreover, users can use it to train new models with their own data.
Jian Zhao 0034, Minhui Zhuang, Meng Zhang 0046, Cong Zeng, Bin Jiang 0001
BMC Bioinform.1