VLDB 2026 Research / reviewers in the wild / expert
Zilin Ren
dblp:340/1343
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-3621-1024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MuFaDDG: a sequence-based multiscale feature fusion framework for protein stability changes predictionabstractMOTIVATION: Predicting the thermodynamic stability of proteins upon single-point mutations is a pivotal step in both protein engineering and medicine. In the study of predicting protein thermodynamic stability, various computational methods, whether they extract features at the local-level or global-level, exhibit their respective advantages and limitations. To leverage the advantages of both features, we developed MuFaDDG, a novel sequence-based method that integrated multiscale feature fusion for improved prediction of protein stability changes (ΔΔG). RESULTS: MuFaDDG achieves comparable performance on the S669 benchmark, demonstrating strong capabilities in stabilizing mutations. Notably, it shows a significant advantage in the ACC metric, with values of 0.75, 0.88, and 0.81 on the direct, reverse, and overall datasets of the CAGI5 Challenge's Frataxin, respectively. Furthermore, our method outperforms leading sequence-based approaches including THPLM, DDGemb, DDGun, and INPS-Seq on protein Myoglobin stability prediction. Additionally, MuFaDDG demonstrates exceptional predictive performance with higher PCC and ACC on the protein ThreeFoil, which is uncurated by FireProtDB and ProThermDB databases. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/PengjiaMa23/MuFaDDG. Jianting Gong, Pengjia Ma, Zilin Ren, Zhiguo Fu, Xiaochen Bo |
Bioinform. | 3 |
| 2025 | PLMTox: Low-Rank Adaptation on Protein Language Models for Peptide Toxicity PredictionabstractPeptides, as biologically active substances between small molecules and proteins, play a crucial role in drug development due to their high specificity. However, peptides may trigger toxic reactions when exerting a therapeutic effect. Recently, deep learning has emerged as a leading paradigm for peptide toxicity prediction. With the development of protein language models (PLMs), their advantages in protein sequence analysis provide a new technological path for peptide toxicity prediction, but their practical application is limited by the high computational resource requirements. To address the above problems, this paper proposes a peptide toxicity prediction model, PLMTox, that combines PLMs and low-rank adaptation (LoRA) technology. We use protein language models to learn evolutionary information and deep semantic features behind peptide sequences. Simultaneously, PLMTox employs low-rank matrix decomposition to retain the understanding ability of the pre-trained model while realizing the efficient task adaptation of parameters, thus reducing the demand for computational resources. For different datasets, PLMTox (AUC$=0.963$,$\text{SP}=0.960$) exhibits good performance and robustness. These experiment results suggest that PLMTox is expected to promote the safety assessment process of peptide drug development. Lingyun Song, Zilin Ren, Xuequn Shang 0001 |
BIBM | 5 |
| 2025 | DeepPFP: a multi-task-aware architecture for protein function predictionabstractDeriving protein function from protein sequences poses a significant challenge due to the intricate relationship between sequence and function. Deep learning has made remarkable strides in predicting sequence-function relationships. However, models tailored for specific tasks or protein types encounter difficulties when using transfer learning across domains. This is attributed to the fact that protein function relies heavily on structural characteristics rather than mere sequence information. Consequently, there is a pressing need for a model capable of capturing shared features among diverse sequence-function mapping tasks to address the generalization issue. In this study, we explore the potential of Model-Agnostic Meta-Learning combined with a protein language model called Evolutionary Scale Modeling to tackle this challenge. Our approach involves training the architecture on five out-domain deep mutational scanning (DMS) datasets and evaluating its performance across four key dimensions. Our findings demonstrate that the proposed architecture exhibits satisfactory performance in terms of generalization and employs an effective few-shot learning strategy. To explain further, Compared to the best results, the Pearson's correlation coefficient (PCC) in the final stage increased by ~0.31%. Furthermore, we leverage the trained architecture to predict binding affinity scores of the DMS dataset of SARS-CoV-2 using transfer learning. Notably, training on a subset of the Ube4b dataset with 500 samples resulted in a notable improvement of 0.11 in the PCC. These results underscore the potential of our conceptual architecture as a promising methodology for multi-task protein function prediction. Zilin Ren, Jinghong Sun, Yongbing Chen, Xiaochen Bo, Jiguo Xue, Jingyang Gao |
Briefings Bioinform. | 2 |
| 2024 | NASTRA: accurate analysis of short tandem repeat markers by nanopore sequencing with repeat-structure-aware algorithmabstractShort-tandem repeats (STRs) are the type of genetic markers extensively utilized in biomedical and forensic applications. Due to sequencing noise in nanopore sequencing, accurate analysis methods are lacking. We developed NASTRA, an innovative tool for Nanopore Autosomal Short Tandem Repeat Analysis, which overcomes traditional database-based methods' limitations and provides a precise germline analysis of STR genetic markers without the need for allele sequence reference. Demonstrating high accuracy in cell line authentication testing and paternity testing, NASTRA significantly surpasses existing methods in both speed and accuracy. This advancement makes it a promising solution for rapid cell line authentication and kinship testing, highlighting the potential of nanopore sequencing for in-field applications. Zilin Ren, Jiarong Zhang, Jiguo Xue, Xiaochen Bo, Jiangwei Yan |
Briefings Bioinform. | 1 |
| 2024 | CodonBERT: a BERT-based architecture tailored for codon optimization using the cross-attention mechanismabstractMOTIVATION: Due to the varying delivery methods of mRNA vaccines, codon optimization plays a critical role in vaccine design to improve the stability and expression of proteins in specific tissues. Considering the many-to-one relationship between synonymous codons and amino acids, the number of mRNA sequences encoding the same amino acid sequence could be enormous. Finding stable and highly expressed mRNA sequences from the vast sequence space using in silico methods can generally be viewed as a path-search problem or a machine translation problem. However, current deep learning-based methods inspired by machine translation may have some limitations, such as recurrent neural networks, which have a weak ability to capture the long-term dependencies of codon preferences. RESULTS: We develop a BERT-based architecture that uses the cross-attention mechanism for codon optimization. In CodonBERT, the codon sequence is randomly masked with each codon serving as a key and a value. In the meantime, the amino acid sequence is used as the query. CodonBERT was trained on high-expression transcripts from Human Protein Atlas mixed with different proportions of high codon adaptation index codon sequences. The result showed that CodonBERT can effectively capture the long-term dependencies between codons and amino acids, suggesting that it can be used as a customized training framework for specific optimization targets. AVAILABILITY AND IMPLEMENTATION: CodonBERT is freely available on https://github.com/FPPGroup/CodonBERT. Zilin Ren, Yaxin Di, Dufei Zhang, Jianli Gong, Jianting Gong, Qiwei Jiang, Zhiguo Fu |
Bioinform. | 1 |
| 2023 | PLMTHP: An Ensemble Framework for Tumor Homing Peptide Prediction based on Protein Language ModelabstractTumor homing peptides (THPs) play significant role in recognizing and specifically binding to tumor cells. Traditional experimental methods can accurately identify THPs but often suffer from high measurement costs and long experimental cycles. In-silico methods can rapidly screen THPs and accelerate the THPs identification process. Existing THPs prediction methods primarily rely on constructing peptide sequence features and machine learning approaches. These tools rely on manual feature extraction, and the combination of features is crucial for both model performance and robustness. In this study, we propose a method called PLMTHP based on protein language model and integrate multiple machine learning models. Compared to existing models, PLMTHP improves predictions by incorporating high-dimensional features encoded through a protein language model. Our method achieved impressive performance on an independent test set, with ACC of 0.915, MCC of 0.831, and AUC of 0.964. PLMTHP can be downloaded from the following website: https://github.com/Chenyb939/PLMTHP Yongbing Chen, Han Qin, Xiaodan Cui, Tianhui Yang, Jianting Gong, Zilin Ren |
BIBM | 7 |
| 2023 | THPLM: a sequence-based deep learning framework for protein stability changes prediction upon point variations using pretrained protein language modelabstractMOTIVATION: Quantitative determination of protein thermodynamic stability is a critical step in protein and drug design. Reliable prediction of protein stability changes caused by point variations contributes to developing-related fields. Over the past decades, dozens of structure-based and sequence-based methods have been proposed, showing good prediction performance. Despite the impressive progress, it is necessary to explore wild-type and variant protein representations to address the problem of how to represent the protein stability change in view of global sequence. With the development of structure prediction using learning-based methods, protein language models (PLMs) have shown accurate and high-quality predictions of protein structure. Because PLM captures the atomic-level structural information, it can help to understand how single-point variations cause functional changes. RESULTS: Here, we proposed THPLM, a sequence-based deep learning model for stability change prediction using Meta's ESM-2. With ESM-2 and a simple convolutional neural network, THPLM achieved comparable or even better performance than most methods, including sequence-based and structure-based methods. Furthermore, the experimental results indicate that the PLM's ability to generate representations of sequence can effectively improve the ability of protein function prediction. AVAILABILITY AND IMPLEMENTATION: The source code of THPLM and the testing data can be accessible through the following links: https://github.com/FPPGroup/THPLM. Jianting Gong, Yongbing Chen, Zhiguo Fu, Zilin Ren, Mingyao Tian |
Bioinform. | 10 |
| 2023 | Model performance and interpretability of semi-supervised generative adversarial networks to predict oncogenic variants with unlabeled dataabstractBACKGROUND: It remains an important challenge to predict the functional consequences or clinical impacts of genetic variants in human diseases, such as cancer. An increasing number of genetic variants in cancer have been discovered and documented in public databases such as COSMIC, but the vast majority of them have no functional or clinical annotations. Some databases, such as CiVIC are available with manual annotation of functional mutations, but the size of the database is small due to the use of human annotation. Since the unlabeled data (millions of variants) typically outnumber labeled data (thousands of variants), computational tools that take advantage of unlabeled data may improve prediction accuracy. RESULT: To leverage unlabeled data to predict functional importance of genetic variants, we introduced a method using semi-supervised generative adversarial networks (SGAN), incorporating features from both labeled and unlabeled data. Our SGAN model incorporated features from clinical guidelines and predictive scores from other computational tools. We also performed comparative analysis to study factors that influence prediction accuracy, such as using different algorithms, types of features, and training sample size, to provide more insights into variant prioritization. We found that SGAN can achieve competitive performances with small labeled training samples by incorporating unlabeled samples, which is a unique advantage compared to traditional machine learning methods. We also found that manually curated samples can achieve a more stable predictive performance than publicly available datasets. CONCLUSIONS: By incorporating much larger samples of unlabeled data, the SGAN method can improve the ability to detect novel oncogenic variants, compared to other machine-learning algorithms that use only labeled datasets. SGAN can be potentially used to predict the pathogenicity of more complex variants such as structural variants or non-coding variants, with the availability of more training samples and informative features. Zilin Ren, Kajia Cao, Marilyn M. Li, Yunyun Zhou |
BMC Bioinform. | 1 |
| 2022 | Correction: Model performance and interpretability of semi-supervised generative adversarial networks to predict oncogenic variants with unlabeled data
Zilin Ren, Kajia Cao, Marilyn M. Li, Yunyun Zhou |
BMC Bioinform. | 1 |