EDBT 2026 Demo / reviewers in the wild / expert
Zhongshen Li
dblp:39/10176
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | METRON: Metabolic Dynamic Perception Kolmogorov-Arnold Network for Biological Age EstimationabstractBiological age is a more direct reflection of physiological status than chronological age, serving as a vital measure to evaluate health risks and aging interventions. While steroid metabolomics offers rich information for exploring aging mechanisms, the complex and nonlinear interactions within metabolic networks remain challenging in modeling. Here, we propose and describe METRON as a deep learning framework to predict biological ages from steroid metabolomics. Specifically, a Metabolite Interaction Perception Module (MIPM) is proposed to capture the interactions. Subsequently, a Group-Rational Kolmogorov-Arnold Network is also integrated to capture intricate dependencies and enhance the representation capability. We demonstrate that METRON achieves promising performance as compared to other machine learning and deep learning methods. Beyond performance, METRON offers interpretability by recovering the established markers such as Dehydroepiandrosterone (DHEA) and identifying 17-hydroxyprogesterone (17-OH-P4) as the key signature linked to hypothalamic-pituitary-adrenal axis dynamics. These results support the capacity of METRON not only to estimate biological age but also to uncover underappreciated metabolic drivers behind aging. Zhongshen Li, Jixiang Yu, Shen You, Hao Liu 0072, Luyang Cai, Yuxuan Deng, Leyi Wei, Junkai Ji, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2026 | UniBreak: A Unified Evolutionary Token-Level Jailbreaking Framework for Large Language ModelsabstractLarge Language Models (LLMs) demonstrate promising capabilities in natural language understanding and reasoning with enormous parameter spaces and vast amounts of training data. These attributes have facilitated their deployment into diverse application domains. However, the underlying parameters implicitly assume decision-making boundaries, resulting in a significant number of decision spaces not covered by training data. This makes them susceptible to adversarial manipulations through carefully crafted inputs. To illuminate the vulnerabilities of LLMs, we propose a unified token-level jailbreaking attack that makes victim models generate responses for potentially harmful queries. Specifically, we propose an evolutionary algorithm to evolve perturbation sets, utilizing gradient-based and crossover-based operators to enhance performance under multiple scenarios. Furthermore, we develop a repository for reusing past perturbations and conduct an analysis of token sensitivity within LLMs, facilitating zero-shot attacks with convergence. Extensive benchmark experiments validate the effectiveness of our method on three different models, achieving increases of 62.36%, 57.89%, and 64.81% in attack success rates compared to baseline methods under white-box scenarios. In addition, the evaluation experiments demonstrate our method is effective for multiple scenarios and different size of models. This research reveals vulnerabilities of LLMs, provides theoretical foundations for developing more robust defense strategies, and contributes to building more reliable AI systems. Shen You, Wei Jiang 0016, Hefei Mei, Danei Gong, Zhongshen Li, Jixiang Yu, Junkai Ji, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE Trans. Evol. Comput. | 5 |
| 2026 | SOPA: Sensitivity-Oriented Poisoning Attack for Self-Supervised Graph Embedding Model via Bilevel Evolutionary OptimizationabstractDespite the popularity of graph neural networks, perturbed graph data is still a serious threat towards its inherent vulnerabilities. Adversarial examples can still easily manipulate the output of graph neural networks across various attack scenarios. Meanwhile, attacks on graph networks also appear to be crucial, as it can help model designers enhance the robustness of their models. In this study, we propose a sensitivity-oriented poisoning attack for self-supervised graph embedding models through bilevel optimization, which employs different optimization methods at each level. In addition, in order to improve attack effectiveness, we analyze graph structure to identify sensitive nodes and edges that guide attack directions, combining gradient-based and query-based methods to target both edge connections and node attributes. Besides, according to the defects of existing graph masked auto-encoders models, we design the feature sensitivity and feature variance to reduce the feature differentiability, which impairs the performance of the downstream model. Ablation studies validate our operator is effective on three citation datasets. And benchmark-based experiments support the effectiveness of our method on three different graph tasks. Specifically, our approach can achieve an average reduction of 3% in the accuracy of node classification compared to existing methods for attacking neural structures alone. For attacking both graph structures and attributes, our model has even achieved an average reduction of 4.5% for the node classification task, outperforming the existing methods. Shen You, Kai Zhou 0001, Zhongshen Li, Kay Chen Tan, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE Trans. Evol. Comput. | 3 |
| 2025 | cis-positional information in regulatory single nucleotide variation prioritizationabstractAbstract Duttke et al. have proved and generalized past observations on the positional preferences of regulatory genomics at multiple functional levels in July 2024 [1]. However, the explicit open-box distribution learning on those positional preferences are under-explored in the existing gene regulation methods including deep learning. Contributing towards such directions, we propose to develop regulatory positional distribution models. The positional distribution models can capture the spatial features of gene transcription which can improve different downstream applications such as rSNV prioritization. Experiments have been conducted to substantiate its claim in deleterious rSNV predictions of ClinVar. In particular, we have collected the ClinVar dataset (i.e., ‘variant_summary.txt.gz’ on 2024-09-18) and retrieved all deleterious SNVs (i.e., labelled as ‘Pathogenic’ and ‘Likely pathogenic’ in the ‘ClinicalSigificance’ column) around all human TSS locations from Ensembl (i.e., Ensembl Genes 112). In particular, it was surprising that, although CADD is already an ensemble approach built upon different state-of-the-arts methods [2], CADD can still be improved with statistical significance after cis-positional information has been incorporated across different situations where evolutionary conservation signals (PhastCons and PhyloP) have been integrated. Based on the results, we propose two future research directions. The first direction is to examine different statistical distributions for position-aware gene regulation modelling while the second direction is to incorporate those distributions into different downstream applications such as eQTL analysis and deleterious rSNV predictions. The outcomes will have broad implications across different downstream gene regulation modelling studies. [1] Duttke S.H. et al. ‘Position-dependent function of human sequence-specific transcription factors.’ Nature 2024;631:891–898. [2] Rentzsch, P. et al. Nucleic acids research 2019:47(D1):D886–D894. Ka-Chun Wong, Zhongyu Yao, Weidun Xie, Zhongshen Li, Tianchi Lu, Cho Ling |
Briefings Bioinform. | 4 |
| 2025 | Elucidating spatiotemporal chromatin dynamics with multi-stage differential variations from Hi-CabstractHigh-throughput sequencing such as Hi-C captures spatiotemporal chromatin interactions, revealing the intricate interplays within transcriptional regulation and chromatin dynamics during cellular reprogramming and developmental processes. However, the forecast on chromatin dynamics in successive developmental stages remains challenging due to the inherent complexity of spatial and temporal patterns in Hi-C data across developmental stages. Towards such a direction, we present StarMie, a deep learning framework that integrates the spatiotemporal-aware module and the multi-stage differential variation module to predict high-throughput chromatin interactions in next developmental stages. Our comprehensive evaluation demonstrates that StarMie outperforms existing methods and sufficiently captures discriminative spatial and temporal dependencies as well as inter-stage-level variations of chromatin interactions. Moreover, the dual importance of both spatial and temporal information in Hi-C data is observed in parameter analysis. Ablation studies also confirm the essential role of each component in StarMie. Furthermore, five cross-species case studies support StarMie's cross-species generalizability and its capability to extract universal chromatin interaction patterns in different developmental stages. In-depth analysis demonstrates that StarMie uncovers conserved genomic logic in cardiac development and disease. Overall, this work paves a new approach for exploring genome reprogramming and development through predictive modeling of Hi-C dynamics. Zhongshen Li, Jixiang Yu, Shen You, Leyi Wei, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
Knowl. Based Syst. | 1 |
| 2024 | StructuralDPPIV: a novel deep learning model based on atom structure for predicting dipeptidyl peptidase-IV inhibitory peptidesabstractMOTIVATION: Diabetes is a chronic metabolic disorder that has been a major cause of blindness, kidney failure, heart attacks, stroke, and lower limb amputation across the world. To alleviate the impact of diabetes, researchers have developed the next generation of anti-diabetic drugs, known as dipeptidyl peptidase IV inhibitory peptides (DPP-IV-IPs). However, the discovery of these promising drugs has been restricted due to the lack of effective peptide-mining tools. RESULTS: Here, we presented StructuralDPPIV, a deep learning model designed for DPP-IV-IP identification, which takes advantage of both molecular graph features in amino acid and sequence information. Experimental results on the independent test dataset and two wet experiment datasets show that our model outperforms the other state-of-art methods. Moreover, to better study what StructuralDPPIV learns, we used CAM technology and perturbation experiment to analyze our model, which yielded interpretable insights into the reasoning behind prediction results. AVAILABILITY AND IMPLEMENTATION: The project code is available at https://github.com/WeiLab-BioChem/Structural-DPP-IV. Junru Jin, Zhongshen Li, Mushuang Fan, Sirui Liang, Ran Su, Leyi Wei |
Bioinform. | 3 |
| 2023 | CoraL: interpretable contrastive meta-learning for the prediction of cancer-associated ncRNA-encoded small peptidesabstractNcRNA-encoded small peptides (ncPEPs) have recently emerged as promising targets and biomarkers for cancer immunotherapy. Therefore, identifying cancer-associated ncPEPs is crucial for cancer research. In this work, we propose CoraL, a novel supervised contrastive meta-learning framework for predicting cancer-associated ncPEPs. Specifically, the proposed meta-learning strategy enables our model to learn meta-knowledge from different types of peptides and train a promising predictive model even with few labeled samples. The results show that our model is capable of making high-confidence predictions on unseen cancer biomarkers with only five samples, potentially accelerating the discovery of novel cancer biomarkers for immunotherapy. Moreover, our approach remarkably outperforms existing deep learning models on 15 cancer-associated ncPEPs datasets, demonstrating its effectiveness and robustness. Interestingly, our model exhibits outstanding performance when extended for the identification of short open reading frames derived from ncPEPs, demonstrating the strong prediction ability of CoraL at the transcriptome level. Importantly, our feature interpretation analysis discovers unique sequential patterns as the fingerprint for each cancer-associated ncPEPs, revealing the relationship among certain cancer biomarkers that are validated by relevant literature and motif comparison. Overall, we expect CoraL to be a useful tool to decipher the pathogenesis of cancer and provide valuable information for cancer research. The dataset and source code of our proposed method can be found at https://github.com/Johnsunnn/CoraL. Zhongshen Li, Junru Jin, Wentao Long, Haoqing Yu, Xin Gao 0001, Kenta Nakai, Quan Zou 0001, Leyi Wei |
Briefings Bioinform. | 1 |
| 2023 | SiameseCPP: a sequence-based Siamese network to predict cell-penetrating peptides by contrastive learningabstractBACKGROUND: Cell-penetrating peptides (CPPs) have received considerable attention as a means of transporting pharmacologically active molecules into living cells without damaging the cell membrane, and thus hold great promise as future therapeutics. Recently, several machine learning-based algorithms have been proposed for predicting CPPs. However, most existing predictive methods do not consider the agreement (disagreement) between similar (dissimilar) CPPs and depend heavily on expert knowledge-based handcrafted features. RESULTS: In this study, we present SiameseCPP, a novel deep learning framework for automated CPPs prediction. SiameseCPP learns discriminative representations of CPPs based on a well-pretrained model and a Siamese neural network consisting of a transformer and gated recurrent units. Contrastive learning is used for the first time to build a CPP predictive model. Comprehensive experiments demonstrate that our proposed SiameseCPP is superior to existing baseline models for predicting CPPs. Moreover, SiameseCPP also achieves good performance on other functional peptide datasets, exhibiting satisfactory generalization ability. Lesong Wei, Xiucai Ye, Saisai Teng, Zhongshen Li, Junru Jin, Min Jae Kim, Tetsuya Sakurai, Li-Zhen Cui 0001, Balachandran Manavalan, Leyi Wei |
Briefings Bioinform. | 6 |
| 2023 | ExamPle: explainable deep learning framework for the prediction of plant small secreted peptidesabstractMOTIVATION: Plant Small Secreted Peptides (SSPs) play an important role in plant growth, development, and plant-microbe interactions. Therefore, the identification of SSPs is essential for revealing the functional mechanisms. Over the last few decades, machine learning-based methods have been developed, accelerating the discovery of SSPs to some extent. However, existing methods highly depend on handcrafted feature engineering, which easily ignores the latent feature representations and impacts the predictive performance. RESULTS: Here, we propose ExamPle, a novel deep learning model using Siamese network and multi-view representation for the explainable prediction of the plant SSPs. Benchmarking comparison results show that our ExamPle performs significantly better than existing methods in the prediction of plant SSPs. Also, our model shows excellent feature extraction ability. Importantly, by utilizing in silicomutagenesis experiment, ExamPle can discover sequential characteristics and identify the contribution of each amino acid for the predictions. The key novel principle learned by our model is that the head region of the peptide and some specific sequential patterns are strongly associated with the SSPs' functions. Thus, ExamPle is expected to be a useful tool for predicting plant SSPs and designing effective plant SSPs. AVAILABILITY AND IMPLEMENTATION: Our codes and datasets are available at https://github.com/Johnsunnn/ExamPle. Zhongshen Li, Junru Jin, Wentao Long, Yuanhao Ding, Leyi Wei |
Bioinform. | 1 |
| 2022 | Accelerating bioactive peptide discovery via mutual information-based meta-learningabstractRecently, machine learning methods have been developed to identify various peptide bio-activities. However, due to the lack of experimentally validated peptides, machine learning methods cannot provide a sufficiently trained model, easily resulting in poor generalizability. Furthermore, there is no generic computational framework to predict the bioactivities of different peptides. Thus, a natural question is whether we can use limited samples to build an effective predictive model for different kinds of peptides. To address this question, we propose Mutual Information Maximization Meta-Learning (MIMML), a novel meta-learning-based predictive model for bioactive peptide discovery. Using few samples from various functional peptides, MIMML can sufficiently learn the discriminative information amongst various functions and characterize functional differences. Experimental results show excellent performance of MIMML though using far fewer training samples as compared to the state-of-the-art methods. We also decipher the latent relationships among different kinds of functions to understand what meta-model learned to improve a specific task. In summary, this study is a pioneering work in the field of functional peptide mining and provides the first-of-its-kind solution for few-sample learning problems in biological sequence analysis, accelerating the new functional peptide discovery. The source codes and datasets are available on https://github.com/TearsWaiting/MIMML. Junru Jin, Zhongshen Li, Jiaojiao Zhao, Balachandran Manavalan, Ran Su, Xin Gao 0001, Leyi Wei |
Briefings Bioinform. | 4 |