VLDB 2026 Research / reviewers in the wild / expert
Li Wang 0145
dblp:58/6810-145
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-2801-7213ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GSToxi: Gated Cross-Modal Modeling With Graph-Sequence Encoders for Peptide Toxicity PredictionabstractPeptide-based therapeutics hold great potential, yet their cytotoxicity remains a key challenge in drug development. Most existing toxicity prediction models rely solely on sequence information, often overlooking the fusion of multimodal submolecular patterns. We propose GSToxi, a multimodal deep learning framework that leverage sequence and molecular graph featurizer to enhance peptide toxicity prediction. A shared gating mechanism is employed to facilitate semantic alignment and cross-modal integration, while a contrastive regularization loss further optimizes latent-space consistency throughout the training process. Furthermore, GSToxi incorporates embeddings from pre-trained protein language models alongside low-level compositional priors, enabling the capture of both global contextual semantics and local structural features. Experimental results show that GSToxi outperforms state-of-the-art baselines across multiple evaluation metrics on an independent test set. Ablation studies underscore the critical contributions of each component, with the molecular graph encoder and pre-trained embeddings proving particularly impactful. This work offers a generalizable and robust framework for peptide toxicity prediction and provides valuable insights for future multimodal modeling of biological molecules. Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai |
BIBM | 1 |
| 2025 | ET-PROTACs: modeling ternary complex interactions using cross-modal learning and ternary attention for accurate PROTAC-induced degradation predictionabstractMOTIVATION: Accurately predicting the degradation capabilities of proteolysis-targeting chimeras (PROTACs) for given target proteins and E3 ligases is important for PROTAC design. The distinctive ternary structure of PROTACs presents a challenge to traditional drug-target interaction prediction methods, necessitating more innovative approaches. While current state-of-the-art (SOTA) methods using graph neural networks (GNNs) can discern the molecular structure of PROTACs and proteins, thus enabling the efficient prediction of PROTACs' degradation capabilities, they rely heavily on limited crystal structure data of the POI-PROTAC-E3 ternary complex. This reliance underutilizes rich PROTAC experimental data and neglects intricate interaction relationships within ternary complexes. RESULTS: In this study, we propose a model based on cross-modal strategy and ternary attention technology, ET-PROTACs, to predict the targeted degradation capabilities of PROTACs. Our model capitalizes on the strengths of cross-modal methods by using equivariant GNN graph neural networks to process the graph structure and spatial coordinates of PROTAC molecules concurrently while utilizing sequence-based methods to learn the protein sequence information. This integration of cross-modal information is cohesively harnessed and channeled into a ternary attention mechanism, specially tailored for the unique structure of PROTACs, enabling the congruent modeling of both PROTAC and protein modalities. Experimental results demonstrate that the ET-PROTACs model outperforms existing SOTA methods. Moreover, visualizing attention scores illuminates crucial residues and atoms pivotal in specific POI-PROTAC-E3 interactions, thus offering invaluable insights and guidance for future pharmaceutical research. AVAILABILITY AND IMPLEMENTATION: The codes of our model are available at https://github.com/GuanyuYue/ET-PROTACs. Guanyu Yue, Li Wang 0145, Quan Zou 0001, Xiangzheng Fu, Dong-Sheng Cao 0001 |
Briefings Bioinform. | 4 |
| 2025 | MOFormer: navigating the antimicrobial peptide design space with Pareto-based multi-objective transformerabstractAntimicrobial peptide (AMP) design through deep learning holds the potential to revolutionize antibiotic development. Despite recent progress in AMP generation, designing peptide antibiotics with multiple optimal properties remains a significant challenge. We present MOFormer, an advanced multi-objective AMP design pipeline capable of optimizing multiple AMP properties simultaneously. By leveraging a conditional Transformer, the model refines the AMP sequence-property landscape for efficient multi-objective generation. It also incorporates regularization techniques to maintain a highly structured space, enabling the sampling of precise and desirable candidates. Comparative analyses reveal that MOFormer achieves the optimal hypervolume in the multi-objective space, surpassing advanced methods in simultaneously maximizing antimicrobial activity (minimum inhibitory concentration) and minimizing hemolysis and toxicity, thereby yielding the most promising and desirable set of candidate peptides. When extended to a tri-objective scenario, MOFormer continues to exhibit remarkable optimization performance. Finally, we execute a hierarchical and rapid ranking of generated candidates based on Pareto fronts. We conducted a comprehensive validation of the physicochemical properties and target attributes of the candidates, while AlphaFold structure predictions revealed notably reliable predicted local distance difference test scores ranging from 70% to 87%. Our findings suggest that MOFormer holds potential to accelerate the discovery of efficacious peptide antibiotics by optimizing multi-objective trade-offs. Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng |
Briefings Bioinform. | 1 |
| 2025 | PKAN: Leveraging Kolmogorov-Arnold Networks and Multi-Modal Learning for Peptide Prediction With Advanced Language ModelsabstractPeptides can offer highly specific biological activities, serving as essential mediators of intercellular signaling, which are critical for advancing precision medicine and drug development. Their primary structure can be depicted either as an amino acid sequence or as a chemical molecules consisting of atoms and chemical bonds. Large language models (LLMs) hold the potential to thoroughly elucidate the intricate intrinsic properties of peptides. Here we present the Peptide Kolmogorov-Arnold Network (PKAN), a framework leveraging multi-modal representations inspired by advanced language models for peptide activity and functionality prediction. Comparative experiments across tasks show that PKAN outperforms state-of-the-art models while maintaining a streamlined design with superior predictive capabilities. The multi-modal feature importance scoring, anchored in global structures and the significant marginal impacts of derived features on the model, coupled with intricate symbolic regression of specific activation functions, further demonstrates the robustness and precision of the PKAN framework in identifying and elucidating key determinants of peptide functionality. This work provides scientific evidence for investigating the complex mechanisms of peptide materials and supports the progression of peptide language paradigms in biology. Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Multi-Objective Molecular Design in Constrained Latent SpaceabstractIn recent times, molecular design has undergone significant advancements, particularly with the integration of artificial intelligence (AI) techniques for discovering molecules with various desired attributes. The use of generative models, especially variational autoencoders (VAEs), has proven to be a potent and efficient method. These models facilitate the rapid identification of new molecules that align with specific research goals. However, sequence-based generative models, like SMILES-based VAEs, often generate invalid molecules, presenting substantial challenges due to the inherent constraints in their latent spaces. To overcome this issue, we introduce a novel multi-objective molecular design approach that incorporates a corrector-based constraint handling technique. This technique employs a transformer as a corrector to convert invalid molecules into valid ones during the search process. Following correction, the latent space is segmented into distinct zones, organized via a spatial partition tree. Utilizing Monte Carlo tree search, we pinpoint the most promising zones for evolutionary-based sampling. Our method applies multi-objective molecular design within these constrained latent spaces. Our experimental results demonstrate that this approach markedly improves the quality of molecules generated from the latent space. Li Wang 0145, Xiangxiang Zeng |
IJCNN | 5 |
| 2024 | An interpretable deep learning model predicts RNA-small molecule binding sites
Wen-Yu Xi, Ruheng Wang, Li Wang 0145, Xiucai Ye, Tetsuya Sakurai |
Future Gener. Comput. Syst. | 3 |
| 2021 | ITP-Pred: an interpretable method for predicting, therapeutic peptides with fused features low-dimension representationabstractThe peptide therapeutics market is providing new opportunities for the biotechnology and pharmaceutical industries. Therefore, identifying therapeutic peptides and exploring their properties are important. Although several studies have proposed different machine learning methods to predict peptides as being therapeutic peptides, most do not explain the decision factors of model in detail. In this work, an Interpretable Therapeutic Peptide Prediction (ITP-Pred) model based on efficient feature fusion was developed. First, we proposed three kinds of feature descriptors based on sequence and physicochemical property encoded, namely amino acid composition (AAC), group AAC and coding autocorrelation, and concatenated them to obtain the feature representation of therapeutic peptide. Then, we input it into the CNN-Bi-directional Long Short-Term Memory (BiLSTM) model to automatically learn recognition of therapeutic peptides. The cross-validation and independent verification experiments results indicated that ITP-Pred has a higher prediction performance on the benchmark dataset than other comparison methods. Finally, we analyzed the output of the model from two aspects: sequence order and physical and chemical properties, mining important features as guidance for the design of better models that can complement existing methods. Li Wang 0145, Xiangzheng Fu, Chenxing Xia, Xiangxiang Zeng, Quan Zou 0001 |
Briefings Bioinform. | 2 |