EDBT 2026 Demo / reviewers in the wild / expert
Chengwei Ai
dblp:263/9521
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-1314-660XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PharmaQA: Prompt-Based Molecular Representation Learning via Pharmacophore-Oriented Question AnsweringabstractMolecular representation plays a central role in computational drug discovery. Pharmacophores, functional groups responsible for molecular bioactivity, have been widely studied in cheminformatics. However, their incorporation into molecular representation learning, particularly in a context reasoning or generalization, remains relatively limited. To address this gap, we propose PharmaQA, a pharmacophore oriented question answering framework that formulates tailored prompts to extract context-aware molecular semantics. Rather than encoding pharmacophore features, PharmaQA learns to answer pharmacophore related queries. This design enables flexible reasoning across diverse tasks, including molecular property prediction, compound-target interaction prediction, and binding affinity estimation. Experimental results on benchmark datasets demonstrate that PharmaQA achieves competitive performance. In a ligand discovery case study using FDA-approved compounds, the framework identified potential inhibitors for three therapeutic targets, with strong docking performance. As a generalizable and modular solution, PharmaQA incorporates pharmacophoric knowledge into molecular embeddings, enhancing both predictive accuracy and interpretability in drug discovery applications. Chengwei Ai, Qiaozhen Meng, Mengwei Sun, Ruihan Dong, Hongpeng Yang, Shiqiang Ma, Cheng Liang 0001, Fei Guo 0001 |
AAAI | 1 |
| 2026 | Refprogen: a reference-guided molecular generation model with protein-ligand joint representation for property-aware drug design
Chengwei Ai, Jijun Tang, Fei Guo 0001 |
Expert Syst. Appl. | 3 |
| 2025 | MolInterAct: Multiscale Cross-Modal Interaction for Robust Molecular Representation LearningabstractMolecular representation learning, which captures the fundamental characteristics of chemical compounds, is crucial for AI-driven drug discovery. Existing methods integrate various modalities (e.g., 2D topology and 3D geometry) to develop robust representations. However, current multi-modal fusion strategies either align embedding space through independent models separately, thereby overlooking complementary information, or bridge modalities at a coarse-grained level, failing to capture inherent correlations. We present MolInterAct, an innovative pretraining framework designed to promote multiscale interactions between 2D and 3D modalities at both atomic-level and moleculelevel. Specifically, we propose a fine-grained fusion module, coupled with a customized complementary masking strategy, to seamlessly integrate information at the atomic-level, mitigating overlap and similarity between 2D and 3D representations. In addition, we introduce a fusion contrastive module, which operates at the molecule level, to further strengthen the fusion of 2D and 3D representations while preserving modality-specific features. Finally, we incorporate an intra-modal reconstruction module to reconstruct the original information, further refining the model's understanding of individual modality. Extensive experiments demonstrate that our model outperforms existing molecular pretraining methods across both 2D and 3D benchmarks, highlighting the effectiveness of multiscale fusion between modalities. Mengwei Sun, Chengwei Ai, Diya Zhang, Qiaozhen Meng, Shiqiang Ma, Fei Guo 0001 |
BIBM | 2 |
| 2025 | MultiPepDec: Decoupled Prompt Learning for Multi-Activity Therapeutic PeptidesabstractTherapeutic peptides demonstrate significant potential in anti-infection, antitumor, and immunomodulation therapies owing to their high specificity and low toxicity. However, existing computational methods are predominantly limited to single-activity design, restricting their clinical applicability. Here, we present MultiPepDec, a novel decoupled prompt learning framework based on protein language model for concurrent generation of multifunctional peptides, which includes antimicrobial, anticancer, toxic, and metabolic activities. Our approach employs: i) Shared-prompts capturing universal therapeutic patterns via adversarial purification; ii) Private-prompts encoding activity-specific knowledge through contrastive learning, ensuring functional decoupling between four activities. Experimental results demonstrate that generated antimicrobial peptides achieve 80.38% predicted efficacy against E. coli, with comparable performance against most clinically relevant pathogens. This confirms robust broad-spectrum capabilities without requiring pathogen-specific training, while maintaining low computational costs. For other therapeutic activities, the designed sequences not only exhibit the intended biological functions but also show significantly improved diversity. This work establishes a new paradigm for efficient multi-activity peptide design, with potential extensions to other biomolecular engineering domains. Xingdan Wang, Diya Zhang, Chengwei Ai, Shiqiang Ma, Qiaozhen Meng, Junwen Duan, Fei Guo 0001 |
BIBM | 3 |
| 2025 | HyperPhS: a pharmacophore-guided multimodal representation framework for metabolic stability prediction through contrastive hypergraph learningabstractMOTIVATION: Metabolic stability is crucial in the early stage of drug discovery and development. Drug candidate screening and optimization can be streamlined through the accurate prediction of stability. Functional groups within drug molecules are known as pharmacophores, which bind directly to receptors or biological macromolecules to produce biological effects, thereby affecting metabolic stability. Therefore, determining metabolic stability via the pharmacophore groups remains a significant challenge. RESULTS: To address these issues, we propose a Pharmacophore-guided Hypergraph representation framework for predicting metabolic Stability (HyperPhS). In this study, we introduce a hypergraph-based method to extract features from metabolic pharmacophores with multi-view representation and contrastive learning. In particular, we introduce a pharmacophore-based contrastive learning encoder that captures the consistency between functional and nonfunctional structures. Our method applies ChatGPT simultaneously to metabolites and heterogeneous encoders and integrates multimodal representations by using attention-driven fusion modules coupled with fully connected neural networks. On the HLM dataset, HyperPhS achieves outstanding performance with 87.6% in AUC and 62.6% in MCC, alongside an external test AUC of 88.3%. In addition, pharmacophore groups studied by HyperPhS are validated for their interpretability through case studies. Overall, HyperPhS is an effective and interpretable tool for determining metabolic stability, identifying critical functional groups, and optimizing compounds. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/xiaoyiliu-usc/HyperPhS. Chenglong Kang, Chengwei Ai, Hongpeng Yang, Jijun Tang, Fei Guo 0001 |
Bioinform. | 4 |
| 2024 | RetroCaptioner: beyond attention in end-to-end retrosynthesis transformer via contrastively captioned learnable graph representationabstractMOTIVATION: Retrosynthesis identifies available precursor molecules for various and novel compounds. With the advancements and practicality of language models, Transformer-based models have increasingly been used to automate this process. However, many existing methods struggle to efficiently capture reaction transformation information, limiting the accuracy and applicability of their predictions. RESULTS: We introduce RetroCaptioner, an advanced end-to-end, Transformer-based framework featuring a Contrastive Reaction Center Captioner. This captioner guides the training of dual-view attention models using a contrastive learning approach. It leverages learned molecular graph representations to capture chemically plausible constraints within a single-step learning process. We integrate the single-encoder, dual-encoder, and encoder-decoder paradigms to effectively fuse information from the sequence and graph representations of molecules. This involves modifying the Transformer encoder into a uni-view sequence encoder and a dual-view module. Furthermore, we enhance the captioning of atomic correspondence between SMILES and graphs. Our proposed method, RetroCaptioner, achieved outstanding performance with 67.2% in top-1 and 93.4% in top-10 exact matched accuracy on the USPTO-50k dataset, alongside an exceptional SMILES validity score of 99.4%. In addition, RetroCaptioner has demonstrated its reliability in generating synthetic routes for the drug protokylol. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/guofei-tju/RetroCaptioner. Chengwei Ai, Hongpeng Yang, Ruihan Dong, Jijun Tang, Shuangjia Zheng, Fei Guo 0001 |
Bioinform. | 2 |
| 2024 | MTMol-GPT: De novo multi-target molecular generation with transformer-based generative adversarial imitation learningabstractDe novo drug design is crucial in advancing drug discovery, which aims to generate new drugs with specific pharmacological properties. Recently, deep generative models have achieved inspiring progress in generating drug-like compounds. However, the models prioritize a single target drug generation for pharmacological intervention, neglecting the complicated inherent mechanisms of diseases, and influenced by multiple factors. Consequently, developing novel multi-target drugs that simultaneously target specific targets can enhance anti-tumor efficacy and address issues related to resistance mechanisms. To address this issue and inspired by Generative Pre-trained Transformers (GPT) models, we propose an upgraded GPT model with generative adversarial imitation learning for multi-target molecular generation called MTMol-GPT. The multi-target molecular generator employs a dual discriminator model using the Inverse Reinforcement Learning (IRL) method for a concurrently multi-target molecular generation. Extensive results show that MTMol-GPT generates various valid, novel, and effective multi-target molecules for various complex diseases, demonstrating robustness and generalization capability. In addition, molecular docking and pharmacophore mapping experiments demonstrate the drug-likeness properties and effectiveness of generated molecules potentially improve neuropsychiatric interventions. Furthermore, our model's generalizability is exemplified by a case study focusing on the multi-targeted drug design for breast cancer. As a broadly applicable solution for multiple targets, MTMol-GPT provides new insight into future directions to enhance potential complex disease therapeutics by generating high-quality multi-target molecules in drug discovery. Chengwei Ai, Hongpeng Yang, Ruihan Dong, Yijie Ding, Fei Guo 0001 |
PLoS Comput. Biol. | 1 |
| 2023 | MVML-MPI: Multi-View Multi-Label Learning for Metabolic Pathway InferenceabstractDevelopment of robust and effective strategies for synthesizing new compounds, drug targeting and constructing GEnome-scale Metabolic models (GEMs) requires a deep understanding of the underlying biological processes. A critical step in achieving this goal is accurately identifying the categories of pathways in which a compound participated. However, current machine learning-based methods often overlook the multifaceted nature of compounds, resulting in inaccurate pathway predictions. Therefore, we present a novel framework on Multi-View Multi-Label Learning for Metabolic Pathway Inference, hereby named MVML-MPI. First, MVML-MPI learns the distinct compound representations in parallel with corresponding compound encoders to fully extract features. Subsequently, we propose an attention-based mechanism that offers a fusion module to complement these multi-view representations. As a result, MVML-MPI accurately represents and effectively captures the complex relationship between compounds and metabolic pathways and distinguishes itself from current machine learning-based methods. In experiments conducted on the Kyoto Encyclopedia of Genes and Genomes pathways dataset, MVML-MPI outperformed state-of-the-art methods, demonstrating the superiority of MVML-MPI and its potential to utilize the field of metabolic pathway design, which can aid in optimizing drug-like compounds and facilitating the development of GEMs. The code and data underlying this article are freely available at https://github.com/guofei-tju/MVML-MPI. Contact: [email protected], [email protected] or [email protected]. Hongpeng Yang, Chengwei Ai, Yijie Ding, Fei Guo 0001, Jijun Tang |
Briefings Bioinform. | 3 |
| 2023 | Low Rank Matrix Factorization Algorithm Based on Multi-Graph Regularization for Detecting Drug-Disease AssociationabstractDetecting potential associations between drugs and diseases plays an indispensable role in drug development, which has also become a research hotspot in recent years. Compared with traditional methods, some computational approaches have the advantages of fast speed and low cost, which greatly accelerate the progress of predicting the drug-disease association. In this study, we propose a novel similarity-based method of low-rank matrix decomposition based on multi-graph regularization. On the basis of low-rank matrix factorization with$L_{2}$regularization, the multi-graph regularization constraint is constructed by combining a variety of similarity matrices from drugs and diseases respectively. In the experiments, we analyze the difference in the combination of different similarities, resulting that combining all the similarity information on drug space is unnecessary, and only a part of the similarity information can achieve the desired performance. Then our method is compared with other existing models on three data sets (Fdataset, Cdataset and LRSSLdataset) and have a good advantage in the evaluation measurement of AUPR. Besides, a case study experiment is conducted and showing that the superior ability for predicting the potential disease-related drugs of our model. Finally, we compare our model with some methods on six real world datasets, and our model has a good performance in detecting real world data. Chengwei Ai, Hongpeng Yang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | Microbe-bridged disease-metabolite associations identification by heterogeneous graph fusionabstractMOTIVATION: Metabolomics has developed rapidly in recent years, and metabolism-related databases are also gradually constructed. Nowadays, more and more studies are being carried out on diverse microbes, metabolites and diseases. However, the logics of various associations among microbes, metabolites and diseases are limited understanding in the biomedicine of gut microbial system. The collection and analysis of relevant microbial bioinformation play an important role in the revelation of microbe-metabolite-disease associations. Therefore, the dataset that integrates multiple relationships and the method based on complex heterogeneous graphs need to be developed. RESULTS: In this study, we integrated some databases and extracted a variety of associations data among microbes, metabolites and diseases. After obtaining the three interconnected bilateral association data (microbe-metabolite, metabolite-disease and disease-microbe), we considered building a heterogeneous graph to describe the association data. In our model, microbes were used as a bridge between diseases and metabolites. In order to fuse the information of disease-microbe-metabolite graph, we used the bipartite graph attention network on the disease-microbe and metabolite-microbe bipartite graph. The experimental results show that our model has good performance in the prediction of various disease-metabolite associations. Through the case study of type 2 diabetes mellitus, Parkinson's disease, inflammatory bowel disease and liver cirrhosis, it is noted that our proposed methodology are valuable for the mining of other associations and the prediction of biomarkers for different human diseases.Availability and implementation: https://github.com/Selenefreeze/DiMiMe.git. Jitong Feng, Shengbo Wu, Hongpeng Yang, Chengwei Ai, Jianjun Qiao, Junhai Xu, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2022 | A multi-layer multi-kernel neural network for determining associations between non-coding RNAs and diseases
Chengwei Ai, Hongpeng Yang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 1 |
| 2022 | Sparse regularized joint projection model for identifying associations of non-coding RNAs and human diseasesabstractCurrent human biomedical research shows that human diseases are closely related to non-coding RNAs, so it is of great significance for human medicine to study the relationship between diseases and non-coding RNAs. Current research has found associations between non-coding RNAs and human diseases through a variety of effective methods, but most of the methods are complex and targeted at a single RNA or disease. Therefore, we urgently need an effective and simple method to discover the associations between non-coding RNAs and human diseases. In this paper, we propose a sparse regularized joint projection model (SRJP) to identify the associations between non-coding RNAs and diseases. First, we extract information through a series of ncRNA similarity matrices and disease similarity matrices and assign average weights to the similarity matrices of the two sides. Then we decompose the similarity matrices of the two spaces into low-rank matrices and put them into SRJP. In SRJP, we innovatively use the projection matrix to combine the ncRNA side and the disease side to identify the associations between ncRNAs and diseases. Finally, the regularization term in SRJP effectively improves the robustness and generalization ability of the model. We test our model on different datasets involving three types of ncRNAs: circRNA, microRNA and long non-coding RNA. The experimental results show that SRJP has superior ability to identify and predict the associations between ncRNAs and diseases. Prayag Tiwari, Junhai Xu, Yuqing Qian, Chengwei Ai, Yijie Ding, Fei Guo 0001 |
Knowl. Based Syst. | 5 |