Qiaozhen Meng

dblp:223/8127 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PharmaQA: Prompt-Based Molecular Representation Learning via Pharmacophore-Oriented Question Answering
abstract
Molecular representation plays a central role in computational drug discovery. Pharmacophores, functional groups responsible for molecular bioactivity, have been widely studied in cheminformatics. However, their incorporation into molecular representation learning, particularly in a context reasoning or generalization, remains relatively limited. To address this gap, we propose PharmaQA, a pharmacophore oriented question answering framework that formulates tailored prompts to extract context-aware molecular semantics. Rather than encoding pharmacophore features, PharmaQA learns to answer pharmacophore related queries. This design enables flexible reasoning across diverse tasks, including molecular property prediction, compound-target interaction prediction, and binding affinity estimation. Experimental results on benchmark datasets demonstrate that PharmaQA achieves competitive performance. In a ligand discovery case study using FDA-approved compounds, the framework identified potential inhibitors for three therapeutic targets, with strong docking performance. As a generalizable and modular solution, PharmaQA incorporates pharmacophoric knowledge into molecular embeddings, enhancing both predictive accuracy and interpretability in drug discovery applications.
Chengwei Ai, Qiaozhen Meng, Mengwei Sun, Ruihan Dong, Hongpeng Yang, Shiqiang Ma, Cheng Liang 0001, Fei Guo 0001
AAAI2
2026 Motif-Aware Graph Attention Networks for Hemolytic Peptide Prediction
Qiaozhen Meng, Bangguo Tan, Fei Guo 0001
ISBRA (1)1
2025 MolInterAct: Multiscale Cross-Modal Interaction for Robust Molecular Representation Learning
abstract
Molecular representation learning, which captures the fundamental characteristics of chemical compounds, is crucial for AI-driven drug discovery. Existing methods integrate various modalities (e.g., 2D topology and 3D geometry) to develop robust representations. However, current multi-modal fusion strategies either align embedding space through independent models separately, thereby overlooking complementary information, or bridge modalities at a coarse-grained level, failing to capture inherent correlations. We present MolInterAct, an innovative pretraining framework designed to promote multiscale interactions between 2D and 3D modalities at both atomic-level and moleculelevel. Specifically, we propose a fine-grained fusion module, coupled with a customized complementary masking strategy, to seamlessly integrate information at the atomic-level, mitigating overlap and similarity between 2D and 3D representations. In addition, we introduce a fusion contrastive module, which operates at the molecule level, to further strengthen the fusion of 2D and 3D representations while preserving modality-specific features. Finally, we incorporate an intra-modal reconstruction module to reconstruct the original information, further refining the model's understanding of individual modality. Extensive experiments demonstrate that our model outperforms existing molecular pretraining methods across both 2D and 3D benchmarks, highlighting the effectiveness of multiscale fusion between modalities.
Mengwei Sun, Chengwei Ai, Diya Zhang, Qiaozhen Meng, Shiqiang Ma, Fei Guo 0001
BIBM5
2025 MultiPepDec: Decoupled Prompt Learning for Multi-Activity Therapeutic Peptides
abstract
Therapeutic peptides demonstrate significant potential in anti-infection, antitumor, and immunomodulation therapies owing to their high specificity and low toxicity. However, existing computational methods are predominantly limited to single-activity design, restricting their clinical applicability. Here, we present MultiPepDec, a novel decoupled prompt learning framework based on protein language model for concurrent generation of multifunctional peptides, which includes antimicrobial, anticancer, toxic, and metabolic activities. Our approach employs: i) Shared-prompts capturing universal therapeutic patterns via adversarial purification; ii) Private-prompts encoding activity-specific knowledge through contrastive learning, ensuring functional decoupling between four activities. Experimental results demonstrate that generated antimicrobial peptides achieve 80.38% predicted efficacy against E. coli, with comparable performance against most clinically relevant pathogens. This confirms robust broad-spectrum capabilities without requiring pathogen-specific training, while maintaining low computational costs. For other therapeutic activities, the designed sequences not only exhibit the intended biological functions but also show significantly improved diversity. This work establishes a new paradigm for efficient multi-activity peptide design, with potential extensions to other biomolecular engineering domains.
Xingdan Wang, Diya Zhang, Chengwei Ai, Shiqiang Ma, Qiaozhen Meng, Junwen Duan, Fei Guo 0001
BIBM5
2025 DynaPhArM: Adaptive and Physics-Constrained Modeling for Target-Drug Complexes with Drug-Specific Adaptations
abstract
Accurately modeling the target-drug complex at atom level presents a significant challenge in the computer-aided drug design. Traditional methods that rely solely on rigid transformations often fail to capture the adaptive interactions between targets and drugs, particularly during substantial conformational changes in targets upon ligand binding, which becomes especially critical when learning target-drug interactions in drug design. Accurately modeling these changes is crucial for understanding target-drug interactions and improving drug efficacy. To address these challenges, we introduce DynaPhArM, an SE(3)-Equivariant Transformer model specifically designed to capture adaptive alterations occurring within target-drug interactions. DynaPhArM utilizes the cooperative scalar-vector representation, drug-specific embeddings, and a diffusion process to effectively model the evolving dynamics of interactions between targets and drugs. Furthermore, we integrate physical information and energetic principles that maintain essential geometric constraints, such as bond lengths, bond angles, van der Waals forces (vdW), within a multi-task learning (MTL) framework to enhance accuracy. Experimental results demonstrate that DynaPhArM achieves state-of-the-art performance with an overall root mean square deviation (RMSD) of 2.01 Å and a sc-RMSD of 0.29 Å while exhibiting higher success rates compared to existing methodologies. Additionally, DynaPhArM shows promise in enhancing drug specificity, thereby simulating how targets adapt to various drugs through precise modeling of atomic-level interactions and conformational flexibility.
Diya Zhang, Mengwei Sun, Xingdan Wang, Cheng Liang 0001, Qiaozhen Meng, Shiqiang Ma, Fei Guo 0001
NeurIPS5
2024 Rapid screening of multi-point mutations for enzyme thermostability modification by utilizing computational tools
Jia Jin, Qiaozhen Meng, Min Zeng 0004, Guihua Duan, Ercheng Wang, Fei Guo 0001
Future Gener. Comput. Syst.2
2024 DMAMP: A Deep-Learning Model for Detecting Antimicrobial Peptides and Their Multi-Activities
abstract
Due to the broad-spectrum and high-efficiency antibacterial activity, antimicrobial peptides (AMPs) and their functions have been studied in the field of drug discovery. Using biological experiments to detect the AMPs and corresponding activities require a high cost, whereas computational technologies do so for much less. Currently, most computational methods solve the identification of AMPs and their activities as two independent tasks, which ignore the relationship between them. Therefore, the combination and sharing of patterns for two tasks is a crucial problem that needs to be addressed. In this study, we propose a deep learning model, called DMAMP, for detecting AMPs and activities simultaneously, which is benefited from multi-task learning. The first stage is to utilize convolutional neural network models and residual blocks to extract the sharing hidden features from two related tasks. The next stage is to use two fully connected layers to learn the distinct information of two tasks. Meanwhile, the original evolutionary features from the peptide sequence are also fed to the predictor of the second task to complement the forgotten information. The experiments on the independent test dataset demonstrate that our method performs better than the single-task model with 4.28% of Matthews Correlation Coefficient (MCC) on the first task, and achieves 0.2627 of an average MCC which is higher than the single-task model and two existing methods for five activities on the second task. To understand whether features derived from the convolutional layers of models capture the differences between target classes, we visualize these high-dimensional features by projecting into 3D space. In addition, we show that our predictor has the ability to identify peptides that achieve activity against Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2). We hope that our proposed method can give new insights into the discovery of novel antiviral peptide drugs.
Qiaozhen Meng, Genlang Chen, Shixin Zheng, Yulai Lin, Jijun Tang, Fei Guo 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2023 Improved structure-related prediction for insufficient homologous proteins using MSA enhancement and pre-trained language model
abstract
In recent years, protein structure problems have become a hotspot for understanding protein folding and function mechanisms. It has been observed that most of the protein structure works rely on and benefit from co-evolutionary information obtained by multiple sequence alignment (MSA). As an example, AlphaFold2 (AF2) is a typical MSA-based protein structure tool which is famous for its high accuracy. As a consequence, these MSA-based methods are limited by the quality of the MSAs. Especially for orphan proteins that have no homologous sequence, AlphaFold2 performs unsatisfactorily as MSA depth decreases, which may pose a barrier to its widespread application in protein mutation and design problems in which there are no rich homologous sequences and rapid prediction is needed. In this paper, we constructed two standard datasets for orphan and de novo proteins which have insufficient/none homology information, called Orphan62 and Design204, respectively, to fairly evaluate the performance of the various methods in this case. Then, depending on whether or not utilizing scarce MSA information, we summarized two approaches, MSA-enhanced and MSA-free methods, to effectively solve the issue without sufficient MSAs. MSA-enhanced model aims to improve poor MSA quality from the data source by knowledge distillation and generation models. MSA-free model directly learns the relationship between residues on enormous protein sequences from pre-trained models, bypassing the step of extracting the residue pair representation from MSA. Next, we evaluated the performance of four MSA-free methods (trRosettaX-Single, TRFold, ESMFold and ProtT5) and MSA-enhanced (Bagging MSA) method compared with a traditional MSA-based method AlphaFold2, in two protein structure-related prediction tasks, respectively. Comparison analyses show that trRosettaX-Single and ESMFold which belong to MSA-free method can achieve fast prediction ($\sim\! 40$s) and comparable performance compared with AF2 in tertiary structure prediction, especially for short peptides, $\alpha $-helical segments and targets with few homologous sequences. Bagging MSA utilizing MSA enhancement improves the accuracy of our trained base model which is an MSA-based method when poor homology information exists in secondary structure prediction. Our study provides biologists an insight of how to select rapid and appropriate prediction tools for enzyme engineering and peptide drug development. CONTACT: [email protected], [email protected].
Qiaozhen Meng, Fei Guo 0001, Jijun Tang
Briefings Bioinform.1
2023 CLIP: accurate prediction of disordered linear interacting peptides from protein sequences using co-evolutionary information
abstract
One of key features of intrinsically disordered regions (IDRs) is facilitation of protein-protein and protein-nucleic acids interactions. These disordered binding regions include molecular recognition features (MoRFs), short linear motifs (SLiMs) and longer binding domains. Vast majority of current predictors of disordered binding regions target MoRFs, with a handful of methods that predict SLiMs and disordered protein-binding domains. A new and broader class of disordered binding regions, linear interacting peptides (LIPs), was introduced recently and applied in the MobiDB resource. LIPs are segments in protein sequences that undergo disorder-to-order transition upon binding to a protein or a nucleic acid, and they cover MoRFs, SLiMs and disordered protein-binding domains. Although current predictors of MoRFs and disordered protein-binding regions could be used to identify some LIPs, there are no dedicated sequence-based predictors of LIPs. To this end, we introduce CLIP, a new predictor of LIPs that utilizes robust logistic regression model to combine three complementary types of inputs: co-evolutionary information derived from multiple sequence alignments, physicochemical profiles and disorder predictions. Ablation analysis suggests that the co-evolutionary information is particularly useful for this prediction and that combining the three inputs provides substantial improvements when compared to using these inputs individually. Comparative empirical assessments using low-similarity test datasets reveal that CLIP secures area under receiver operating characteristic curve (AUC) of 0.8 and substantially improves over the results produced by the closest current tools that predict MoRFs and disordered protein-binding regions. The webserver of CLIP is freely available at http://biomine.cs.vcu.edu/servers/CLIP/ and the standalone code can be downloaded from http://yanglab.qd.sdu.edu.cn/download/CLIP/.
Zhen-Ling Peng, Zixia Li, Qiaozhen Meng, Bi Zhao, Lukasz A. Kurgan
Briefings Bioinform.3
2022 HDIContact: a novel predictor of residue-residue contacts on hetero-dimer interfaces via sequential information and transfer learning strategy
abstract
Proteins maintain the functional order of cell in life by interacting with other proteins. Determination of protein complex structural information gives biological insights for the research of diseases and drugs. Recently, a breakthrough has been made in protein monomer structure prediction. However, due to the limited number of the known protein structure and homologous sequences of complexes, the prediction of residue-residue contacts on hetero-dimer interfaces is still a challenge. In this study, we have developed a deep learning framework for inferring inter-protein residue contacts from sequential information, called HDIContact. We utilized transfer learning strategy to produce Multiple Sequence Alignment (MSA) two-dimensional (2D) embedding based on patterns of concatenated MSA, which could reduce the influence of noise on MSA caused by mismatched sequences or less homology. For MSA 2D embedding, HDIContact took advantage of Bi-directional Long Short-Term Memory (BiLSTM) with two-channel to capture 2D context of residue pairs. Our comprehensive assessment on the Escherichia coli (E. coli) test dataset showed that HDIContact outperformed other state-of-the-art methods, with top precision of 65.96%, the Area Under the Receiver Operating Characteristic curve (AUROC) of 83.08% and the Area Under the Precision Recall curve (AUPR) of 25.02%. In addition, we analyzed the potential of HDIContact for human-virus protein-protein complexes, by achieving top five precision of 80% on O75475-P04584 related to Human Immunodeficiency Virus. All experiments indicated that our method was a valuable technical tool for predicting inter-protein residue contacts, which would be helpful for understanding protein-protein interaction mechanisms.
Qiaozhen Meng, Jianxin Wang 0001, Fei Guo 0001
Briefings Bioinform.2
2021 Multi-AMP: detecting the antimicrobial peptides and their activities using the multi-task learning
abstract
Recently due to the broad-spectrum and high-efficiency antibacterial activity, antimicrobial peptides (AMPs) have become the best alternative to antibiotics. With the rapid increase of the antibacterial peptides, many computational methods have been developed to identify the AMPs and their specific antibacterial activities. However, most existing methods regard these two problems as independent sub-problems and ignore the correlation between tasks. In this paper, we propose a method, Multi-AMP, which utilizes multi-task learning and solves two tasks simultaneously: 1) whether a given peptide is AMP, 2) which activities it performs. The two tasks share the parameters at the bottom layers of the model and learn the specific information at the top layers. Experiments indicate that our multi-task model performs better than single-task models and two existing predictors, which can give insights to the drug discovery process.
Qiaozhen Meng, Jijun Tang, Fei Guo 0001
BIBM1
2021 Exploring effectiveness of ab-initio protein-protein docking methods on a novel antibacterial protein complex dataset
abstract
Diseases caused by bacterial infections become a critical problem in public heath. Antibiotic, the traditional treatment, gradually loses their effectiveness due to the resistance. Meanwhile, antibacterial proteins attract more attention because of broad spectrum and little harm to host cells. Therefore, exploring new effective antibacterial proteins is urgent and necessary. In this paper, we are committed to evaluating the effectiveness of ab-initio docking methods in antibacterial protein-protein docking. For this purpose, we constructed a three-dimensional (3D) structure dataset of antibacterial protein complex, called APCset, which contained $19$ protein complexes whose receptors or ligands are homologous to antibacterial peptides from Antimicrobial Peptide Database. Then we selected five representative ab-initio protein-protein docking tools including ZDOCK3.0.2, FRODOCK3.0, ATTRACT, PatchDock and Rosetta to identify these complexes' structure, whose performance differences were obtained by analyzing from five aspects, including top/best pose, first hit, success rate, average hit count and running time. Finally, according to different requirements, we assessed and recommended relatively efficient protein-protein docking tools. In terms of computational efficiency and performance, ZDOCK was more suitable as preferred computational tool, with average running time of $6.144$ minutes, average Fnat of best pose of $0.953$ and average rank of best pose of $4.158$. Meanwhile, ZDOCK still yielded better performance on Benchmark 5.0, which proved ZDOCK was effective in performing docking on large-scale dataset. Our survey can offer insights into the research on the treatment of bacterial infections by utilizing the appropriate docking methods.
Qiaozhen Meng, Jijun Tang, Fei Guo 0001
Briefings Bioinform.2
2021 Granular multiple kernel learning for identifying RNA-binding protein residues via integrating sequence and structure information
Yijie Ding, Qiaozhen Meng, Jijun Tang, Fei Guo 0001
Neural Comput. Appl.3
2018 CoABind: a novel algorithm for Coenzyme A (CoA)- and CoA derivatives-binding residues prediction
abstract
Motivation: Coenzyme A (CoA)-protein binding plays an important role in various cellular functions and metabolic pathways. However, no computational methods can be employed for CoA-binding residues prediction. Results: We developed three methods for the prediction of CoA- and CoA derivatives-binding residues, including an ab initio method SVMpred, a template-based method TemPred and a consensus-based method CoABind. In SVMpred, a comprehensive set of features are designed from two complementary sequence profiles and the predicted secondary structure and solvent accessibility. The engine for classification in SVMpred is selected as the support vector machine. For TemPred, the prediction is transferred from homologous templates in the training set, which are detected by the program HHsearch. The assessment on an independent test set consisting of 73 proteins shows that SVMpred and TemPred achieve Matthews correlation coefficient (MCC) of 0.438 and 0.481, respectively. Analysis on the predictions by SVMpred and TemPred shows that these two methods are complementary to each other. Therefore, we combined them together, forming the third method CoABind, which further improves the MCC to 0.489 on the same set. Experiments demonstrate that the proposed methods significantly outperform the state-of-the-art general-purpose ligand-binding residues prediction algorithm COACH. As the first-of-its-kind method, we anticipate CoABind to be helpful for studying CoA-protein interaction. Availability and implementation: http://yanglab.nankai.edu.cn/CoABind. Supplementary information: Supplementary data are available at Bioinformatics online.
Qiaozhen Meng, Zhen-Ling Peng, Jianyi Yang 0002
Bioinform.1