EDBT 2026 Demo / reviewers in the wild / expert
Yang Li 0130
dblp:37/4190-130
· DBLP profile ↗
20ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0002-0403-7287ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identification and characterization of lncRNA-stemness-immune regulatory patternsabstractLong noncoding RNAs (lncRNAs) play critical roles in regulating stemness signature genes (SSGs) and tumor immunity, thereby shaping the tumor microenvironment and antitumor immune responses. Increasing evidence suggests that cancer stem cell traits are closely associated with immune evasion and therapeutic resistance, underscoring the need to systematically characterize the pan-cancer interplay among SSGs, lncRNAs, and tumor immunity. Here, we developed an integrative analytical framework that combines network-based modeling with Bayesian network inference to identify core regulatory triplets (STEM-LncCRTs), each consisting of an lncRNA, an SSG, and an immune gene. We demonstrate that specific stemness-related lncRNAs can distinguish cancer subtypes, and that common stemness-related lncRNAs correlate significantly with immune cell infiltration. Notably, the ATAD5/PRR11-AS1/SKP2 triplet exhibits favorable prognostic potential across multiple cancers and consistently outperforms individual gene markers in predicting 1-, 3-, and 5-year overall survival. Furthermore, using four machine learning algorithms across three independent immunotherapy cohorts, we validate the predictive value of STEM-LncCRTs for immune checkpoint inhibitor response. Importantly, integrating STEM-LncCRTs with tumor mutation burden further improves predictive accuracy. Collectively, this study provides a systems-level view of stemness-related lncRNA regulation in tumor immunity and offers practical biomarkers for predicting immunotherapy efficacy. Zhipeng Qian, Chunlong Zhang, Guohua Wang 0001, Chunyu Wang 0002, Yang Li 0130 |
Briefings Bioinform. | 6 |
| 2025 | Relational similarity-based graph contrastive learning for DTI predictionabstractAs part of the drug repurposing process, it is imperative to predict the interactions between drugs and target proteins in an accurate and efficient manner. With the introduction of contrastive learning into drug-target prediction, the accuracy of drug repurposing will be further improved. However, a large part of DTI prediction methods based on deep learning either focus only on the structural features of proteins and drugs extracted using GNN or CNN, or focus only on their relational features extracted using heterogeneous graph neural networks on a DTI heterogeneous graph. Since the structural and relational features of proteins and drugs describe their attribute information from different perspectives, their combination can improve DTI prediction performance. We propose a relational similarity-based graph contrastive learning for DTI prediction (RSGCL-DTI), which combines the structural and relational features of drugs and proteins to enhance the accuracy of DTI predictions. In our proposed method, the inter-protein relational features and inter-drug relational features are extracted from the heterogeneous drug-protein interaction network through graph contrastive learning, respectively. The results demonstrate that combining the relational features obtained by graph contrastive learning with the structural ones extracted by D-MPNN and CNN enhances feature representation ability, thereby improving DTI prediction performance. Our proposed RSGCL-DTI outperforms eight SOTA baseline models on the four benchmark datasets, performs well on the imbalanced dataset, and also shows excellent generalization ability on unseen drug-protein pairs. Jilong Bian, Limin Wei, Yang Li 0130, Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2025 | MolPrompt: improving multi-modal molecular pre-training with knowledge promptsabstractMOTIVATION: Molecular pre-training has emerged as a foundational approach in computational drug discovery, enabling the extraction of expressive molecular representations from large-scale unlabeled datasets. However, existing methods largely focus on topological or structural features, often neglecting critical physicochemical attributes embedded in molecular systems. RESULT: We present MolPrompt, a knowledge-enhanced multimodal pre-training framework that integrates molecular graphs and textual descriptions via contrastive learning. MolPrompt employs a dual-encoder architecture consisting of Graphormer for graph encoding and BERT for textual encoding, and introduces knowledge prompts, semantic embeddings constructed by converting molecular descriptors into natural language, into the graph encoder to guide structure-aware representation learning. Across tasks including molecular property prediction, toxicity estimation, cross-modal retrieval, and anticancer inhibitor identification, MolPrompt consistently surpasses state-of-the-art baselines. These results highlight the value of embedding domain knowledge into structural learning to improve the depth, interpretability, and transferability of molecular representations. AVAILABILITY AND IMPLEMENTATION: The source code of MolPrompt is available at: https://github.com/catly/MolPrompt. Yang Li 0130, Chang Liu 0082, Xin Gao 0001, Guohua Wang 0001 |
Bioinform. | 1 |
| 2025 | PLiSAGE: enhancing protein-ligand interaction prediction with multimodal surface and geometry encodingabstractMOTIVATION: Accurately predicting protein-ligand interactions is fundamental to elucidating molecular recognition and has far-reaching implications in drug discovery, gene regulation, and signal transduction. Conventional methods predominantly rely on internal structural or sequence-based protein representations. While these approaches have improved predictive performance, their dependence on limited labeled data restricts the capacity to learn expressive features from structural inputs. Moreover, they often neglect the intricate geometric and chemical context encoded on protein surfaces, limiting interpretability, and hindering mechanistic insights into binding interactions. RESULT: Here, we present PLiSAGE, a multimodal framework that integrates 3D structural and surface geometric embeddings to enable accurate prediction of protein-ligand interactions. Central to our approach is the joint pretraining of structural and surface encoders through unsupervised contrastive learning and point cloud reconstruction. Protein surfaces are represented as segmented point cloud patches, allowing the model to capture fine-grained geometric and chemical cues. A Transformer-based encoder further captures both local and global spatial dependencies across patches. The incorporation of spatial topological information during pretraining facilitates the learning of stable, discriminative, and multi-scale protein representations, enhancing the expressive capacity of both modalities. An adaptive fusion module dynamically integrates structural and surface embeddings to yield complete and robust protein representations. PLiSAGE demonstrates superior performance over competitive baselines in binding affinity prediction and interaction classification tasks. Extensive ablation studies underscore the critical contributions of surface features and the pretraining strategy to the model's generalization capabilities. AVAILABILITY AND IMPLEMENTATION: The source code of PLiSAGE is available at: https://github.com/catly/PLiSAGE. Guanyu Qiao, Guohua Wang 0001, Yang Li 0130 |
Bioinform. | 4 |
| 2025 | Fine-grained multimodal molecular pretraining via prompt learning
Yang Li 0130, Zhengxin Wei, Chang Liu 0082, Guohua Wang 0001 |
Knowl. Based Syst. | 1 |
| 2024 | Diffusion Based Counterfactual Augmentation for Dual Sentiment ClassificationabstractState-of-the-art NLP models have demonstrated exceptional performance across various tasks, including sentiment analysis. However, concerns have been raised about their robustness and susceptibility to systematic biases in both training and test data, which may lead to performance challenges when these models encounter out-of-distribution data in real-world applications. Although various data augmentation and adversarial perturbation techniques have shown promise in tackling these issues, prior methods such as word embedding perturbation or synonymous sentence expansion have failed to mitigate the spurious association problem inherent in the original data. Recent counterfactual augmentation methods have attempted to tackle this issue, but they have been limited by rigid rules, resulting in inconsistent context and disrupted semantics. In response to these challenges, we introduce a diffusion-based counterfactual data augmentation (DCA) framework. It utilizes an antonymous paradigm to guide the continuous diffusion model and employs reinforcement learning in combination with contrastive learning to optimize algorithms for generating counterfactual samples with high diversity and quality. Furthermore, we use a dual sentiment classifier to validate the generated antonymous samples and subsequently perform sentiment classification. Our experiments on four benchmark datasets demonstrate that DCA achieves state-of-the-art performance in sentiment classification tasks. Dancheng Xin, Yang Li 0130 |
LREC/COLING | 3 |
| 2024 | Causal enhanced drug-target interaction prediction based on graph generation and multi-source information fusionabstractMOTIVATION: The prediction of drug-target interaction is a vital task in the biomedical field, aiding in the discovery of potential molecular targets of drugs and the development of targeted therapy methods with higher efficacy and fewer side effects. Although there are various methods for drug-target interaction (DTI) prediction based on heterogeneous information networks, these methods face challenges in capturing the fundamental interaction between drugs and targets and ensuring the interpretability of the model. Moreover, they need to construct meta-paths artificially or a lot of feature engineering (prior knowledge), and graph generation can fuse information more flexibly without meta-path selection. RESULTS: We propose a causal enhanced method for drug-target interaction (CE-DTI) prediction that integrates graph generation and multi-source information fusion. First, we represent drugs and targets by modeling the fusion of their multi-source information through automatic graph generation. Once drugs and targets are combined, a network of drug-target pairs is constructed, transforming the prediction of drug-target interactions into a node classification problem. Specifically, the influence of surrounding nodes on the central node is separated into two groups: causal and non-causal variable nodes. Causal variable nodes significantly impact the central node's classification, while non-causal variable nodes do not. Causal invariance is then used to enhance the contrastive learning of the drug-target pairs network. Our method demonstrates excellent performance compared with other competitive benchmark methods across multiple datasets. At the same time, the experimental results also show that the causal enhancement strategy can explore the potential causal effects between DTPs, and discover new potential targets. Additionally, case studies demonstrate that this method can identify potential drug targets. AVAILABILITY AND IMPLEMENTATION: The source code of AdaDR is available at: https://github.com/catly/CE-DTI. Guanyu Qiao, Guohua Wang 0001, Yang Li 0130 |
Bioinform. | 3 |
| 2024 | Improving ncRNA family prediction using multi-modal contrastive learning of sequence and structureabstractMOTIVATION: Recent advancements in high-throughput sequencing technology have significantly increased the focus on non-coding RNA (ncRNA) research within the life sciences. Despite this, the functions of many ncRNAs remain poorly understood. Research suggests that ncRNAs within the same family typically share similar functions, underlining the importance of understanding their roles. There are two primary methods for predicting ncRNA families: biological and computational. Traditional biological methods are not suitable for large-scale data prediction due to the significant human and resource requirements. Concurrently, most existing computational methods either rely solely on ncRNA sequence data or are exclusively based on the secondary structure of ncRNA molecules. These methods fail to fully utilize the rich multimodal information available from ncRNAs, thereby preventing them from learning more comprehensive and in-depth feature representations. RESULTS: To tackle these problems, we proposed MM-ncRNAFP, a multi-modal contrastive learning framework for ncRNA family prediction. We first used a pre-trained language model to encode the primary sequences of a large mammalian ncRNA dataset. Then, we adopted a contrastive learning framework with an attention mechanism to fuse the secondary structure information obtained by graph neural networks. The MM-ncRNAFP method can effectively fuse multi-modal information. Experimental comparisons with several competitive baselines demonstrated that MM-ncRNAFP can achieve more comprehensive representations of ncRNA features by integrating both sequence and structural information. This integration significantly enhances the performance of ncRNA family prediction. Ablation experiments and qualitative analyses were performed to verify the effectiveness of each component in our model. Moreover, since our model is pre-trained on a large amount of ncRNA data, it has the potential to bring significant improvements to other ncRNA-related tasks. AVAILABILITY AND IMPLEMENTATION: MM-ncRNAFP and the datasets are available at https://github.com/xuruiting2/MM-ncRNAFP. Ruiting Xu, Guohua Wang 0001, Yang Li 0130 |
Bioinform. | 5 |
| 2024 | DAPM-CDR: A domain adaptation prompting model for drug response prediction
Youhan Sun, Guanyu Qiao, Yang Li 0130 |
Future Gener. Comput. Syst. | 4 |
| 2023 | Graph Pruning and Representation Learning for Stance DetectionabstractStance detection has attracted wide attention, especially on social media platforms, which is particularly challenging. It can be regarded as a short text classification task, which aims to identify the author's stance (Favor, Against, or None) expressed within the text towards a specific target. Most of the existing research based on the pre-training language model ignores global word co-occurrence with discontinuous and long-distance semantics in the corpus itself. To overcome such drawbacks, we propose a graph neural network framework that integrates semantic features and structural features to improve the performance of stance detection. Through the fine-tuning of the pre-training model on the task of stance detection, and graph pruning and sampling in graph pre-training model, the representation of text nodes takes into account both global structure and semantic information. Experimental results on the public dataset SemEval-2016 show that our model outperforms other baseline methods. In addition, the ablation experiments demonstrates the effectiveness of each part of our proposed method. Yang Li 0130 |
IJCNN | 2 |
| 2023 | End-to-end interpretable disease-gene association predictionabstractIdentifying disease-gene associations is a fundamental and critical biomedical task towards understanding molecular mechanisms, the diagnosis and treatment of diseases. It is time-consuming and expensive to experimentally verify causal links between diseases and genes. Recently, deep learning methods have achieved tremendous success in identifying candidate genes for genetic diseases. The gene prediction problem can be modeled as a link prediction problem based on the features of nodes and edges of the gene-disease graph. However, most existing researches either build homogeneous networks based on one single data source or heterogeneous networks based on multi-source data, and artificially define meta-paths, so as to learn the network representation of diseases and genes. The former cannot make use of abundant multi-source heterogeneous information, while the latter needs domain knowledge and experience when defining meta-paths, and the accuracy of the model largely depends on the definition of meta-paths. To address the aforementioned challenges above bottlenecks, we propose an end-to-end disease-gene association prediction model with parallel graph transformer network (DGP-PGTN), which deeply integrates the heterogeneous information of diseases, genes, ontologies and phenotypes. DGP-PGTN can automatically and comprehensively capture the multiple latent interactions between diseases and genes, discover the causal relationship between them and is fully interpretable at the same time. We conduct comprehensive experiments and show that DGP-PGTN outperforms the state-of-the-art methods significantly on the task of disease-gene association prediction. Furthermore, DGP-PGTN can automatically learn the implicit relationship between diseases and genes without manually defining meta paths. Yang Li 0130, Zihou Guo, Xin Gao 0001, Guohua Wang 0001 |
Briefings Bioinform. | 1 |
| 2023 | MMCL-CDR: enhancing cancer drug response prediction with multi-omics and morphology images contrastive representation learningabstractMOTIVATION: Cancer is a complex disease that results in a significant number of global fatalities. Treatment strategies can vary among patients, even if they have the same type of cancer. The application of precision medicine in cancer shows promise for treating different types of cancer, reducing healthcare expenses, and improving recovery rates. To achieve personalized cancer treatment, machine learning models have been developed to predict drug responses based on tumor and drug characteristics. However, current studies either focus on constructing homogeneous networks from single data source or heterogeneous networks from multiomics data. While multiomics data have shown potential in predicting drug responses in cancer cell lines, there is still a lack of research that effectively utilizes insights from different modalities. Furthermore, effectively utilizing the multimodal knowledge of cancer cell lines poses a challenge due to the heterogeneity inherent in these modalities. RESULTS: To address these challenges, we introduce MMCL-CDR (Multimodal Contrastive Learning for Cancer Drug Responses), a multimodal approach for cancer drug response prediction that integrates copy number variation, gene expression, morphology images of cell lines, and chemical structure of drugs. The objective of MMCL-CDR is to align cancer cell lines across different data modalities by learning cell line representations from omic and image data, and combined with structural drug representations to enhance the prediction of cancer drug responses (CDR). We have carried out comprehensive experiments and show that our model significantly outperforms other state-of-the-art methods in CDR prediction. The experimental results also prove that the model can learn more accurate cell line representation by integrating multiomics and morphological data from cell lines, thereby improving the accuracy of CDR prediction. In addition, the ablation study and qualitative analysis also confirm the effectiveness of each part of our proposed model. Last but not least, MMCL-CDR opens up a new dimension for cancer drug response prediction through multimodal contrastive learning, pioneering a novel approach that integrates multiomics and multimodal drug and cell line modeling. AVAILABILITY AND IMPLEMENTATION: MMCL-CDR is available at https://github.com/catly/MMCL-CDR. Yang Li 0130, Zihou Guo, Xin Gao 0001, Guohua Wang 0001 |
Bioinform. | 1 |
| 2023 | KEIC: A tag recommendation framework with knowledge enhancement and interclass correlation
Yang Li 0130, Weipeng Jing 0001 |
Inf. Sci. | 2 |
| 2022 | Generative Data Augmentation with Contrastive Learning for Zero-Shot Stance DetectionabstractStance detection aims to identify whether the author of an opinionated text is in favor of, against, or neutral towards a given target.Remarkable success has been achieved when sufficient labeled training data is available.However, it is labor-intensive to annotate sufficient data and train the model for every new target.Therefore, zero-shot stance detection, aiming at identifying stances of unseen targets with seen targets, has gradually attracted attention.Among them, one of the important challenges is to reduce the domain transfer between seen and unseen targets.To tackle this problem, we propose a generative data augmentation approach to generate training samples containing targets and stances for testing data, and map the real samples and generated synthetic samples into the same embedding space with contrastive learning, then perform the final classification based on the augmented data.We evaluate our proposed model on two benchmark datasets.Experimental results show that our approach achieves state of-the-art performance on most topics in the task of zero-shot stance detection. Yang Li 0130 |
EMNLP | 1 |
| 2022 | Drug-target interaction predication via multi-channel graph neural networksabstractDrug-target interaction (DTI) is an important step in drug discovery. Although there are many methods for predicting drug targets, these methods have limitations in using discrete or manual feature representations. In recent years, deep learning methods have been used to predict DTIs to improve these defects. However, most of the existing deep learning methods lack the fusion of topological structure and semantic information in DPP representation learning process. Besides, when learning the DPP node representation in the DPP network, the different influences between neighboring nodes are ignored. In this paper, a new model DTI-MGNN based on multi-channel graph convolutional network and graph attention is proposed for DTI prediction. We use two independent graph attention networks to learn the different interactions between nodes for the topology graph and feature graph with different strengths. At the same time, we use a graph convolutional network with shared weight matrices to learn the common information of the two graphs. The DTI-MGNN model combines topological structure and semantic features to improve the representation learning ability of DPPs, and obtain the state-of-the-art results on public datasets. Specifically, DTI-MGNN has achieved a high accuracy in identifying DTIs (the area under the receiver operating characteristic curve is 0.9665). Yang Li 0130, Guanyu Qiao, Guohua Wang 0001 |
Briefings Bioinform. | 1 |
| 2022 | Supervised graph co-contrastive learning for drug-target interaction predictionabstractMOTIVATION: Identification of Drug-Target Interactions (DTIs) is an essential step in drug discovery and repositioning. DTI prediction based on biological experiments is time-consuming and expensive. In recent years, graph learning-based methods have aroused widespread interest and shown certain advantages on this task, where the DTI prediction is often modeled as a binary classification problem of the nodes composed of drug and protein pairs (DPPs). Nevertheless, in many real applications, labeled data are very limited and expensive to obtain. With only a few thousand labeled data, models could hardly recognize comprehensive patterns of DPP node representations, and are unable to capture enough commonsense knowledge, which is required in DTI prediction. Supervised contrastive learning gives an aligned representation of DPP node representations with the same class label. In embedding space, DPP node representations with the same label are pulled together, and those with different labels are pushed apart. RESULTS: We propose an end-to-end supervised graph co-contrastive learning model for DTI prediction directly from heterogeneous networks. By contrasting the topology structures and semantic features of the drug-protein-pair network, as well as the new selection strategy of positive and negative samples, SGCL-DTI generates a contrastive loss to guide the model optimization in a supervised manner. Comprehensive experiments on three public datasets demonstrate that our model outperforms the SOTA methods significantly on the task of DTI prediction, especially in the case of cold start. Furthermore, SGCL-DTI provides a new research perspective of contrastive learning for DTI prediction. AVAILABILITY AND IMPLEMENTATION: The research shows that this method has certain applicability in the discovery of drugs, the identification of drug-target pairs and so on. Yang Li 0130, Guanyu Qiao, Xin Gao 0001, Guohua Wang 0001 |
Bioinform. | 1 |
| 2021 | Evaluating disease similarity based on gene network reconstruction and representationabstractMOTIVATION: Quantifying the associations between diseases is of great significance in increasing our understanding of disease biology, improving disease diagnosis, re-positioning and developing drugs. Therefore, in recent years, the research of disease similarity has received a lot of attention in the field of bioinformatics. Previous work has shown that the combination of the ontology (such as disease ontology and gene ontology) and disease-gene interactions are worthy to be regarded to elucidate diseases and disease associations. However, most of them are either based on the overlap between disease-related gene sets or distance within the ontology's hierarchy. The diseases in these methods are represented by discrete or sparse feature vectors, which cannot grasp the deep semantic information of diseases. Recently, deep representation learning has been widely studied and gradually applied to various fields of bioinformatics. Based on the hypothesis that disease representation depends on its related gene representations, we propose a disease representation model using two most representative gene resources HumanNet and Gene Ontology to construct a new gene network and learn gene (disease) representations. The similarity between two diseases is computed by the cosine similarity of their corresponding representations. RESULTS: We propose a novel approach to compute disease similarity, which integrates two important factors disease-related genes and gene ontology hierarchy to learn disease representation based on deep representation learning. Under the same experimental settings, the AUC value of our method is 0.8074, which improves the most competitive baseline method by 10.1%. The quantitative and qualitative experimental results show that our model can learn effective disease representations and improve the accuracy of disease similarity computation significantly. AVAILABILITY AND IMPLEMENTATION: The research shows that this method has certain applicability in the prediction of gene-related diseases, the migration of disease treatment methods, drug development and so on. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Li 0130, Guohua Wang 0001 |
Bioinform. | 1 |
| 2019 | Topical Co-Attention Networks for hashtag recommendation on microblogs
Yang Li 0130, Ting Liu 0001, Jingwen Hu 0002, Jing Jiang 0001 |
Neurocomputing | 1 |
| 2017 | Personalized Microtopic Recommendation on MicroblogsabstractMicroblogging services such as Sina Weibo and Twitter allow users to create tags explicitly indicated by the # symbol. In Sina Weibo, these tags are called microtopics , and in Twitter, they are called hashtags . In Sina Weibo, each microtopic has a designate page and can be directly visited or commented on. Recommending these microtopics to users based on their interests can help users efficiently acquire information. However, it is non-trivial to recommend microtopics to users to satisfy their information needs. In this article, we investigate the task of personalized microtopic recommendation, which exhibits two challenges. First, users usually do not give explicit ratings to microtopics. Second, there exists rich information about users and microtopics, for example, users' published content and biographical information, but it is not clear how to best utilize such information. To address the above two challenges, we propose a joint probabilistic latent factor model to integrate rich information into a matrix factorization-based solution to microtopic recommendation. Our model builds on top of collaborative filtering, content analysis, and feature regression. Using two real-world datasets, we evaluate our model with different kinds of content and contextual information. Experimental results show that our model significantly outperforms a few competitive baseline methods, especially in the circumstance where users have few adoption behaviors. Yang Li 0130, Jing Jiang 0001, Ting Liu 0001, Minghui Qiu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2016 | Hashtag Recommendation with Topical Attention-Based LSTMabstractMicroblogging services allow users to create hashtags to categorize their posts. In recent years, the task of recommending hashtags for microblogs has been given increasing attention. However, most of existing methods depend on hand-crafted features. Motivated by the successful use of long short-term memory (LSTM) for many natural language processing tasks, in this paper, we adopt LSTM to learn the representation of a microblog post. Observing that hashtags indicate the primary topics of microblog posts, we propose a novel attention-based LSTM model which incorporates topic modeling into the LSTM architecture through an attention mechanism. We evaluate our model using a large real-world dataset. Experimental results show that our model significantly outperforms various competitive baseline methods. Furthermore, the incorporation of topical attention mechanism gives more than 7.4% improvement in F1 score compared with standard LSTM method. Yang Li 0130, Ting Liu 0001, Jing Jiang 0001 |
COLING | 1 |