Ning Yu 0004

dblp:24/3024-4 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-1385-6882ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Tissue-Aware Prototype Learning Model for Predicting Anticancer Drug Response in Patients
abstract
Cancer treatments often yield different results from patient to patient due to the genomic heterogeneity of tumors. Accurately predicting a patient's response to an anticancer drug is challenging, especially when using traditional machine learning models trained on cell lines and applied to patient data. These models struggle with domain shift (out-of-distribution data due to differences between cell line and patient data), loss of tissue specificity, and data imbalance across domains. To address these issues, we developed the Tissue-Aware and Prototype Learning for Drug Response Prediction (TAPL-DRP) model. This model predicts anticancer drug responses using a two-stage process. At the stage of tissue-aware cross-domain feature extraction, we integrate a Variational Autoencoder (VAE) and a Generative Adversarial Network (GAN) to extract features from both cell line and patient data. This process incorporates tissue prototypes, InfoNCE loss, and class-balance loss to ensure the features are not only domain-invariant but also biologically meaningful and tissue-specific. At the drug response prediction stage, the model combines these extracted features with molecular drug graph features. It then uses the tissue prototypes to guide the training of the classifier, which further improves the prediction accuracy. We tested TAPL-DRP on the TCGA clinical dataset and the PDTC in vitro dataset. The results show that our model significantly outperforms other methods in key metrics like AUC, AUPRC, ACC, and MCC. This demonstrates its effectiveness in handling domain shift and maintaining tissue specificity. Further analysis confirmed that the tissue prototypes, mutual information loss, and class-balance strategies are all crucial components of the model. In summary, TAPL-DRP offers an effective and precise solution for predicting anticancer drug responses in personalized medicine.The source code is available at https://github.com/weiba/TAPL-DRP.
Wei Peng 0004, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Ning Yu 0004
BIBM7
2025 Predicting Anti-Cancer Drug Response Based on Hypergraph Representation Learning
abstract
Accurate prediction of drug responses is critical for advancing personalized cancer therapies. Although current graph neural network (GNN)-based approaches predominantly focus on pairwise interactions between cell lines and drugs, they often neglect the potential of higher-order interactions. In this study, we present HRLCDR, a novel computational framework that utilizes Hypergraph Representation Learning to predict Cancer Drug Responses. HRLCDR begins by constructing hypergraphs for both cell lines and drugs and then processes through low-pass and high-pass hypergraph convolutions, allowing the model to extract both common and different features from the complex higher-order interactions between cell lines and drugs. After that, HRLCDR constructs a heterogeneous graph using known cell line responses to drugs. Parallel heterogeneous graph convolution operations are then employed to extract primary interaction features between cell lines and drugs from these associations. Finally, HRLCDR integrates the features learned from both the hypergraphs and the heterogeneous graph, predicting drug response via Classifiers. We evaluated HRLCDR's performance on two major cancer drug response datasets: the Cancer Drug Sensitivity Data (GDSC) and the Cancer Cell Line Encyclopedia (CCLE). The results demonstrate that HRLCDR outperforms current state-of-the-art methods, underscoring its potential to enhance the accuracy and reliability of cancer drug response predictions.
Wei Peng 0004, Jiangzhen Lin, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Ning Yu 0004
IEEE Trans. Comput. Biol. Bioinform.9
2025 Predicting Clinical Anticancer Drug Response of Patients by Using Domain Alignment and Prototypical Learning
abstract
Anticancer drug response prediction is crucial in developing personalized treatment plans for cancer patients. However, High-quality patient anticancer drug response data are scarce and cell line data and patient data have different distributions, models trained solely on cell line data perform poorly. Some existing methods predict anticancer drug response by transferring knowledge from the cell line domain to the patient domain using transfer learning. However, the robustness of these classifiers is affected by anomalies in the cell line data, and they do not utilize the knowledge in the unlabeled target domain data. To this end, we proposed a model called DAPL to predict patient responses to anticancer drugs. The model extracts domain-invariant features from cell lines and patients by constructing multiple VAEs and extracts drug features using GNNs. These features are then combined for prototypical learning to train a classifier, resulting in better predictions of patient anticancer drug response. We used the cell line datasets CCLE and GDSC as source domains and the patient datasets TCGA and PDTC as target domains and conducted experiments. The results indicate that DAPL shows excellent performance in predicting patient anticancer drug response compared to other state-of-the-art methods.
Wei Peng 0004, Chuyue Chen, Wei Dai 0012, Ning Yu 0004, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics4
2025 Hierarchical Graph Representation Learning With Multi-Granularity Features for Anti-Cancer Drug Response Prediction
abstract
Patients with the same type of cancer often respond differently to identical drug treatments due to unique genomic traits. Accurately predicting a patient's response to drug is crucial in guiding treatment decisions, alleviating patient suffering, and improving cancer prognosis. Current computational methods utilize deep learning models trained on extensive drug screening data to predict anti-cancer drug responses based on features of cell lines and drugs. However, the interaction between cell lines and drugs is a complex biological process involving interactions across various levels, from internal cellular and drug structures to the external interactions among different molecules.To address this complexity, we propose a novel Hierarchical graph representation Learning with Multi-Granularity features (HLMG) algorithm for predicting anti-cancer drug responses. The HLMG algorithm combines features at two granularities: the overall gene expression and pathway substructures of cell lines, and the overall molecular fingerprints and substructures of drugs. Subsequently, it constructs a heterogeneous graph including cell lines, drugs, known cell line-drug responses, and the associations between similar cell lines and similar drugs. Through a graph convolutional network model, the HLMG learns the final cell line and drug representations by aggregating features of their multi-level neighbor in the heterogeneous graph. The multi-level neighbors consist of the node self, directly related drugs/cell lines, and indirectly related similar drugs/cell lines. Finally, a linear correlation coefficient decoder is employed to reconstruct the cell line-drug correlation matrix to predict anti-cancer drug responses. Our model was tested on the Genomics of Drug Sensitivity in Cancer (GDSC) and the Cancer Cell Line Encyclopedia (CCLE) databases. Results indicate that HLMG outperforms other state-of-the-art methods in accurately predicting anti-cancer drug responses.
Wei Peng 0004, Jiangzhen Lin, Wei Dai 0012, Ning Yu 0004, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics4
2024 Sparse Attention-based Hierarchical Node Representation for Spatial Domain Identification
abstract
Using deep learning models on spatial transcriptomics data to identify the spatial domain is crucial for uncovering the spatial distribution of cells and gene expression patterns within tissues, essential for understanding complex biological processes and disease mechanisms. Existing methods for spatial domain partitioning often rely on predefined adjacency relationships at a single scale, overlooking the hierarchical structure and functional characteristics of biological tissues. In this paper, we propose SpaNFM, a novel method that leverages sparse attention-based hierarchical node representation and multi-view contrastive learning for spatial domain identification in spatial transcriptomics data. The SpaNFM first treats each spot as a node and constructs two views using different data augmentation techniques based on tissue image information, gene expression profiles, and spatial coordinates of cells. Subsequently, SpaNFM utilizes a sparse attention-based hierarchical node fusion module to generate coarse-grained node representations. This fine-to-coarse hierarchical structure integrates complementary information from multi-granularity node features and reduces model complexity due to the decreased node size. The model parameters are updated using gene expression reconstruction loss and contrastive loss on the coarse-grained node representations from the two views. Finally, the learned node features are subjected to downstream clustering using the Leiden algorithm. We tested SpaNFM on the human dorsolateral prefrontal cortex dataset. The results demonstrate that SpaNFM outperforms other state-of-the-art methods in most cases. The data and code are available at: https://github.com/weiba/SpaNFM
Wei Peng 0004, Zhihao Ping, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Ning Yu 0004
BIBM7
2024 A Dual-Approach Framework for Enhancing Network Traffic Analysis (DAFENTA): Leveraging NumeroLogic LLM Embeddings and Transformer Models for Intrusion Detection
abstract
In cybersecurity, network traffic analysis is essential for identifying abnormal patterns that may indicate cyberattacks, often beyond the capabilities of human detection. While Machine Learning (ML) has proven effective in this domain, traditional ML algorithms face significant challenges when dealing with large-scale data and class imbalances, which are commonly found in network traffic logs. These challenges can compromise model accuracy and reliability. To address them, we propose a novel Dual-Approach Framework for Enhancing Network Traffic Analysis (DAFENTA): integrating Large Language Model (LLM) embeddings with NumeroLogic encoding to enhance feature representation and employing a Transformer encoder model for direct binary classification of network traffic. By embedding network traffic data using the sentence-transformers model, we improved feature contextualization for typical ML classifiers like Random Forest, AdaBoost, Gradient Boost, Extra Trees, Logistic Regression, and KNN. In addition, the numerical reasoning is enhanced in LLM by applying the NumeroLogic encoding. At the same time, we leveraged Transformer models to capture complex feature dependencies in data through attention mechanisms. This framework was evaluated through the KDD and ISCX traffic flow datasets. Our results have shown that LLM embeddings, when integrated with NumeroLogic encoding method, significantly enhanced the performance of traditional ML models by improving feature generalization, especially in handling out-of-distribution samples. Additionally, the Transformer model demonstrated efficiency in managing large-scale data and addressing class imbalance, leading to marked improvements in accuracy. This dual-approach framework offers a powerful method for advancing anomaly detection in network traffic analysis, leveraging the capabilities of Large Language Models.
Ning Yu 0004, Liam Davies
IEEE Big Data1
2024 LGCDA: Predicting CircRNA-Disease Association Based on Fusion of Local and Global Features
abstract
CircRNA has been shown to be involved in the occurrence of many diseases. Several computational frameworks have been proposed to identify circRNA-disease associations. Despite the existing computational methods have obtained considerable successes, these methods still require to be improved as their performance may degrade due to the sparsity of the data and the problem of memory overflow. We develop a novel computational framework called LGCDA to predict circRNA-disease associations by fusing local and global features to solve the above mentioned problems. First, we construct closed local subgraphs by using k-hop closed subgraph and label the subgraphs to obtain rich graph pattern information. Then, the local features are extracted by using graph neural network (GNN). In addition, we fuse Gaussian interaction profile (GIP) kernel and cosine similarity to obtain global features. Finally, the score of circRNA-disease associations is predicted by using the multilayer perceptron (MLP) based on local and global features. We perform five-fold cross validation on five datasets for model evaluation and our model surpasses other advanced methods.
Wei Lan 0001, Qingfeng Chen, Ning Yu 0004, Yi Pan 0001, Yu Zheng 0013, Yi-Ping Phoebe Chen
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 A multi-view comparative learning method for spatial transcriptomics data clustering
abstract
Clustering individual cells or spots based on their gene expression profiles in a spatial context is a powerful approach to uncovering the underlying biological diversity and relationships among cells. The intricate information within spatial transcriptomics data demands sophisticated algorithms that effectively integrate gene expression, cell position, and tissue image data for accurate cell or spot clustering. This work proposes a Multi-View Comparative Learning method for clustering Spatial Transcriptomics data (MVCLST). MVCLST first builds on two data views using gene expression profiles, cell space coordinates, and image features. Then it employs four different encoders to capture the common and private features of the two views. The model employs a contrastive learning loss to encourage effective interaction between the two views and ensure feature consistency. The shared and private features from both views are fused using corresponding decoders. Finally, the model employs the Leiden algorithm for downstream clustering of the learned features. We test the MVCLST method on a human dorsolateral prefrontal cortex dataset. The results show that MVCLST outperforms other state-of-the-art methods in most cases. Additionally, the clusters identified by MVCLST align closely with manual annotations and established neuroscience definitions.
Wei Peng 0004, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Ning Yu 0004
BIBM7
2023 Identifying cancer driver genes based on multi-view heterogeneous graph convolutional network and self-attention mechanism
abstract
BACKGROUND: Correctly identifying the driver genes that promote cell growth can significantly assist drug design, cancer diagnosis and treatment. The recent large-scale cancer genomics projects have revealed multi-omics data from thousands of cancer patients, which requires to design effective models to unlock the hidden knowledge within the valuable data and discover cancer drivers contributing to tumorigenesis. RESULTS: In this work, we propose a graph convolution network-based method called MRNGCN that integrates multiple gene relationship networks to identify cancer driver genes. First, we constructed three gene relationship networks, including the gene-gene, gene-outlying gene and gene-miRNA networks. Then, genes learnt feature presentations from the three networks through three sharing-parameter heterogeneous graph convolution network (HGCN) models with the self-attention mechanism. After that, these gene features pass a convolution layer to generate fused features. Finally, we utilized the fused features and the original feature to optimize the model by minimizing the node and link prediction losses. Meanwhile, we combined the fused features, the original features and the three features learned from every network through a logistic regression model to predict cancer driver genes. CONCLUSIONS: We applied the MRNGCN to predict pan-cancer and cancer type-specific driver genes. Experimental results show that our model performs well in terms of the area under the ROC curve (AUC) and the area under the precision-recall curve (AUPRC) compared to state-of-the-art methods. Ablation experimental results show that our model successfully improved the cancer driver identification by integrating multiple gene relationship networks.
Wei Peng 0004, Wei Dai 0012, Ning Yu 0004
BMC Bioinform.4
2022 Predicting cancer drug response using parallel heterogeneous graph convolutional networks with neighborhood interactions
abstract
MOTIVATION: Due to cancer heterogeneity, the therapeutic effect may not be the same when a cohort of patients of the same cancer type receive the same treatment. The anticancer drug response prediction may help develop personalized therapy regimens to increase survival and reduce patients' expenses. Recently, graph neural network-based methods have aroused widespread interest and achieved impressive results on the drug response prediction task. However, most of them apply graph convolution to process cell line-drug bipartite graphs while ignoring the intrinsic differences between cell lines and drug nodes. Moreover, most of these methods aggregate node-wise neighbor features but fail to consider the element-wise interaction between cell lines and drugs. RESULTS: This work proposes a neighborhood interaction (NI)-based heterogeneous graph convolution network method, namely NIHGCN, for anticancer drug response prediction in an end-to-end way. Firstly, it constructs a heterogeneous network consisting of drugs, cell lines and the known drug response information. Cell line gene expression and drug molecular fingerprints are linearly transformed and input as node attributes into an interaction model. The interaction module consists of a parallel graph convolution network layer and a NI layer, which aggregates node-level features from their neighbors through graph convolution operation and considers the element-level of interactions with their neighbors in the NI layer. Finally, the drug response predictions are made by calculating the linear correlation coefficients of feature representations of cell lines and drugs. We have conducted extensive experiments to assess the effectiveness of our model on Cancer Drug Sensitivity Data (GDSC) and Cancer Cell Line Encyclopedia (CCLE) datasets. It has achieved the best performance compared with the state-of-the-art algorithms, especially in predicting drug responses for new cell lines, new drugs and targeted drugs. Furthermore, our model that was well trained on the GDSC dataset can be successfully applied to predict samples of PDX and TCGA, which verified the transferability of our model from cell line in vitro to the datasets in vivo. AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/weiba/NIHGCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wei Peng 0004, Hancheng Liu, Wei Dai 0012, Ning Yu 0004, Jianxin Wang 0001
Bioinform.4
2021 A comprehensive survey on computational methods of non-coding RNA and disease association prediction
abstract
The studies on relationships between non-coding RNAs and diseases are widely carried out in recent years. A large number of experimental methods and technologies of producing biological data have also been developed. However, due to their high labor cost and production time, nowadays, calculation-based methods, especially machine learning and deep learning methods, have received a lot of attention and been used commonly to solve these problems. From a computational point of view, this survey mainly introduces three common non-coding RNAs, i.e. miRNAs, lncRNAs and circRNAs, and the related computational methods for predicting their association with diseases. First, the mainstream databases of above three non-coding RNAs are introduced in detail. Then, we present several methods for RNA similarity and disease similarity calculations. Later, we investigate ncRNA-disease prediction methods in details and classify these methods into five types: network propagating, recommend system, matrix completion, machine learning and deep learning. Furthermore, we provide a summary of the applications of these five types of computational methods in predicting the associations between diseases and miRNAs, lncRNAs and circRNAs, respectively. Finally, the advantages and limitations of various methods are identified, and future researches and challenges are also discussed.
Xiujuan Lei, Thosini Bamunu Mudiyanselage, Yuchen Zhang 0003, Chen Bian, Wei Lan 0001, Ning Yu 0004, Yi Pan 0001
Briefings Bioinform.6
2019 Reconstruction of Hidden Representation for Robust Feature Extraction
abstract
This article aims to develop a new and robust approach to feature representation. Motivated by the success of Auto-Encoders, we first theoretically analyze and summarize the general properties of all algorithms that are based on traditional Auto-Encoders: (1) The reconstruction error of the input cannot be lower than a lower bound, which can be viewed as a guiding principle for reconstructing the input. Additionally, when the input is corrupted with noises, the reconstruction error of the corrupted input also cannot be lower than a lower bound. (2) The reconstruction of a hidden representation achieving its ideal situation is the necessary condition for the reconstruction of the input to reach the ideal state. (3) Minimizing the Frobenius norm of the Jacobian matrix of the hidden representation has a deficiency and may result in a much worse local optimum value. We believe that minimizing the reconstruction error of the hidden representation is more robust than minimizing the Frobenius norm of the Jacobian matrix of the hidden representation. Based on the above analysis, we propose a new model termedDouble Denoising Auto-Encoders(DDAEs), which uses corruption and reconstruction on both the input and the hidden representation. We demonstrate that the proposed model is highly flexible and extensible and has a potentially better capability to learn invariant and robust feature representations. We also show that our model is more robust than Denoising Auto-Encoders (DAEs) for dealing with noises or inessential features. Furthermore, we detail how to train DDAEs with two different pretraining methods by optimizing the objective function in a combined and separate manner, respectively. Comparative experiments illustrate that the proposed model is significantly better for representation learning than the state-of-the-art models.
Zeng Yu 0001, Tianrui Li 0001, Ning Yu 0004, Yi Pan 0001, Hongmei Chen 0001, Bing Liu 0001
ACM Trans. Intell. Syst. Technol.3
2018 Convolutional networks with cross-layer neurons for image recognition
Zeng Yu 0001, Tianrui Li 0001, Guangchun Luo, Hamido Fujita, Ning Yu 0004, Yi Pan 0001
Inf. Sci.5
2017 Evaluating the Impact of Encoding Schemes on Deep Auto-Encoders for DNA Annotation
Ning Yu 0004, Zeng Yu 0001, Feng Gu 0001, Yi Pan 0001
ISBRA1
2017 A deep learning method for lincRNA detection using auto-encoder algorithm
abstract
BACKGROUND: RNA sequencing technique (RNA-seq) enables scientists to develop novel data-driven methods for discovering more unidentified lincRNAs. Meantime, knowledge-based technologies are experiencing a potential revolution ignited by the new deep learning methods. By scanning the newly found data set from RNA-seq, scientists have found that: (1) the expression of lincRNAs appears to be regulated, that is, the relevance exists along the DNA sequences; (2) lincRNAs contain some conversed patterns/motifs tethered together by non-conserved regions. The two evidences give the reasoning for adopting knowledge-based deep learning methods in lincRNA detection. Similar to coding region transcription, non-coding regions are split at transcriptional sites. However, regulatory RNAs rather than message RNAs are generated. That is, the transcribed RNAs participate the biological process as regulatory units instead of generating proteins. Identifying these transcriptional regions from non-coding regions is the first step towards lincRNA recognition. RESULTS: The auto-encoder method achieves 100% and 92.4% prediction accuracy on transcription sites over the putative data sets. The experimental results also show the excellent performance of predictive deep neural network on the lincRNA data sets compared with support vector machine and traditional neural network. In addition, it is validated through the newly discovered lincRNA data set and one unreported transcription site is found by feeding the whole annotated sequences through the deep learning machine, which indicates that deep learning method has the extensive ability for lincRNA prediction. CONCLUSIONS: The transcriptional sequences of lincRNAs are collected from the annotated human DNA genome data. Subsequently, a two-layer deep neural network is developed for the lincRNA detection, which adopts the auto-encoder algorithm and utilizes different encoding schemes to obtain the best performance over intergenic DNA sequence data. Driven by those newly annotated lincRNA data, deep learning methods based on auto-encoder algorithm can exert their capability in knowledge learning in order to capture the useful features and the information correlation along DNA genome sequences for lincRNA detection. As our knowledge, this is the first application to adopt the deep learning techniques for identifying lincRNA transcription sequences.
Ning Yu 0004, Zeng Yu 0001, Yi Pan 0001
BMC Bioinform.1
2016 J2M: a Java to MapReduce translator for cloud computing
Bing Li 0012, Junbo Zhang 0004, Ning Yu 0004, Yi Pan 0001
J. Supercomput.3
2015 DNA AS X: An Information-Coding-Based Model to Improve the Sensitivity in Comparative Gene Analysis
Ning Yu 0004, Xuan Guo 0004, Feng Gu 0001, Yi Pan 0001
ISBRA1
2014 Cloud computing for detecting high-order genome-wide epistatic interaction via dynamic clustering
abstract
BACKGROUND: Taking the advantage of high-throughput single nucleotide polymorphism (SNP) genotyping technology, large genome-wide association studies (GWASs) have been considered to hold promise for unravelling complex relationships between genotype and phenotype. At present, traditional single-locus-based methods are insufficient to detect interactions consisting of multiple-locus, which are broadly existing in complex traits. In addition, statistic tests for high order epistatic interactions with more than 2 SNPs propose computational and analytical challenges because the computation increases exponentially as the cardinality of SNPs combinations gets larger. RESULTS: In this paper, we provide a simple, fast and powerful method using dynamic clustering and cloud computing to detect genome-wide multi-locus epistatic interactions. We have constructed systematic experiments to compare powers performance against some recently proposed algorithms, including TEAM, SNPRuler, EDCF and BOOST. Furthermore, we have applied our method on two real GWAS datasets, Age-related macular degeneration (AMD) and Rheumatoid arthritis (RA) datasets, where we find some novel potential disease-related genetic factors which are not shown up in detections of 2-loci epistatic interactions. CONCLUSIONS: Experimental results on simulated data demonstrate that our method is more powerful than some recently proposed methods on both two- and three-locus disease models. Our method has discovered many novel high-order associations that are significantly enriched in cases from two real GWAS datasets. Moreover, the running time of the cloud implementation for our method on AMD dataset and RA dataset are roughly 2 hours and 50 hours on a cluster with forty small virtual machines for detecting two-locus interactions, respectively. Therefore, we believe that our method is suitable and effective for the full-scale analysis of multiple-locus epistatic interactions in GWAS.
Xuan Guo 0004, Meng Yu 0001, Ning Yu 0004, Yi Pan 0001
BMC Bioinform.3