Yuanyuan Zhang 0008

dblp:23/6185-8 · also Yuan-Yuan Zhang 0008 · DBLP profile ↗
← Back
29ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0003-3935-3201ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Micro ribonucleic acids-drug sensitivity prediction by variational graph auto-encoder and collaborative matrix factorization
Yunyin Li, Yuanyuan Zhang 0008, Chuanru Ren, Tiyao Liu, Yingye Liu
Eng. Appl. Artif. Intell.3
2026 LoHi-SSL: A Multi-Level Synergistic Learning Model for Integrating Single-Cell Multi-Omics Data via Low- and High-Order Information Fusion
abstract
Recent advancements in single-cell sequencing technologies have enabled researchers to identify cell subpopulations and their functional states with greater accuracy, thereby uncovering cellular heterogeneity. However, due to the heterogeneity across different single-cell multi-omics datasets and the intrinsic variability among cells, effectively integrating data from multiple molecular layers remains a significant challenge. To address this issue, a Single-cell Multi-level Synergistic Learning model (LoHi-SSL) is proposed, which integrates low-order and high-order information to achieve efficient multi-omics data fusion. LoHi-SSL consists of three key modules: low-order information learning, high-order information learning, and feature integration. The low-order information learning module focuses on addressing intra-omics cellular heterogeneity. It first extracts features from each omics dataset, then constructs a graph structure to capture intercellular relationships. A Graph Autoencoder is employed to extract local neighborhood information, effectively preserving intra-omics cellular similarity. The high-order information learning module is designed to eliminate cross-omics heterogeneity and align data in a unified latent representation space. To achieve this, multi-omics hypergraph learning is introduced to model complex cellular relationships across different omics, enhancing feature interactions.In the feature integration module, contrastive learning is utilized to guide the model in learning more discriminative feature representations by constructing positive and negative sample pairs. For different omics data from the same cell, LoHi-SSL encourages feature alignment within a shared latent space, reducing cross-omics heterogeneity. Meanwhile, for different cell types, a contrastive loss function is applied to increase the separation between their representations, thereby enhancing cellular distinguishability and achieving efficient single-cell multi-omics integration. Experimental results demonstrate that LoHi-SSL outperforms existing methods on six publicly available datasets, achieving superior performance in clustering tasks, particularly in terms of NMI (Normalized Mutual Information), ARI (Adjusted Rand Index), AMI (Adjusted Mutual Information), and ACC (Clustering Accuracy). Furthermore, robustness analysis shows that LoHi-SSL exhibits strong resistance to noise. Additionally, cell trajectory analysis using the latent representations learned by LoHi-SSL accurately reflects biological evolutionary pathways. In summary, LoHi-SSL provides an efficient and robust approach for single-cell multi-omics data integration, offering a powerful tool for studying cellular state transitions, heterogeneity, and regulatory mechanisms.
Xiaoyun Xiong, Kaihao Zhang, Chengdong Zhang, Yuanyuan Zhang 0008
IEEE Trans. Comput. Biol. Bioinform.4
2026 Synthetic Assessment of Transfer Learning Method From Cell Lines to Single-Cell on Drug Response Prediction
abstract
Drug response prediction is of critical importance in precision medicine and novel drug development, yet it remains highly challenging due to the complexity, high cost, and low success rate of the process. With the rapid advancement of single-cell sequencing technologies, researchers are now able to gain deeper insights into intratumoral clonal heterogeneity and drug resistance mechanisms, thereby the necessity of drug response prediction is underscored at the single-cell level. However, despite the growing availability of single-cell expression data, corresponding drug sensitivity annotations remain extremely scarce, significantly hindering the development and generalization of supervised learning models. To bridge this data gap, transfer learning has emerged as a key strategy, enabling the migration of drug response knowledge learned from bulk cell line data to single-cell contexts. However, most existing studies focus on model development without conducting systematic, comparative evaluations. To address this gap, this study presents the first comprehensive assessment of representative single-cell drug response prediction models based on transfer learning. We perform a multidimensional analysis of these methods-including transfer mechanisms, feature alignment strategies, and predictive performance-using publicly available datasets for empirical benchmarking. It provides valuable methodological guidance for the future selection and design of predictive models, and establishes a foundation for advancing toward clinically actionable single-cell pharmacogenomics.
Yuanyuan Zhang 0008, Wenying Li, Zhennuo Wang, Shuang Du
IEEE Trans. Comput. Biol. Bioinform.1
2026 A Semantic Conditional Diffusion Model for Enhanced Personal Privacy Preservation in Medical Images
abstract
Deep learning has significantly advanced medical image processing, yet the inherent inclusion of personally identifiable information (PII) within medical images-such as facial features, distinctive anatomical structures, rare lesions, or specific textural patterns-poses a critical risk to patient privacy during data transmission. To mitigate this risk, we introduce the Medical Semantic Diffusion Model (MSDM), a novel framework designed to synthesize medical images guided by semantic information, synthesis images with the same distribution as the original data, which effectively removes the PPI of the original data to ensure robust privacy protection. Unlike conventional techniques that combine semantic and noisy images for denoising, MSDM integrates Adaptive Batch Normalization (AdaBN) to encode semantic information into high-dimensional latent space, embedding it directly within the denoising neural network. This approach enhances image quality and semantic accuracy while ensuring that the synthetic and original images belong to the same distribution. In addition, to further accelerate synthesis and reduce dependency on manually crafted semantic masks, we propose the Spread Algorithm, which automatically generates these masks. Extensive experiments conducted on the BraTS 2021, MSD Lung, DSB18, and FIVES datasets confirm the efficacy of MSDM, yielding state-of-the-art results across several performance metrics. Augmenting datasets with MSDM-generated images in nnUNet segmentation experiments led to Dice scores of 0.6243, 0.9531, 0.9406, and 0.9562 underscoring its potential for enhancing both image quality and privacy-preserving data augmentation.
Zhiyuan Zhao 0003, Yawu Zhao, Yuanyuan Zhang 0008, Jiehuan Wang, Sibo Qiao, Zhihan Lyu
IEEE J. Biomed. Health Informatics5
2025 DBAANet: Dual-Branch Attention Aggregation Network for Medical Image Segmentation
abstract
Recent hybrid architectures combining Transformer encoders with U-Net have advanced medical image segmentation by modeling global dependencies, but they often suffer from semantic misalignment between local CNN features and global Transformer representations, leading to inefficient multi-scale fusion and boundary detail loss in resourceconstrained clinical settings. To overcome these limitations, we propose DBAANet, a novel and efficient architecture featuring a dual-branch encoder that synergistically combines an enhanced Vision Transformer (ViT) with multi-scale convolutional blocks to optimize feature extraction. In the decoder, we employ channel and spatial attention mechanisms to capture complex inter-feature relationships and introduce a gated attention mechanism to fuse multi-source features from different stages, thereby fully leveraging diverse information sources. To evaluate its feasibility, we conducted extensive experiments on two 3D medical image segmentation tasks and five polyp segmentation datasets. The results demonstrate that DBAANet achieves strong performance while maintaining excellent computational efficiency, providing an accurate and efficient solution for medical image segmentation.
Qihao Wang, Guiling Shi, Kaihao Zhang, Yuanyuan Zhang 0008
BIBM6
2025 Dynamic scheduling in flexible and hybrid disassembly systems with manual and automated workstations using reward-shaping enhanced reinforcement learning
Jinlong Wang 0002, Qihuiyang Liang, Zelin Qu, Yuanyuan Zhang 0008
Eng. Appl. Artif. Intell.5
2025 Risk identification of listed companies violation by integrating knowledge graph and multi-source risk factors
Jinlong Wang 0002, Pengjun Li, Yingmin Liu, Xiaoyun Xiong, Yuanyuan Zhang 0008, Zhihan Lyu
Eng. Appl. Artif. Intell.5
2025 Multi-echelon inventory optimization of waste electrical and electronic equipment closed-loop supply chain based on reinforcement learning under carbon tax policy
Jinlong Wang 0002, Shangzhuo Zhou, Guanyu Ren, Xianquan Ren, Xiaoyun Xiong, Yuanyuan Zhang 0008
Eng. Appl. Artif. Intell.7
2025 Deciphering circRNA-drug sensitivity associations via global-local heterogeneous matrix factorization and hypergraph contrastive learning
Tiyao Liu, Yuanyuan Zhang 0008, Wenjing Yin, Yingye Liu
Expert Syst. Appl.3
2025 MOHGCN: A trustworthy multi-omics data integration framework based on specificity-aware heterogeneous graph convolutional neural networks for disease diagnosis
Yuanyuan Zhang 0008, Kuijie Zhang, Wenjing Yin
Expert Syst. Appl.3
2025 MPSO-CD: A Multi-Objective Particle Swarm Optimization Community Detection Method for Identifying Disease Modules
abstract
The dysfunction of biological systems caused by disease-related genes is one of the inducements of complex diseases. To understand molecular mechanisms of complex diseases, the identification of disease-related gene modules in biological networks through community detection is emerging as a promising approach. However, most community detection methods are not suitable for biological networks because their topological structures are complex and the scale of biologically relevant modules are small. In this paper, a novel community detection method called MPSO-CD was proposed based on multi-objective particle swarm optimization, in which negative ratio association and ratio cut were employed as objective functions. Highlights of MPSO-CD are a mutation strategy based on clustering coefficient and the procedure of disease module screening referring to the internal connection density and functional similarity. Experimental results of social and synthetic complex networks indicate that MPSO-CD is comparable and often superior to four compared methods. Eventually, MPSO-CD is applied to the asthma gene co-expression network for identifying potential disease modules that provide the molecular mechanism information about asthma. Most of the captured modules have been proven to be associated with asthma through Gene Ontology and pathway enrichment analysis.
Xuhui Zhu, Mingyuan Bi, Junliang Shang, Feng Li 0033, Yuanyuan Zhang 0008, Ling-Yun Dai, Shengjun Li, Jin-Xing Liu 0001
IEEE Trans. Comput. Biol. Bioinform.6
2025 Convolution Bridge: An Effective Algorithmic Migration Strategy From CNNs to GNNs
abstract
Graph neural networks (GNNs), as a rising star in machine learning, are widely used in relational data models and have achieved outstanding performance in graph tasks. GNN continuously takes inspiration from mature models in other domains such as computer vision and natural language processing to motivate the development of graph algorithms. However, due to the various data structures from different domains, the cross-domain migration of models has to go through a long period of disassembly and reconstruction, which may not yield the desired results. To preserve the excellent properties of convolution and optimize the migration process from convolutional neural networks (CNNs) to GNNs, we propose a convolution bridge. The convolution bridge realizes the data alignment from CNN to GNN, so that the CNN-based model can be efficiently migrated to the graph structure model. To demonstrate the effectiveness of our migration strategy, we migrated the inception module and U-Net architecture from CNNs to GNNs, named GraInc and GraU-Net, for the node-level task and the graph-level task, respectively. Experimental results show that GraInc and GraU-Net are highly competitive compared to the current state-of-the-art models, particularly on dense graph datasets.
Kuijie Zhang, Huahui Yang, Yuanyuan Zhang 0008, Hengxiao Li, Jerry Chun-Wei Lin
IEEE Trans. Neural Networks Learn. Syst.4
2024 A multi-objective genetic algorithm based on neighborhood coevolution for community detection
abstract
Community detection has attracted growing interest, with multi-objective evolutionary algorithms proving to be highly competitive in this area. In this paper, a community detection method based on a multi-objective neighborhood coevolution genetic algorithm, NCMOGA, is proposed. To improve the computational efficiency in large-scale networks, NCMOGA introduces a network processing strategy to simplify the network before and during evolution. A neighborhood coevolution strategy is proposed, in which the corresponding subpopulation is formed according to the neighborhood of each individual. A series of operations such as crossover, mutation and update are performed in the subpopulation, emphasizing the synergy between individuals and their neighbors. Mating selection and crossover operations are performed based on the center selection idea of density peak clustering, and the most important nodes are selected to generate offspring. The effectiveness of NCMOGA is verified on synthetic networks and real-world networks. In addition, the results in guiding the classification of disease and healthy samples demonstrate the high quality of the modules detected by NCMOGA.
Mingyuan Bi, Junliang Shang, Xiaotong Kong, Feng Li 0033, Yuanyuan Zhang 0008, Jin-Xing Liu 0001
BIBM5
2024 IEARACDA: Predicting circRNA-disease associations based on information enhancement and arctangent rank approximation
abstract
Circular RNAs (circRNAs) are non-coding RNA molecules that play a significant role in cell regulation and disease occurrence. In recent years, the use of computational methods to predict circRNAs associated with diseases has become a research focus. However, these methods are often constrained by the noise interference and an insufficient exploration of the underlying information in known associations. Therefore, this paper proposes an information enhancement and arctangent rank approximation based method to predict potential circRNA-disease associations (IEARACDA). Firstly, information enhancement methods are used to make circRNA and disease similarity information more complete and to reduce the sparsity of known association. Subsequently, the arctangent rank approximation method is employed to mitigate the potential bias issues that may arise from the nuclear norm, and optimization is carried out using the inexact augmented Lagrange multiplier (IALM) method. Finally, potential associations are predicted through heterogeneous graph inference method. The experimental results demonstrate that IEARACDA exhibits superior performance compared to existing methods. Furthermore, case studies for hepatocellular carcinoma and lung cancer provide additional confirmation of the accuracy and practical significance of the model predictions.
Zheqi Song, Tiyao Liu, Yuanyuan Zhang 0008
BIBM5
2024 A Particle Swarm Optimization Algorithm Based on Multi-Population Mutual Learning for SNP-SNP Interaction Detection
abstract
Single nucleotide polymorphism (SNPs) data have become abundant thanks to the quick advancement of high-throughput sequencing technology, which provides convenience for genome-wide association studies. Single SNPs have been proven to be the cause of some diseases, and the emergence of complex diseases is often thought to be the result of the interaction of multiple SNPs. However, the possible interaction of millions of SNPs imposes a heavy computational burden for uncovering complex disease mechanisms. The existing SNP-SNP interaction detection algorithms frequently have flaws including high computation complexity and poor optimization effectiveness. In this study, a particle swarm optimization algorithm based on multi-population mutual learning (PSOMPML) is proposed to detect SNP-SNP interactions. In this algorithm, the mutual learning strategy is introduced to deal with different particles in different sub-populations to facilitate knowledge exchange. In addition, the elite preservation mechanism is incorporated into PSOMPML, to better preserve the good SNPs in the elite particles. The promising region local search strategy searches the optimal solution along the target solution and its near space to increase the convergence speed of the proposed algorithm. Experiments on simulated data sets and real data also demonstrate the effectiveness of the proposed algorithm.
Linqian Zhao, Yahan Li, Junliang Shang, Qianqian Ren, Yuanyuan Zhang 0008, Jin-Xing Liu 0001
BIBM5
2024 CPSORCL: A Cooperative Particle Swarm Optimization Method with Random Contrastive Learning for Interactive Feature Selection
Junliang Shang, Yahan Li, Feng Li 0033, Yuanyuan Zhang 0008, Jin-Xing Liu 0001
ISBRA (2)5
2024 Graph attention autoencoder model with dual decoder for clustering single-cell RNA sequencing data
Yu Zhang 0265, Yuanyuan Zhang 0008, Jionglong Su, Yingye Liu
Appl. Intell.3
2024 Collaborative Construction Method of Biomedical Knowledge Graph Based on Multi-Blockchain
abstract
The collaborative construction method of knowledge graph based on blockchain is crucial for the safe and efficient construction of biomedical knowledge graph (BioKG), and has received widespread attention from the academic and medical circles. At present, the data sources of BioKG are very wide. While using multi-source data to collaboratively build a knowledge graph, data redundancy problems often occur. Once data from a large number of sources is available, it becomes difficult for the system to ensure data consistency in a reliable way. In addition, the increase in inconsistent and redundant data also reduces the efficiency of collaboration in the system. The traditional blockchain-based collaborative construction method of BioKG is difficult to effectively solve these problems. Multi-blockchain technology provides a new solution for the construction of collaborative knowledge graphs in the biomedical field, which is characterized by good scalability and diversified applications. To this end, we propose a multi-person collaboration construction method for BioKG based on multi-blockchain, named collaborative construction method chain (CCMC). CCMC uses the multi-chain structure and clock cycle control method of the blockchain to separate user operations and system verification, and uses the read-write separation function to improve the consistency of the collaborative version. In addition, we propose a semantic fusion method that effectively reduces data redundancy in the collaborative process and improves the efficiency of BioKG collaborative construction. Performance testing and comparative evaluation confirmed that compared with previous methods, CCMC significantly improves the creation of collaborative knowledge graphs and meets the requirements of multi-person collaborative construction.
Jinlong Wang 0002, Zhenxi Xie, Hui Xin, Pengjun Li, Yuanyuan Zhang 0008, Xiaoyun Xiong, Jerry Chun-Wei Lin
Distributed Ledger Technol. Res. Pract.5
2024 Multi-resolution sequence and structure feature extraction for binding site prediction
Wenjing Yin, Sibo Qiao, Yuanyuan Zhang 0008
Eng. Appl. Artif. Intell.4
2024 DeFuseDTI: Interpretable drug target interaction prediction model with dual-branch encoder and multiview fusion
Baoming Feng, Yuanyuan Zhang 0008, Niu-Wang-Jie Niu, Hao-Yu Zheng, Jinlong Wang 0002
Future Gener. Comput. Syst.2
2024 TSCNet: Topology and semantic co-mining node representation learning based on direct perception strategy
Kuijie Zhang, Yuanyuan Zhang 0008, Xiao He 0012, Haiyuan Gui
Knowl. Based Syst.3
2024 MOSGAT: Uniting Specificity-Aware GATs and Cross Modal-Attention to Integrate Multi-Omics Data for Disease Diagnosis
abstract
With the advancement of sequencing methodologies, the acquisition of vast amounts of multi-omics data presents a significant opportunity for comprehending the intricate biological mechanisms underlying diseases and achieving precise diagnosis and treatment for complex disorders. However, as diverse omics data are integrated, extracting sample-specific features within each omics modality and exploring potential correlations among different modalities while avoiding mutual interference becomes a critical challenge in multi-omics data integration research. In the context of this study, we proposed a framework that unites specificity-aware GATs and cross-modal attention to integrate different omics data (MOSGAT). To be specific, we devise Graph Attention Networks (GATs) tailored for each omics modality data to perform feature extraction on samples. Additionally, an adaptive confidence attention weighting technique is incorporated to enhance the confidence in the extracted features. Finally, a cross-modal attention mechanism was devised based on multi-head self-attention, thoroughly uncovering potential correlations between different omics data. Extensive experiments were conducted on four publicly available medical datasets, highlighting the superiority of the proposed framework when compared to state-of-the-art methodologies, particularly in the realm of classification tasks. The experimental results underscore MOSGAT's effectiveness in extracting features and exploring potential inter-omics associations.
Yuanyuan Zhang 0008, Wenjing Yin, Yawu Zhao
IEEE J. Biomed. Health Informatics3
2023 Predicting potential small molecule-miRNA associations utilizing truncated schatten p-norm
abstract
MicroRNAs (miRNAs) have significant implications in diverse human diseases and have proven to be effectively targeted by small molecules (SMs) for therapeutic interventions. However, current SM-miRNA association prediction models do not adequately capture SM/miRNA similarity. Matrix completion is an effective method for association prediction, but existing models use nuclear norm instead of rank function, which has some drawbacks. Therefore, we proposed a new approach for predicting SM-miRNA associations by utilizing the truncated schatten p-norm (TSPN). First, the SM/miRNA similarity was preprocessed by incorporating the Gaussian interaction profile kernel similarity method. This identified more SM/miRNA similarities and significantly improved the SM-miRNA prediction accuracy. Next, we constructed a heterogeneous SM-miRNA network by combining biological information from three matrices and represented the network with its adjacency matrix. Finally, we constructed the prediction model by minimizing the truncated schatten p-norm of this adjacency matrix and we developed an efficient iterative algorithmic framework to solve the model. In this framework, we also used a weighted singular value shrinkage algorithm to avoid the problem of excessive singular value shrinkage. The truncated schatten p-norm approximates the rank function more closely than the nuclear norm, so the predictions are more accurate. We performed four different cross-validation experiments on two separate datasets, and TSPN outperformed various most advanced methods. In addition, public literature confirms a large number of predictive associations of TSPN in four case studies. Therefore, TSPN is a reliable model for SM-miRNA association prediction.
Tiyao Liu, Chuanru Ren, Zhiyuan Zhao 0003, Yuanyuan Zhang 0008
Briefings Bioinform.7
2023 Generative Adversarial Matrix Completion Network based on Multi-Source Data Fusion for miRNA-Disease Associations Prediction
abstract
Numerous biological studies have shown that considering disease-associated micro RNAs (miRNAs) as potential biomarkers or therapeutic targets offers new avenues for the diagnosis of complex diseases. Computational methods have gradually been introduced to reveal disease-related miRNAs. Considering that previous models have not fused sufficiently diverse similarities, that their inappropriate fusion methods may lead to poor quality of the comprehensive similarity network and that their results are often limited by insufficiently known associations, we propose a computational model called Generative Adversarial Matrix Completion Network based on Multi-source Data Fusion (GAMCNMDF) for miRNA-disease association prediction. We create a diverse network connecting miRNAs and diseases, which is then represented using a matrix. The main task of GAMCNMDF is to complete the matrix and obtain the predicted results. The main innovations of GAMCNMDF are reflected in two aspects: GAMCNMDF integrates diverse data sources and employs a nonlinear fusion approach to update the similarity networks of miRNAs and diseases. Also, some additional information is provided to GAMCNMDF in the form of a 'hint' so that GAMCNMDF can work successfully even when complete data are not available. Compared with other methods, the outcomes of 10-fold cross-validation on two distinct databases validate the superior performance of GAMCNMDF with statistically significant results. It is worth mentioning that we apply GAMCNMDF in the identification of underlying small molecule-related miRNAs, yielding outstanding performance results in this specific domain. In addition, two case studies about two important neoplasms show that GAMCNMDF is a promising prediction method.
Yunyin Li, Yuanyuan Zhang 0008, Sibo Qiao, Yu Zhang 0265, Fuyu Wang 0003
Briefings Bioinform.3
2023 VGAEDTI: drug-target interaction prediction based on variational inference and graph autoencoder
abstract
MOTIVATION: Accurate identification of Drug-Target Interactions (DTIs) plays a crucial role in many stages of drug development and drug repurposing. (i) Traditional methods do not consider the use of multi-source data and do not consider the complex relationship between data sources. (ii) How to better mine the hidden features of drug and target space from high-dimensional data, and better solve the accuracy and robustness of the model. RESULTS: To solve the above problems, a novel prediction model named VGAEDTI is proposed in this paper. We constructed a heterogeneous network with multiple sources of information using multiple types of drug and target dataIn order to obtain deeper features of drugs and targets, we use two different autoencoders. One is variational graph autoencoder (VGAE) which is used to infer feature representations from drug and target spaces. The second is graph autoencoder (GAE) propagating labels between known DTIs. Experimental results on two public datasets show that the prediction accuracy of VGAEDTI is better than that of six DTIs prediction methods. These results indicate that model can predict new DTIs and provide an effective tool for accelerating drug development and repurposing.
Yuanyuan Zhang 0008, Yinfei Feng, Zengqian Deng
BMC Bioinform.1
2023 AF-GCN: Completing various graph tasks efficiently via adaptive quadratic frequency response function in graph spectral domain
Kuijie Zhang, Gan Wang, Jerry Chun-Wei Lin, Fuyu Wang 0003, Yuanyuan Zhang 0008
Inf. Sci.8
2021 HGDD: A Drug-Disease High-Order Association Information Extraction Method for Drug Repurposing via Hypergraph
Kuijie Zhang, Yuanyuan Zhang 0008, Sibo Qiao
ISBRA4
2021 TagSNP-set selection for genotyping using integrated data
abstract
Single-nucleotide polymorphisms (SNPs) are vital in identifying genetic level variations in complex disease. It was found that the information of SNPs on adjacent or identical genes can be represented by a few tagSNPs (called tag SNP-set or tagSNP-set). In this work, we propose a novel method called TagSNP-set Selection by Optimal Iteration with Linkage Disequilibrium (TSOILD) and develop a quantificationally analytical tagSNP-set prediction method called Physical Distance-Linkage Disequilibrium Prediction Method (PDLDPM). To verify the validity of TSOILD method and PDLDPM, a large amount of test data is generated by simulation software HAPGEN2. According to the experimental results, the prediction accuracy of TSOILD is improved by 6.73%, 3.19%, 6.52% and 1.72% over the Random Sampling, Genetic Algorithm (GA) , Greedy Algorithm and TagSNP-Set Selection Method with Maximum Information (TSMI) respectively. In addition, PDLDPM, Linkage Coverage and selection of tag SNPs to maximize prediction accuracy (STAMPA) are used to evaluate the tagSNP-set selected by Random Sampling, GA, Greedy Algorithm and TSMI. Results show that the PDLDPM performs better than the other two methods. These methods provide effective assistance for the study of genetic level variation of complex diseases.
Gaowei Liu, Xinzeng Wang, Yuanyuan Zhang 0008
Future Gener. Comput. Syst.4
2019 PEIS: a novel approach of tumor purity estimation by identifying information sites through integrating signal based on DNA methylation data
abstract
BACKGROUND: Tumor purity plays an important role in understanding the pathogenic mechanism of tumors. The purity of tumor samples is highly sensitive to tumor heterogeneity. Due to Intratumoral heterogeneity of genetic and epigenetic data, it is suitable to study the purity of tumors. Among them, there are many purity estimation methods based on copy number variation, gene expression and other data, while few use DNA methylation data and often based on selected information sites. Consequently, how to choose methylation sites as information sites has an important influence on the purity estimation results. At present, the selection of information sites was often based on the differentially methylated sites that only consider the mean signal, without considering other possible signals and the strong correlation among adjacent sites. RESULTS: Considering integrating multi-signals and strong correlation among adjacent sites, we propose an approach, PEIS, to estimate the purity of tumor samples by selecting informative differential methylation sites. Application to 12 publicly available tumor datasets, it is shown that PEIS provides accurate results in the estimation of tumor purity which has a high consistency with other existing methods. Also, through comparing the results of different information sites selection methods in the evaluation of tumor purity, it shows the PEIS is superior to other methods. CONCLUSIONS: A new method to estimate the purity of tumor samples is proposed. This approach integrates multi-signals of the CpG sites and the correlation between the sites. Experimental analysis shows that this method is in good agreement with other existing methods for estimating tumor purity.
Yuanyuan Zhang 0008, Xinzeng Wang
BMC Bioinform.3