VLDB 2026 Research / reviewers in the wild / expert
Yufang Qin
dblp:76/10146
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-8906-8727ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StackAge: an ensemble-based clock for precise quantification of biological age using multi-omics dataabstractAccurate quantification of biological age is essential for early risk stratification and intervention of chronic diseases. Here, we present StackAge, an ensemble-based biological aging clock that integrates large-scale plasma proteomic and metabolomic profiles from 30 376 participants in the UK Biobank. StackAge demonstrated high accuracy in age prediction (Pearson r ≈ 0.93 with chronological age) and substantially enhanced risk prediction for 12 chronic diseases, achieving AUCs exceeding 0.90 for type 2 diabetes, Alzheimer's disease, and chronic kidney disease. Notably, the incorporation of estimated aging rates consistently improved disease prediction beyond conventional omics and demographic features. Feature interpretation and pathway enrichment analyses revealed that aging-associated biomarkers were enriched in inflammation, metabolic stress, and extracellular matrix remodeling pathways. Mediation analysis further indicated that modifiable lifestyle factors may accelerate biological aging, thereby increasing susceptibility to cardiovascular, neurological, immune, and musculoskeletal disorders. Together, these findings establish a robust multi-omics framework for quantifying individual aging trajectories and highlight biological age as a clinically actionable indicator for precision prevention and health management of age-related diseases. Yingyi Jiang, Yuan Fei, Xiaoqi Zheng, Yufang Qin |
Briefings Bioinform. | 6 |
| 2025 | Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniquesabstractIndividualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO). Yufang Qin |
PLoS Comput. Biol. | 3 |
| 2024 | Prediction of anticancer drug sensitivity using an interpretable model guided by deep learningabstractBACKGROUND: The prediction of drug sensitivity plays a crucial role in improving the therapeutic effect of drugs. However, testing the effectiveness of drugs is challenging due to the complex mechanism of drug reactions and the lack of interpretability in most machine learning and deep learning methods. Therefore, it is imperative to establish an interpretable model that receives various cell line and drug feature data to learn drug response mechanisms and achieve stable predictions between available datasets. RESULTS: This study proposes a new and interpretable deep learning model, DrugGene, which integrates gene expression, gene mutation, gene copy number variation of cancer cells, and chemical characteristics of anticancer drugs to predict their sensitivity. This model comprises two different branches of neural networks, where the first involves a hierarchical structure of biological subsystems that uses the biological processes of human cells to form a visual neural network (VNN) and an interpretable deep neural network for human cancer cells. DrugGene receives genotype input from the cell line and detects changes in the subsystem states. We also employ a traditional artificial neural network (ANN) to capture the chemical structural features of drugs. DrugGene generates final drug response predictions by combining VNN and ANN and integrating their outputs into a fully connected layer. The experimental results using drug sensitivity data extracted from the Cancer Drug Sensitivity Genome Database and the Cancer Treatment Response Portal v2 reveal that the proposed model is better than existing prediction methods. Therefore, our model achieves higher accuracy, learns the reaction mechanisms between anticancer drugs and cell lines from various features, and interprets the model's predicted results. CONCLUSIONS: Our method utilizes biological pathways to construct neural networks, which can use genotypes to monitor changes in the state of network subsystems, thereby interpreting the prediction results in the model and achieving satisfactory prediction accuracy. This will help explore new directions in cancer treatment. More available code resources can be downloaded for free from GitHub ( https://github.com/pangweixiong/DrugGene ). Weixiong Pang, Yufang Qin |
BMC Bioinform. | 3 |
| 2022 | Deconvolution of tumor composition using partially available DNA methylation dataabstractBACKGROUND: Deciphering proportions of constitutional cell types in tumor tissues is a crucial step for the analysis of tumor heterogeneity and the prediction of response to immunotherapy. In the process of measuring cell population proportions, traditional experimental methods have been greatly hampered by the cost and extensive dropout events. At present, the public availability of large amounts of DNA methylation data makes it possible to use computational methods to predict proportions. RESULTS: In this paper, we proposed PRMeth, a method to deconvolve tumor mixtures using partially available DNA methylation data. By adopting an iteratively optimized non-negative matrix factorization framework, PRMeth took DNA methylation profiles of a portion of the cell types in the tissue mixtures (including blood and solid tumors) as input to estimate the proportions of all cell types as well as the methylation profiles of unknown cell types simultaneously. We compared PRMeth with five different methods through three benchmark datasets and the results show that PRMeth could infer the proportions of all cell types and recover the methylation profiles of unknown cell types effectively. Then, applying PRMeth to four types of tumors from The Cancer Genome Atlas (TCGA) database, we found that the immune cell proportions estimated by PRMeth were largely consistent with previous studies and met biological significance. CONCLUSIONS: Our method can circumvent the difficulty of obtaining complete DNA methylation reference data and obtain satisfactory deconvolution accuracy, which will be conducive to exploring the new directions of cancer immunotherapy. PRMeth is implemented in R and is freely available from GitHub ( https://github.com/hedingqin/PRMeth ). Dingqin He, Chunhui Song, Yufang Qin |
BMC Bioinform. | 5 |
| 2021 | Drug-induced cell viability prediction from LINCS-L1000 through WRFEN-XGBoost algorithmabstractBACKGROUND: Predicting the drug response of the cancer diseases through the cellular perturbation signatures under the action of specific compounds is very important in personalized medicine. In the process of testing drug responses to the cancer, traditional experimental methods have been greatly hampered by the cost and sample size. At present, the public availability of large amounts of gene expression data makes it a challenging task to use machine learning methods to predict the drug sensitivity. RESULTS: In this study, we introduced the WRFEN-XGBoost cell viability prediction algorithm based on LINCS-L1000 cell signatures. We integrated the LINCS-L1000, CTRP and Achilles datasets and adopted a weighted fusion algorithm based on random forest and elastic net for key gene selection. Then the FEBPSO algorithm was introduced into XGBoost learning algorithm to predict the cell viability induced by the drugs. The proposed method was compared with some new methods, and it was found that our model achieved good results with 0.83 Pearson correlation. At the same time, we completed the drug sensitivity validation on the NCI60 and CCLE datasets, which further demonstrated the effectiveness of our method. CONCLUSIONS: The results showed that our method was conducive to the elucidation of disease mechanisms and the exploration of new therapies, which greatly promoted the progress of clinical medicine. Yufang Qin |
BMC Bioinform. | 3 |
| 2020 | Deconvolution of heterogeneous tumor samples using partial reference signalsabstractDeconvolution of heterogeneous bulk tumor samples into distinct cellular populations is an important yet challenging problem, particularly when only partial references are available. A common approach to dealing with this problem is to deconvolve the mixed signals using available references and leverage the remaining signal as a new cell component. However, as indicated in our simulation, such an approach tends to over-estimate the proportions of known cell types and fails to detect novel cell types. Here, we propose PREDE, a partial reference-based deconvolution method using an iterative non-negative matrix factorization algorithm. Our method is verified to be effective in estimating cell proportions and expression profiles of unknown cell types based on simulated datasets at a variety of parameter settings. Applying our method to TCGA tumor samples, we found that proportions of pure cancer cells better indicate different subtypes of tumor samples. We also detected several cell types for each cancer type whose proportions successfully predicted patient survival. Our method makes a significant contribution to deconvolution of heterogeneous tumor samples and could be widely applied to varieties of high throughput bulk data. PREDE is implemented in R and is freely available from GitHub (https://xiaoqizheng.github.io/PREDE). Yufang Qin, Siwei Nan, Nana Wei, Hua-Jun Wu, Xiaoqi Zheng |
PLoS Comput. Biol. | 1 |
| 2016 | Prediction the Substrate Specificities of Membrane Transport Proteins Based on Support Vector Machine and Hybrid FeaturesabstractMembrane transport proteins and their substrate specificities play crucial roles in a variety of cellular functions. Identifying the substrate specificities of membrane transport proteins is closely related to the protein-target interaction prediction, drug design, membrane recruitment, and dysregulation analysis. However, experimental methods to this aim are time consuming, labor intensive, and costly. Therefore, we proposed a novel method basing on support vector machine (SVM) to predict substrate specificities of membrane transport proteins by integrating features from position-specific score matrix (PSSM), PROFEAT, and Gene Ontology (GO). Finally, jackknife cross-validation tests were adopted on a benchmark and independent datasets to measure the performance of the proposed method. The overall accuracy of 96.16 and 80.45 percent were obtained for two datasets, which are higher (from 2.12 to 20.44 percent) than that by the state-of-the-art tool. Comparison results indicate that the proposed model is more reliable and efficient for accurate prediction the substrate specificities of membrane transport proteins. Liqi Li, Yongsheng Li 0003, Yufang Qin, Shiwen Zhou 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2013 | Prioritization of candidate disease genes by topological similarity between disease and protein diffusion profilesabstractBACKGROUND: Identification of gene-phenotype relationships is a fundamental challenge in human health clinic. Based on the observation that genes causing the same or similar phenotypes tend to correlate with each other in the protein-protein interaction network, a lot of network-based approaches were proposed based on different underlying models. A recent comparative study showed that diffusion-based methods achieve the state-of-the-art predictive performance. RESULTS: In this paper, a new diffusion-based method was proposed to prioritize candidate disease genes. Diffusion profile of a disease was defined as the stationary distribution of candidate genes given a random walk with restart where similarities between phenotypes are incorporated. Then, candidate disease genes are prioritized by comparing their diffusion profiles with that of the disease. Finally, the effectiveness of our method was demonstrated through the leave-one-out cross-validation against control genes from artificial linkage intervals and randomly chosen genes. Comparative study showed that our method achieves improved performance compared to some classical diffusion-based methods. To further illustrate our method, we used our algorithm to predict new causing genes of 16 multifactorial diseases including Prostate cancer and Alzheimer's disease, and the top predictions were in good consistent with literature reports. CONCLUSIONS: Our study indicates that integration of multiple information sources, especially the phenotype similarity profile data, and introduction of global similarity measure between disease and gene diffusion profiles are helpful for prioritizing candidate disease genes. AVAILABILITY: Programs and data are available upon request. Yufang Qin, Taigang Liu, Xiaoqi Zheng |
BMC Bioinform. | 2 |