VLDB 2026 Research / reviewers in the wild / expert
Yunyan Gu
dblp:16/8249
· DBLP profile ↗
13ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0001-5693-4126ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deciphering the cardiac neuron landscape in heart failure patientsabstractNeurons exert a pivotal role in the preservation of cardiac physiological function. However, there is a lack of explanation about the mechanism of cardiac neurons in the pathogenesis of cardiac dysfunction. Here, we generated a cardiac neuron landscape including 11,026 neuronal cells based on the integration of published single-nucleus RNA sequencing data from 75 patients with heart failure and 45 healthy donors. We determined ten distinct neuronal cell subsets differing in abundances, compositions, and biological functions in the heart. In particular, N4-ALK neurons were significantly enriched in failing hearts relative to healthy controls, and their abundance was associated with the response to left ventricular assist device implantation. RXRG, a transcription factor highly expressed in neuronal cells, participated in the transcriptional regulatory network of N4-ALK neurons and showed a positive correlation with the expression of their marker genes. Notably, in heart failure, the PTN-PTPRZ1 axis mediated specific crosstalk between cardiac fibroblasts and N4-ALK neurons. Finally, we used N4-ALK-related features to develop an optimized prediction model for identifying individuals with heart failure. Overall, our integrative cardiac neuron atlas comprehensively characterizes the molecular and functional diversity of neuronal cells, providing a new perspective for further exploration of the regulatory function of neurons in heart failure. Shuping Zhuang, Xiuqi Yang, Jiangqi Liu, Kaidong Liu, Huiming Han, Songmei Zhai, Haihai Liang, Yunyan Gu, Yanjie Lu |
PLoS Comput. Biol. | 10 |
| 2025 | Unveiling a novel cancer hallmark by evaluation of neural infiltration in cancerabstractCancer cells acquire necessary functional capabilities for malignancy through the influence of the nervous system. We evaluate the extent of neural infiltration within the tumor microenvironment (TME) across multiple cancer types, highlighting its role as a cancer hallmark. We identify cancer-related neural genes using 40 bulk RNA-seq datasets across 10 cancer types, developing a predictive score for cancer-related neural infiltration (C-Neural score). Cancer samples with elevated C-Neural scores exhibit perineural invasion, recurrence, metastasis, higher stage or grade, or poor prognosis. Epithelial cells show the highest C-Neural scores among all cell types in 55 single-cell RNA sequencing datasets. The epithelial cells with high C-Neural scores (epi-highCNs) characterized by increased copy number variation, reduced cell differentiation, higher epithelial-mesenchymal transition scores, and elevated metabolic level. Epi-highCNs frequently communicate with Schwann cells by FN1 signaling pathway. The co-culture experiment indicates that Schwann cells may facilitate cancer progression through upregulation of VDAC1. Moreover, C-Neural scores positively correlate with the infiltration of antitumor immune cells, indicating potential response for immunotherapy. Melanoma patients with high C-Neural scores may benefit from trametinib. These analyses illuminate the extent of neural influence within TME, suggesting potential role as a cancer hallmark and offering implications for effective therapeutic strategies against cancer. Lingxue Ren, Kaidong Liu, Linzhu Wang, Shaocong Sang, Yang Hui, Haihai Liang, Yunyan Gu |
Briefings Bioinform. | 16 |
| 2021 | Reference genome and annotation updates lead to contradictory prognostic predictions in gene expression signatures: a case study of resected stage I lung adenocarcinomaabstractRNA-sequencing enables accurate and low-cost transcriptome-wide detection. However, expression estimates vary as reference genomes and gene annotations are updated, confounding existing expression-based prognostic signatures. Herein, prognostic 9-gene pair signature (GPS) was applied to 197 patients with stage I lung adenocarcinoma derived from previous and latest data from The Cancer Genome Atlas (TCGA) processed with different reference genomes and annotations. For 9-GPS, 6.6% of patients exhibited discordant risk classifications between the two TCGA versions. Similar results were observed for other prognostic signatures, including IRGPI, 15-gene and ORACLE. We found that conflicting annotations for gene length and overlap were the major cause of their discordant risk classification. Therefore, we constructed a prognostic 40-GPS based on stable genes across GENCODE v20-v30 and validated it using public data of 471 stage I samples (log-rank P < 0.0010). Risk classification was still stable in RNA-sequencing data processed with the newest GENCODE v32 versus GENCODE v20-v30. Specifically, 40-GPS could predict survival for 30 stage I samples with formalin-fixed paraffin-embedded tissues (log-rank P = 0.0177). In conclusion, this method overcomes the vulnerability of existing prognostic signatures due to reference genome and annotation updates. 40-GPS may offer individualized clinical applications due to its prognostic accuracy and classification stability. Zheyang Zhang, Zhangxiang Zhao, Changjing Chen, Juxuan Zhang, Mengyue Li, Zixin Wei, Wenbin Jiang 0008, Ying Li 0028, Yingyue Cao, Wenyuan Zhao, Yunyan Gu, Qingwei Meng, Lishuang Qi |
Briefings Bioinform. | 15 |
| 2021 | An absolute human stemness index associated with oncogenic dedifferentiationabstractThe progression of cancer is accompanied by the acquisition of stemness features. Many stemness evaluation methods based on transcriptional profiles have been presented to reveal the relationship between stemness and cancer. However, instead of absolute stemness index values-the values with certain range-these methods gave the values without range, which makes them unable to intuitively evaluate the stemness. Besides, these indices were based on the absolute expression values of genes, which were found to be seriously influenced by batch effects and the composition of samples in the dataset. Recently, we have showed that the signatures based on the relative expression orderings (REOs) of gene pairs within a sample were highly robust against these factors, which makes that the REO-based signatures have been stably applied in the evaluations of the continuous scores with certain range. Here, we provided an absolute REO-based stemness index to evaluate the stemness. We found that this stemness index had higher correlation with the culture time of the differentiated stem cells than the previous stemness index. When applied to the cancer and normal tissue samples, the stemness index showed its significant difference between cancers and normal tissues and its ability to reveal the intratumor heterogeneity at stemness level. Importantly, higher stemness index was associated with poorer prognosis and greater oncogenic dedifferentiation reflected by histological grade. All results showed the capability of the REO-based stemness index to assist the assignment of tumor grade and its potential therapeutic and diagnostic implications. Hailong Zheng, Yelin Fu, Tianyi You, Wenbing Guo, Liangliang Jin, Yunyan Gu, Lishuang Qi, Wenyuan Zhao |
Briefings Bioinform. | 9 |
| 2019 | A rank-based algorithm of differential expression analysis for small cell line data with statistical controlabstractTo detect differentially expressed genes (DEGs) in small-scale cell line experiments, usually with only two or three technical replicates for each state, the commonly used statistical methods such as significance analysis of microarrays (SAM), limma and RankProd (RP) lack statistical power, while the fold change method lacks any statistical control. In this study, we demonstrated that the within-sample relative expression orderings (REOs) of gene pairs were highly stable among technical replicates of a cell line but often widely disrupted after certain treatments such like gene knockdown, gene transfection and drug treatment. Based on this finding, we customized the RankComp algorithm, previously designed for individualized differential expression analysis through REO comparison, to identify DEGs with certain statistical control for small-scale cell line data. In both simulated and real data, the new algorithm, named CellComp, exhibited high precision with much higher sensitivity than the original RankComp, SAM, limma and RP methods. Therefore, CellComp provides an efficient tool for analyzing small-scale cell line data. Xianlong Wang 0002, Lu Ao, You Guo, Yunyan Gu, Lishuang Qi, Qingzhou Guan, Zheng Guo 0002 |
Briefings Bioinform. | 7 |
| 2019 | Link synthetic lethality to drug sensitivity of cancer cellsabstractSynthetic lethal (SL) interactions occur when alterations in two genes lead to cell death but alteration in only one of them is not lethal. SL interactions provide a new strategy for molecular-targeted cancer therapy. Currently, there are few drugs targeting SL interactions that entered into clinical trials. Therefore, it is necessary to investigate the link between SL interactions and drug sensitivity of cancer cells systematically for drug development purpose. We identified SL interactions by integrating the high-throughput data from The Cancer Genome Atlas, small hairpin RNA data and genetic interactions of yeast. By integrating SL interactions from other studies, we tested whether the SL pairs that consist of drug target genes and the genes with genomic alterations are related with drug sensitivity of cancer cells. We found that only 6.26%∼34.61% of SL interactions showed the expected significant drug sensitivity using the pooled cancer cell line data from different tissues, but the proportion increased significantly to approximately 90% using the cancer cell line data for each specific tissue. From an independent pharmacogenomics data of 41 breast cancer cell lines, we found three SL interactions (ABL1-IFI16, ABL1-SLC50A1 and ABL1-SYT11) showed significantly better prognosis for the patients with both genes being altered than the patients with only one gene being altered, which partially supports the SL effect between the gene pairs. Our study not only provides a new way for unraveling the complex mechanisms of drug sensitivity but also suggests numerous potentially important drug targets for cancer therapy. Ruiping Wang 0002, Zhangxiang Zhao, Xianlong Wang 0002, Lishuang Qi, Wenyuan Zhao, Zheng Guo 0002, Yunyan Gu |
Briefings Bioinform. | 11 |
| 2018 | A landscape of synthetic viable interactions in cancerabstractSynthetic viability, which is defined as the combination of gene alterations that can rescue the lethal effects of a single gene alteration, may represent a mechanism by which cancer cells resist targeted drugs. Approaches to detect synthetic viable (SV) interactions in cancer genome to investigate drug resistance are still scarce. Here, we present a computational method to detect synthetic viability-induced drug resistance (SVDR) by integrating the multidimensional data sets, including copy number alteration, whole-exome mutation, expression profile and clinical data. SVDR comprehensively characterized the landscape of SV interactions across 8580 tumors in 32 cancer types by integrating The Cancer Genome Atlas data, small hairpin RNA-based functional experimental data and yeast genetic interaction data. We revealed that the SV interactions are favorable to cells and can predict clinical prognosis for cancer patients, which were robustly observed in an independent data set. By integrating the cancer pharmacogenomics data sets from Cancer Cell Line Encyclopedia (CCLE) and Broad Cancer Therapeutics Response Portal, we have demonstrated that SVDR enables drug resistance prediction and exhibits high reliability between two databases. To our knowledge, SVDR is the first genome-scale data-driven approach for the identification of SV interactions related to drug resistance in cancer cells. This data-driven approach lays the foundation for identifying the genomic markers to predict drug resistance and successfully infers the potential drug combination for anti-cancer therapy. Yunyan Gu, Ruiping Wang 0002, Zhangxiang Zhao, Fuduan Peng, Haihai Liang, Lishuang Qi, Wenyuan Zhao, Da Yang 0003, Zheng Guo 0002 |
Briefings Bioinform. | 1 |
| 2018 | Individualized analysis of differentially expressed miRNAs with application to the identification of miRNAs deregulated commonly in lung cancer tissuesabstractIdentifying differentially expressed microRNAs (DE miRNAs) between cancer samples and normal controls is a common way to investigate carcinogenesis mechanisms. However, for a DE miRNA detected at the population-level, we do not know whether it is DE in a particular cancer sample. Here, based on the finding that the within-sample relative expression orderings of miRNA pairs are highly stable in a particular type of normal tissues but widely disrupted in the corresponding cancer tissues, we proposed a method, called RankMiRNA, to identify DE miRNAs in each cancer tissue compared with its own normal state. Evaluated with pair-matched miRNA expression profiles of cancer tissues and adjacent normal tissues for lung and liver cancers, RankMiRNA exhibited excellent performance. Finally, we exemplified an application of the individual-level differential expression analysis by finding miRNAs DE in at least 90% lung cancer tissues, defined as common DE miRNAs of lung cancer. After identifying DE miRNAs for each of 991 lung cancer samples from The Cancer Genome Atlas with RankMiRNA, we found that hsa-mir-210 was upregulated, while hsa-mir-490 and hsa-mir-486 were downregulated in > 90% of the 991 lung cancer samples. These common DE miRNAs were validated in independent pair-matched samples of cancer tissues and adjacent normal tissues measured with different platforms. In conclusion, RankMiRNA provides us a novel tool to find common and subtype-specific miRNAs for a type of cancer, allowing us to study cancer mechanisms in a novel way. Haidan Yan, Qingzhou Guan, You Guo, Yunyan Gu, Lishuang Qi, Zheng Guo 0002 |
Briefings Bioinform. | 10 |
| 2016 | Critical limitations of prognostic signatures based on risk scores summarized from gene expression levels: a case study for resected stage I non-small-cell lung cancerabstractMost of current gene expression signatures for cancer prognosis are based on risk scores, usually calculated as some summaries of expression levels of the signature genes, whose applications require presetting risk score thresholds and data normalization. In this study, we demonstrate the critical limitations of such type of signatures that the risk scores of samples will change greatly when they are normalized together with different samples, which would induce spurious risk classification and difficulty in clinical settings, and the risk scores of independent samples are incomparable if data normalization is not adopted. To overcome these limitations, we propose a rank-based method to extract a prognostic gene pair signature for overall survival of stage I non-small-cell lung cancer. The prognostic gene pair signature is verified in three integrated data sets detected by different laboratories with different microarray platforms. We conclude that, different from the type of signatures based on risk scores summarized from gene expression levels, the rank-based signatures could be robustly applied at the individualized level to independent clinical samples assessed in different laboratories. Lishuang Qi, Libin Chen, Yang Li 0023, Rufei Pan, Wenyuan Zhao, Yunyan Gu, Hongwei Wang 0003, Ruiping Wang 0002, Xiangqi Chen, Zheng Guo 0002 |
Briefings Bioinform. | 7 |
| 2016 | Individualized identification of disease-associated pathways with disrupted coordination of gene expressionabstractCurrent pathway analysis approaches are primarily dedicated to capturing deregulated pathways at the population level and cannot provide patient-specific pathway deregulation information. In this article, the authors present a simple approach, called individPath, to detect pathways with significantly disrupted intra-pathway relative expression orderings for each disease sample compared with the stable, normal intra-pathway relative expression orderings pre-determined in previously accumulated normal samples. Through the analysis of multiple microarray data sets for lung and breast cancer, the authors demonstrate individPath's effectiveness for detecting cancer-associated pathways with disrupted relative expression orderings at the individual level and dissecting the heterogeneity of pathway deregulation among different patients. The portable use of this simple approach in clinical contexts is exemplified by the identification of prognostic intra-pathway gene pair signatures to predict overall survival of resected early-stage lung adenocarcinoma patients and signatures to predict relapse-free survival of estrogen receptor-positive breast cancer patients after tamoxifen treatment. Hongwei Wang 0003, Lu Ao, Haidan Yan, Wenyuan Zhao, Lishuang Qi, Yunyan Gu, Zheng Guo 0002 |
Briefings Bioinform. | 7 |
| 2015 | Individual-level analysis of differential expression of genes and pathways for personalized medicineabstractMOTIVATION: The differential expression analysis focusing on inter-group comparison can capture only differentially expressed genes (DE genes) at the population level, which may mask the heterogeneity of differential expression in individuals. Thus, to provide patient-specific information for personalized medicine, it is necessary to conduct differential expression analysis at the individual level. RESULTS: We proposed a method to detect DE genes in individual disease samples by using the disrupted ordering in individual disease samples. In both simulated data and real paired cancer-normal sample data, this method showed excellent performance. It was found to be insensitive to experimental batch effects and data normalization. The landscape of stable gene pairs in a particular type of normal tissue could be predetermined using previously accumulated data, based on which dysregulated genes and pathways for any disease sample can be readily detected. The usefulness of the RankComp method in clinical settings was exemplified by the identification and application of prognostic markers for lung cancer. AVAILABILITY AND IMPLEMENTATION: RankComp is implemented in R script that is freely available from Supplementary Materials. Hongwei Wang 0003, Wenyuan Zhao, Lishuang Qi, Yunyan Gu, Pengfei Li 0002, Yang Li 0023, Zheng Guo 0002 |
Bioinform. | 5 |
| 2012 | GO-function: deriving biologically relevant functions from statistically significant functionsabstractIn high-throughput studies of diseases, terms enriched with disease-related genes based on Gene Ontology (GO) are routinely found. However, most current algorithms used to find significant GO terms cannot handle the redundancy that results from the dependencies of GO terms. Simply based on some numerical considerations, current algorithms developed for reducing this redundancy may produce results that do not account for biologically interesting cases. In this article, we present several rules used to design a tool called GO-function for extracting biologically relevant terms from statistically significant GO terms for a disease. Using one gene expression profile for colorectal cancer, we compared GO-function with four algorithms designed to treat redundancy. Then, we validated results obtained in this data set by GO-function using another data set for colorectal cancer. Our analysis showed that GO-function can identify disease-related terms that are more statistically and biologically meaningful than those found by the other four algorithms. Jing Wang 0004, Xianxiao Zhou, Jing Zhu 0004, Yunyan Gu, Wenyuan Zhao, Jinfeng Zou, Zheng Guo 0002 |
Briefings Bioinform. | 4 |
| 2010 | Extracting consistent knowledge from highly inconsistent cancer gene data sourcesabstractBACKGROUND: Hundreds of genes that are causally implicated in oncogenesis have been found and collected in various databases. For efficient application of these abundant but diverse data sources, it is of fundamental importance to evaluate their consistency. RESULTS: First, we showed that the lists of cancer genes from some major data sources were highly inconsistent in terms of overlapping genes. In particular, most cancer genes accumulated in previous small-scale studies could not be rediscovered in current high-throughput genome screening studies. Then, based on a metric proposed in this study, we showed that most cancer gene lists from different data sources were highly functionally consistent. Finally, we extracted functionally consistent cancer genes from various data sources and collected them in our database F-Census. CONCLUSIONS: Although they have very low gene overlapping, most cancer gene data sources are highly consistent at the functional level, which indicates that they can separately capture partial genes in a few key pathways associated with cancer. Our results suggest that the sample sizes currently used for cancer studies might be inadequate for consistently capturing individual cancer genes, but could be sufficient for finding a number of cancer genes that could represent functionally most cancer genes. The F-Census database provides biologists with a useful tool for browsing and extracting functionally consistent cancer genes from various data sources. Ruihong Wu, Yuannv Zhang, Wenyuan Zhao, Lixin Cheng, Yunyan Gu, Lin Zhang 0057, Jing Wang 0004, Jing Zhu 0004, Zheng Guo 0002 |
BMC Bioinform. | 6 |