VLDB 2026 Research / reviewers in the wild / expert
Hongyan Cao
dblp:205/9707
· DBLP profile ↗
11ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CEDR: robust consensus cancer subtyping with multi-omics data via ensemble dimensionality reductionabstractCancer is a highly heterogeneous disease underpinned by complex molecular alterations. Accurate subtyping is critical for guiding personalized treatment and improving clinical outcomes. However, multi-omics data are high-dimensional, noisy, and heterogeneous across platforms, posing major challenges for reliable subtyping. To address this, dimensionality reduction is necessary to capture underlying molecular patterns in a low-dimensional space, facilitating both computational efficiency and biological interpretation. We present Consensus subtyping method with Ensemble Dimensionality Reduction for multi-omics data integration (CEDR), a consensus subtyping framework that integrates complementary linear and nonlinear dimensionality reduction methods with robust clustering and probabilistic ensemble modeling. Different from existing dimensionality reduction techniques, our framework adopts an ensemble learning framework that integrates multiple dimensionality reduction techniques with robust clustering to achieve reliable consensus cancer subtyping. We apply Optimally Tuned Robust Improper Maximum Likelihood Estimator to the concatenated low-dimensional matrix for robust subtyping, and ensemble the result with the Mixture Model for Clustering Ensembles to identify stable subtypes. Across extensive simulations, CEDR consistently outperformed conventional dimensionality reduction-based clustering, the Cluster Of Clusters Analysis (COCA) ensemble strategy, and state-of-the-art multi-omics integration algorithms (SNF and CIMLR) in both accuracy and robustness. Application to clear cell renal cell carcinoma and lower-grade glioma revealed biologically interpretable subtypes characterized by distinctive survival outcomes, pathway activities, and immune infiltration patterns. These findings demonstrate that CEDR provides a powerful and reliable strategy for multi-omics data integration and cancer subtyping, with strong potential for broader applications in high-dimensional multimodal data analysis. Hongyan Cao, Zhaoyang Xu, Shilong Lin, Gang Du, Tong Wang 0019, Juping Wang, Ruiling Fang, Ping Zeng, Hongmei Yu, Yuehua Cui |
Briefings Bioinform. | 1 |
| 2026 | An integrative association analysis for complex diseases in underrepresented groups by leveraging the trans-ethnic genetic similarityabstractGenome-wide association studies (GWASs) have been conducted primarily in European (EUR) populations, limiting insights into underrepresented groups such as East Asian (EAS), but cross-ancestry GWASs have demonstrated high trans-ethnic genetic similarity between EUR and non-EUR populations. To enhance association analysis power in EAS populations, we propose tranScore, a novel summary-statistics-based transfer learning method that leverages trans-ethnic genetic similarity through hierarchical modeling. By considering EUR as auxiliary population, tranScore performs joint testing of genetic effects in auxiliary and target populations via well-established P-value combination procedures. Simulations demonstrate that tranScore maintains control of type I error rates and provides substantial power gains for diverse genetic architectures, showing robustness against various challenges including incomplete SNP overlap and effect heterogeneity. In the real-data application of eight diseases from the China Kadoorie Biobank (CKB), after incorporating the genetic information of the EUR population, tranScore identified significantly more genes than the traditional score test which ignored such information. Approximately 41.9% of discovered genes were replicated in the Biobank Japan cohort. Overall, tranScore represents a flexible and powerful statistical approach for association analysis of complex diseases and traits through transfer learning of shared genetic similarities between the auxiliary and target populations. Jike Qi, Hongyan Cao, Ping Zeng |
Briefings Bioinform. | 7 |
| 2025 | Multi-omics data integration for enhanced cancer subtyping via interactive multi-kernel learningabstractCancer is a highly heterogeneous disease characterized by complex molecular changes. Subtypes identified through multi-omics data hold significant promise for improving prognosis and facilitating personalized precision treatment. Recent multi-omics integration methods have mostly focused on capturing complementary information from different data types, often overlooking potential interactions between omics data. Here we develop a novel method named interactive multi-kernel learning (iMKL), which incorporates omics-omics interactions alongside heterogeneous data types under the unsupervised multi-kernel learning framework, to improve subtype identification. Using the sample-similarity kernel for each dataset, we propose a joint Hadamard product strategy to capture higher-order interactive effects from different omics data types. We applied iMKL to two renal cell carcinoma (RCC) datasets-clear renal cell carcinoma (ccRCC) and type II papillary renal cell carcinoma (type II pRCC)-both including miRNA expression, mRNA expression, and DNA methylation data. Stability analysis through random sampling of patients or features demonstrated that iMKL exhibits strong robustness and accuracy in identifying patient subtypes. The identified subtypes revealed dramatic differences in patient survival, with both ccRCC and type II pRCC classified into three distinct subtypes. The findings in the real application highlight potential biomarkers associated with adverse patient outcomes and demonstrate substantial advancement in cancer subtype identification. The iMKL method effectively identifies tumor molecular subtypes that are strongly associated with clinical features and survival rates, providing valuable insights for accurate cancer subtyping, clinical decision-making, and the realization of personalized treatment strategies. Hongyan Cao, Tong Wang 0019, Zhaoyang Xu, Gaiqin Liu, Ruiling Fang, Ping Zeng, Hongmei Yu, Yuehua Cui |
Briefings Bioinform. | 1 |
| 2024 | A metagene based similarity network fusion approach for multi-omics data integration identified novel subtypes in renal cell carcinomaabstractRenal cell carcinoma (RCC) ranks among the most prevalent cancers worldwide, with both incidence and mortality rates increasing annually. The heterogeneity among RCC patients presents considerable challenges for developing universally effective treatment strategies, emphasizing the necessity of in-depth research into RCC's molecular mechanisms, understanding the variations among RCC patients and further identifying distinct molecular subtypes for precise treatment. We proposed a metagene-based similarity network fusion (Meta-SNF) method for RCC subtype identification with multi-omics data, using a non-negative matrix factorization technique to capture alternative structures inherent in the dataset as metagenes. These latent metagenes were then integrated to construct a fused network under the Similarity Network Fusion (SNF) framework for more precise subtyping. We conducted simulation studies and analyzed real-world data from two RCC datasets, namely kidney renal clear cell carcinoma (KIRC) and kidney renal papillary cell carcinoma (KIRP) to demonstrate the utility of Meta-SNF. The simulation studies indicated that Meta-SNF achieved higher accuracy in subtype identification compared with the original SNF and other state-of-the-art methods. In analyses of real data, Meta-SNF produced more distinct and well-separated clusters, classifying both KIRC and KIRP into four subtypes with significant differences in survival outcomes. Subsequently, we performed comprehensive bioinformatics analyses focused on subtypes with poor prognoses in KIRC and KIRP and identified several potential biomarkers. Meta-SNF offers a novel strategy for subtype identification using multi-omics data, and its application to RCC datasets has yielded diverse biological insights which are highly valuable for informing clinical decision-making processes in the treatment of RCC. Congcong Jia, Tong Wang 0019, Dingtong Cui, Yaxin Tian, Gaiqin Liu, Zhaoyang Xu, Ruiling Fang, Hongmei Yu, Yuehua Cui, Hongyan Cao |
Briefings Bioinform. | 12 |
| 2024 | A high-dimensional omnibus test for set-based association analysisabstractSet-based association analysis is a valuable tool in studying the etiology of complex diseases in genome-wide association studies, as it allows for the joint testing of variants in a region or group. Two common types of single nucleotide polymorphism (SNP)-disease functional models are recognized when evaluating the joint function of a set of SNP: the cumulative weak signal model, in which multiple functional variants with small effects contribute to disease risk, and the dominating strong signal model, in which a few functional variants with large effects contribute to disease risk. However, existing methods have two main limitations that reduce their power. Firstly, they typically only consider one disease-SNP association model, which can result in significant power loss if the model is misspecified. Secondly, they do not account for the high-dimensional nature of SNPs, leading to low power or high false positives. In this study, we propose a solution to these challenges by using a high-dimensional inference procedure that involves simultaneously fitting many SNPs in a regression model. We also propose an omnibus testing procedure that employs a robust and powerful P-value combination method to enhance the power of SNP-set association. Our results from extensive simulation studies and a real data analysis demonstrate that our set-based high-dimensional inference strategy is both flexible and computationally efficient and can substantially improve the power of SNP-set association analysis. Application to a real dataset further demonstrates the utility of the testing strategy. Fuzhao Chen, Hongyan Cao, Lina Yan, Xia Gao, Yuehua Cui |
Briefings Bioinform. | 5 |
| 2024 | DAE-CFR: detecting microRNA-disease associations using deep autoencoder and combined feature representationabstractBACKGROUND: MicroRNA (miRNA) has been shown to play a key role in the occurrence and progression of diseases, making uncovering miRNA-disease associations vital for disease prevention and therapy. However, traditional laboratory methods for detecting these associations are slow, strenuous, expensive, and uncertain. Although numerous advanced algorithms have emerged, it is still a challenge to develop more effective methods to explore underlying miRNA-disease associations. RESULTS: In the study, we designed a novel approach on the basis of deep autoencoder and combined feature representation (DAE-CFR) to predict possible miRNA-disease associations. We began by creating integrated similarity matrices of miRNAs and diseases, performing a logistic function transformation, balancing positive and negative samples with k-means clustering, and constructing training samples. Then, deep autoencoder was used to extract low-dimensional feature from two kinds of feature representations for miRNAs and diseases, namely, original association information-based and similarity information-based. Next, we combined the resulting features for each miRNA-disease pair and used a logistic regression (LR) classifier to infer all unknown miRNA-disease interactions. Under five and tenfold cross-validation (CV) frameworks, DAE-CFR not only outperformed six popular algorithms and nine classifiers, but also demonstrated superior performance on an additional dataset. Furthermore, case studies on three diseases (myocardial infarction, hypertension and stroke) confirmed the validity of DAE-CFR in practice. CONCLUSIONS: DAE-CFR achieved outstanding performance in predicting miRNA-disease associations and can provide evidence to inform biological experiments and clinical therapy. Ruiyan Zhang, Xiaojing Dong, Hongyan Cao |
BMC Bioinform. | 6 |
| 2023 | Cancer subtyping with heterogeneous multi-omics data via hierarchical multi-kernel learningabstractDifferentiating cancer subtypes is crucial to guide personalized treatment and improve the prognosis for patients. Integrating multi-omics data can offer a comprehensive landscape of cancer biological process and provide promising ways for cancer diagnosis and treatment. Taking the heterogeneity of different omics data types into account, we propose a hierarchical multi-kernel learning (hMKL) approach, a novel cancer molecular subtyping method to identify cancer subtypes by adopting a two-stage kernel learning strategy. In stage 1, we obtain a composite kernel borrowing the cancer integration via multi-kernel learning (CIMLR) idea by optimizing the kernel parameters for individual omics data type. In stage 2, we obtain a final fused kernel through a weighted linear combination of individual kernels learned from stage 1 using an unsupervised multiple kernel learning method. Based on the final fusion kernel, k-means clustering is applied to identify cancer subtypes. Simulation studies show that hMKL outperforms the one-stage CIMLR method when there is data heterogeneity. hMKL can estimate the number of clusters correctly, which is the key challenge in subtyping. Application to two real data sets shows that hMKL identified meaningful subtypes and key cancer-associated biomarkers. The proposed method provides a novel toolkit for heterogeneous multi-omics data integration and cancer subtypes identification. Yifang Wei, Lingmei Li, Jian Sa, Hongyan Cao, Yuehua Cui |
Briefings Bioinform. | 6 |
| 2021 | Gene-based mediation analysis in epigenetic studiesabstractMediation analysis has been a useful tool for investigating the effect of mediators that lie in the path from the independent variable to the outcome. With the increasing dimensionality of mediators such as in (epi)genomics studies, high-dimensional mediation model is needed. In this work, we focus on epigenetic studies with the goal to identify important DNA methylations that act as mediators between an exposure disease outcome. Specifically, we focus on gene-based high-dimensional mediation analysis implemented with kernel principal component analysis to capture potential nonlinear mediation effect. We first review the current high-dimensional mediation models and then propose two gene-based analytical approaches: gene-based high-dimensional mediation analysis based on linearity assumption between mediators and outcome (gHMA-L) and gene-based high-dimensional mediation analysis based on nonlinearity assumption (gHMA-NL). Since the underlying true mediation relationship is unknown in practice, we further propose an omnibus test of gene-based high-dimensional mediation analysis (gHMA-O) by combing gHMA-L and gHMA-NL. Extensive simulation studies show that gHMA-L performs better under the model linear assumption and gHMA-NL does better under the model nonlinear assumption, while gHMA-O is a more powerful and robust method by combining the two. We apply the proposed methods to two datasets to investigate genes whose methylation levels act as important mediators in the relationship: (1) between alcohol consumption and epithelial ovarian cancer risk using data from the Mayo Clinic Ovarian Cancer Case-Control Study and (2) between childhood maltreatment and comorbid post-traumatic stress disorder and depression in adulthood using data from the Gray Trauma Project. Ruiling Fang, Yuzhao Gao, Hongyan Cao, Ellen L. Goode, Yuehua Cui |
Briefings Bioinform. | 4 |
| 2021 | Identifying complex gene-gene interactions: a mixed kernel omnibus testing approachabstractGenes do not function independently; rather, they interact with each other to fulfill their joint tasks. Identification of gene-gene interactions has been critically important in elucidating the molecular mechanisms responsible for the variation of a phenotype. Regression models are commonly used to model the interaction between two genes with a linear product term. The interaction effect of two genes can be linear or nonlinear, depending on the true nature of the data. When nonlinear interactions exist, the linear interaction model may not be able to detect such interactions; hence, it suffers from substantial power loss. While the true interaction mechanism (linear or nonlinear) is generally unknown in practice, it is critical to develop statistical methods that can be flexible to capture the underlying interaction mechanism without assuming a specific model assumption. In this study, we develop a mixed kernel function which combines both linear and Gaussian kernels with different weights to capture the linear or nonlinear interaction of two genes. Instead of optimizing the weight function, we propose a grid search strategy and use a Cauchy transformation of the P-values obtained under different weights to aggregate the P-values. We further extend the two-gene interaction model to a high-dimensional setup using a de-biased LASSO algorithm. Extensive simulation studies are conducted to verify the performance of the proposed method. Application to two case studies further demonstrates the utility of the model. Our method provides a flexible and computationally efficient tool for disentangling complex gene-gene interactions associated with complex traits. Yan Liu 0093, Yuzhao Gao, Ruiling Fang, Hongyan Cao, Jian Sa, Jianrong Wang, Hongqi Liu, Tong Wang 0019, Yuehua Cui |
Briefings Bioinform. | 4 |
| 2020 | Multilevel heterogeneous omics data integration with kernel fusionabstractHigh-throughput omics data are generated almost with no limit nowadays. It becomes increasingly important to integrate different omics data types to disentangle the molecular machinery of complex diseases with the hope for better disease prevention and treatment. Since the relationship among different omics data features are typically unknown, a supervised learning model assuming a particular distribution with a specific structure will not serve the purpose to capture the underlying complex relationship between multiple features and a disease phenotype. In this work, we briefly reviewed methods for kernel fusion (KF) based on support vector machine and kernel partial least squares (KPLS) algorithms. We then proposed a fused KPLS (fKPLS) model for disease classification and prediction with multilevel omics data. The fused kernel can deal with effect heterogeneity in which different omic data types may have different effect contribution to the trait of interest, with the purpose to improve the prediction performance. We proposed to optimize the kernel parameters and kernel weights with the genetic algorithm (GA). The proposed GA-fKPLS model can substantially improve disease classification performance by integrating multiple omics data types, demonstrated via extensive simulations and real data analysis. With properly defined fitness functions during GA optimization, the proposed KF method can be extended to other kernel-based analyses such as in kernel association analysis with common or rare variants. Hongyan Cao, Tong Wang 0019, Yuehua Cui |
Briefings Bioinform. | 2 |
| 2017 | Predicting disease trait with genomic data: a composite kernel approachabstractWith the advancement of biotechniques, a vast amount of genomic data is generated with no limit. Predicting a disease trait based on these data offers a cost-effective and time-efficient way for early disease screening. Here we proposed a composite kernel partial least squares (CKPLS) regression model for quantitative disease trait prediction focusing on genomic data. It can efficiently capture nonlinear relationships among features compared with linear learning algorithms such as Least Absolute Shrinkage and Selection Operator or ridge regression. We proposed to optimize the kernel parameters and kernel weights with the genetic algorithm (GA). In addition to improved performance for parameter optimization, the proposed GA-CKPLS approach also has better learning capacity and generalization ability compared with single kernel-based KPLS method as well as other nonlinear prediction models such as the support vector regression. Extensive simulation studies demonstrated that GA-CKPLS had better prediction performance than its counterparts under different scenarios. The utility of the method was further demonstrated through two case studies. Our method provides an efficient quantitative platform for disease trait prediction based on increasing volume of omics data. Shaoyu Li, Hongyan Cao, Chichen Zhang, Yuehua Cui |
Briefings Bioinform. | 3 |