VLDB 2026 Research / reviewers in the wild / expert
Ruiling Fang
dblp:295/1978
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0003-5915-2287ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CEDR: robust consensus cancer subtyping with multi-omics data via ensemble dimensionality reductionabstractCancer is a highly heterogeneous disease underpinned by complex molecular alterations. Accurate subtyping is critical for guiding personalized treatment and improving clinical outcomes. However, multi-omics data are high-dimensional, noisy, and heterogeneous across platforms, posing major challenges for reliable subtyping. To address this, dimensionality reduction is necessary to capture underlying molecular patterns in a low-dimensional space, facilitating both computational efficiency and biological interpretation. We present Consensus subtyping method with Ensemble Dimensionality Reduction for multi-omics data integration (CEDR), a consensus subtyping framework that integrates complementary linear and nonlinear dimensionality reduction methods with robust clustering and probabilistic ensemble modeling. Different from existing dimensionality reduction techniques, our framework adopts an ensemble learning framework that integrates multiple dimensionality reduction techniques with robust clustering to achieve reliable consensus cancer subtyping. We apply Optimally Tuned Robust Improper Maximum Likelihood Estimator to the concatenated low-dimensional matrix for robust subtyping, and ensemble the result with the Mixture Model for Clustering Ensembles to identify stable subtypes. Across extensive simulations, CEDR consistently outperformed conventional dimensionality reduction-based clustering, the Cluster Of Clusters Analysis (COCA) ensemble strategy, and state-of-the-art multi-omics integration algorithms (SNF and CIMLR) in both accuracy and robustness. Application to clear cell renal cell carcinoma and lower-grade glioma revealed biologically interpretable subtypes characterized by distinctive survival outcomes, pathway activities, and immune infiltration patterns. These findings demonstrate that CEDR provides a powerful and reliable strategy for multi-omics data integration and cancer subtyping, with strong potential for broader applications in high-dimensional multimodal data analysis. Hongyan Cao, Zhaoyang Xu, Shilong Lin, Gang Du, Tong Wang 0019, Juping Wang, Ruiling Fang, Ping Zeng, Hongmei Yu, Yuehua Cui |
Briefings Bioinform. | 8 |
| 2025 | Multi-omics data integration for enhanced cancer subtyping via interactive multi-kernel learningabstractCancer is a highly heterogeneous disease characterized by complex molecular changes. Subtypes identified through multi-omics data hold significant promise for improving prognosis and facilitating personalized precision treatment. Recent multi-omics integration methods have mostly focused on capturing complementary information from different data types, often overlooking potential interactions between omics data. Here we develop a novel method named interactive multi-kernel learning (iMKL), which incorporates omics-omics interactions alongside heterogeneous data types under the unsupervised multi-kernel learning framework, to improve subtype identification. Using the sample-similarity kernel for each dataset, we propose a joint Hadamard product strategy to capture higher-order interactive effects from different omics data types. We applied iMKL to two renal cell carcinoma (RCC) datasets-clear renal cell carcinoma (ccRCC) and type II papillary renal cell carcinoma (type II pRCC)-both including miRNA expression, mRNA expression, and DNA methylation data. Stability analysis through random sampling of patients or features demonstrated that iMKL exhibits strong robustness and accuracy in identifying patient subtypes. The identified subtypes revealed dramatic differences in patient survival, with both ccRCC and type II pRCC classified into three distinct subtypes. The findings in the real application highlight potential biomarkers associated with adverse patient outcomes and demonstrate substantial advancement in cancer subtype identification. The iMKL method effectively identifies tumor molecular subtypes that are strongly associated with clinical features and survival rates, providing valuable insights for accurate cancer subtyping, clinical decision-making, and the realization of personalized treatment strategies. Hongyan Cao, Tong Wang 0019, Zhaoyang Xu, Gaiqin Liu, Ruiling Fang, Ping Zeng, Hongmei Yu, Yuehua Cui |
Briefings Bioinform. | 7 |
| 2024 | A metagene based similarity network fusion approach for multi-omics data integration identified novel subtypes in renal cell carcinomaabstractRenal cell carcinoma (RCC) ranks among the most prevalent cancers worldwide, with both incidence and mortality rates increasing annually. The heterogeneity among RCC patients presents considerable challenges for developing universally effective treatment strategies, emphasizing the necessity of in-depth research into RCC's molecular mechanisms, understanding the variations among RCC patients and further identifying distinct molecular subtypes for precise treatment. We proposed a metagene-based similarity network fusion (Meta-SNF) method for RCC subtype identification with multi-omics data, using a non-negative matrix factorization technique to capture alternative structures inherent in the dataset as metagenes. These latent metagenes were then integrated to construct a fused network under the Similarity Network Fusion (SNF) framework for more precise subtyping. We conducted simulation studies and analyzed real-world data from two RCC datasets, namely kidney renal clear cell carcinoma (KIRC) and kidney renal papillary cell carcinoma (KIRP) to demonstrate the utility of Meta-SNF. The simulation studies indicated that Meta-SNF achieved higher accuracy in subtype identification compared with the original SNF and other state-of-the-art methods. In analyses of real data, Meta-SNF produced more distinct and well-separated clusters, classifying both KIRC and KIRP into four subtypes with significant differences in survival outcomes. Subsequently, we performed comprehensive bioinformatics analyses focused on subtypes with poor prognoses in KIRC and KIRP and identified several potential biomarkers. Meta-SNF offers a novel strategy for subtype identification using multi-omics data, and its application to RCC datasets has yielded diverse biological insights which are highly valuable for informing clinical decision-making processes in the treatment of RCC. Congcong Jia, Tong Wang 0019, Dingtong Cui, Yaxin Tian, Gaiqin Liu, Zhaoyang Xu, Ruiling Fang, Hongmei Yu, Yuehua Cui, Hongyan Cao |
Briefings Bioinform. | 8 |
| 2021 | Gene-based mediation analysis in epigenetic studiesabstractMediation analysis has been a useful tool for investigating the effect of mediators that lie in the path from the independent variable to the outcome. With the increasing dimensionality of mediators such as in (epi)genomics studies, high-dimensional mediation model is needed. In this work, we focus on epigenetic studies with the goal to identify important DNA methylations that act as mediators between an exposure disease outcome. Specifically, we focus on gene-based high-dimensional mediation analysis implemented with kernel principal component analysis to capture potential nonlinear mediation effect. We first review the current high-dimensional mediation models and then propose two gene-based analytical approaches: gene-based high-dimensional mediation analysis based on linearity assumption between mediators and outcome (gHMA-L) and gene-based high-dimensional mediation analysis based on nonlinearity assumption (gHMA-NL). Since the underlying true mediation relationship is unknown in practice, we further propose an omnibus test of gene-based high-dimensional mediation analysis (gHMA-O) by combing gHMA-L and gHMA-NL. Extensive simulation studies show that gHMA-L performs better under the model linear assumption and gHMA-NL does better under the model nonlinear assumption, while gHMA-O is a more powerful and robust method by combining the two. We apply the proposed methods to two datasets to investigate genes whose methylation levels act as important mediators in the relationship: (1) between alcohol consumption and epithelial ovarian cancer risk using data from the Mayo Clinic Ovarian Cancer Case-Control Study and (2) between childhood maltreatment and comorbid post-traumatic stress disorder and depression in adulthood using data from the Gray Trauma Project. Ruiling Fang, Yuzhao Gao, Hongyan Cao, Ellen L. Goode, Yuehua Cui |
Briefings Bioinform. | 1 |
| 2021 | Identifying complex gene-gene interactions: a mixed kernel omnibus testing approachabstractGenes do not function independently; rather, they interact with each other to fulfill their joint tasks. Identification of gene-gene interactions has been critically important in elucidating the molecular mechanisms responsible for the variation of a phenotype. Regression models are commonly used to model the interaction between two genes with a linear product term. The interaction effect of two genes can be linear or nonlinear, depending on the true nature of the data. When nonlinear interactions exist, the linear interaction model may not be able to detect such interactions; hence, it suffers from substantial power loss. While the true interaction mechanism (linear or nonlinear) is generally unknown in practice, it is critical to develop statistical methods that can be flexible to capture the underlying interaction mechanism without assuming a specific model assumption. In this study, we develop a mixed kernel function which combines both linear and Gaussian kernels with different weights to capture the linear or nonlinear interaction of two genes. Instead of optimizing the weight function, we propose a grid search strategy and use a Cauchy transformation of the P-values obtained under different weights to aggregate the P-values. We further extend the two-gene interaction model to a high-dimensional setup using a de-biased LASSO algorithm. Extensive simulation studies are conducted to verify the performance of the proposed method. Application to two case studies further demonstrates the utility of the model. Our method provides a flexible and computationally efficient tool for disentangling complex gene-gene interactions associated with complex traits. Yan Liu 0093, Yuzhao Gao, Ruiling Fang, Hongyan Cao, Jian Sa, Jianrong Wang, Hongqi Liu, Tong Wang 0019, Yuehua Cui |
Briefings Bioinform. | 3 |