EDBT 2026 Demo / reviewers in the wild / expert
Yuehua Cui
dblp:80/6905
· DBLP profile ↗
23ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CEDR: robust consensus cancer subtyping with multi-omics data via ensemble dimensionality reductionabstractCancer is a highly heterogeneous disease underpinned by complex molecular alterations. Accurate subtyping is critical for guiding personalized treatment and improving clinical outcomes. However, multi-omics data are high-dimensional, noisy, and heterogeneous across platforms, posing major challenges for reliable subtyping. To address this, dimensionality reduction is necessary to capture underlying molecular patterns in a low-dimensional space, facilitating both computational efficiency and biological interpretation. We present Consensus subtyping method with Ensemble Dimensionality Reduction for multi-omics data integration (CEDR), a consensus subtyping framework that integrates complementary linear and nonlinear dimensionality reduction methods with robust clustering and probabilistic ensemble modeling. Different from existing dimensionality reduction techniques, our framework adopts an ensemble learning framework that integrates multiple dimensionality reduction techniques with robust clustering to achieve reliable consensus cancer subtyping. We apply Optimally Tuned Robust Improper Maximum Likelihood Estimator to the concatenated low-dimensional matrix for robust subtyping, and ensemble the result with the Mixture Model for Clustering Ensembles to identify stable subtypes. Across extensive simulations, CEDR consistently outperformed conventional dimensionality reduction-based clustering, the Cluster Of Clusters Analysis (COCA) ensemble strategy, and state-of-the-art multi-omics integration algorithms (SNF and CIMLR) in both accuracy and robustness. Application to clear cell renal cell carcinoma and lower-grade glioma revealed biologically interpretable subtypes characterized by distinctive survival outcomes, pathway activities, and immune infiltration patterns. These findings demonstrate that CEDR provides a powerful and reliable strategy for multi-omics data integration and cancer subtyping, with strong potential for broader applications in high-dimensional multimodal data analysis. Hongyan Cao, Zhaoyang Xu, Shilong Lin, Gang Du, Tong Wang 0019, Juping Wang, Ruiling Fang, Ping Zeng, Hongmei Yu, Yuehua Cui |
Briefings Bioinform. | 13 |
| 2026 | Nonlinear kernel-based high-dimensional inference for set-based genetic association studiesabstractNonlinear genetic architectures, including epistasis and threshold effects, are increasingly recognized as contributors to complex disease risk, yet most existing SNP-set association tests rely on linear modeling assumptions, resulting in reduced power and unstable inference when genetic effects are nonlinear or heterogeneously distributed across variants. To address this limitation, we propose a nonlinear high-dimensional inference framework for set-based genetic association analysis that integrates scalable kernel representations with valid statistical inference. The framework combines distance correlation-based sure independence screening to reduce ultra-high dimensional predictors, kernel principal component analysis with Nyström approximation for nonlinear feature extraction, and de-sparsified LASSO to enable asymptotically valid hypothesis testing in high dimensions, together with a two-stage omnibus testing strategy that adaptively aggregates evidence across complementary signal models. Extensive simulation studies demonstrate that the proposed method maintains well-calibrated Type I error and consistently achieves higher power than established set-based approaches, including Sequence Kernel Association Test and adaptive Sum of Powered Score test, particularly under nonlinear and heterogeneous genetic effect scenarios, while remaining competitive in linear settings. Application to Alzheimer's Disease Neuroimaging Initiative data identifies gene-level associations with brain regional volumes that converge on neuronal excitability, calcium signaling, and cytoskeletal regulation, biological processes centrally implicated in neurodegeneration. Together, this work provides a robust and scalable framework for nonlinear set-based inference in genome-wide studies, expanding the analytical toolbox for dissecting complex genetic contributions to disease. Meilin Zhu, Fuzhao Chen, Yuehua Cui |
Briefings Bioinform. | 7 |
| 2025 | Multi-omics data integration for enhanced cancer subtyping via interactive multi-kernel learningabstractCancer is a highly heterogeneous disease characterized by complex molecular changes. Subtypes identified through multi-omics data hold significant promise for improving prognosis and facilitating personalized precision treatment. Recent multi-omics integration methods have mostly focused on capturing complementary information from different data types, often overlooking potential interactions between omics data. Here we develop a novel method named interactive multi-kernel learning (iMKL), which incorporates omics-omics interactions alongside heterogeneous data types under the unsupervised multi-kernel learning framework, to improve subtype identification. Using the sample-similarity kernel for each dataset, we propose a joint Hadamard product strategy to capture higher-order interactive effects from different omics data types. We applied iMKL to two renal cell carcinoma (RCC) datasets-clear renal cell carcinoma (ccRCC) and type II papillary renal cell carcinoma (type II pRCC)-both including miRNA expression, mRNA expression, and DNA methylation data. Stability analysis through random sampling of patients or features demonstrated that iMKL exhibits strong robustness and accuracy in identifying patient subtypes. The identified subtypes revealed dramatic differences in patient survival, with both ccRCC and type II pRCC classified into three distinct subtypes. The findings in the real application highlight potential biomarkers associated with adverse patient outcomes and demonstrate substantial advancement in cancer subtype identification. The iMKL method effectively identifies tumor molecular subtypes that are strongly associated with clinical features and survival rates, providing valuable insights for accurate cancer subtyping, clinical decision-making, and the realization of personalized treatment strategies. Hongyan Cao, Tong Wang 0019, Zhaoyang Xu, Gaiqin Liu, Ruiling Fang, Ping Zeng, Hongmei Yu, Yuehua Cui |
Briefings Bioinform. | 12 |
| 2024 | BayesKAT: bayesian optimal kernel-based test for genetic association studies reveals joint genetic effects in complex diseasesabstractGenome-wide Association Studies (GWAS) methods have identified individual single-nucleotide polymorphisms (SNPs) significantly associated with specific phenotypes. Nonetheless, many complex diseases are polygenic and are controlled by multiple genetic variants that are usually non-linearly dependent. These genetic variants are marginally less effective and remain undetected in GWAS analysis. Kernel-based tests (KBT), which evaluate the joint effect of a group of genetic variants, are therefore critical for complex disease analysis. However, choosing different kernel functions in KBT can significantly influence the type I error control and power, and selecting the optimal kernel remains a statistically challenging task. A few existing methods suffer from inflated type 1 errors, limited scalability, inferior power or issues of ambiguous conclusions. Here, we present a new Bayesian framework, BayesKAT (https://github.com/wangjr03/BayesKAT), which overcomes these kernel specification issues by selecting the optimal composite kernel adaptively from the data while testing genetic associations simultaneously. Furthermore, BayesKAT implements a scalable computational strategy to boost its applicability, especially for high-dimensional cases where other methods become less effective. Based on a series of performance comparisons using both simulated and real large-scale genetics data, BayesKAT outperforms the available methods in detecting complex group-level associations and controlling type I errors simultaneously. Applied on a variety of groups of functionally related genetic variants based on biological pathways, co-expression gene modules and protein complexes, BayesKAT deciphers the complex genetic basis and provides mechanistic insights into human diseases. Sikta Das Adhikari, Yuehua Cui, Jianrong Wang |
Briefings Bioinform. | 2 |
| 2024 | A metagene based similarity network fusion approach for multi-omics data integration identified novel subtypes in renal cell carcinomaabstractRenal cell carcinoma (RCC) ranks among the most prevalent cancers worldwide, with both incidence and mortality rates increasing annually. The heterogeneity among RCC patients presents considerable challenges for developing universally effective treatment strategies, emphasizing the necessity of in-depth research into RCC's molecular mechanisms, understanding the variations among RCC patients and further identifying distinct molecular subtypes for precise treatment. We proposed a metagene-based similarity network fusion (Meta-SNF) method for RCC subtype identification with multi-omics data, using a non-negative matrix factorization technique to capture alternative structures inherent in the dataset as metagenes. These latent metagenes were then integrated to construct a fused network under the Similarity Network Fusion (SNF) framework for more precise subtyping. We conducted simulation studies and analyzed real-world data from two RCC datasets, namely kidney renal clear cell carcinoma (KIRC) and kidney renal papillary cell carcinoma (KIRP) to demonstrate the utility of Meta-SNF. The simulation studies indicated that Meta-SNF achieved higher accuracy in subtype identification compared with the original SNF and other state-of-the-art methods. In analyses of real data, Meta-SNF produced more distinct and well-separated clusters, classifying both KIRC and KIRP into four subtypes with significant differences in survival outcomes. Subsequently, we performed comprehensive bioinformatics analyses focused on subtypes with poor prognoses in KIRC and KIRP and identified several potential biomarkers. Meta-SNF offers a novel strategy for subtype identification using multi-omics data, and its application to RCC datasets has yielded diverse biological insights which are highly valuable for informing clinical decision-making processes in the treatment of RCC. Congcong Jia, Tong Wang 0019, Dingtong Cui, Yaxin Tian, Gaiqin Liu, Zhaoyang Xu, Ruiling Fang, Hongmei Yu, Yuehua Cui, Hongyan Cao |
Briefings Bioinform. | 11 |
| 2024 | A high-dimensional omnibus test for set-based association analysisabstractSet-based association analysis is a valuable tool in studying the etiology of complex diseases in genome-wide association studies, as it allows for the joint testing of variants in a region or group. Two common types of single nucleotide polymorphism (SNP)-disease functional models are recognized when evaluating the joint function of a set of SNP: the cumulative weak signal model, in which multiple functional variants with small effects contribute to disease risk, and the dominating strong signal model, in which a few functional variants with large effects contribute to disease risk. However, existing methods have two main limitations that reduce their power. Firstly, they typically only consider one disease-SNP association model, which can result in significant power loss if the model is misspecified. Secondly, they do not account for the high-dimensional nature of SNPs, leading to low power or high false positives. In this study, we propose a solution to these challenges by using a high-dimensional inference procedure that involves simultaneously fitting many SNPs in a regression model. We also propose an omnibus testing procedure that employs a robust and powerful P-value combination method to enhance the power of SNP-set association. Our results from extensive simulation studies and a real data analysis demonstrate that our set-based high-dimensional inference strategy is both flexible and computationally efficient and can substantially improve the power of SNP-set association analysis. Application to a real dataset further demonstrates the utility of the testing strategy. Fuzhao Chen, Hongyan Cao, Lina Yan, Xia Gao, Yuehua Cui |
Briefings Bioinform. | 9 |
| 2024 | Uncertainty quantification in high-dimensional linear models incorporating graphical structures with applications to gene set analysisabstractMOTIVATION: The functions of genes in networks are typically correlated due to their functional connectivity. Variable selection methods have been developed to select important genes associated with a trait while incorporating network graphical information. However, no method has been proposed to quantify the uncertainty of individual genes under such settings. RESULTS: In this paper, we construct confidence intervals (CIs) and provide P-values for parameters of a high-dimensional linear model incorporating graphical structures where the number of variables p diverges with the number of observations. For combining the graphical information, we propose a graph-constrained desparsified LASSO (least absolute shrinkage and selection operator) (GCDL) estimator, which reduces dramatically the influence of high correlation of predictors and enjoys the advantage of faster computation and higher accuracy compared with the desparsified LASSO. Theoretical results show that the GCDL estimator achieves asymptotic normality. The asymptotic property of the uniform convergence is established, with which an explicit expression of the uniform CI can be derived. Extensive numerical results indicate that the GCDL estimator and its (uniform) CI perform well even when predictors are highly correlated. AVAILABILITY AND IMPLEMENTATION: An R package implementing the proposed method is available at https://github.com/XiaoZhangryy/gcdl. Xiangyong Tan, Yuehua Cui, Xu Liu 0024 |
Bioinform. | 3 |
| 2023 | Cancer subtyping with heterogeneous multi-omics data via hierarchical multi-kernel learningabstractDifferentiating cancer subtypes is crucial to guide personalized treatment and improve the prognosis for patients. Integrating multi-omics data can offer a comprehensive landscape of cancer biological process and provide promising ways for cancer diagnosis and treatment. Taking the heterogeneity of different omics data types into account, we propose a hierarchical multi-kernel learning (hMKL) approach, a novel cancer molecular subtyping method to identify cancer subtypes by adopting a two-stage kernel learning strategy. In stage 1, we obtain a composite kernel borrowing the cancer integration via multi-kernel learning (CIMLR) idea by optimizing the kernel parameters for individual omics data type. In stage 2, we obtain a final fused kernel through a weighted linear combination of individual kernels learned from stage 1 using an unsupervised multiple kernel learning method. Based on the final fusion kernel, k-means clustering is applied to identify cancer subtypes. Simulation studies show that hMKL outperforms the one-stage CIMLR method when there is data heterogeneity. hMKL can estimate the number of clusters correctly, which is the key challenge in subtyping. Application to two real data sets shows that hMKL identified meaningful subtypes and key cancer-associated biomarkers. The proposed method provides a novel toolkit for heterogeneous multi-omics data integration and cancer subtypes identification. Yifang Wei, Lingmei Li, Jian Sa, Hongyan Cao, Yuehua Cui |
Briefings Bioinform. | 7 |
| 2022 | Gene set analysis with graph-embedded kernel association testabstractMOTIVATION: Kernel-based association test (KAT) has been a popular approach to evaluate the association of expressions of a gene set (e.g. pathway) with a phenotypic trait. KATs rely on kernel functions which capture the sample similarity across multiple features, to capture potential linear or non-linear relationship among features in a gene set. When calculating the kernel functions, no network graphical information about the features is considered. While genes in a functional group (e.g. a pathway) are not independent in general due to regulatory interactions, incorporating regulatory network (or graph) information can potentially increase the power of KAT. In this work, we propose a graph-embedded kernel association test, termed gKAT. gKAT incorporates prior pathway knowledge when constructing a kernel function into hypothesis testing. RESULTS: We apply a diffusion kernel to capture any graph structures in a gene set, then incorporate such information to build a kernel function for further association test. We illustrate the geometric meaning of the approach. Through extensive simulation studies, we show that the proposed gKAT algorithm can improve testing power compared to the one without considering graph structures. Application to a real dataset further demonstrate the utility of the method. AVAILABILITY AND IMPLEMENTATION: The R code used for the analysis can be accessed at https://github.com/JialinQu/gKAT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jialin Qu, Yuehua Cui |
Bioinform. | 2 |
| 2021 | Gene-based mediation analysis in epigenetic studiesabstractMediation analysis has been a useful tool for investigating the effect of mediators that lie in the path from the independent variable to the outcome. With the increasing dimensionality of mediators such as in (epi)genomics studies, high-dimensional mediation model is needed. In this work, we focus on epigenetic studies with the goal to identify important DNA methylations that act as mediators between an exposure disease outcome. Specifically, we focus on gene-based high-dimensional mediation analysis implemented with kernel principal component analysis to capture potential nonlinear mediation effect. We first review the current high-dimensional mediation models and then propose two gene-based analytical approaches: gene-based high-dimensional mediation analysis based on linearity assumption between mediators and outcome (gHMA-L) and gene-based high-dimensional mediation analysis based on nonlinearity assumption (gHMA-NL). Since the underlying true mediation relationship is unknown in practice, we further propose an omnibus test of gene-based high-dimensional mediation analysis (gHMA-O) by combing gHMA-L and gHMA-NL. Extensive simulation studies show that gHMA-L performs better under the model linear assumption and gHMA-NL does better under the model nonlinear assumption, while gHMA-O is a more powerful and robust method by combining the two. We apply the proposed methods to two datasets to investigate genes whose methylation levels act as important mediators in the relationship: (1) between alcohol consumption and epithelial ovarian cancer risk using data from the Mayo Clinic Ovarian Cancer Case-Control Study and (2) between childhood maltreatment and comorbid post-traumatic stress disorder and depression in adulthood using data from the Gray Trauma Project. Ruiling Fang, Yuzhao Gao, Hongyan Cao, Ellen L. Goode, Yuehua Cui |
Briefings Bioinform. | 6 |
| 2021 | Identifying complex gene-gene interactions: a mixed kernel omnibus testing approachabstractGenes do not function independently; rather, they interact with each other to fulfill their joint tasks. Identification of gene-gene interactions has been critically important in elucidating the molecular mechanisms responsible for the variation of a phenotype. Regression models are commonly used to model the interaction between two genes with a linear product term. The interaction effect of two genes can be linear or nonlinear, depending on the true nature of the data. When nonlinear interactions exist, the linear interaction model may not be able to detect such interactions; hence, it suffers from substantial power loss. While the true interaction mechanism (linear or nonlinear) is generally unknown in practice, it is critical to develop statistical methods that can be flexible to capture the underlying interaction mechanism without assuming a specific model assumption. In this study, we develop a mixed kernel function which combines both linear and Gaussian kernels with different weights to capture the linear or nonlinear interaction of two genes. Instead of optimizing the weight function, we propose a grid search strategy and use a Cauchy transformation of the P-values obtained under different weights to aggregate the P-values. We further extend the two-gene interaction model to a high-dimensional setup using a de-biased LASSO algorithm. Extensive simulation studies are conducted to verify the performance of the proposed method. Application to two case studies further demonstrates the utility of the model. Our method provides a flexible and computationally efficient tool for disentangling complex gene-gene interactions associated with complex traits. Yan Liu 0093, Yuzhao Gao, Ruiling Fang, Hongyan Cao, Jian Sa, Jianrong Wang, Hongqi Liu, Tong Wang 0019, Yuehua Cui |
Briefings Bioinform. | 9 |
| 2020 | A Novel Quality Enhanced Low Complexity Rate Control Algorithm for HEVCabstractRate control (RC) is a key technology in video coding which is mainly responsible for adapting the compressed video quality as much as possible under limited bandwidth. Typical RC methods consist of initial quantization parameters of I-frames decision and control parameters updating procedures. However, on the one hand, I-frames quantization parameter (QP) is decided by the sum of absolute transformed difference (SATD) computation, which is time consuming for low delay applications. On the other hand, the RC parameters are only updated ignoring the distortion characteristics for inter frame, which cannot achieve optimal rate distortion (RD) performance. Therefore, a novel RC method which considers distortion characteristics for model parameters updating for quality enhancement is proposed in this paper, including a low complexity based I-frame QP decision strategy of low delay applications. For which estimated distortion characteristics and the previous QPs are employed, respectively. Through the experimental results, with more accurate RC precision and negligible I-frame QP computation time, the proposed rate control scheme improves the performance gain of 2.6% bitrate savings for the whole test sequences on average in HM 16.9. ChungWen Ku, Guoqing Xiang, Huizhu Jia, Yuehua Cui, Yuan Li 0014 |
VCIP | 5 |
| 2020 | Multilevel heterogeneous omics data integration with kernel fusionabstractHigh-throughput omics data are generated almost with no limit nowadays. It becomes increasingly important to integrate different omics data types to disentangle the molecular machinery of complex diseases with the hope for better disease prevention and treatment. Since the relationship among different omics data features are typically unknown, a supervised learning model assuming a particular distribution with a specific structure will not serve the purpose to capture the underlying complex relationship between multiple features and a disease phenotype. In this work, we briefly reviewed methods for kernel fusion (KF) based on support vector machine and kernel partial least squares (KPLS) algorithms. We then proposed a fused KPLS (fKPLS) model for disease classification and prediction with multilevel omics data. The fused kernel can deal with effect heterogeneity in which different omic data types may have different effect contribution to the trait of interest, with the purpose to improve the prediction performance. We proposed to optimize the kernel parameters and kernel weights with the genetic algorithm (GA). The proposed GA-fKPLS model can substantially improve disease classification performance by integrating multiple omics data types, demonstrated via extensive simulations and real data analysis. With properly defined fitness functions during GA optimization, the proposed KF method can be extended to other kernel-based analyses such as in kernel association analysis with common or rare variants. Hongyan Cao, Tong Wang 0019, Yuehua Cui |
Briefings Bioinform. | 5 |
| 2020 | Comparison of methods for the detection of outliers and associated biomarkers in mislabeled omics dataabstractBACKGROUND: Previous studies have reported that labeling errors are not uncommon in omics data. Potential outliers may severely undermine the correct classification of patients and the identification of reliable biomarkers for a particular disease. Three methods have been proposed to address the problem: sparse label-noise-robust logistic regression (Rlogreg), robust elastic net based on the least trimmed square (enetLTS), and Ensemble. Ensemble is an ensembled classification based on distinct feature selection and modeling strategies. The accuracy of biomarker selection and outlier detection of these methods needs to be evaluated and compared so that the appropriate method can be chosen. RESULTS: The accuracy of variable selection, outlier identification, and prediction of three methods (Ensemble, enetLTS, Rlogreg) were compared for simulated and an RNA-seq dataset. On simulated datasets, Ensemble had the highest variable selection accuracy, as measured by a comprehensive index, and lowest false discovery rate among the three methods. When the sample size was large and the proportion of outliers was ≤5%, the positive selection rate of Ensemble was similar to that of enetLTS. However, when the proportion of outliers was 10% or 15%, Ensemble missed some variables that affected the response variables. Overall, enetLTS had the best outlier detection accuracy with false positive rates < 0.05 and high sensitivity, and enetLTS still performed well when the proportion of outliers was relatively large. With 1% or 2% outliers, Ensemble showed high outlier detection accuracy, but with higher proportions of outliers Ensemble missed many mislabeled samples. Rlogreg and Ensemble were less accurate in identifying outliers than enetLTS. The prediction accuracy of enetLTS was better than that of Rlogreg. Running Ensemble on a subset of data after removing the outliers identified by enetLTS improved the variable selection accuracy of Ensemble. CONCLUSIONS: When the proportion of outliers is ≤5%, Ensemble can be used for variable selection. When the proportion of outliers is > 5%, Ensemble can be used for variable selection on a subset after removing outliers identified by enetLTS. For outlier identification, enetLTS is the recommended method. In practice, the proportion of outliers can be estimated according to the inaccuracy of the diagnostic methods used. Yuehua Cui, Tong Wang 0019 |
BMC Bioinform. | 2 |
| 2018 | Learning Collaborative Model for Visual TrackingabstractThis paper proposes a robust visual tracking method by designing a collaborative model. The collaborative model employs a two-stage tracker and a HOG-based detector, which exploits both holistic and local information of the target. The two-stage tracker learns a linear classifier from the patches of original images and the HOG-based detector trains a linear discriminant analysis classifier with the object exemplar. Finally, a result decision making strategy is developed by considering both the original template and the appearance variations, making the tracker and the detector collaborate with each other. The proposed method has been evaluated on OTB-50, OTB-100 and Temple-Color datasets, and results demonstrate that the proposed method is able to effectively address the challenging cases such as scale variation and out-of-view and gets better performance than the state-of-the-art trackers. Ding Ma 0001, Wei Bu, Yuehua Cui, Xiangqian Wu 0002 |
ICPR | 3 |
| 2018 | Segmentation-Guided Tracking with Prior Map DecisionabstractFor visual tracking, the target object is represented by an appearance model and the location of the target is estimated in each frame. Numerous tracking algorithms model the appearance of the target with a confidence score and rarely take into account the semantic information of the target. In this paper, we propose an efficient tracking algorithm that models the appearance of the target based on semantic segmentation. The overall architecture consists of two parts: the segmentation part and the tracking part. In the segmentation part, an attention model is employed, providing spatial highlights of the candidate region of the target. In the tracking part, the tracker is constructed by an online updated convolutional neural networks to identify the target in subsequent frames, taking advantage of the segmentation information of the target from the segmentation part. To enhance the performance of this architecture, we design an incremental updated prior map taking both the segmentation signal and the tracking signal into consideration. Extensive experiments on two benchmarks including OTB-50, OTB-100, and Temple-Color, show that the proposed method outperforms other trackers. Ding Ma 0001, Wei Bu, Yuehua Cui, Xiangqian Wu 0002 |
ICPR | 4 |
| 2017 | Predicting disease trait with genomic data: a composite kernel approachabstractWith the advancement of biotechniques, a vast amount of genomic data is generated with no limit. Predicting a disease trait based on these data offers a cost-effective and time-efficient way for early disease screening. Here we proposed a composite kernel partial least squares (CKPLS) regression model for quantitative disease trait prediction focusing on genomic data. It can efficiently capture nonlinear relationships among features compared with linear learning algorithms such as Least Absolute Shrinkage and Selection Operator or ridge regression. We proposed to optimize the kernel parameters and kernel weights with the genetic algorithm (GA). In addition to improved performance for parameter optimization, the proposed GA-CKPLS approach also has better learning capacity and generalization ability compared with single kernel-based KPLS method as well as other nonlinear prediction models such as the support vector regression. Extensive simulation studies demonstrated that GA-CKPLS had better prediction performance than its counterparts under different scenarios. The utility of the method was further demonstrated through two case studies. Our method provides an efficient quantitative platform for disease trait prediction based on increasing volume of omics data. Shaoyu Li, Hongyan Cao, Chichen Zhang, Yuehua Cui |
Briefings Bioinform. | 5 |
| 2015 | Learning directed acyclic graphical structures with genetical genomics dataabstractMOTIVATION: Large amount of research efforts have been focused on estimating gene networks based on gene expression data to understand the functional basis of a living organism. Such networks are often obtained by considering pairwise correlations between genes, thus may not reflect the true connectivity between genes. By treating gene expressions as quantitative traits while considering genetic markers, genetical genomics analysis has shown its power in enhancing the understanding of gene regulations. Previous works have shown the improved performance on estimating the undirected network graphical structure by incorporating genetic markers as covariates. Knowing that gene expressions are often due to directed regulations, it is more meaningful to estimate the directed graphical network. RESULTS: In this article, we introduce a covariate-adjusted Gaussian graphical model to estimate the Markov equivalence class of the directed acyclic graphs (DAGs) in a genetical genomics analysis framework. We develop a two-stage estimation procedure to first estimate the regression coefficient matrix by [Formula: see text] penalization. The estimated coefficient matrix is then used to estimate the mean values in our multi-response Gaussian model to estimate the regulatory networks of gene expressions using PC-algorithm. The estimation consistency for high dimensional sparse DAGs is established. Simulations are conducted to demonstrate our theoretical results. The method is applied to a human Alzheimer's disease dataset in which differential DAGs are identified between cases and controls. R code for implementing the method can be downloaded at http://www.stt.msu.edu/∼cui. AVAILABILITY AND IMPLEMENTATION: R code for implementing the method is freely available at http://www.stt.msu.edu/∼cui/software.html. Yuehua Cui |
Bioinform. | 2 |
| 2014 | Boosting signals in gene-based association studies via efficient SNP selectionabstractSet-based association studies based on genes or pathways have shown great promise in interpreting association signals associated with complex diseases. These approaches are particularly useful when variants in a set have moderate effects and are difficult to be detected with single marker analysis, especially when variants function jointly in a complicated manner. The set-based analyses use a summary statistic such as the maximum or average of individual signal (e.g. a chi-square statistic) over all variants in a set, or consider their joint distribution to assess the significance of the set. The signal obtained with this treatment, however, could be potentially diluted when noisy variants are not taken good care of, leading to either inflated false negatives or false positives. Thus, the selection of disease informative single-nucleotide polymorphism (diSNPs) plays a crucial role in improving the power of the set-based association study. In this work, we propose an efficient diSNP selection method based on the information theory. We select diSNP variants by considering their relative information contribution to a disease status, which is different from the usual tag SNP selection. The relative merit of pre-selecting diSNPs in a set-based association analysis is demonstrated through extensive simulation studies and real data analysis. Cen Wu, Yuehua Cui |
Briefings Bioinform. | 2 |
| 2014 | A new set-valued system identification approach to identifying rare genetic variants for ordered categorical phenotypeabstractabstract Wenjian Bi, Guolian Kang, Yuehua Cui, Christine Hartford, Wing Leung, Ji-Feng Zhang |
BMC Bioinform. | 3 |
| 2012 | Bayesian inference for genomic imprinting underlying developmental characteristicsabstractThe identification of imprinted genes is becoming a standard procedure in searching for quantitative trait loci (QTL) underlying complex traits. When a developmental characteristic such as growth or drug response is observed at multiple time points, understanding the dynamics of gene function governing the underlying feature should provide more biological information regarding the genetic control of an organism. Recognizing that differential imprinting can be development-specific, mapping imprinted genes considering the dynamic imprinting effect can provide additional biological insights into the epigenetic control of a complex trait. In this study, we proposed a Bayesian imprinted QTL (iQTL) mapping framework considering the dynamics of imprinting effects and model multiple iQTLs with an efficient Bayesian model selection procedure. The method overcomes the limitation of likelihood-based mapping procedure, and can simultaneously identify multiple iQTLs with different gene action modes across the whole genome with high computational efficiency. An inference procedure using Bayes factors to distinguish different imprinting patterns of iQTL was proposed. Monte Carlo simulations were conducted to evaluate the performance of the method. The utility of the approach was illustrated through an analysis of a body weight growth data set in an F(2) family derived from LG/J and SM/J mouse stains. The proposed Bayesian mapping method provides an efficient and computationally feasible framework for genome-wide multiple iQTL inference with complex developmental traits. Runqing Yang, Yuehua Cui |
Briefings Bioinform. | 3 |
| 2011 | Varying coefficient model for gene-environment interaction: a non-linear lookabstractMOTIVATION: The genetic basis of complex traits often involves the function of multiple genetic factors, their interactions and the interaction between the genetic and environmental factors. Gene-environment (G×E) interaction is considered pivotal in determining trait variations and susceptibility of many genetic disorders such as neurodegenerative diseases or mental disorders. Regression-based methods assuming a linear relationship between a disease response and the genetic and environmental factors as well as their interaction is the commonly used approach in detecting G×E interaction. The linearity assumption, however, could be easily violated due to non-linear genetic penetrance which induces non-linear G×E interaction. RESULTS: In this work, we propose to relax the linear G×E assumption and allow for non-linear G×E interaction under a varying coefficient model framework. We propose to estimate the varying coefficients with regression spline technique. The model allows one to assess the non-linear penetrance of a genetic variant under different environmental stimuli, therefore help us to gain novel insights into the etiology of a complex disease. Several statistical tests are proposed for a complete dissection of G×E interaction. A wild bootstrap method is adopted to assess the statistical significance. Both simulation and real data analysis demonstrate the power and utility of the proposed method. Our method provides a powerful and testable framework for assessing non-linear G×E interaction. Shujie Ma, Roberto Romero, Yuehua Cui |
Bioinform. | 4 |
| 2005 | Mapping genome-genome epistasis: a high-dimensional modelabstractMOTIVATION: The proper development of any organ or tissue requires the coordinated expression of its underlying genes that can be located on different genomes present in an organism. For instance, each step in the development of seed for a higher plant is the consequence of gene interactions from the maternal, embryo and endosperm genomes. RESULTS: We present a multivariate statistical model for mapping quantitative trait loci (QTL) by incorporating two important aspects of seed development in plants-QTL interactions derived from different genomes, the maternal, embryo and endosperm, and genetic correlations among phenotypic traits expressed in different genome-specific tissues. This model, which has a high dimensionality, is constructed within the maximum-likelihood context based on a finite mixture model. The implementation of the expectation-maximization algorithm allows for the efficient estimation of QTL positions, their action and interaction effects and pleiotropic effects. The application of this high-dimensional model to a real rice dataset has validated its usefulness. CONCLUSIONS: Our model was derived for self-pollinated plants, but it can be extended to cross-pollinated plants and to animals. With the burgeoning of genetic and genomic data, this high-dimensional model will have many implications for agricultural and evolutionary genetic research. AVAILABILITY: A package of software will be provided from the corresponding author upon request. Yuehua Cui, Rongling Wu |
Bioinform. | 1 |