Guoyou Qin

dblp:14/8929 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-1413-3117ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
YearPublicationVenuePosition
2025 Transfer learning reveals the mediating mechanisms of cross-ethnic lipid metabolic pathways in the association between APOE gene and Alzheimer's disease
abstract
Lipid-mediated effects play a crucial role in elucidating the pathological mechanisms linking the ε4 allele of the apolipoprotein E gene (APOE ε4) to Alzheimer's disease (AD). However, traditional mediation analysis methods often suffer from insufficient statistical power in studies involving minority populations due to limited sample sizes. This study innovatively develops a high-dimensional mediation analysis model (TransHDM) based on a transfer learning framework. By leveraging information from source data with large-scale samples, it significantly enhances the ability to identify potential mediators in small sample target data. The method first constructs a high-dimensional regression model using aggregated data from the source data and target data, then applies transfer regularization to adjust for heterogeneity between the source and target domains, correcting for estimation bias in high-dimensional Lasso. Ultimately, it achieves parameter transfer across domains, addressing statistical bias and inferential uncertainty caused by small sample sizes. Simulation results demonstrate that, compared to traditional methods, this approach significantly improves the power in identifying true mediator variables while effectively controlling the family-wise error rate in multiple testing. When applied to the Alzheimer's Disease Neuroimaging Initiative cohort, TransHDM transferred large-scale data from white and other ethnic groups, identifying additional lipid metabolic pathways mediating the influence of the APOE ε4 allele on AD pathological progression in African American populations compared to pre-transfer analysis. These pathways include glycerophospholipid metabolism, glycerolipid metabolism, sphingolipid metabolism, and ether lipid metabolism (false discovery rate < 0.05). The TransHDM framework not only provides a powerful methodological tool for small sample population research but also offers valuable insights for future research in exploring disease mechanisms and developing biomarkers for disease prediction.
Lulu Pan, Yahang Liu, Ruilang Lin, Yongfu Yu, Guoyou Qin
Briefings Bioinform.6
2025 Deconfounded and debiased estimation for high-dimensional linear regression under hidden confounding with application to omics data
abstract
MOTIVATION: A critical challenge in observational studies arises from the presence of hidden confounders in high-dimensional data. This leads to biases in causal effect estimation due to both hidden confounding and high-dimensional estimation. Some classical deconfounding methods are inadequate for high-dimensional scenarios and typically require prior information on hidden confounders. We propose a two-step deconfounded and debiased estimation for high-dimensional linear regression with hidden confounding. RESULTS: First, we reduce hidden confounding via spectral transformation. Second, we correct bias from the weighted ℓ1 penalty, commonly used in high-dimensional estimation, by inverting the Karush-Kuhn-Tucker conditions and solving convex optimization programs. This deconfounding technique by spectral transformation requires no prior knowledge of hidden confounders. This novel debiasing approach improves over recent work by not assuming a sparse precision matrix, making it more suitable for cases with intrinsic covariate correlations. Simulations show that the proposed method corrects both biases and provides more precise coefficient estimates than existing approaches. We also apply the proposed method to a deoxyribonucleic acid methylation dataset from the Alzheimer's disease (AD) neuroimaging initiative database to investigate the association between cerebrospinal fluid tau protein levels and AD severity. AVAILABILITY AND IMPLEMENTATION: The code for the proposed method is available on GitHub (https://github.com/Li-Zhaoy/Dec-Deb.git) and archived on Zenodo (DOI: https://10.5281/zenodo.15478745).
Yahang Liu, Kecheng Wei, Yongfu Yu, Guoyou Qin, Zhongyi Zhu
Bioinform.5
2025 Debiased machine learning for ultra-high dimensional mediation analysis
abstract
MOTIVATION: In ultra-high dimensional mediation analysis, confounding variables can influence both mediators and outcomes through complex functional forms. While machine learning (ML) approaches are effective at modeling such complex relationships, they can introduce bias when estimating mediation effects. In this article, we propose a debiased ML framework that mitigates this bias, enabling accurate identification of key mediators and precise estimation and inference of their respective contributions. RESULTS: We construct an orthogonalized score function and use cross-fitting to reduce bias introduced by ML. To tackle ultra-high dimensional potential mediators, we implement screening and regularization techniques for variable selection and effect estimation. For statistical inference of the mediators' contributions, we use an adjusted Sobel-type test. Simulation results demonstrate the superior performance of the proposed method in handling complex confounding. Applying this method to Alzheimer's Disease Neuroimaging Initiative data, we identify several cytosine-phosphate-guanine sites where DNA methylation mediates the effect of body mass index on Alzheimer's Disease. AVAILABILITY AND IMPLEMENTATION: The R function DML_HDMA implementing the proposed methods is available online at https://github.com/Wei-Kecheng/DML_HDMA.
Kecheng Wei, Yahang Liu, Ruilang Lin, Yongfu Yu, Guoyou Qin
Bioinform.6
2025 A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data
abstract
Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.
Lulu Pan, Kecheng Wei, Yongfu Yu, Guoyou Qin, Tong Wang 0019
PLoS Comput. Biol.5
2024 High-dimensional generalized median adaptive lasso with application to omics data
abstract
Recently, there has been a growing interest in variable selection for causal inference within the context of high-dimensional data. However, when the outcome exhibits a skewed distribution, ensuring the accuracy of variable selection and causal effect estimation might be challenging. Here, we introduce the generalized median adaptive lasso (GMAL) for covariate selection to achieve an accurate estimation of causal effect even when the outcome follows skewed distributions. A distinctive feature of our proposed method is that we utilize a linear median regression model for constructing penalty weights, thereby maintaining the accuracy of variable selection and causal effect estimation even when the outcome presents extremely skewed distributions. Simulation results showed that our proposed method performs comparably to existing methods in variable selection when the outcome follows a symmetric distribution. Besides, the proposed method exhibited obvious superiority over the existing methods when the outcome follows a skewed distribution. Meanwhile, our proposed method consistently outperformed the existing methods in causal estimation, as indicated by smaller root-mean-square error. We also utilized the GMAL method on a deoxyribonucleic acid methylation dataset from the Alzheimer's disease (AD) neuroimaging initiative database to investigate the association between cerebrospinal fluid tau protein levels and the severity of AD.
Yahang Liu, Kecheng Wei, Yongfu Yu, Guoyou Qin, Tong Wang 0019
Briefings Bioinform.7
2024 Robust double machine learning model with application to omics data
abstract
Recently, there has been a growing interest in combining causal inference with machine learning algorithms. Double machine learning model (DML), as an implementation of this combination, has received widespread attention for their expertise in estimating causal effects within high-dimensional complex data. However, the DML model is sensitive to the presence of outliers and heavy-tailed noise in the outcome variable. In this paper, we propose the robust double machine learning (RDML) model to achieve a robust estimation of causal effects when the distribution of the outcome is contaminated by outliers or exhibits symmetrically heavy-tailed characteristics. In the modelling of RDML model, we employed median machine learning algorithms to achieve robust predictions for the treatment and outcome variables. Subsequently, we established a median regression model for the prediction residuals. These two steps ensure robust causal effect estimation. Simulation study show that the RDML model is comparable to the existing DML model when the data follow normal distribution, while the RDML model has obvious superiority when the data follow mixed normal distribution and t-distribution, which is manifested by having a smaller RMSE. Meanwhile, we also apply the RDML model to the deoxyribonucleic acid methylation dataset from the Alzheimer’s disease (AD) neuroimaging initiative database with the aim of investigating the impact of Cerebrospinal Fluid Amyloid $$\upbeta$$ 42 (CSF A $$\upbeta$$ 42) on AD severity. These findings illustrate that the RDML model is capable of robustly estimating causal effect, even when the outcome distribution is affected by outliers or displays symmetrically heavy-tailed properties.
Xuqing Wang, Yahang Liu, Guoyou Qin, Yongfu Yu
BMC Bioinform.3