VLDB 2026 Research / reviewers in the wild / expert
Kuangnan Fang
dblp:13/10049
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GE-IA-NAM: gene-environment interaction analysis via imaging-assisted neural additive modelabstractMOTIVATION: Gene-environment (G-E) interaction analysis is crucial in cancer research, offering insights into how genetic and environmental factors jointly influence cancer outcomes. Most existing G-E interaction methods are regression-based, which may lack flexibility to capture complex data patterns. Recent advances have investigated deep neural network-based G-E models. However, these methods may be more vulnerable to information deficiency due to challenges such as limited sample size and high dimensionality. Apart from genetic and environmental data, pathological images have emerged as a widely accessible and informative resource for cancer modeling, presenting its potential to enhance G-E modeling. RESULTS: We propose the pathological imaging-assisted neural additive model for G-E analysis (GE-IA-NAM). The flexible and interpretable additive network architecture is adopted to account for individualized effects associated with genetic factors, environmental factors, and their interactions. To improve G-E modeling, an assisted-learning strategy is investigated, which adopts a joint analysis to integrate information from pathological images. Simulations and the analysis of lung and skin cancer datasets from The Cancer Genome Atlas demonstrate the competitive performance of the proposed method. AVAILABILITY AND IMPLEMENTATION: Python code implementing the proposed method is available at https://github.com/Mr-maoge/NAM-IA-GE. The data that support the findings in this article are openly available in TCGA (The Cancer Genome Atlas) at https://portal.gdc.cancer.gov/. Jingmao Li, Yaqing Xu, Shuangge Ma, Kuangnan Fang |
Bioinform. | 4 |
| 2025 | Fraud Detection by Integrating Multisource Heterogeneous Presence-Only DataabstractIn credit fraud detection practice, certain fraudulent transactions often evade detection because of the hidden nature of fraudulent behavior. To address this issue, an increasing number of positive-unlabeled (PU) learning techniques have been employed by more and more financial institutions. However, most of these methods are designed for single data sets and do not take into account the heterogeneity of data when they are collected from different sources. In this paper, we propose an integrative PU learning method (I-PU) for pooling information from multiple heterogeneous PU data sets. A novel approach that penalizes group differences is developed to explicitly and automatically identify the cluster structures of coefficients across different data sets, thus offering a plausible interpretation of heterogeneity. Furthermore, we apply a bilevel selection method to detect the sparse structure at both the group level and within-group level. Theoretically, we show that our proposed estimator has the oracle property. Computationally, we design an expectation-maximization (EM) algorithm framework and propose an alternating direction method of multipliers (ADMM) algorithm to solve it. Simulation results show that our proposed method has better numerical performance in terms of variable selection, parameter estimation, and prediction ability. Finally, a real-world application showcases the effectiveness of our method in identifying distinct coefficient clusters and its superior prediction performance compared with direct data merging or separate modeling. This result also offers valuable insights for financial institutions in developing targeted fraud detection systems. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: This work was supported by the National Natural Science Foundation of China [Grants 72071169, 72231005, 72233002, and 72471169], the Fundamental Research Funds for the Central Universities of China [Grant 20720231060], the National Social Science Fund of China [Grant 21&ZD146], and Shuimu Tsinghua Scholar Program. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0366 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0366 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Yongqin Qiu, Yuanxing Chen, Kan Fang, Lean Yu, Kuangnan Fang |
INFORMS J. Comput. | 5 |
| 2025 | Joint modeling of mixed outcomes using a rank-based sparse neural network
Jiajing Xue, Yaqing Xu, Jingmao Li, Shuangge Ma, Kuangnan Fang |
J. Biomed. Informatics | 5 |
| 2024 | Heterogeneity-aware Clustered Distributed Learning for Multi-source Data AnalysisabstractIn diverse fields ranging from finance to omics, it is increasingly common that data is distributed with multiple individual sources (referred to as “clients” in some studies). Integrating raw data, although powerful, is often not feasible, for example, when there are considerations on privacy protection. Distributed learning techniques have been developed to integrate summary statistics as opposed to raw data. In many existing distributed learning studies, it is stringently assumed that all the clients have the same model. To accommodate data heterogeneity, some federated learning methods allow for client-specific models. In this article, we consider the scenario that clients form clusters, those in the same cluster have the same model, and different clusters have different models. Further considering the clustering structure can lead to a better understanding of the “interconnections” among clients and reduce the number of parameters. To this end, we develop a novel penalization approach. Specifically, group penalization is imposed for regularized estimation and selection of important variables, and fusion penalization is imposed to automatically cluster clients. An effective ADMM algorithm is developed, and the estimation, selection, and clustering consistency properties are established under mild conditions. Simulation and data analysis further demonstrate the practical utility and superiority of the proposed approach. Yuanxing Chen, Qingzhao Zhang 0002, Shuangge Ma, Kuangnan Fang |
J. Mach. Learn. Res. | 4 |
| 2023 | FunctanSNP: an R package for functional analysis of dense SNP data (with interactions)abstractSUMMARY: Densely measured SNP data are routinely analyzed but face challenges due to its high dimensionality, especially when gene-environment interactions are incorporated. In recent literature, a functional analysis strategy has been developed, which treats dense SNP measurements as a realization of a genetic function and can 'bypass' the dimensionality challenge. However, there is a lack of portable and friendly software, which hinders practical utilization of these functional methods. We fill this knowledge gap and develop the R package FunctanSNP. This comprehensive package encompasses estimation, identification, and visualization tools and has undergone extensive testing using both simulated and real data, confirming its reliability. FunctanSNP can serve as a convenient and reliable tool for analyzing SNP and other densely measured data. AVAILABILITY AND IMPLEMENTATION: The package is available at https://CRAN.R-project.org/package=FunctanSNP. Kuangnan Fang, Qingzhao Zhang 0002, Shuangge Ma |
Bioinform. | 2 |
| 2022 | iSFun: an R package for integrative dimension reduction analysisabstractSUMMARY: In the analysis of high-dimensional omics data, dimension reduction techniques-including principal component analysis (PCA), partial least squares (PLS) and canonical correlation analysis (CCA)-have been extensively used. When there are multiple datasets generated by independent studies with compatible designs, integrative analysis has been developed and shown to outperform meta-analysis, other multidatasets analysis, and individual-data analysis. To facilitate integrative dimension reduction analysis in daily practice, we develop the R package iSFun, which can comprehensively conduct integrative sparse PCA, PLS and CCA, as well as meta-analysis and stacked analysis. The package can conduct analysis under the homogeneity and heterogeneity models and with the magnitude- and sign-based contrasted penalties. As a 'byproduct', this article is the first to develop integrative analysis built on the CCA technique, further expanding the scope of integrative analysis. AVAILABILITY AND IMPLEMENTATION: The package is available at https://CRAN.R-project.org/package=iSFun. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. Kuangnan Fang, Qingzhao Zhang 0002, Shuangge Ma |
Bioinform. | 1 |