EDBT 2026 Demo / reviewers in the wild / expert
Canhong Wen
dblp:200/9874
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-0220-9986ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Theory of computation · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Association Recovery in High Dimensions by Parallel Learning
Ruipeng Dong, Canhong Wen |
INFORMS J. Comput. | 2 |
| 2025 | An Efficient Pruner for Large Language Model with Theoretical GuaranteeabstractLarge Language Models (LLMs) have showcased remarkable performance across a range of tasks but are hindered by their massive parameter sizes, which impose significant computational and storage demands. Pruning has emerged as an effective solution to reduce model size, but traditional methods often involve inefficient retraining or rely on heuristic-based one-shot approaches that lack theoretical guarantees. In this paper, we reformulate the pruning problem as an $\ell_0$-penalized optimization problem and propose a monotone accelerated Iterative Hard Thresholding (mAIHT) method. Our approach combines solid theoretical foundations with practical effectiveness, offering a detailed theoretical analysis that covers convergence, convergence rates, and risk upper bounds. Through extensive experiments, we demonstrate that mAIHT outperforms state-of-the-art pruning techniques by effectively pruning the LLaMA-7B model across various evaluation metrics. Canhong Wen, Yihong Zuo, Wenliang Pan |
ICML | 1 |
| 2024 | SWGP: Semi-supervised clustering via Wasserstein generative adversarial network with gradient penalty for uncovering brain disease heterogeneity from medical imagesabstractDisease heterogeneity poses a significant challenge for the accurate diagnosis and treatment in brain diseases such as Alzheimer’s disease. Previous attempts to address this problem using machine learning methods have proven to be either insufficiently In order to overcom these limitations, we propose a novel method called Semi-supervised clustering via WGANs with Gradiant Penalty (SWGP) to identify distinct subgroups among patient cohorts through the application of semi-supervised clustering techniques. Essentially, SWGP is a Generative Adversarial Network (GAN) model enhanced with a gradient penalty, consisting of three key modules: the generator, the critic and the clustering. The generator module learns a transformation function by generating pseudo patient data based on healthy individuals, with latent mapping variables to characterize the transformation directions. The critic module evaluates and assigns different scores to the generated pseudo data and the real patient data. Moreover, the clustering module is trained interactively with the generator by assigning patients to their respective subtype memberships. To validate the efficacy of our proposed model, extensive experiments have been conducted on simulated and synthetic datasets. The results highlight that the proposed model outperforms other state-of-the-art solutions in terms of both accuracy and computational efficiency. We further apply our proposal to a publicly available dataset related to Alzheimer’s disease, and demonstrate its potential in capturing biomedical imaging patterns associated with Alzheimer’s disease. Canhong Wen, Haizhu Tan, Chiyu Wei |
IJCNN | 2 |
| 2024 | Subset Selection in Support Vector MachineabstractSupport Vector Machines (SVMs) face a significant challenge when dealing with high-dimensional datasets, as their performance can be severely compromised by the inclusion of numerous redundant variables. To address this challenge, we propose a novel approach that integrates ℓ0and ℓ2penalties to enhance the performance of SVM in high-dimensional settings. The ℓ0penalty is introduced for best subset selection, addressing the issue of redundancy. In addition, we incorporate an ℓ2penalty in our proposal, which effectively reduces the impact of noise, thereby improving the stability and generalization capability of the model. To solve the underlying optimization problem, we develop an efficient algorithm that utilizes primaldual hybrid gradient method and the concept of splicing to reach a stable solution. Through simulation studies, we demonstrate the superior performance of our proposal compared to the state-of-the-art methods in terms of prediction accuracy and variables selection. To further validate the effectiveness of our method, we conduct experiments using three classification datasets from the NIPS 2003 Feature Selection Challenge. The results show the benefits of our proposed method in selecting variables and achieving more accurate predictions. Jiangshuo Zhao, Canhong Wen |
IJCNN | 2 |
| 2023 | Simultaneous Dimension Reduction and Variable Selection for Multinomial Logistic RegressionabstractMultinomial logistic regression is a useful model for predicting the probabilities of multiclass outcomes. Because of the complexity and high dimensionality of some data, it is challenging to fit a valid model with high accuracy and interpretability. We propose a novel sparse reduced-rank multinomial logistic regression model to jointly select variables and reduce the dimension via a nonconvex row constraint. We develop a block-wise iterative algorithm with a majorizing surrogate function to efficiently solve the optimization problem. From an algorithmic aspect, we show that the output estimator enjoys consistency in estimation and sparsity recovery even in a high-dimensional setting. The finite sample performance of the proposed method is investigated via simulation studies and two real image data sets. The results show that our proposal has competitive performance in both estimation accuracy and computation time. History: Accepted by Andrea Lodi, Area Editor for Design & Analysis of Algorithms–Discrete. Funding: This work was supported by the National Natural Science Foundation of China [Grants 71991474, 12171449, 11801540, and 12071494] and the Natural Science Foundation of Anhui Province [Grant BJ2040170017]. Supplemental Material: The online appendix is available at https://doi.org/10.1287/ijoc.2022.0132 . Canhong Wen, Zhenduo Li, Ruipeng Dong, Yijin Ni, Wenliang Pan |
INFORMS J. Comput. | 1 |
| 2023 | ℓ0 Trend FilteringabstractThe [Formula: see text] trend filtering ([Formula: see text]-TF) is a new effective tool for nonparametric regression with the power of automatic knot detection in function values or derivatives. It overcomes the drawback of [Formula: see text]-TF that is known to have bias issues. To solve the [Formula: see text]-TF problem, we propose an alternating minimization induced active set (AMIAS) search method based on the necessary optimality conditions derived from an augmented Lagrangian framework. The proposed method takes full advantage of the primal and dual variables with complementary supports, and decouples the high-dimensional problem into two subsystems on the active and inactive sets, respectively. A sequential AMIAS algorithm with warm start initialization is developed for efficient determination of the cardinality parameter, along with the output of solution paths. Theoretically, the oracle estimator of [Formula: see text]-TF is justified to behave like regression splines under the continuous time setting with mild conditions. Our numerical experiments include simulation studies for comparing [Formula: see text]-TF to [Formula: see text]-TF and free-knot splines on several synthetic examples, and a real data application of time series segmentation on Hong Kong PM2.5 indexes. History: Accepted by Antonio Frangioni, Area Editor for Design & Analysis of Algorithms – Continuous. Funding: This work was supported in part by Hong Kong General Research Fund [No. 17306519]. C. Wen’s research is partially supported by National Science Foundation of China [12171449] and Fundamental Research Funds for the Central Universities [WK3470000027, YD2040002019]. X. Wang’s research is partially supported by National Natural Science Foundation of China [Grants 72171216, 12231017, 71921001, and 71991474], and the National Key R&D Program of China [No. 2022YFA1003803]. Supplemental Material: The e-companion is available at https://doi.org/10.1287/ijoc.2021.0313 . The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2021.0313 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2021.0313 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Canhong Wen, Aijun Zhang |
INFORMS J. Comput. | 1 |
| 2021 | Co-sparse reduced-rank regression for association analysis between imaging phenotypes and genetic variantsabstractMOTIVATION: The association analysis between genetic variants and imaging phenotypes must be carried out to understand the inherited neuropsychiatric disorders via imaging genetic studies. Given the high dimensionality in imaging and genetic data, traditional methods based on massive univariate regression entail large computational cost and disregard many-to-many correlations between phenotypes and genetic variants. Several multivariate imaging genetic methods have been proposed to alleviate the above problems. However, most of these methods are based on the l1 penalty, which might cause the over-selection of variables and thus mislead scientists in analyzing data from the field of neuroimaging genetics. RESULTS: To address these challenges in both statistics and computation, we propose a novel co-sparse reduced-rank regression model that identifies complex correlations in a dimensional reduction manner. We developed an iterative algorithm based on a group primal dual-active set formulation to detect simultaneously important genetic variants and imaging phenotypes efficiently and precisely via non-convex penalty. The simulation studies showed that our method achieved accurate and stable performance in parameter estimation and variable selection. In real application, the proposed approach successfully detected several novel Alzheimer's disease-related genetic variants and regions of interest, which indicate that our method may be a valuable statistical toolbox for imaging genetic studies. AVAILABILITY AND IMPLEMENTATION: The R package csrrr, and the code for experiments in this article is available in Github: https://github.com/hailongba/csrrr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Canhong Wen, Hailong Ba, Wenliang Pan, Meiyan Huang |
Bioinform. | 1 |
| 2020 | Image denoising via K-SVD with primal-dual active set algorithmabstractK-SVD algorithm has been successfully applied to image denoising tasks dozens of years but the big bottleneck in speed and accuracy still needs attention to break. For the sparse coding stage in K-SVD, which involves ℓ0constraint, prevailing methods usually seek approximate solutions greedily but are less effective once the noise level is high. The alternative ℓ1optimization is proved to be powerful than ℓ0, however, the time consumption prevents it from the implementation. In this paper, we propose a new K-SVD framework called K-SVDPby applying the Primal-dual active set (PDAS) algorithm to it. Different from the greedy algorithms based K-SVD, the K-SVDPalgorithm develops a selection strategy motivated by KKT (Karush-Kuhn-Tucker) condition and yields to an efficient update in the sparse coding stage. Since the K-SVDPalgorithm seeks for an equivalent solution to the dual problem iteratively with simple explicit expression in this denoising problem, speed and quality of denoising can be reached simultaneously. Experiments are carried out and demonstrate the comparable denoising performance of our K-SVDPwith state-of-the-art methods. Quan Xiao, Canhong Wen, Zirui Yan |
WACV | 2 |
| 2020 | Genome-wide association studies of brain imaging data via weighted distance correlationabstractMOTIVATION: Imaging genetics is mainly used to reveal the pathogenesis of neuropsychiatric risk genes and understand the relationship between human brain structure, functional and individual differences. Increasingly, the brain-wide imaging phenotypes in voxels are available to test the association with genetic markers. A challenge with analyzing such data is their high dimensionality and complex relationships. RESULTS: To tackle this challenge, we introduce a weighed distance correlation (wdCor) that can assess the association between genetic markers and voxel-based imaging data. Importantly, the wdCor test takes the voxel-based data as a whole multivariate phenotype, which preserves the spatial continuity and might enhance the power. Besides, an adaptive permutation procedure is introduced to determine the P-values of the wdCor test and also alleviate the computational burden in GWAS. In extensive simulation studies, wdCor achieves much better performances compared to the original distance correlation. We also successfully apply wdCor to conduct a large-scale analysis on data from the Alzheimer's disease neuroimaging project (ADNI). AVAILABILITY AND IMPLEMENTATION: Our wdCor method provides new research directions and ideas for multivariate analysis of high-dimensional data, it can also be used as a tool for scientific analysis of imaging genetics research in practical applications. The R package wdcor, and the code for reproducing all results in this article is available in Github: https://github.com/yangyuhui0129/wdcor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Canhong Wen, Yuhui Yang, Quan Xiao, Meiyan Huang, Wenliang Pan |
Bioinform. | 1 |