Yuwen Gu

dblp:211/4174 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 QuanDA: Quantile-Based Discriminant Analysis for High-Dimensional Imbalanced Classification
abstract
Binary classification with imbalanced classes is a common and fundamental task, where standard machine learning methods often struggle to provide reliable predictive performance. Although numerous approaches have been proposed to address this issue, classification in low-sample-size and high-dimensional settings still remains particularly challenging. The abundance of noisy features in high-dimensional data limits the effectiveness of classical methods due to overfitting, and the minority class is even difficult to detect because of its severe underrepresentation with low sample size. To address this challenge, we introduce Quantile-based Discriminant Analysis (QuanDA), which builds upon a novel connection with quantile regression and naturally accounts for class imbalance through appropriately chosen quantile levels. We provide comprehensive theoretical analysis to validate QuanDA in ultra-high dimensional settings. Through extensive simulation studies and high-dimensional benchmark data analysis, we demonstrate that QuanDA overall outperforms existing classification methods for imbalanced data, including cost-sensitive large-margin classifiers, random forests, and SMOTE.
Yuwen Gu
NeurIPS2
2023 Density-Convoluted Support Vector Machines for High-Dimensional Classification
abstract
The support vector machine (SVM) is a popular classification method which enjoys good performance in many real applications. The SVM can be viewed as a penalized minimization problem in which the objective function is the expectation of hinge loss function with respect to the standard non-smooth empirical measure corresponding to the true underlying measure. We further extend this viewpoint and propose a smoothed SVM by substituting a kernel density estimator for the measure in the expectation calculation. The resulting method is called density convoluted support vector machine (DCSVM). We argue that the DCSVM is particularly more interesting than the standard SVM in the context of high-dimensional classification. We systematically study the rate of convergence of the elastic-net penalized DCSVM under general random design setting. We further develop novel efficient algorithm for computing elastic-net penalized DCSVM. Simulation studies and ten benchmark datasets are used to demonstrate the superior classification performance of elastic-net DCSVM over other competitors, and it is demonstrated in these numerical studies that the computation of DCSVM can be more than 100 times faster than that of the SVM.
Yuwen Gu
IEEE Trans. Inf. Theory3
2020 Sparse Composite Quantile Regression in Ultrahigh Dimensions With Tuning Parameter Calibration
abstract
When estimating coefficients in a linear model, the (sparse) composite quantile regression was first proposed in Zou and Yuan (2008) as an efficient alternative to the (sparse) least squares to handle arbitrary error distribution. The highly nonsmooth nature of the composite loss in the sparse composite quantile regression makes its theoretical analysis as well as numerical computation much more challenging than the least squares method. The theory in Zou and Yuan (2008) was proven under fixed-dimension asymptotics and the estimator was computed via linear programming that does not scale well with high dimensions. In this paper, we study the sparse composite quantile regression under ultrahigh dimensionality and make three contributions. First, we provide a non-asymptotic analysis of both the lasso and the folded concave penalized composite quantile regression, which reveals a practical way of achieving the oracle estimator. Second, we construct a novel information criterion for selecting the regularization parameter in the folded concave penalized composite quantile regression and prove its selection consistency. Third, we exploit the structure of the composite loss and design a specialized optimization algorithm for computing the penalized composite quantile regression via the alternating direction method of multipliers. We conduct extensive simulations to illustrate the theoretical results. Our analysis provides a unified treatment of the concentration inequalities involving the composite loss. Those inequalities could be of independent interest.
Yuwen Gu
IEEE Trans. Inf. Theory1
2017 Flu MODELO 1.0: A simulation model and graphic interface for training and decision support for influenza management
abstract
There are many models for simulating infectious disease outbreaks that assist health policy officials in making better decisions about mitigation strategies in the event of an epidemic. However, none of the existing models address the need to train students and policy makers in the concepts and use of such models. In this paper, we present Flu MODELO 1.0, an implementation of a model that simulates an outbreak of influenza that is mitigated with a variety of strategies. Flu MODELO 1.0 is coupled with a GUI for running, replicating, and analyzing simulations. A novel addition in the GUI is the module for educating users about the concepts used in flu outbreak simulation. Flu MODELO 1.0 includes a parallelized version for use in simulating large populations.
Greg Ostroy, Diana Prieto, Yuwen Gu, Elise de Doncker, Rajib Paul
BIBM3