Yachong Yang

dblp:268/8038 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0001-9780-4918ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning theory · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
high-dimensional statistics
0.612022
The Price of Competition: Effect Size Heterogeneity Matters in High Dimensions · IEEE Trans. Inf. Theory 2022
Machine learning › Learning theory
model selection
0.612022
The Price of Competition: Effect Size Heterogeneity Matters in High Dimensions · IEEE Trans. Inf. Theory 2022
Machine learning and data management › statistical learning
high-dimensional regression
0.612022
The Price of Competition: Effect Size Heterogeneity Matters in High Dimensions · IEEE Trans. Inf. Theory 2022
Machine learning and data management › statistical learning
lasso
0.612022
The Price of Competition: Effect Size Heterogeneity Matters in High Dimensions · IEEE Trans. Inf. Theory 2022
Machine learning › Learning theory
high-dimensional regression
0.412020
The Complete Lasso Tradeoff Diagram · NeurIPS 2020
Machine learning › Learning theory › model selection
variable selection
0.412020
The Complete Lasso Tradeoff Diagram · NeurIPS 2020
Mathematical optimization › statistical estimation › regression › sparse regression
lasso
0.412020
The Complete Lasso Tradeoff Diagram · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

approximate message passing · 1.1simulation · 0.9asymptotic analysis · 0.9
YearPublicationVenuePosition
2022 The Price of Competition: Effect Size Heterogeneity Matters in High Dimensions
abstract
In high-dimensional sparse regression, would increasing the signal-to-noise ratio while fixing the sparsity level always lead to better model selection? For high-dimensional sparse regression problems, surprisingly, in this paper we answer this question in thenegativein the regime of linear sparsity for the Lasso method, relying on a new concept we termeffect size heterogeneity. Roughly speaking, a regression coefficient vector has high effect size heterogeneity if its nonzero entries have significantly different magnitudes. From the viewpoint of this new measure, we prove that the false and true positive rates achieve the optimal trade-offuniformlyalong the Lasso path when this measure is maximal in a certain sense, and the worst trade-off is achieved when it is minimal in the sense that all nonzero effect sizes are roughly equal. Moreover, we demonstrate that the first false selection occurs much earlier when effect size heterogeneity is minimal than when it is maximal. The underlying cause of these two phenomena is, metaphorically speaking, the “competition” among variables with effect sizes of the same magnitude in entering the model. Taken together, our findings suggest that effect size heterogeneity shall serve as an important complementary measure to the sparsity of regression coefficients in the analysis of high-dimensional regression problems. Our proofs use techniques from approximate message passing theory as well as a novel technique for estimating the rank of the first false variable.
Yachong Yang, Weijie J. Su
IEEE Trans. Inf. Theory2
2020 The Complete Lasso Tradeoff Diagram
abstract
A fundamental problem in high-dimensional regression is to understand the tradeoff between type I and type II errors or, equivalently, false discovery rate (FDR) and power in variable selection. To address this important problem, we offer the first complete diagram that distinguishes all pairs of FDR and power that can be asymptotically realized by the Lasso from the remaining pairs, in a regime of linear sparsity under random designs. The tradeoff between the FDR and power characterized by our diagram holds no matter how strong the signals are. In particular, our results complete the earlier Lasso tradeoff diagram in previous literature by recognizing two simple constraints on the pairs of FDR and power. The improvement is more substantial when the regression problem is above the Donoho-Tanner phase transition. Finally, we present extensive simulation studies to confirm the sharpness of the complete Lasso tradeoff diagram.
Yachong Yang, Zhiqi Bu, Weijie J. Su
NeurIPS2