VLDB 2026 Research / reviewers in the wild / expert
Fan Yang 0106
dblp:29/3081-106
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0001-6972-0784ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Asymptotic Theory of Eigenvectors for Latent Embeddings With Generalized Laplacian MatricesabstractLaplacian matrices are widely used in practice to capture latent structural information in data, ranging from graphs to manifolds. Their normalization forms naturally induce dependencies among matrix entries, and such dependencies are known to pose significant challenges for advances in random matrix theory (RMT). Motivated by this, we introduce a general class of generalized Laplacian matrices, which includes both the standard Laplacian and random adjacency matrices as special cases, and we develop a new framework—Asymptotic Theory of Eigenvectors for latent embeddings with Generalized Laplacian matrices (ATE-GL)—for studying their spectral properties. Our theory is driven by two key ingredients: the use of generalized quadratic vector equations to handle dependency in RMT, and refined high-order asymptotic expansions for empirical spiked eigenvectors and eigenvalues based on local laws. These results lead to asymptotic normality for both spiked eigenvectors and eigenvalues, enabling precise statistical inference and uncertainty quantification for a broad class of applications involving generalized Laplacian matrices. We also discuss two motivating applications of the ATE-GL framework and demonstrate its effectiveness through numerical examples. Jianqing Fan, Jinchi Lv, Fan Yang 0106, Diwen Yu |
IEEE Trans. Inf. Theory | 4 |
| 2025 | Precise High-Dimensional Asymptotics for Quantifying Heterogeneous TransfersabstractThe problem of learning one task using samples from another task is central to transfer learning. In this paper, we focus on answering the following question: when does combining the samples from two related tasks perform better than learning with one target task alone? This question is motivated by an empirical phenomenon known as negative transfer often observed in transfer learning practice. While the transfer effect from one task to another depends on factors such as their sample sizes and the spectrum of their covariance matrices, precisely quantifying this dependence has remained a challenging problem. In order to compare a transfer learning estimator to single-task learning, one needs to compare the risks between the two estimators precisely. Further, the comparison depends on the distribution shifts between the two tasks. This paper applies recent developments of random matrix theory to tackle this challenge in a high-dimensional linear regression setting with two tasks. We provide precise high-dimensional asymptotics for the bias and variance of a classical hard parameter sharing (HPS) estimator in the proportional limit, when the sample sizes of both tasks increase proportionally with dimension at fixed ratios. The precise asymptotics apply to various types of distribution shifts, including covariate shifts, model shifts, and combinations of both. We illustrate these results in a random-effects model to mathematically prove a phase transition from positive to negative transfer as the number of source task samples increases. One insight from the analysis is that a rebalanced HPS estimator, which downsizes the source task when the model shift is high, achieves the minimax optimal rate. The finding regarding phase transition also applies to multiple tasks when feature covariates are shared across all tasks. Simulations validate the accuracy of the high-dimensional asymptotics for finite dimensions. Fan Yang 0106, Hongyang R. Zhang, Sen Wu 0002, Christopher Ré, Weijie J. Su |
J. Mach. Learn. Res. | 1 |
| 2022 | Tracy-Widom Distribution for Heterogeneous Gram Matrices With Applications in Signal DetectionabstractDetection of the number of signals corrupted by high-dimensional noise is a fundamental problem in signal processing and statistics. This paper focuses on a general setting where the high-dimensional noise has an unknown complicated heterogeneous variance structure. We propose a sequential test which utilizes the edge singular values (i.e., the largest few singular values) of the data matrix. It also naturally leads to a consistent sequential testing estimate of the number of signals. We describe the asymptotic distribution of the test statistic in terms of the Tracy-Widom distribution. The test is shown to be accurate and have full power against the alternative, both theoretically and numerically. The theoretical analysis relies on establishing the Tracy-Widom law for a large class of Gram type random matrices with non-zero means and completely arbitrary variance profiles, which can be of independent interest. Xiucai Ding, Fan Yang 0106 |
IEEE Trans. Inf. Theory | 2 |