VLDB 2026 Research / reviewers in the wild / expert
Yi Yu 0016
dblp:99/111-16
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0002-3395-7744ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
6 papers |
Mathematical optimization · 43% Information theory · 41% Graph algorithms and graph theory · 7% | |
| Artificial intelligence
1 paper |
Transfer learning and domain adaptation · 50% Reinforcement learning · 50% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information theory › hypothesis testing
change-point detection |
1.8 | 3 | 2023 | Change point detection and inference in multivariate non-parametric models under mixing conditions · NeurIPS 2023 Optimal Nonparametric Multivariate Change Point Detection and Localization · IEEE Trans. Inf. Theory 2022 Change-point Detection for Sparse and Dense Functional Data in General Dimensions · NeurIPS 2022 |
Mathematical optimization
statistical estimation |
1.7 | 3 | 2023 | Change point detection and inference in multivariate non-parametric models under mixing conditions · NeurIPS 2023 Change-point Detection for Sparse and Dense Functional Data in General Dimensions · NeurIPS 2022 Lattice partition recovery with dyadic CART · NeurIPS 2021 |
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift |
0.9 | 1 | 2025 | Transfer Learning for Nonparametric Contextual Dynamic Pricing · ICML 2025 |
Machine learning › Reinforcement learning
regret minimization |
0.9 | 1 | 2025 | Transfer Learning for Nonparametric Contextual Dynamic Pricing · ICML 2025 |
Data mining
anomaly detection |
0.6 | 1 | 2022 | Change point localization in dependent dynamic nonparametric random dot product graphs · J. Mach. Learn. Res. 2022 |
Data mining › time series analysis
change point detection |
0.6 | 1 | 2022 | Change point localization in dependent dynamic nonparametric random dot product graphs · J. Mach. Learn. Res. 2022 |
Mathematical optimization › regularization
group lasso |
0.6 | 1 | 2022 | Functional Linear Regression with Mixed Predictors · J. Mach. Learn. Res. 2022 |
Mathematical optimization › statistical learning theory
high-dimensional regression |
0.6 | 1 | 2022 | Functional Linear Regression with Mixed Predictors · J. Mach. Learn. Res. 2022 |
Computational geometry › geometric inference › geometric estimation
localization |
0.6 | 1 | 2022 | Optimal Nonparametric Multivariate Change Point Detection and Localization · IEEE Trans. Inf. Theory 2022 |
Information theory › estimation theory › minimax estimation
minimax lower bounds |
0.6 | 1 | 2022 | Optimal Nonparametric Multivariate Change Point Detection and Localization · IEEE Trans. Inf. Theory 2022 |
Information theory › statistical inference
nonparametric statistics |
0.6 | 1 | 2022 | Optimal Nonparametric Multivariate Change Point Detection and Localization · IEEE Trans. Inf. Theory 2022 |
Graph algorithms and graph theory › random graph models
random dot product graph |
0.6 | 1 | 2022 | Change point localization in dependent dynamic nonparametric random dot product graphs · J. Mach. Learn. Res. 2022 |
Mathematical optimization › statistical estimation › high-dimensional estimation
sparse estimation |
0.6 | 1 | 2022 | Functional Linear Regression with Mixed Predictors · J. Mach. Learn. Res. 2022 |
Information theory › signal processing
denoising |
0.5 | 1 | 2021 | Lattice partition recovery with dyadic CART · NeurIPS 2021 |
Computational finance and economics › pricing
dynamic pricing |
0.3 | 1 | 2025 | Transfer Learning for Nonparametric Contextual Dynamic Pricing · ICML 2025 |
Computational finance and economics › mechanism design
revenue maximization |
0.3 | 1 | 2025 | Transfer Learning for Nonparametric Contextual Dynamic Pricing · ICML 2025 |
Mathematical optimization
functional data analysis |
0.2 | 1 | 2022 | Change-point Detection for Sparse and Dense Functional Data in General Dimensions · NeurIPS 2022 |
Algorithms and data structures › decision tree › decision tree learning
classification and regression tree |
0.1 | 1 | 2021 | Lattice partition recovery with dyadic CART · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
minimax analysis · 3.4transfer learning · 1.7lipschitz condition · 1.7nonparametric estimation · 1.1latent position estimation · 1.1CUSUM statistic · 1.1mixing conditions · 0.7long-run variance estimation · 0.7penalized least squares · 0.6kernel methods · 0.6coordinate descent · 0.6binary segmentation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Transfer Learning for Nonparametric Contextual Dynamic PricingabstractDynamic pricing strategies are crucial for firms to maximize revenue by adjusting prices based on market conditions and customer characteristics. However, designing optimal pricing strategies becomes challenging when historical data are limited, as is often the case when launching new products or entering new markets. One promising approach to overcome this limitation is to leverage information from related products or markets to inform the focal pricing decisions. In this paper, we explore transfer learning for nonparametric contextual dynamic pricing under a covariate shift model, where the marginal distributions of covariates differ between source and target domains while the reward functions remain the same. We propose a novel Transfer Learning for Dynamic Pricing (TLDP) algorithm that can effectively leverage pre-collected data from a source domain to enhance pricing decisions in the target domain. The regret upper bound of TLDP is established under a simple Lipschitz condition on the reward function. To establish the optimality of TLDP, we further derive a matching minimax lower bound, which includes the target-only scenario as a special case and is presented for the first time in the literature. Extensive numerical experiments validate our approach, demonstrating its superiority over existing methods and highlighting its practical utility in real-world applications. Feiyu Jiang, Zifeng Zhao, Yi Yu 0016 |
ICML | 4 |
| 2024 | Quickest Detection in High-Dimensional Linear Regression Models via Implicit RegularizationabstractIn this paper, we consider the quickest detection problem in high-dimensional streaming data, where the unknown regression coefficients might change at some unknown time. We propose a quickest detection algorithm based on the implicit regularization algorithm via gradient descent, and provide theoretical guarantees on the average run length to false alarm and detection delay. Numerical studies are conducted to validate the theoretical results. Qunzhi Xu, Yi Yu 0016, Yajun Mei |
ISIT | 2 |
| 2023 | Change point detection and inference in multivariate non-parametric models under mixing conditionsabstractThis paper addresses the problem of localizing and inferring multiple change points, in non-parametric multivariate time series settings. Specifically, we consider a multivariate time series with potentially short-range dependence, whose underlying distributions have Hölder smooth densities and can change over time in a piecewise-constant manner. The change points, which correspond to the times when the distribution changes, are unknown.
We present the limiting distributions of the change point estimators under the scenarios where the minimal jump size vanishes or remains constant. Such results have not been revealed in the literature in non-parametric change point settings. As byproducts, we develop a sharp estimator that can accurately localize the change points in multivariate non-parametric time series, and a consistent block-type long-run variance estimator. Numerical studies are provided to complement our theoretical findings. Carlos Misael Madrid Padilla, Daren Wang, Oscar Hernan Madrid Padilla, Yi Yu 0016 |
NeurIPS | 5 |
| 2022 | Denoising and change point localisation in piecewise-constant high-dimensional regression coefficientsabstractWe study the theoretical properties of the fused lasso procedure originally proposed by Tibshirani et al. (2005) in the context of a linear regression model in which the regression coefficient are totally ordered and assumed to be sparse and piecewise constant. Despite its popularity, to the best of our knowledge, estimation error bounds in high-dimensional settings have only been obtained for the simple case in which the design matrix is the identity matrix. We formulate a novel restricted isometry condition on the design matrix that is tailored to the fused lasso estimator and derive estimation bounds for both the constrained version of the fused lasso assuming dense coefficients and for its penalised version. We observe that the estimation error can be dominated by either the lasso or the fused lasso rate, depending on whether the number of non-zero coefficient is larger than the number of piece-wise constant segments. Finally, we devise a post-processing procedure to recover the piecewise-constant pattern of the coefficients. Extensive numerical experiments support our theoretical findings. Oscar Hernan Madrid Padilla, Yi Yu 0016, Alessandro Rinaldo |
AISTATS | 3 |
| 2022 | Optimal partition recovery in general graphsabstractWe consider a graph-structured change point problem in which we observe a random vector with piece-wise constant but otherwise unknown mean and whose independent, sub-Gaussian coordinates correspond to the $n$ nodes of a fixed graph. We are interested in the localisation task of recovering the partition of the nodes associated to the constancy regions of the mean vector or, equivalently, of estimating the cut separating the sub-graphs over which the mean remains constant. Although graph-valued signals of this type have been previously studied in the literature for the different tasks of testing for the presence of an anomalous cluster and of estimating the mean vector, no localisation results are known outside the classical case of chain graphs. When the partition $\mathcal{S}$ consists of only two elements, we characterise the difficulty of the localisation problem in terms of four key parameters: the maximal noise variance $\sigma^2$, the size $\Delta$ of the smaller element of the partition, the magnitude $\kappa$ of the difference in the signal values across contiguous elements of the partition and the sum of the effective resistance edge weights $|\partial_r(\mathcal{S})|$ of the corresponding cut – a graph theoretic quantity quantifying the size of the partition boundary. In particular, we demonstrate an information theoretical lower bound implying that, in the low signal-to-noise ratio regime $\kappa^2 \Delta \sigma^{-2} |\partial_r(\mathcal{S})|^{-1} \lesssim 1$, no consistent estimator of the true partition exists. On the other hand, when $\kappa^2 \Delta \sigma^{-2} |\partial_r(\mathcal{S})|^{-1} \gtrsim \zeta_n \log\{r(|E|)\}$, with $r(|E|)$ being the sum of effective resistance weighted edges and $\zeta_n$ being any diverging sequence in $n$, we show that a polynomial-time, approximate $\ell_0$-penalised least squared estimator delivers a localisation error – measured by the symmetric difference between the true and estimated partition – of order $ \kappa^{-2} \sigma^2 |\partial_r(\mathcal{S})| \log\{r(|E|)\}$. Aside from the $\log\{r(|E|)\}$ term, this rate is minimax optimal. Finally, we provide discussions on the localisation error for more general partitions of unknown sizes. Yi Yu 0016, Oscar Hernan Madrid Padilla, Alessandro Rinaldo |
AISTATS | 1 |
| 2022 | Change-point Detection for Sparse and Dense Functional Data in General DimensionsabstractWe study the problem of change-point detection and localisation for functional data sequentially observed on a general $d$-dimensional space, where we allow the functional curves to be either sparsely or densely sampled. Data of this form naturally arise in a wide range of applications such as biology, neuroscience, climatology and finance. To achieve such a task, we propose a kernel-based algorithm named functional seeded binary segmentation (FSBS). FSBS is computationally efficient, can handle discretely observed functional data, and is theoretically sound for heavy-tailed and temporally-dependent observations. Moreover, FSBS works for a general $d$-dimensional domain, which is the first in the literature of change-point estimation for functional data. We show the consistency of FSBS for multiple change-point estimation and further provide a sharp localisation error rate, which reveals an interesting phase transition phenomenon depending on the number of functional curves observed and the sampling frequency for each curve. Extensive numerical experiments illustrate the effectiveness of FSBS and its advantage over existing methods in the literature under various settings. A real data application is further conducted, where FSBS localises change-points of sea surface temperature patterns in the south Pacific attributed to El Ni\~{n}o. Carlos Misael Madrid Padilla, Daren Wang, Zifeng Zhao, Yi Yu 0016 |
NeurIPS | 4 |
| 2022 | Change point localization in dependent dynamic nonparametric random dot product graphsabstractIn this paper, we study the offline change point localization problem in a sequence of dependent nonparametric random dot product graphs. To be specific, assume that at every time point, a network is generated from a nonparametric random dot product graph model (see e.g. Athreya et al., 2018), where the latent positions are generated from unknown underlying distributions. The underlying distributions are piecewise constant in time and change at unknown locations, called change points. Most importantly, we allow for dependence among networks generated between two consecutive change points. This setting incorporates edge-dependence within networks and temporal dependence between networks, which is the most flexible setting in the published literature. To accomplish the task of consistently localizing change points, we propose a novel change point detection algorithm, consisting of two steps. First, we estimate the latent positions of the random dot product model, our theoretical result being a refined version of the state-of-the-art results, allowing the dimension of the latent positions to diverge. Subsequently, we construct a nonparametric version of the CUSUM statistic (e.g. Page, 1954; Padilla et al., 2019a) that allows for temporal dependence. Consistent localization is proved theoretically and supported by extensive numerical experiments, which illustrate state-of-the-art performance. We also provide in depth discussion of possible extensions to give more understanding and insights. Oscar Hernan Madrid Padilla, Yi Yu 0016, Carey E. Priebe |
J. Mach. Learn. Res. | 2 |
| 2022 | Functional Linear Regression with Mixed PredictorsabstractWe study a functional linear regression model that deals with functional responses and allows for both functional covariates and high-dimensional vector covariates. The proposed model is flexible and nests several functional regression models in the literature as special cases. Based on the theory of reproducing kernel Hilbert spaces (RKHS), we propose a penalized least squares estimator that can accommodate functional variables observed on discrete sample points. Besides a conventional smoothness penalty, a group Lasso-type penalty is further imposed to induce sparsity in the high-dimensional vector predictors. We derive finite sample theoretical guarantees and show that the excess prediction risk of our estimator is minimax optimal. Furthermore, our analysis reveals an interesting phase transition phenomenon that the optimal excess risk is determined jointly by the smoothness and the sparsity of the functional regression coefficients. A novel efficient optimization algorithm based on iterative coordinate descent is devised to handle the smoothness and group penalties simultaneously. Simulation studies and real data applications illustrate the promising performance of the proposed approach compared to the state-of-the-art methods in the literature. Daren Wang, Zifeng Zhao, Yi Yu 0016, Rebecca Willett |
J. Mach. Learn. Res. | 3 |
| 2022 | Optimal Nonparametric Multivariate Change Point Detection and LocalizationabstractWe study the multivariate nonparametric change point detection problem, where the data are a sequence of independent$p$-dimensional random vectors whose distributions are piecewise-constant with Lipschitz densities changing at unknown times, called change points. We quantify the size of the distributional change at any change point with the supremum norm of the difference between the corresponding densities. We are concerned with the localization task of estimating the positions of the change points. In our analysis, we allow for the model parameters to vary with the total number of time points, including the minimal spacing between consecutive change points and the magnitude of the smallest distributional change. We provide information-theoretic lower bounds on both the localization rate and the minimal signal-to-noise ratio required to guarantee consistent localization. We formulate a novel algorithm based on kernel density estimation that nearly achieves the minimax lower bound, save possibly for logarithm factors. We have provided extensive numerical evidence to support our theoretical findings. Oscar Hernan Madrid Padilla, Yi Yu 0016, Daren Wang, Alessandro Rinaldo |
IEEE Trans. Inf. Theory | 2 |
| 2021 | Localizing Changes in High-Dimensional Regression ModelsabstractThis paper addresses the problem of localizing change points in high-dimensional linear regression models with piecewise constant regression coefficients. We develop a dynamic programming approach to estimate the locations of the change points whose performance improves upon the current state-of-the-art, even as the dimension, the sparsity of the regression coefficients, the temporal spacing between two consecutive change points, and the magnitude of the difference of two consecutive regression coefficient vectors are allowed to vary with the sample size. Furthermore, we devise a computationally-efficient refinement procedure that provably reduces the localization error of preliminary estimates of the change points. We demonstrate minimax lower bounds on the localization error that nearly match the upper bound on the localization error of our methodology and show that the signal-to-noise condition we impose is essentially the weakest possible based on information-theoretic arguments. Extensive numerical results support our theoretical findings, and experiments on real air quality data reveal change points supported by historical information not used by the algorithm. Alessandro Rinaldo, Daren Wang, Qin Wen, Rebecca Willett, Yi Yu 0016 |
AISTATS | 5 |
| 2021 | Lattice partition recovery with dyadic CARTabstractWe study piece-wise constant signals corrupted by additive Gaussian noise over a $d$-dimensional lattice. Data of this form naturally arise in a host of applications, and the tasks of signal detection or testing, de-noising and estimation have been studied extensively in the statistical and signal processing literature. In this paper we consider instead the problem of partition recovery, i.e.~of estimating the partition of the lattice induced by the constancy regions of the unknown signal, using the computationally-efficient dyadic classification and regression tree (DCART) methodology proposed by \citep{donoho1997cart}. We prove that, under appropriate regularity conditions on the shape of the partition elements, a DCART-based procedure consistently estimates the underlying partition at a rate of order $\sigma^2 k^* \log (N)/\kappa^2$, where $k^*$ is the minimal number of rectangular sub-graphs obtained using recursive dyadic partitions supporting the signal partition, $\sigma^2$ is the noise variance, $\kappa$ is the minimal magnitude of the signal difference among contiguous elements of the partition and $N$ is the size of the lattice. Furthermore, under stronger assumptions, our method attains a sharper estimation error of order $\sigma^2\log(N)/\kappa^2$, independent of $k^*$, which we show to be minimax rate optimal. Our theoretical guarantees further extend to the partition estimator based on the optimal regression tree estimator (ORT) of \cite{chatterjee2019adaptive} and to the one obtained through an NP-hard exhaustive search method. We corroborate our theoretical findings and the effectiveness of DCART for partition recovery in simulations. Oscar Hernan Madrid Padilla, Yi Yu 0016, Alessandro Rinaldo |
NeurIPS | 2 |