Haoshu Xu

dblp:388/2403 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning theory · 75% Probabilistic and Bayesian machine learning · 25%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 67% Recommender systems · 33%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › nonparametric regression
fréchet regression
0.912025
Wasserstein F-tests for Frechet regression on Bures-Wasserstein manifolds · J. Mach. Learn. Res. 2025
Machine learning › Learning theory
generalization bounds
0.912025
A Practical Theory of Generalization in Selectivity Learning · Proc. VLDB Endow. 2025
Machine learning › Learning theory
hypothesis testing
0.912025
Wasserstein F-tests for Frechet regression on Bures-Wasserstein manifolds · J. Mach. Learn. Res. 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
non-parametric methods
0.912025
Wasserstein F-tests for Frechet regression on Bures-Wasserstein manifolds · J. Mach. Learn. Res. 2025
Query processing and optimization
cardinality estimation
0.912025
A Practical Theory of Generalization in Selectivity Learning · Proc. VLDB Endow. 2025
Query processing and optimization › selectivity estimation
learned selectivity estimation
0.912025
A Practical Theory of Generalization in Selectivity Learning · Proc. VLDB Endow. 2025
Recommender systems › trustworthy recommendation › robust recommendation
out-of-distribution generalization
0.912025
A Practical Theory of Generalization in Selectivity Learning · Proc. VLDB Endow. 2025

Methods — techniques the papers use, named apart from their topics

signed measures · 1.7PAC learning · 1.7wasserstein distance · 0.9asymptotic theory · 0.9
YearPublicationVenuePosition
2025 Wasserstein F-tests for Frechet regression on Bures-Wasserstein manifolds
abstract
This paper addresses regression analysis for covariance matrix-valued outcomes with Euclidean covariates, motivated by applications in single-cell genomics and neuroscience where covariance matrices are observed across many samples. Our analysis leverages Fréchet regression on the Bures-Wasserstein manifold to estimate the conditional Fréchet mean given covariates $x$. We establish a non-asymptotic uniform $\sqrt{n}$-rate of convergence (up to logarithmic factors) over covariates with $\|x\| \lesssim \sqrt{\log n}$ and derive a pointwise central limit theorem to enable statistical inference. For testing covariate effects, we devise a novel test whose null distribution converges to a weighted sum of independent chi-square distributions, with power guarantees against a sequence of contiguous alternatives. Simulations validate the accuracy of the asymptotic theory. Finally, we apply our methods to a single-cell gene expression dataset, revealing age-related changes in gene co-expression networks.
Haoshu Xu, Hongzhe Li
J. Mach. Learn. Res.1
2025 A Practical Theory of Generalization in Selectivity Learning
abstract
Query-driven machine learning models have emerged as a promising estimation technique for query selectivities. Yet, surprisingly little is known about the efficacy of these techniques from a theoretical perspective, as there exist substantial gaps between practical solutions and state-of-the-art (SOTA) theory based on the Probably Approximately Correct (PAC) learning framework. In this paper, we aim to bridge the gaps between theory and practice. First, we demonstrate that selectivity predictors induced by signed measures are learnable, which relaxes the reliance on probability measures in SOTA theory. More importantly, beyond the PAC learning framework (which only allows us to characterize how the model behaves when both training and test workloads are drawn from the same distribution), we establish, under mild assumptions, that selectivity predictors from this class exhibit favorable out-of-distribution (OOD) generalization error bounds. These theoretical advances provide us with a better understanding of both the in-distribution and OOD generalization capabilities of query-driven selectivity learning, and facilitate the design of two general strategies to improve OOD generalization for existing query-driven selectivity models. We empirically verify that our techniques help query-driven selectivity models generalize significantly better to OOD queries both in terms of prediction accuracy and query latency performance, while maintaining their superior in-distribution generalization performance.
Peizhi Wu, Haoshu Xu, Ryan Marcus, Zachary G. Ives
Proc. VLDB Endow.2