EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyi Mai
dblp:201/7061
· DBLP profile ↗
7ranked-venue papers
7as first author
3since 2021 · last 2025
0009-0004-3121-2708ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Learning theory · 69% Learning paradigms · 21% Graph learning · 9% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
empirical risk minimization |
0.9 | 1 | 2025 | The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor Mixtures · ICLR 2025 |
Machine learning › Learning theory › high-dimensional statistics › high-dimensional asymptotics
gaussian universality |
0.9 | 1 | 2025 | The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor Mixtures · ICLR 2025 |
Machine learning › Learning theory
high-dimensional statistics |
0.9 | 1 | 2025 | The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor Mixtures · ICLR 2025 |
Machine learning › Learning paradigms
semi-supervised learning |
0.8 | 2 | 2021 | Consistent Semi-Supervised Graph Regularization for High Dimensional Data · J. Mach. Learn. Res. 2021 A Random Matrix Analysis and Improvement of Semi-Supervised Learning for Large Dimensional Data · J. Mach. Learn. Res. 2018 |
Machine learning › Graph learning
graph regularization |
0.5 | 1 | 2021 | Consistent Semi-Supervised Graph Regularization for High Dimensional Data · J. Mach. Learn. Res. 2021 |
Machine learning › Learning theory › statistical learning theory
asymptotic analysis |
0.3 | 1 | 2018 | A Random Matrix Analysis and Improvement of Semi-Supervised Learning for Large Dimensional Data · J. Mach. Learn. Res. 2018 |
Machine learning › Learning paradigms › semi-supervised learning
graph-based semi-supervised learning |
0.3 | 1 | 2018 | A Random Matrix Analysis and Improvement of Semi-Supervised Learning for Large Dimensional Data · J. Mach. Learn. Res. 2018 |
Machine learning › Learning theory
random matrix theory |
0.3 | 1 | 2018 | A Random Matrix Analysis and Improvement of Semi-Supervised Learning for Large Dimensional Data · J. Mach. Learn. Res. 2018 |
Algorithms and data structures
classification |
0.3 | 1 | 2025 | The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor Mixtures · ICLR 2025 |
Data mining
clustering |
0.1 | 1 | 2021 | Consistent Semi-Supervised Graph Regularization for High Dimensional Data · J. Mach. Learn. Res. 2021 |
Data mining › clustering
spectral clustering |
0.1 | 1 | 2021 | Consistent Semi-Supervised Graph Regularization for High Dimensional Data · J. Mach. Learn. Res. 2021 |
Data mining › predictive modeling
classification |
0.1 | 1 | 2018 | A Random Matrix Analysis and Improvement of Semi-Supervised Learning for Large Dimensional Data · J. Mach. Learn. Res. 2018 |
Data mining › predictive modeling › classification › pattern classification
high-dimensional classification |
0.1 | 1 | 2018 | A Random Matrix Analysis and Improvement of Semi-Supervised Learning for Large Dimensional Data · J. Mach. Learn. Res. 2018 |
Methods — techniques the papers use, named apart from their topics
random matrix theory · 2.4asymptotic analysis · 1.7laplacian regularization · 1.0centering · 1.0graph regularization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor MixturesabstractThe assumption of Gaussian or Gaussian mixture data has been extensively exploited in a long series of precise performance analyses of machine learning (ML) methods, on large datasets having comparably numerous samples and features.
To relax this restrictive assumption, subsequent efforts have been devoted to establish "Gaussian equivalent principles" by studying scenarios of Gaussian universality where the asymptotic performance of ML methods on non-Gaussian data remains unchanged when replaced with Gaussian data having the *same mean and covariance*.
Beyond the realm of Gaussian universality, there are few exact results on how the data distribution affects the learning performance.
In this article, we provide a precise high-dimensional characterization of empirical risk minimization, for classification under a general mixture data setting of *linear factor models* that extends Gaussian mixtures.
The Gaussian universality is shown to break down under this setting, in the sense that the asymptotic learning performance depends on the data distribution *beyond* the class means and covariances.
To clarify the limitations of Gaussian universality in the classification of mixture data and to understand the impact of its breakdown, we specify conditions for Gaussian universality and discuss their implications for the choice of loss function. Xiaoyi Mai, Zhenyu Liao 0001 |
ICLR | 1 |
| 2022 | On The Effectiveness of Active Learning by Uncertainty Sampling in Classification of High-Dimensional Gaussian Mixture DataabstractActive learning aims to reduce the cost of labeling through selective sampling. Despite reported empirical success over passive learning, many popular active learning heuristics such as uncertainty sampling still lack satisfying theoretical guarantees. Towards closing the gap between practical use and theoretical understanding in active learning, we propose to characterize the exact behavior of uncertainty sampling for high-dimensional Gaussian mixture data, in a modern regime of big data where the numbers of samples and features are commensurately large. Through a sharp characterization of the learning results, our analysis sheds light on the important question of when uncertainty sampling works better than passive learning. Our results show that the effectiveness of uncertainty sampling is not always ensured. In fact it depends crucially on the choice of i) an adequate initial classifier used to start the active sampling process and ii) a proper loss function that allows an adaptive treatment of samples queried at various steps. Xiaoyi Mai, Amir Salman Avestimehr, Antonio Ortega, Mahdi Soltanolkotabi |
ICASSP | 1 |
| 2021 | Consistent Semi-Supervised Graph Regularization for High Dimensional DataabstractSemi-supervised Laplacian regularization, a standard graph-based approach for learning from both labelled and unlabelled data, was recently demonstrated to have an insignificant high dimensional learning efficiency with respect to unlabelled data, causing it to be outperformed by its unsupervised counterpart, spectral clustering, given sufficient unlabelled data. Following a detailed discussion on the origin of this inconsistency problem, a novel regularization approach involving centering operation is proposed as solution, supported by both theoretical analysis and empirical results. Xiaoyi Mai, Romain Couillet |
J. Mach. Learn. Res. | 1 |
| 2019 | Revisiting and Improving Semi-supervised Learning: A Large Dimensional ApproachabstractThe recent work [1] shows that in the big data regime (i.e., numerous high dimensional data), the popular semi-supervised graph regularization, known as semi-supervised Laplacian regularization, fails to effectively extract information from unlabelled data. In response to this problem, we propose in this article an improved approach based on a simple yet fundamental update of the classical method. The effectiveness of the former is supported by both asymptotic results and simulations on finite data samples. Xiaoyi Mai, Romain Couillet |
ICASSP | 1 |
| 2019 | A Large Scale Analysis of Logistic Regression: Asymptotic Performance and New InsightsabstractLogistic regression, one of the most popular machine learning binary classification methods, has been long believed to be unbiased. In this paper, we consider the "hard" classification problem of separating high dimensional Gaussian vectors, where the data dimension p and the sample size n are both large. Based on recent advances in random matrix theory (RMT) and high dimensional statistics, we evaluate the asymptotic distribution of the logistic regression classifier and consequently, provide the associated classification performance. This brings new insights into the internal mechanism of logistic regression classifier, including a possible bias in the separating hyperplane, as well as on practical issues such as hyper-parameter tuning, thereby opening the door to novel RMT-inspired improvements. Xiaoyi Mai, Zhenyu Liao 0001, Romain Couillet |
ICASSP | 1 |
| 2018 | A Random Matrix Analysis and Improvement of Semi-Supervised Learning for Large Dimensional DataabstractThis article provides an original understanding of the behavior of a class of graph-oriented semi-supervised learning algorithms in the limit of large and numerous data. It is demonstrated that the intuition at the root of these methods collapses in this limit and that, as a result, most of them become inconsistent. Corrective measures and a new data-driven parametrization scheme are proposed along with a theoretical analysis of the asymptotic performances of the resulting approach. A surprisingly close behavior between theoretical performances on Gaussian mixture models and on real data sets is also illustrated throughout the article, thereby suggesting the importance of the proposed analysis for dealing with practical data. As a result, significant performance gains are observed on practical data classification using the proposed parametrization. Xiaoyi Mai |
J. Mach. Learn. Res. | 1 |
| 2017 | The counterintuitive mechanism of graph-based semi-supervised learning in the big data regimeabstractIn this article, a new approach is proposed to study the performance of graph-based semi-supervised learning methods, under the assumptions that the dimension of data p and their number n grow large at the same rate and that the data arise from a Gaussian mixture model. Unlike small dimensional systems, the large dimensions allow for a Taylor expansion to linearize the weight (or kernel) matrix W, thereby providing in closed form the limiting performance of semi-supervised learning algorithms. This notably allows to predict the classification error rate as a function of the normalization parameters and of the choice of the kernel function. Despite the Gaussian assumption for the data, the theoretical findings match closely the performance achieved with real datasets, particularly here on the popular MNIST database. Xiaoyi Mai, Romain Couillet |
ICASSP | 1 |