VLDB 2026 Research / reviewers in the wild / expert
Georgios Batzolis
dblp:287/8984
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Representation and self-supervised learning · 44% Generative modeling · 30% Deep learning architectures and training · 11% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning |
1.6 | 2 | 2025 | Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows · ICML 2025 Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024 |
Machine learning › Generative modeling › diffusion model
score-based generative model |
1.6 | 2 | 2025 | Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows · ICML 2025 Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024 |
Machine learning › Deep learning architectures and training
autoencoder |
0.9 | 1 | 2025 | Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows · ICML 2025 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
riemannian manifold |
0.9 | 1 | 2025 | Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › intrinsic dimension
intrinsic dimension estimation |
0.8 | 1 | 2024 | Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.6 | 1 | 2022 | How to Distribute Data across Tasks for Meta-Learning? · AAAI 2022 |
Machine learning › Learning theory
sample complexity |
0.6 | 1 | 2022 | How to Distribute Data across Tasks for Meta-Learning? · AAAI 2022 |
Machine learning › Representation and self-supervised learning
data manifold |
0.2 | 1 | 2024 | Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
isometry regularization · 0.9anisotropic normalizing flow · 0.9score matching · 0.8mixed linear regression · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic FlowsabstractData-driven Riemannian geometry has emerged as a powerful tool for interpretable representation learning, offering improved efficiency in downstream tasks. Moving forward, it is crucial to balance cheap manifold mappings with efficient training algorithms. In this work, we integrate concepts from pullback Riemannian geometry and generative models to propose a framework for data-driven Riemannian geometry that is scalable in both geometry and learning: score-based pullback Riemannian geometry. Focusing on unimodal distributions as a first step, we propose a score-based Riemannian structure with closed-form geodesics that pass through the data probability density. With this structure, we construct a Riemannian autoencoder (RAE) with error bounds for discovering the correct data manifold dimension. This framework can naturally be used with anisotropic normalizing flows by adopting isometry regularization during training. Through numerical experiments on diverse datasets, including image data, we demonstrate that the proposed framework produces high-quality geodesics passing through the data support, reliably estimates the intrinsic dimension of the data manifold, and provides a global chart of the manifold. To the best of our knowledge, this is the first scalable framework for extracting the complete geometry of the data manifold. Willem Diepeveen, Georgios Batzolis, Zakhar Shumaylov, Carola-Bibiane Schönlieb |
ICML | 2 |
| 2024 | Diffusion Models Encode the Intrinsic Dimension of Data ManifoldsabstractIn this work, we provide a mathematical proof that diffusion models encode data manifolds by approximating their normal bundles. Based on this observation we propose a novel method for extracting the intrinsic dimension of the data manifold from a trained diffusion model. Our insights are based on the fact that a diffusion model approximates the score function i.e. the gradient of the log density of a noise-corrupted version of the target distribution for varying levels of corruption. We prove that as the level of corruption decreases, the score function points towards the manifold, as this direction becomes the direction of maximal likelihood increase. Therefore, at low noise levels, the diffusion model provides us with an approximation of the manifold’s normal bundle, allowing for an estimation of the manifold’s intrinsic dimension. To the best of our knowledge our method is the first estimator of intrinsic dimension based on diffusion models and it outperforms well established estimators in controlled experiments on both Euclidean and image data. Jan Stanczuk, Georgios Batzolis, Teo Deveney, Carola-Bibiane Schönlieb |
ICML | 2 |
| 2022 | How to Distribute Data across Tasks for Meta-Learning?abstractMeta-learning models transfer the knowledge acquired from previous tasks to quickly learn new ones. They are trained on benchmarks with a fixed number of data points per task. This number is usually arbitrary and it is unknown how it affects performance at testing. Since labelling of data is expensive, finding the optimal allocation of labels across training tasks may reduce costs. Given a fixed budget of labels, should we use a small number of highly labelled tasks, or many tasks with few labels each? Should we allocate more labels to some tasks and less to others? We show that: 1) If tasks are homogeneous, there is a uniform optimal allocation, whereby all tasks get the same amount of data; 2) At fixed budget, there is a trade-off between number of tasks and number of data points per task, with a unique solution for the optimum; 3) When trained separately, harder task should get more data, at the cost of a smaller number of tasks; 4) When training on a mixture of easy and hard tasks, more data should be allocated to easy tasks. Interestingly, Neuroscience experiments have shown that human visual skills also transfer better from easy tasks. We prove these results mathematically on mixed linear regression, and we show empirically that the same results hold for few-shot image classification on CIFAR-FS and mini-ImageNet. Our results provide guidance for allocating labels across tasks when collecting data for meta-learning. Alexandru Cioba, Michael Bromberg, Ritwik Niyogi, Georgios Batzolis, Jezabel R. Garcia, Da-Shan Shiu, Alberto Bernacchia |
AAAI | 5 |