Georgios Batzolis

dblp:287/8984 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 44% Generative modeling · 30% Deep learning architectures and training · 11%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
1.622025
Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows · ICML 2025
Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024
Machine learning › Generative modeling › diffusion model
score-based generative model
1.622025
Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows · ICML 2025
Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024
Machine learning › Deep learning architectures and training
autoencoder
0.912025
Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
riemannian manifold
0.912025
Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows · ICML 2025
Machine learning › Generative modeling
diffusion model
0.812024
Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › intrinsic dimension
intrinsic dimension estimation
0.812024
Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024
Machine learning › Transfer learning and domain adaptation
meta-learning
0.612022
How to Distribute Data across Tasks for Meta-Learning? · AAAI 2022
Machine learning › Learning theory
sample complexity
0.612022
How to Distribute Data across Tasks for Meta-Learning? · AAAI 2022
Machine learning › Representation and self-supervised learning
data manifold
0.212024
Diffusion Models Encode the Intrinsic Dimension of Data Manifolds · ICML 2024

Methods — techniques the papers use, named apart from their topics

isometry regularization · 0.9anisotropic normalizing flow · 0.9score matching · 0.8mixed linear regression · 0.6
YearPublicationVenuePosition
2025 Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows
abstract
Data-driven Riemannian geometry has emerged as a powerful tool for interpretable representation learning, offering improved efficiency in downstream tasks. Moving forward, it is crucial to balance cheap manifold mappings with efficient training algorithms. In this work, we integrate concepts from pullback Riemannian geometry and generative models to propose a framework for data-driven Riemannian geometry that is scalable in both geometry and learning: score-based pullback Riemannian geometry. Focusing on unimodal distributions as a first step, we propose a score-based Riemannian structure with closed-form geodesics that pass through the data probability density. With this structure, we construct a Riemannian autoencoder (RAE) with error bounds for discovering the correct data manifold dimension. This framework can naturally be used with anisotropic normalizing flows by adopting isometry regularization during training. Through numerical experiments on diverse datasets, including image data, we demonstrate that the proposed framework produces high-quality geodesics passing through the data support, reliably estimates the intrinsic dimension of the data manifold, and provides a global chart of the manifold. To the best of our knowledge, this is the first scalable framework for extracting the complete geometry of the data manifold.
Willem Diepeveen, Georgios Batzolis, Zakhar Shumaylov, Carola-Bibiane Schönlieb
ICML2
2024 Diffusion Models Encode the Intrinsic Dimension of Data Manifolds
abstract
In this work, we provide a mathematical proof that diffusion models encode data manifolds by approximating their normal bundles. Based on this observation we propose a novel method for extracting the intrinsic dimension of the data manifold from a trained diffusion model. Our insights are based on the fact that a diffusion model approximates the score function i.e. the gradient of the log density of a noise-corrupted version of the target distribution for varying levels of corruption. We prove that as the level of corruption decreases, the score function points towards the manifold, as this direction becomes the direction of maximal likelihood increase. Therefore, at low noise levels, the diffusion model provides us with an approximation of the manifold’s normal bundle, allowing for an estimation of the manifold’s intrinsic dimension. To the best of our knowledge our method is the first estimator of intrinsic dimension based on diffusion models and it outperforms well established estimators in controlled experiments on both Euclidean and image data.
Jan Stanczuk, Georgios Batzolis, Teo Deveney, Carola-Bibiane Schönlieb
ICML2
2022 How to Distribute Data across Tasks for Meta-Learning?
abstract
Meta-learning models transfer the knowledge acquired from previous tasks to quickly learn new ones. They are trained on benchmarks with a fixed number of data points per task. This number is usually arbitrary and it is unknown how it affects performance at testing. Since labelling of data is expensive, finding the optimal allocation of labels across training tasks may reduce costs. Given a fixed budget of labels, should we use a small number of highly labelled tasks, or many tasks with few labels each? Should we allocate more labels to some tasks and less to others? We show that: 1) If tasks are homogeneous, there is a uniform optimal allocation, whereby all tasks get the same amount of data; 2) At fixed budget, there is a trade-off between number of tasks and number of data points per task, with a unique solution for the optimum; 3) When trained separately, harder task should get more data, at the cost of a smaller number of tasks; 4) When training on a mixture of easy and hard tasks, more data should be allocated to easy tasks. Interestingly, Neuroscience experiments have shown that human visual skills also transfer better from easy tasks. We prove these results mathematically on mixed linear regression, and we show empirically that the same results hold for few-shot image classification on CIFAR-FS and mini-ImageNet. Our results provide guidance for allocating labels across tasks when collecting data for meta-learning.
Alexandru Cioba, Michael Bromberg, Ritwik Niyogi, Georgios Batzolis, Jezabel R. Garcia, Da-Shan Shiu, Alberto Bernacchia
AAAI5