Jan E. Gerken

dblp:293/9373 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-0172-7944ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 45% Trustworthy machine learning · 16% Learning theory · 9%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
equivariant neural network
1.722025
Learning Chern Numbers of Multiband Topological Insulators with Gauge Equivariant Neural Networks · NeurIPS 2025
Equivariant Neural Tangent Kernels · ICML 2025
Machine learning › Deep learning architectures and training › equivariant neural network
group convolutional network
0.912025
Equivariant Neural Tangent Kernels · ICML 2025
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.912025
Equivariant Neural Tangent Kernels · ICML 2025
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
0.812024
Diffeomorphic Counterfactuals With Generative Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.812024
Emergent Equivariance in Deep Ensembles · ICML 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Diffeomorphic Counterfactuals With Generative Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.812024
HEAL-SWIN: A Vision Transformer on the Sphere · CVPR 2024
Computer vision › Image recognition and object detection
image classification
0.612022
Equivariance versus Augmentation for Spherical Images · ICML 2022
Machine learning › Deep learning architectures and training › equivariant neural network
rotation equivariance
0.612022
Equivariance versus Augmentation for Spherical Images · ICML 2022
Computer vision › Segmentation and scene understanding
semantic segmentation
0.422024
HEAL-SWIN: A Vision Transformer on the Sphere · CVPR 2024
Equivariance versus Augmentation for Spherical Images · ICML 2022
Machine learning › Deep learning architectures and training
data augmentation
0.312025
Equivariant Neural Tangent Kernels · ICML 2025
Machine learning › Representation and self-supervised learning
equivariance
0.312025
Equivariant Neural Tangent Kernels · ICML 2025
Computer vision › 3D vision › depth estimation
depth regression
0.212024
HEAL-SWIN: A Vision Transformer on the Sphere · CVPR 2024
Machine learning › Generative modeling
generative model
0.212024
Diffeomorphic Counterfactuals With Generative Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › 3D vision › 3d shape representation
spherical representation
0.212024
HEAL-SWIN: A Vision Transformer on the Sphere · CVPR 2024

Methods — techniques the papers use, named apart from their topics

gauge equivariant normalization · 1.7data augmentation · 1.3universal approximation theorems · 0.9universal approximation theorem · 0.9kernel regression · 0.9group convolution · 0.9shifted-window attention · 0.8neural tangent kernel · 0.8gradient ascent · 0.8diffeomorphic coordinate transformation · 0.8HEALPix grid · 0.8
YearPublicationVenuePosition
2025 Equivariant Neural Tangent Kernels
abstract
Little is known about the training dynamics of equivariant neural networks, in particular how it compares to data augmented training of their non-equivariant counterparts. Recently, neural tangent kernels (NTKs) have emerged as a powerful tool to analytically study the training dynamics of wide neural networks. In this work, we take an important step towards a theoretical understanding of training dynamics of equivariant models by deriving neural tangent kernels for a broad class of equivariant architectures based on group convolutions. As a demonstration of the capabilities of our framework, we show an interesting relationship between data augmentation and group convolutional networks. Specifically, we prove that they share the same expected prediction over initializations at all training times and even off the data manifold. In this sense, they have the same training dynamics. We demonstrate in numerical experiments that this still holds approximately for finite-width ensembles. By implementing equivariant NTKs for roto-translations in the plane ($G=C_{n}\ltimes\mathbb{R}^{2}$) and 3d rotations ($G=\mathrm{SO}(3)$), we show that equivariant NTKs outperform their non-equivariant counterparts as kernel predictors for histological image classification and quantum mechanical property prediction.
Philipp Misof, Pan Kessel, Jan E. Gerken
ICML3
2025 Learning Chern Numbers of Multiband Topological Insulators with Gauge Equivariant Neural Networks
abstract
Equivariant network architectures are a well-established tool for predicting invariant or equivariant quantities. However, almost all learning problems considered in this context feature a global symmetry, i.e. each point of the underlying space is transformed with the same group element, as opposed to a local *gauge* symmetry, where each point is transformed with a different group element, exponentially enlarging the size of the symmetry group. We use gauge equivariant networks to predict topological invariants (Chern numbers) of multiband topological insulators for the first time. The gauge symmetry of the network guarantees that the predicted quantity is a topological invariant. A major technical challenge is that the relevant gauge equivariant networks are plagued by instabilities in their training, severely limiting their usefulness. In particular, for larger gauge groups the instabilities make training impossible. We resolve this problem by introducing a novel gauge equivariant normalization layer which stabilizes the training. Furthermore, we prove a universal approximation theorem for our model. We train on samples with trivial Chern number only but show that our model generalizes to samples with non-trivial Chern number and provide various ablations of our setup.
Longde Huang, Oleksandr Balabanov, Hampus Linander, Mats Granath, Daniel Persson, Jan E. Gerken
NeurIPS6
2024 HEAL-SWIN: A Vision Transformer on the Sphere
abstract
High-resolution wide-angle fisheye images are becoming more and more important for robotics applications such as autonomous driving. However, using ordinary convolutional neural networks or vision transformers on this data is problematic due to projection and distortion losses introduced when projecting to a rectangular grid on the plane. We introduce the HEAL-SWIN transformer, which combines the highly uniform Hierarchi-cal Equal Area iso-Latitude Pixelation (HEALPix) grid used in astrophysics and cosmology with the Hierarchical Shifted-Window (SWIN) transformer to yield an efficient and flexible model capable of training on high-resolution, distortion-free spherical data. In HEAL-SWIN, the nested structure of the HEALPix grid is used to perform the patching and windowing operations of the SWIN transformer, enabling the network to process spherical representations with minimal computational overhead. We demonstrate the superior performance of our model on both synthetic and real automotive datasets, as well as a selection of other image datasets, for semantic segmentation, depth regression and classification tasks. Our code is publicly available11https://github.com/JanEGerken/HEAL-SWIN.
Oscar Carlsson, Jan E. Gerken, Hampus Linander, Heiner Spieß, Fredrik Ohlsson, Christoffer Petersson, Daniel Persson
CVPR2
2024 Emergent Equivariance in Deep Ensembles
abstract
We show that deep ensembles become equivariant for all inputs and at all training times by simply using data augmentation. Crucially, equivariance holds off-manifold and for any architecture in the infinite width limit. The equivariance is emergent in the sense that predictions of individual ensemble members are not equivariant but their collective prediction is. Neural tangent kernel theory is used to derive this result and we verify our theoretical insights using detailed numerical experiments.
Jan E. Gerken, Pan Kessel
ICML1
2024 Diffeomorphic Counterfactuals With Generative Models
abstract
Counterfactuals can explain classification decisions of neural networks in a human interpretable way. We propose a simple but effective method to generate such counterfactuals. More specifically, we perform a suitable diffeomorphic coordinate transformation and then perform gradient ascent in these coordinates to find counterfactuals which are classified with great confidence as a specified target class. We propose two methods to leverage generative models to construct such suitable coordinate systems that are either exactly or approximately diffeomorphic. We analyze the generation process theoretically using Riemannian differential geometry and validate the quality of the generated counterfactuals using various qualitative and quantitative measures.
Ann-Kathrin Dombrowski, Jan E. Gerken, Klaus-Robert Müller, Pan Kessel
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Equivariance versus Augmentation for Spherical Images
abstract
We analyze the role of rotational equivariance in convolutional neural networks (CNNs) applied to spherical images. We compare the performance of the group equivariant networks known as S2CNNs and standard non-equivariant CNNs trained with an increasing amount of data augmentation. The chosen architectures can be considered baseline references for the respective design paradigms. Our models are trained and evaluated on single or multiple items from the MNIST- or FashionMNIST dataset projected onto the sphere. For the task of image classification, which is inherently rotationally invariant, we find that by considerably increasing the amount of data augmentation and the size of the networks, it is possible for the standard CNNs to reach at least the same performance as the equivariant network. In contrast, for the inherently equivariant task of semantic segmentation, the non-equivariant networks are consistently outperformed by the equivariant networks with significantly fewer parameters. We also analyze and compare the inference latency and training times of the different networks, enabling detailed tradeoff considerations between equivariant architectures and data augmentation for practical problems.
Jan E. Gerken, Oscar Carlsson, Hampus Linander, Fredrik Ohlsson, Christoffer Petersson, Daniel Persson
ICML1