Mario Geiger

dblp:206/7093 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-5433-0900ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Deep learning architectures and training · 36% Generative modeling · 32% Learning theory · 14%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 67% Computational science and engineering · 33%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 50% GPUs and heterogeneous computing · 50%

Topics — the 25 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
equivariant neural network
1.842023
A General Framework for Equivariant Neural Networks on Reductive Lie Groups · NeurIPS 2023
SE(3)-equivariant prediction of molecular wavefunctions and electronic densities · NeurIPS 2021
3D Steerable CNNs: Learning Rotationally Equivariant Features in Volumetric Data · NeurIPS 2018
Bioinformatics and computational biology › molecular informatics › molecular modeling
molecular conformation generation
1.622025
Efficient Molecular Conformer Generation with SO(3)-Averaged Flow Matching and Reflow · ICML 2025
Symphony: Symmetry-Equivariant Point-Centered Spherical Harmonics for 3D Molecule Generation · ICLR 2024
Machine learning › Deep learning architectures and training
training dynamics
1.022023
Dissecting the Effects of SGD Noise in Distinct Regimes of Deep Learning · ICML 2023
Comparing Dynamics: Deep Neural Networks versus Glassy Systems · ICML 2018
Machine learning › Generative modeling
flow matching
0.912025
Efficient Molecular Conformer Generation with SO(3)-Averaged Flow Matching and Reflow · ICML 2025
Machine learning › Generative modeling › diffusion model
molecular conformation generation
0.912025
Efficient Molecular Conformer Generation with SO(3)-Averaged Flow Matching and Reflow · ICML 2025
Machine learning › Generative modeling
normalizing flow
0.912025
Proteina: Scaling Flow-based Protein Structure Generative Models · ICLR 2025
Machine learning › Generative modeling › protein design
protein structure generation
0.912025
Proteina: Scaling Flow-based Protein Structure Generative Models · ICLR 2025
Computational science and engineering
computational chemistry
0.912025
Efficient Molecular Conformer Generation with SO(3)-Averaged Flow Matching and Reflow · ICML 2025
Bioinformatics and computational biology
protein design
0.912025
Proteina: Scaling Flow-based Protein Structure Generative Models · ICLR 2025
GPUs and heterogeneous computing
GPU training
0.912025
Optimizing Data Distribution and Kernel Performance for Efficient Training of Chemistry Foundation Models: A Case Study with MACE · HPDC 2025
Machine learning › Generative modeling
autoregressive model
0.812024
Symphony: Symmetry-Equivariant Point-Centered Spherical Harmonics for 3D Molecule Generation · ICLR 2024
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecule generation
0.812024
Symphony: Symmetry-Equivariant Point-Centered Spherical Harmonics for 3D Molecule Generation · ICLR 2024
Machine learning › Graph learning › graph neural network › geometric graph neural network
equivariant message passing
0.712023
A General Framework for Equivariant Neural Networks on Reductive Lie Groups · NeurIPS 2023
Machine learning › Learning theory
generalization
0.712023
Dissecting the Effects of SGD Noise in Distinct Regimes of Deep Learning · ICML 2023
Machine learning › Graph learning › graph neural network
message passing
0.712023
A General Framework for Equivariant Neural Networks on Reductive Lie Groups · NeurIPS 2023
Machine learning › Optimization for machine learning
optimization landscape
0.712023
Dissecting the Effects of SGD Noise in Distinct Regimes of Deep Learning · ICML 2023
Machine learning › Learning theory
phase transition
0.712023
Dissecting the Effects of SGD Noise in Distinct Regimes of Deep Learning · ICML 2023
Machine learning › Learning theory
generalization error
0.512021
Relative stability toward diffeomorphisms indicates performance in deep nets · NeurIPS 2021
Computational science and engineering › computational chemistry
quantum chemistry
0.512021
SE(3)-equivariant prediction of molecular wavefunctions and electronic densities · NeurIPS 2021
Machine learning › Deep learning architectures and training › equivariant neural network
group equivariant CNN
0.412019
A General Theory of Equivariant CNNs on Homogeneous Spaces · NeurIPS 2019
Machine learning › Deep learning architectures and training › loss landscape
loss landscape geometry
0.312018
Comparing Dynamics: Deep Neural Networks versus Glassy Systems · ICML 2018
Computer vision › 3D vision › geometric deep learning
rotation-equivariant learning
0.312018
3D Steerable CNNs: Learning Rotationally Equivariant Features in Volumetric Data · NeurIPS 2018
Machine learning › Deep learning architectures and training › equivariant neural network
spherical CNN
0.312018
Spherical CNNs · ICLR 2018
Machine learning › Deep learning architectures and training › equivariant neural network
steerable CNN
0.312018
3D Steerable CNNs: Learning Rotationally Equivariant Features in Volumetric Data · NeurIPS 2018
Mathematical optimization
nonconvex optimization
0.112018
Comparing Dynamics: Deep Neural Networks versus Glassy Systems · ICML 2018

Methods — techniques the papers use, named apart from their topics

flow matching · 3.5transformer · 1.7symmetric tensor contraction optimization · 1.7reflow · 1.7multi-objective bin packing · 1.7distillation · 1.7classifier-free guidance · 1.7SO(3)-averaged flow · 1.7LoRA · 1.7e(3) equivariance · 1.5spherical harmonics · 0.8message passing · 0.8statistical physics · 0.3mean-field theory · 0.3
YearPublicationVenuePosition
2025 Optimizing Data Distribution and Kernel Performance for Efficient Training of Chemistry Foundation Models: A Case Study with MACE
abstract
Chemistry Foundation Models (CFMs) that leverage Graph Neural Networks (GNNs) operating on 3D molecular graph structures are becoming indispensable tools for computational chemists and materials scientists. These models facilitate the understanding of matter and the discovery of new molecules and materials. In contrast to GNNs operating on large homogeneous graphs, GNNs used by CFMs process a large number of geometric graphs of varying sizes, requiring different optimization strategies than those developed for large homogeneous GNNs. This paper presents optimizations for two critical phases of CFM training: data distribution and model training, targeting MACE - a state-of-the-art CFM. We address the challenge of load balancing in data distribution by formulating it as a multi-objective bin packing problem. We propose an iterative algorithm that provides a highly effective, fast, and practical solution, ensuring efficient data distribution. For the training phase, we identify symmetric tensor contraction as the key computational kernel in MACE and optimize this kernel to improve the overall performance. Our combined approach of balanced data distribution and kernel optimization significantly enhances the training process of MACE. Experimental results demonstrate a substantial speedup, reducing per-epoch execution time for training from 12 to 2 minutes on 740 GPUs with a 2.6M sample dataset.
Jesun Sahariar Firoz, Franco Pellegrini, Mario Geiger, Darren Hsu, Jenna A. Bilbrey, Han-Yi Chou, Maximilian Stadler, Markus Höhnerbach, Tingyu Wang 0001, Dejun Lin, Emine Küçükbenli, Henry Sprueill, Ilyes Batatia, Sotiris S. Xantheas, MalSoon Lee, Christopher J. Mundy, Gábor Csányi, Justin S. Smith, P. Sadayappan, Sutanay Choudhury
HPDC3
2025 Proteina: Scaling Flow-based Protein Structure Generative Models
abstract
Recently, diffusion- and flow-based generative models of protein structures have emerged as a powerful tool for de novo protein design. Here, we develop *Proteina*, a new large-scale flow-based protein backbone generator that utilizes hierarchical fold class labels for conditioning and relies on a tailored scalable transformer architecture with up to $5\times$ as many parameters as previous models. To meaningfully quantify performance, we introduce a new set of metrics that directly measure the distributional similarity of generated proteins with reference sets, complementing existing metrics. We further explore scaling training data to millions of synthetic protein structures and explore improved training and sampling recipes adapted to protein backbone generation. This includes fine-tuning strategies like LoRA for protein backbones, new guidance methods like classifier-free guidance and autoguidance for protein backbones, and new adjusted training objectives. Proteina achieves state-of-the-art performance on de novo protein backbone design and produces diverse and designable proteins at unprecedented length, up to 800 residues. The hierarchical conditioning offers novel control, enabling high-level secondary-structure guidance as well as low-level fold-specific generation.
Tomas Geffner, Kieran Didi, Zuobai Zhang, Danny Reidenbach, Zhonglin Cao, Jason Yim, Mario Geiger, Christian Dallago, Emine Küçükbenli, Arash Vahdat, Karsten Kreis
ICLR7
2025 Efficient Molecular Conformer Generation with SO(3)-Averaged Flow Matching and Reflow
abstract
Fast and accurate generation of molecular conformers is desired for downstream computational chemistry and drug discovery tasks. Currently, training and sampling state-of-the-art diffusion or flow-based models for conformer generation require significant computational resources. In this work, we build upon flow-matching and propose two mechanisms for accelerating training and inference of generative models for 3D molecular conformer generation. For fast training, we introduce the SO(3)-*Averaged Flow* training objective, which leads to faster convergence to better generation quality compared to conditional optimal transport flow or Kabsch-aligned flow. We demonstrate that models trained using SO(3)-*Averaged Flow* can reach state-of-the-art conformer generation quality. For fast inference, we show that the reflow and distillation methods of flow-based models enable few-steps or even one-step molecular conformer generation with high quality. The training techniques proposed in this work show a path towards highly efficient molecular conformer generation with flow-based models.
Zhonglin Cao, Mario Geiger, Allan dos Santos Costa, Danny Reidenbach, Karsten Kreis, Tomas Geffner, Franco Pellegrini, Emine Küçükbenli
ICML2
2024 Symphony: Symmetry-Equivariant Point-Centered Spherical Harmonics for 3D Molecule Generation
abstract
We present Symphony, an $E(3)$ equivariant autoregressive generative model for 3D molecular geometries that iteratively builds a molecule from molecular fragments. Existing autoregressive models such as G-SchNet and G-SphereNet for molecules utilize rotationally invariant features to respect the 3D symmetries of molecules. In contrast, Symphony uses message-passing with higher-degree $E(3)$-equivariant features. This allows a novel representation of probability distributions via spherical harmonic signals to efficiently model the 3D geometry of molecules. We show that Symphony is able to accurately generate small molecules from the QM9 dataset, outperforming existing autoregressive models and approaching the performance of diffusion models. Our code is available at https://github.com/atomicarchitects/symphony.
Ameya Daigavane, Song Kim, Mario Geiger, Tess E. Smidt
ICLR3
2023 Dissecting the Effects of SGD Noise in Distinct Regimes of Deep Learning
abstract
Understanding when the noise in stochastic gradient descent (SGD) affects generalization of deep neural networks remains a challenge, complicated by the fact that networks can operate in distinct training regimes. Here we study how the magnitude of this noise $T$ affects performance as the size of the training set $P$ and the scale of initialization $\alpha$ are varied. For gradient descent, $\alpha$ is a key parameter that controls if the network is lazy’ ($\alpha\gg1$) or instead learns features ($\alpha\ll1$). For classification of MNIST and CIFAR10 images, our central results are: *(i)* obtaining phase diagrams for performance in the $(\alpha,T)$ plane. They show that SGD noise can be detrimental or instead useful depending on the training regime. Moreover, although increasing $T$ or decreasing $\alpha$ both allow the net to escape the lazy regime, these changes can have opposite effects on performance. *(ii)* Most importantly, we find that the characteristic temperature $T_c$ where the noise of SGD starts affecting the trained model (and eventually performance) is a power law of $P$. We relate this finding with the observation that key dynamical quantities, such as the total variation of weights during training, depend on both $T$ and $P$ as power laws. These results indicate that a key effect of SGD noise occurs late in training, by affecting the stopping process whereby all data are fitted. Indeed, we argue that due to SGD noise, nets must develop a strongersignal’, i.e. larger informative weights, to fit the data, leading to a longer training time. A stronger signal and a longer training time are also required when the size of the training set $P$ increases. We confirm these views in the perceptron model, where signal and noise can be precisely measured. Interestingly, exponents characterizing the effect of SGD depend on the density of data near the decision boundary, as we explain.
Antonio Sclocchi, Mario Geiger, Matthieu Wyart
ICML2
2023 A General Framework for Equivariant Neural Networks on Reductive Lie Groups
abstract
Reductive Lie Groups, such as the orthogonal groups, the Lorentz group, or the unitary groups, play essential roles across scientific fields as diverse as high energy physics, quantum mechanics, quantum chromodynamics, molecular dynamics, computer vision, and imaging. In this paper, we present a general Equivariant Neural Network architecture capable of respecting the symmetries of the finite-dimensional representations of any reductive Lie Group. Our approach generalizes the successful ACE and MACE architectures for atomistic point clouds to any data equivariant to a reductive Lie group action. We also introduce the lie-nn software library, which provides all the necessary tools to develop and implement such general G-equivariant neural networks. It implements routines for the reduction of generic tensor products of representations into irreducible representations, making it easy to apply our architecture to a wide range of problems and groups. The generality and performance of our approach are demonstrated by applying it to the tasks of top quark decay tagging (Lorentz group) and shape recognition (orthogonal group).
Ilyes Batatia, Mario Geiger, Jose M. Munoz, Tess E. Smidt, Lior Silberman, Christoph Ortner
NeurIPS2
2021 Relative stability toward diffeomorphisms indicates performance in deep nets
abstract
Understanding why deep nets can classify data in large dimensions remains a challenge. It has been proposed that they do so by becoming stable to diffeomorphisms, yet existing empirical measurements support that it is often not the case. We revisit this question by defining a maximum-entropy distribution on diffeomorphisms, that allows to study typical diffeomorphisms of a given norm. We confirm that stability toward diffeomorphisms does not strongly correlate to performance on benchmark data sets of images. By contrast, we find that the stability toward diffeomorphisms relative to that of generic transformations $R_f$ correlates remarkably with the test error $\epsilon_t$. It is of order unity at initialization but decreases by several decades during training for state-of-the-art architectures. For CIFAR10 and 15 known architectures, we find $\epsilon_t\approx 0.2\sqrt{R_f}$, suggesting that obtaining a small $R_f$ is important to achieve good performance. We study how $R_f$ depends on the size of the training set and compare it to a simple model of invariant learning.
Leonardo Petrini, Alessandro Favero, Mario Geiger, Matthieu Wyart
NeurIPS3
2021 SE(3)-equivariant prediction of molecular wavefunctions and electronic densities
abstract
Machine learning has enabled the prediction of quantum chemical properties with high accuracy and efficiency, allowing to bypass computationally costly ab initio calculations. Instead of training on a fixed set of properties, more recent approaches attempt to learn the electronic wavefunction (or density) as a central quantity of atomistic systems, from which all other observables can be derived. This is complicated by the fact that wavefunctions transform non-trivially under molecular rotations, which makes them a challenging prediction target. To solve this issue, we introduce general SE(3)-equivariant operations and building blocks for constructing deep learning architectures for geometric point cloud data and apply them to reconstruct wavefunctions of atomistic systems with unprecedented accuracy. Our model achieves speedups of over three orders of magnitude compared to ab initio methods and reduces prediction errors by up to two orders of magnitude compared to the previous state-of-the-art. This accuracy makes it possible to derive properties such as energies and forces directly from the wavefunction in an end-to-end manner. We demonstrate the potential of our approach in a transfer learning application, where a model trained on low accuracy reference wavefunctions implicitly learns to correct for electronic many-body interactions from observables computed at a higher level of theory. Such machine-learned wavefunction surrogates pave the way towards novel semi-empirical methods, offering resolution at an electronic level while drastically decreasing computational cost. Additionally, the predicted wavefunctions can serve as initial guess in conventional ab initio methods, decreasing the number of iterations required to arrive at a converged solution, thus leading to significant speedups without any loss of accuracy or robustness. While we focus on physics applications in this contribution, the proposed equivariant framework for deep learning on point clouds is promising also beyond, say, in computer vision or graphics.
Oliver T. Unke, Mihail Bogojeski, Michael Gastegger, Mario Geiger, Tess E. Smidt, Klaus-Robert Müller
NeurIPS4
2019 A General Theory of Equivariant CNNs on Homogeneous Spaces
abstract
We present a general theory of Group equivariant Convolutional Neural Networks (G-CNNs) on homogeneous spaces such as Euclidean space and the sphere. Feature maps in these networks represent fields on a homogeneous base space, and layers are equivariant maps between spaces of fields. The theory enables a systematic classification of all existing G-CNNs in terms of their symmetry group, base space, and field type. We also answer a fundamental question: what is the most general kind of equivariant linear map between feature spaces (fields) of given types? We show that such maps correspond one-to-one with generalized convolutions with an equivariant kernel, and characterize the space of such kernels.
Taco Cohen, Mario Geiger, Maurice Weiler
NeurIPS2
2018 Spherical CNNs
Taco Cohen, Mario Geiger, Jonas Köhler 0001, Max Welling
ICLR2
2018 Comparing Dynamics: Deep Neural Networks versus Glassy Systems
abstract
We analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are the complexity of the loss-landscape and of the dynamics within it, and to what extent DNNs share similarities with glassy systems. Our findings, obtained for different architectures and data-sets, suggest that during the training process the dynamics slows down because of an increasingly large number of flat directions. At large times, when the loss is approaching zero, the system diffuses at the bottom of the landscape. Despite some similarities with the dynamics of mean-field glassy systems, in particular, the absence of barrier crossing, we find distinctive dynamical behaviors in the two cases, thus showing that the statistical properties of the corresponding loss and energy landscapes are different. In contrast, when the network is under-parametrized we observe a typical glassy behavior, thus suggesting the existence of different phases depending on whether the network is under-parametrized or over-parametrized.
Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gérard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, Giulio Biroli
ICML3
2018 3D Steerable CNNs: Learning Rotationally Equivariant Features in Volumetric Data
abstract
We present a convolutional network that is equivariant to rigid body motions. The model uses scalar-, vector-, and tensor fields over 3D Euclidean space to represent data, and equivariant convolutions to map between such representations. These SE(3)-equivariant convolutions utilize kernels which are parameterized as a linear combination of a complete steerable kernel basis, which is derived analytically in this paper. We prove that equivariant convolutions are the most general equivariant linear maps between fields over R^3. Our experimental results confirm the effectiveness of 3D Steerable CNNs for the problem of amino acid propensity prediction and protein structure classification, both of which have inherent SE(3) symmetry.
Maurice Weiler, Mario Geiger, Max Welling, Wouter Boomsma, Taco Cohen
NeurIPS2