Heiner Kremer

dblp:319/3506 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 37% Deep learning architectures and training · 27% Learning theory · 19%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
recurrent neural network
0.912025
Implicit Language Models are RNNs: Balancing Parallelization and Expressivity · ICML 2025
Machine learning › Deep learning architectures and training
state space model
0.912025
Implicit Language Models are RNNs: Balancing Parallelization and Expressivity · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
conditional moment restrictions
0.822023
Functional Generalized Empirical Likelihood Estimation for Conditional Moment Restrictions · ICML 2022
Estimation Beyond Data Reweighting: Kernel Method of Moments · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference
instrumental variable regression
0.812024
Geometry-Aware Instrumental Variable Regression · ICML 2024
Machine learning › Trustworthy machine learning › adversarial machine learning
robustness to adversarial attacks
0.812024
Geometry-Aware Instrumental Variable Regression · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
method of moments
0.712023
Estimation Beyond Data Reweighting: Kernel Method of Moments · ICML 2023
Machine learning › Efficient and distributed learning › distributed training
parallelization
0.312025
Implicit Language Models are RNNs: Balancing Parallelization and Expressivity · ICML 2025
Mathematical optimization
optimal transport
0.212024
Geometry-Aware Instrumental Variable Regression · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.212022
Functional Generalized Empirical Likelihood Estimation for Conditional Moment Restrictions · ICML 2022

Methods — techniques the papers use, named apart from their topics

generalized method of moments · 2.2sinkhorn divergence · 1.5optimal transport · 1.5implicit models · 0.9fixed-point iteration · 0.9maximum mean discrepancy · 0.7empirical likelihood · 0.7neural network · 0.6kernel methods · 0.6generalized empirical likelihood · 0.6
YearPublicationVenuePosition
2025 Implicit Language Models are RNNs: Balancing Parallelization and Expressivity
abstract
State-space models (SSMs) and transformers dominate the language modeling landscape. However, they are constrained to a lower computational complexity than classical recurrent neural networks (RNNs), limiting their expressivity. In contrast, RNNs lack parallelization during training, raising fundamental questions about the trade off between parallelization and expressivity. We propose implicit SSMs, which iterate a transformation until convergence to a fixed point. Theoretically, we show that implicit SSMs implement the non-linear state-transitions of RNNs. Empirically, we find that only approximate fixed-point convergence suffices, enabling the design of a scalable training curriculum that largely retains parallelization, with full convergence required only for a small subset of tokens. Our approach demonstrates superior state-tracking capabilities on regular languages, surpassing transformers and SSMs. We further scale implicit SSMs to natural language reasoning tasks and pretraining of large-scale language models up to 1.3B parameters on 207B tokens - representing, to our knowledge, the largest implicit model trained to date. Notably, our implicit models outperform their explicit counterparts on standard benchmarks. Our code is publicly available at github.com/microsoft/implicit_languagemodels
Mark Schöne, Babak Rahmani, Heiner Kremer, Fabian Falck, Hitesh Ballani, Jannes Gladrow
ICML3
2024 Geometry-Aware Instrumental Variable Regression
abstract
Instrumental variable (IV) regression can be approached through its formulation in terms of conditional moment restrictions (CMR). Building on variants of the generalized method of moments, most CMR estimators are implicitly based on approximating the population data distribution via reweightings of the empirical sample. While for large sample sizes, in the independent identically distributed (IID) setting, reweightings can provide sufficient flexibility, they might fail to capture the relevant information in presence of corrupted data or data prone to adversarial attacks. To address these shortcomings, we propose the Sinkhorn Method of Moments, an optimal transport-based IV estimator that takes into account the geometry of the data manifold through data-derivative information. We provide a simple plug-and-play implementation of our method that performs on par with related estimators in standard settings but improves robustness against data corruption and adversarial attacks.
Heiner Kremer, Bernhard Schölkopf
ICML1
2023 Glare Removal for Astronomical Images with High Local Dynamic Range
abstract
Veiling glare is a common imaging corruption, describing the phenomenon of light spreading from bright regions into otherwise dark regions. This leads to a reduction of contrast, locally limiting high dynamic range imaging capabilities in addition to what the sensor’s pixels are able to record. In astronomical imaging, which motivates the present study, high local dynamic ranges are ubiquitous, an extreme example being the contrast between the moon and the sun’s corona during a total solar eclipse. While this corruption phenomenon has been known for a long time, the field lacks a specialized method for this problem. Building on the structure of the glare-corruption operator, we devise an efficient framework which can be combined with off-the-shelf optimization methods to address the problem of glare removal. We also present a novel evaluation set consisting of pairs of images with and without glare. We show state-of-the art glare removal performance on this benchmark and other real world images, including telescope images of a total solar eclipse.
Max-Olivier Van Bastelaer, Heiner Kremer, Valentin Volchkov, Jean-Claude Passy, Bernhard Schölkopf
ICCP2
2023 Estimation Beyond Data Reweighting: Kernel Method of Moments
abstract
Moment restrictions and their conditional counterparts emerge in many areas of machine learning and statistics ranging from causal inference to reinforcement learning. Estimators for these tasks, generally called methods of moments, include the prominent generalized method of moments (GMM) which has recently gained attention in causal inference. GMM is a special case of the broader family of empirical likelihood estimators which are based on approximating a population distribution by means of minimizing a $\varphi$-divergence to an empirical distribution. However, the use of $\varphi$-divergences effectively limits the candidate distributions to reweightings of the data samples. We lift this long-standing limitation and provide a method of moments that goes beyond data reweighting. This is achieved by defining an empirical likelihood estimator based on maximum mean discrepancy which we term the kernel method of moments (KMM). We provide a variant of our estimator for conditional moment restrictions and show that it is asymptotically first-order optimal for such problems. Finally, we show that our method achieves competitive performance on several conditional moment restriction tasks.
Heiner Kremer, Yassine Nemmour, Bernhard Schölkopf, Jia-Jie Zhu
ICML1
2022 Functional Generalized Empirical Likelihood Estimation for Conditional Moment Restrictions
abstract
Important problems in causal inference, economics, and, more generally, robust machine learning can be expressed as conditional moment restrictions, but estimation becomes challenging as it requires solving a continuum of unconditional moment restrictions. Previous works addressed this problem by extending the generalized method of moments (GMM) to continuum moment restrictions. In contrast, generalized empirical likelihood (GEL) provides a more general framework and has been shown to enjoy favorable small-sample properties compared to GMM-based estimators. To benefit from recent developments in machine learning, we provide a functional reformulation of GEL in which arbitrary models can be leveraged. Motivated by a dual formulation of the resulting infinite dimensional optimization problem, we devise a practical method and explore its asymptotic properties. Finally, we provide kernel- and neural network-based implementations of the estimator, which achieve state-of-the-art empirical performance on two conditional moment restriction problems.
Heiner Kremer, Jia-Jie Zhu, Krikamol Muandet, Bernhard Schölkopf
ICML1