VLDB 2026 Research / reviewers in the wild / expert
Heiner Kremer
dblp:319/3506
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Probabilistic and Bayesian machine learning · 37% Deep learning architectures and training · 27% Learning theory · 19% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
recurrent neural network |
0.9 | 1 | 2025 | Implicit Language Models are RNNs: Balancing Parallelization and Expressivity · ICML 2025 |
Machine learning › Deep learning architectures and training
state space model |
0.9 | 1 | 2025 | Implicit Language Models are RNNs: Balancing Parallelization and Expressivity · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
conditional moment restrictions |
0.8 | 2 | 2023 | Functional Generalized Empirical Likelihood Estimation for Conditional Moment Restrictions · ICML 2022 Estimation Beyond Data Reweighting: Kernel Method of Moments · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
instrumental variable regression |
0.8 | 1 | 2024 | Geometry-Aware Instrumental Variable Regression · ICML 2024 |
Machine learning › Trustworthy machine learning › adversarial machine learning
robustness to adversarial attacks |
0.8 | 1 | 2024 | Geometry-Aware Instrumental Variable Regression · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
method of moments |
0.7 | 1 | 2023 | Estimation Beyond Data Reweighting: Kernel Method of Moments · ICML 2023 |
Machine learning › Efficient and distributed learning › distributed training
parallelization |
0.3 | 1 | 2025 | Implicit Language Models are RNNs: Balancing Parallelization and Expressivity · ICML 2025 |
Mathematical optimization
optimal transport |
0.2 | 1 | 2024 | Geometry-Aware Instrumental Variable Regression · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.2 | 1 | 2022 | Functional Generalized Empirical Likelihood Estimation for Conditional Moment Restrictions · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
generalized method of moments · 2.2sinkhorn divergence · 1.5optimal transport · 1.5implicit models · 0.9fixed-point iteration · 0.9maximum mean discrepancy · 0.7empirical likelihood · 0.7neural network · 0.6kernel methods · 0.6generalized empirical likelihood · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Implicit Language Models are RNNs: Balancing Parallelization and ExpressivityabstractState-space models (SSMs) and transformers dominate the language modeling landscape. However, they are constrained to a lower computational complexity than classical recurrent neural networks (RNNs), limiting their expressivity. In contrast, RNNs lack parallelization during training, raising fundamental questions about the trade off between parallelization and expressivity. We propose implicit SSMs, which iterate a transformation until convergence to a fixed point. Theoretically, we show that implicit SSMs implement the non-linear state-transitions of RNNs. Empirically, we find that only approximate fixed-point convergence suffices, enabling the design of a scalable training curriculum that largely retains parallelization, with full convergence required only for a small subset of tokens. Our approach demonstrates superior state-tracking capabilities on regular languages, surpassing transformers and SSMs. We further scale implicit SSMs to natural language reasoning tasks and pretraining of large-scale language models up to 1.3B parameters on 207B tokens - representing, to our knowledge, the largest implicit model trained to date. Notably, our implicit models outperform their explicit counterparts on standard benchmarks. Our code is publicly available at github.com/microsoft/implicit_languagemodels Mark Schöne, Babak Rahmani, Heiner Kremer, Fabian Falck, Hitesh Ballani, Jannes Gladrow |
ICML | 3 |
| 2024 | Geometry-Aware Instrumental Variable RegressionabstractInstrumental variable (IV) regression can be approached through its formulation in terms of conditional moment restrictions (CMR). Building on variants of the generalized method of moments, most CMR estimators are implicitly based on approximating the population data distribution via reweightings of the empirical sample. While for large sample sizes, in the independent identically distributed (IID) setting, reweightings can provide sufficient flexibility, they might fail to capture the relevant information in presence of corrupted data or data prone to adversarial attacks. To address these shortcomings, we propose the Sinkhorn Method of Moments, an optimal transport-based IV estimator that takes into account the geometry of the data manifold through data-derivative information. We provide a simple plug-and-play implementation of our method that performs on par with related estimators in standard settings but improves robustness against data corruption and adversarial attacks. Heiner Kremer, Bernhard Schölkopf |
ICML | 1 |
| 2023 | Glare Removal for Astronomical Images with High Local Dynamic RangeabstractVeiling glare is a common imaging corruption, describing the phenomenon of light spreading from bright regions into otherwise dark regions. This leads to a reduction of contrast, locally limiting high dynamic range imaging capabilities in addition to what the sensor’s pixels are able to record. In astronomical imaging, which motivates the present study, high local dynamic ranges are ubiquitous, an extreme example being the contrast between the moon and the sun’s corona during a total solar eclipse. While this corruption phenomenon has been known for a long time, the field lacks a specialized method for this problem. Building on the structure of the glare-corruption operator, we devise an efficient framework which can be combined with off-the-shelf optimization methods to address the problem of glare removal. We also present a novel evaluation set consisting of pairs of images with and without glare. We show state-of-the art glare removal performance on this benchmark and other real world images, including telescope images of a total solar eclipse. Max-Olivier Van Bastelaer, Heiner Kremer, Valentin Volchkov, Jean-Claude Passy, Bernhard Schölkopf |
ICCP | 2 |
| 2023 | Estimation Beyond Data Reweighting: Kernel Method of MomentsabstractMoment restrictions and their conditional counterparts emerge in many areas of machine learning and statistics ranging from causal inference to reinforcement learning. Estimators for these tasks, generally called methods of moments, include the prominent generalized method of moments (GMM) which has recently gained attention in causal inference. GMM is a special case of the broader family of empirical likelihood estimators which are based on approximating a population distribution by means of minimizing a $\varphi$-divergence to an empirical distribution. However, the use of $\varphi$-divergences effectively limits the candidate distributions to reweightings of the data samples. We lift this long-standing limitation and provide a method of moments that goes beyond data reweighting. This is achieved by defining an empirical likelihood estimator based on maximum mean discrepancy which we term the kernel method of moments (KMM). We provide a variant of our estimator for conditional moment restrictions and show that it is asymptotically first-order optimal for such problems. Finally, we show that our method achieves competitive performance on several conditional moment restriction tasks. Heiner Kremer, Yassine Nemmour, Bernhard Schölkopf, Jia-Jie Zhu |
ICML | 1 |
| 2022 | Functional Generalized Empirical Likelihood Estimation for Conditional Moment RestrictionsabstractImportant problems in causal inference, economics, and, more generally, robust machine learning can be expressed as conditional moment restrictions, but estimation becomes challenging as it requires solving a continuum of unconditional moment restrictions. Previous works addressed this problem by extending the generalized method of moments (GMM) to continuum moment restrictions. In contrast, generalized empirical likelihood (GEL) provides a more general framework and has been shown to enjoy favorable small-sample properties compared to GMM-based estimators. To benefit from recent developments in machine learning, we provide a functional reformulation of GEL in which arbitrary models can be leveraged. Motivated by a dual formulation of the resulting infinite dimensional optimization problem, we devise a practical method and explore its asymptotic properties. Finally, we provide kernel- and neural network-based implementations of the estimator, which achieve state-of-the-art empirical performance on two conditional moment restriction problems. Heiner Kremer, Jia-Jie Zhu, Krikamol Muandet, Bernhard Schölkopf |
ICML | 1 |