EDBT 2026 Demo / reviewers in the wild / expert
Wolfgang Lehrach
dblp:190/7782
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Probabilistic and Bayesian machine learning · 40% Reinforcement learning · 34% Robot navigation and mapping · 6% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
model-based reinforcement learning |
1.6 | 2 | 2025 | Improving Transformer World Models for Data-Efficient RL · ICML 2025 Learning Cognitive Maps from Transformer Representations for Efficient Planning in Partially Observed Environments · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference |
1.3 | 2 | 2024 | PGMax: Factor Graphs for Discrete Probabilistic Graphical Models and Loopy Belief Propagation in JAX · J. Mach. Learn. Res. 2024 Query Training: Learning a Worse Model to Infer Better Marginals in Undirected Graphical Models with Hidden Variables · AAAI 2021 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
1.3 | 2 | 2024 | PGMax: Factor Graphs for Discrete Probabilistic Graphical Models and Loopy Belief Propagation in JAX · J. Mach. Learn. Res. 2024 Query Training: Learning a Worse Model to Infer Better Marginals in Undirected Graphical Models with Hidden Variables · AAAI 2021 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.9 | 1 | 2025 | Improving Transformer World Models for Data-Efficient RL · ICML 2025 |
Robotics › Robot navigation and mapping › robot mapping
cognitive map |
0.8 | 1 | 2024 | Learning Cognitive Maps from Transformer Representations for Efficient Planning in Partially Observed Environments · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
factor graphs |
0.8 | 1 | 2024 | PGMax: Factor Graphs for Discrete Probabilistic Graphical Models and Loopy Belief Propagation in JAX · J. Mach. Learn. Res. 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
loopy belief propagation |
0.8 | 1 | 2024 | PGMax: Factor Graphs for Discrete Probabilistic Graphical Models and Loopy Belief Propagation in JAX · J. Mach. Learn. Res. 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | DMC-VB: A Benchmark for Representation Learning for Control with Visual Distractors · NeurIPS 2024 |
Robotics › Motion planning and robot control
path planning |
0.8 | 1 | 2024 | Learning Cognitive Maps from Transformer Representations for Efficient Planning in Partially Observed Environments · ICML 2024 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.8 | 1 | 2024 | Learning Cognitive Maps from Transformer Representations for Efficient Planning in Partially Observed Environments · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
marginal inference |
0.5 | 1 | 2021 | Query Training: Learning a Worse Model to Infer Better Marginals in Undirected Graphical Models with Hidden Variables · AAAI 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation |
0.5 | 1 | 2021 | Query Training: Learning a Worse Model to Infer Better Marginals in Undirected Graphical Models with Hidden Variables · AAAI 2021 |
Computer vision › Segmentation and scene understanding › image segmentation › document image segmentation
character segmentation |
0.2 | 1 | 2016 | Generative Shape Models: Joint Text Recognition and Segmentation with Very Little Training Data · NIPS 2016 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.2 | 1 | 2016 | Generative Shape Models: Joint Text Recognition and Segmentation with Very Little Training Data · NIPS 2016 |
Machine learning › Generative modeling › 3d generative model
mesh generative model |
0.2 | 1 | 2016 | Generative Shape Models: Joint Text Recognition and Segmentation with Very Little Training Data · NIPS 2016 |
Computer vision › Image recognition and object detection
scene text recognition |
0.2 | 1 | 2016 | Generative Shape Models: Joint Text Recognition and Segmentation with Very Little Training Data · NIPS 2016 |
Methods — techniques the papers use, named apart from their topics
teacher forcing · 0.9nearest neighbor tokenizer · 0.9dyna · 0.9RNN · 0.9CNN · 0.9transformer · 0.8representation pretraining · 0.8discrete bottleneck · 0.8behavioral cloning · 0.8LSTM · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Transformer World Models for Data-Efficient RLabstractWe present an approach to model-based RL that achieves a new state of the art performance on the challenging Craftax-classic benchmark, an open-world 2D survival game that requires agents to exhibit a wide range of general abilities---such as strong generalization, deep exploration, and long-term reasoning. With a series of careful design choices aimed at improving sample efficiency, our MBRL algorithm achieves a reward of 69.66% after only 1M environment steps, significantly outperforming DreamerV3, which achieves $53.2\%$, and, for the first time, exceeds human performance of 65.0%. Our method starts by constructing a SOTA model-free baseline, using a novel policy architecture that combines CNNs and RNNs.
We then add three improvements to the standard MBRL setup: (a) "Dyna with warmup", which trains the policy on real and imaginary data, (b) "nearest neighbor tokenizer" on image patches, which improves the scheme to create the transformer world model (TWM) inputs, and (c) "block teacher forcing", which allows the TWM to reason jointly about the future tokens of the next timestep. Antoine Dedieu, Joseph Ortiz, Xinghua Lou, Carter Wendelken, J. Swaroop Guntupalli, Wolfgang Lehrach, Miguel Lázaro-Gredilla, Kevin Murphy 0002 |
ICML | 6 |
| 2024 | Learning Cognitive Maps from Transformer Representations for Efficient Planning in Partially Observed EnvironmentsabstractDespite their stellar performance on a wide range of tasks, including in-context tasks only revealed during inference, vanilla transformers and variants trained for next-token predictions (a) do not learn an explicit world model of their environment which can be flexibly queried and (b) cannot be used for planning or navigation. In this paper, we consider partially observed environments (POEs), where an agent receives perceptually aliased observations as it navigates, which makes path planning hard. We introduce a transformer with (multiple) discrete bottleneck(s), TDB, whose latent codes learn a compressed representation of the history of observations and actions. After training a TDB to predict the future observation(s) given the history, we extract interpretable cognitive maps of the environment from its active bottleneck(s) indices. These maps are then paired with an external solver to solve (constrained) path planning problems. First, we show that a TDB trained on POEs (a) retains the near-perfect predictive performance of a vanilla transformer or an LSTM while (b) solving shortest path problems exponentially faster. Second, a TDB extracts interpretable representations from text datasets, while reaching higher in-context accuracy than vanilla sequence models. Finally, in new POEs, a TDB (a) reaches near-perfect in-context accuracy, (b) learns accurate in-context cognitive maps (c) solves in-context path planning problems. Antoine Dedieu, Wolfgang Lehrach, Dileep George, Miguel Lázaro-Gredilla |
ICML | 2 |
| 2024 | DMC-VB: A Benchmark for Representation Learning for Control with Visual DistractorsabstractLearning from previously collected data via behavioral cloning or offline reinforcement learning (RL) is a powerful recipe for scaling generalist agents by avoiding the need for expensive online learning. Despite strong generalization in some respects, agents are often remarkably brittle to minor visual variations in control-irrelevant factors such as the background or camera viewpoint. In this paper, we present theDeepMind Control Visual Benchmark (DMC-VB), a dataset collected in the DeepMind Control Suite to evaluate the robustness of offline RL agents for solving continuous control tasks from visual input in the presence of visual distractors. In contrast to prior works, our dataset (a) combines locomotion and navigation tasks of varying difficulties, (b) includes static and dynamic visual variations, (c) considers data generated by policies with different skill levels, (d) systematically returns pairs of state and pixel observation, (e) is an order of magnitude larger, and (f) includes tasks with hidden goals. Accompanying our dataset, we propose three benchmarks to evaluate representation learning methods for pretraining, and carry out experiments on several recently proposed methods. First, we find that pretrained representations do not help policy learning on DMC-VB, and we highlight a large representation gap between policies learned on pixel observations and on states. Second, we demonstrate when expert data is limited, policy learning can benefit from representations pretrained on (a) suboptimal data, and (b) tasks with stochastic hidden goals. Our dataset and benchmark code to train and evaluate agents are available at https://github.com/google-deepmind/dmcvisionbenchmark. Joseph Ortiz, Antoine Dedieu, Wolfgang Lehrach, J. Swaroop Guntupalli, Carter Wendelken, Ahmad Humayun, Sivaramakrishnan Swaminathan, Miguel Lázaro-Gredilla, Kevin Murphy 0002 |
NeurIPS | 3 |
| 2024 | PGMax: Factor Graphs for Discrete Probabilistic Graphical Models and Loopy Belief Propagation in JAXabstractPGMax is an open-source Python/ JAX package for (a) easily specifying discrete Probabilistic Graphical Models (PGMs) as factor graphs; and (b) automatically running efficient and scalable differentiable Loopy Belief Propagation (LBP). PGMax supports general factor graphs with tractable factors, and leverages modern accelerators like GPUs for inference. Compared with alternative libraries, PGMax obtains higher-quality inference results with up to three orders-of-magnitude inference time speedups. PGMax interacts seamlessly with the growing JAX ecosystem, opening up new research possibilities. Our source code, examples and documentation are available at https://github.com/google-deepmind/PGMax Antoine Dedieu, Nishanth Kumar, Wolfgang Lehrach, Shrinu Kushagra, Dileep George, Miguel Lázaro-Gredilla |
J. Mach. Learn. Res. | 4 |
| 2021 | Query Training: Learning a Worse Model to Infer Better Marginals in Undirected Graphical Models with Hidden VariablesabstractProbabilistic graphical models (PGMs) provide a compact representation of knowledge that can be queried in a flexible way: after learning the parameters of a graphical model once, new probabilistic queries can be answered at test time without retraining. However, when using undirected PGMS with hidden variables, two sources of error typically compound in all but the simplest models (a) learning error (both computing the partition function and integrating out the hidden variables is intractable); and (b) prediction error (exact inference is also intractable). Here we introduce query training (QT), a mechanism to learn a PGM that is optimized for the approximate inference algorithm that will be paired with it. The resulting PGM is a worse model of the data (as measured by the likelihood), but it is tuned to produce better marginals for a given inference algorithm. Unlike prior works, our approach preserves the querying flexibility of the original PGM: at test time, we can estimate the marginal of any variable given any partial evidence. We demonstrate experimentally that QT can be used to learn a challenging 8-connected grid Markov random field with hidden variables and that it consistently outperforms the state-of-the-art AdVIL when tested on three undirected models across multiple datasets. Miguel Lázaro-Gredilla, Wolfgang Lehrach, Nishad Gothoskar, Antoine Dedieu, Dileep George |
AAAI | 2 |
| 2016 | Generative Shape Models: Joint Text Recognition and Segmentation with Very Little Training DataabstractWe demonstrate that a generative model for object shapes can achieve state of the art results on challenging scene text recognition tasks, and with orders of magnitude fewer training images than required for competing discriminative methods. In addition to transcribing text from challenging images, our method performs fine-grained instance segmentation of characters. We show that our model is more robust to both affine transformations and non-affine deformations compared to previous approaches. Xinghua Lou, Ken Kansky, Wolfgang Lehrach, C. C. Laan, Bhaskara Marthi, D. Scott Phoenix, Dileep George |
NIPS | 3 |