EDBT 2026 Demo / reviewers in the wild / expert
Oleg Arenz
dblp:190/8525
· DBLP profile ↗
9ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-9470-2833ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 57% Probabilistic and Bayesian machine learning · 43% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation |
0.9 | 1 | 2025 | Maximum Total Correlation Reinforcement Learning · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.8 | 2 | 2020 | Trust-Region Variational Inference with Gaussian Mixture Models · J. Mach. Learn. Res. 2020 Efficient Gradient-Free Variational Inference using Policy Search · ICML 2018 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.7 | 1 | 2023 | LS-IQ: Implicit Reward Regularization for Inverse Reinforcement Learning · ICLR 2023 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.7 | 1 | 2023 | LS-IQ: Implicit Reward Regularization for Inverse Reinforcement Learning · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture simplification |
0.4 | 1 | 2020 | Trust-Region Variational Inference with Gaussian Mixture Models · J. Mach. Learn. Res. 2020 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
mixture density estimation |
0.4 | 1 | 2020 | Expected Information Maximization: Using the I-Projection for Mixture Density Estimation · ICLR 2020 |
Machine learning › Reinforcement learning › policy optimization
trust region methods |
0.4 | 1 | 2020 | Trust-Region Variational Inference with Gaussian Mixture Models · J. Mach. Learn. Res. 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model |
0.3 | 1 | 2018 | Efficient Gradient-Free Variational Inference using Policy Search · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model |
0.3 | 1 | 2018 | Efficient Gradient-Free Variational Inference using Policy Search · ICML 2018 |
Machine learning › Reinforcement learning
policy search |
0.1 | 1 | 2018 | Efficient Gradient-Free Variational Inference using Policy Search · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
total correlation · 0.9regularization · 0.9lower bound approximation · 0.9policy search · 0.8inverse reinforcement learning · 0.7implicit reward regularization · 0.7trust region · 0.4information geometry · 0.4i-projection · 0.4gaussian mixture model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Maximum Total Correlation Reinforcement LearningabstractSimplicity is a powerful inductive bias. In reinforcement learning, regularization is used for simpler policies, data augmentation for simpler representations, and sparse reward functions for simpler objectives, all that, with the underlying motivation to increase generalizability and robustness by focusing on the essentials. Supplementary to these techniques, we investigate how to promote simple behavior throughout the episode. To that end, we introduce a modification of the reinforcement learning problem that additionally maximizes the total correlation within the induced trajectories. We propose a practical algorithm that optimizes all models, including policy and state representation, based on a lower-bound approximation. In simulated robot environments, our method naturally generates policies that induce periodic and compressible trajectories, and that exhibit superior robustness to noise and changes in dynamics compared to baseline methods, while also improving performance in the original tasks. Bang You, Puze Liu, Jan Peters 0001, Oleg Arenz |
ICML | 5 |
| 2025 | Context-Aware Deep Lagrangian Networks for Model Predictive ControlabstractControlling a robot based on physics-consistent dynamic models, such as Deep Lagrangian Networks (DeLaN), can improve the generalizability and interpretability of the resulting behavior. However, in complex environments, the number of objects to potentially interact with is vast, and their physical properties are often uncertain. This complexity makes it infeasible to employ a single global model. Therefore, we need to resort to online system identification of context-aware models that capture only the currently relevant aspects of the environment. While physical principles such as the conservation of energy may not hold across varying contexts, ensuring physical plausibility for any individual context-aware model can still be highly desirable, particularly when using it for receding horizon control methods such as model predictive control (MPC). Hence, in this work, we extend DeLaN to make it context-aware, combine it with a recurrent network for online system identification, and integrate it with an MPC for adaptive, physics-consistent control. We also combine DeLaN with a residual dynamics model to leverage the fact that a nominal model of the robot is typically available. We evaluate our method on a 7-DOF robot arm for trajectory tracking under varying loads. Our method reduces the end-effector tracking error by 39%, compared to a 21% improvement achieved by a baseline that uses an extended Kalman filter. Lucas Schulze, Jan Peters 0001, Oleg Arenz |
IROS | 3 |
| 2023 | LS-IQ: Implicit Reward Regularization for Inverse Reinforcement Learning
Firas Al-Hafez, Davide Tateo, Oleg Arenz, Guoping Zhao, Jan Peters 0001 |
ICLR | 3 |
| 2022 | Integrating contrastive learning with dynamic models for reinforcement learning from images
Bang You, Oleg Arenz, Youping Chen, Jan Peters 0001 |
Neurocomputing | 2 |
| 2020 | Expected Information Maximization: Using the I-Projection for Mixture Density Estimation
Philipp Becker, Oleg Arenz, Gerhard Neumann |
ICLR | 2 |
| 2020 | Deep Adversarial Reinforcement Learning for Object DisentanglingabstractDeep learning in combination with improved training techniques and high computational power has led to recent advances in the field of reinforcement learning (RL) and to successful robotic RL applications such as in-hand manipulation. However, most robotic RL relies on a well known initial state distribution. In real-world tasks, this information is however often not available. For example, when disentangling waste objects the actual position of the robot w.r.t. the objects may not match the positions the RL policy was trained for. To solve this problem, we present a novel adversarial reinforcement learning (ARL) framework. The ARL framework utilizes an adversary, which is trained to steer the original agent, the protagonist, to challenging states. We train the protagonist and the adversary jointly to allow them to adapt to the changing policy of their opponent. We show that our method can generalize from training to test scenarios by training an end-to-end system for robot control to solve a challenging object disentangling task. Experiments with a KUKA LBR+ 7-DOF robot arm show that our approach outperforms the baseline method in disentangling when starting from different initial states than provided during training. Melvin Laux, Oleg Arenz, Jan Peters 0001, Joni Pajarinen |
IROS | 2 |
| 2020 | Trust-Region Variational Inference with Gaussian Mixture ModelsabstractMany methods for machine learning rely on approximate inference from intractable probability distributions. Variational inference approximates such distributions by tractable models that can be subsequently used for approximate inference. Learning sufficiently accurate approximations requires a rich model family and careful exploration of the relevant modes of the target distribution. We propose a method for learning accurate GMM approximations of intractable probability distributions based on insights from policy search by using information-geometric trust regions for principled exploration. For efficient improvement of the GMM approximation, we derive a lower bound on the corresponding optimization objective enabling us to update the components independently. Our use of the lower bound ensures convergence to a stationary point of the original objective. The number of components is adapted online by adding new components in promising regions and by deleting components with negligible weight. We demonstrate on several domains that we can learn approximations of complex, multimodal distributions with a quality that is unmet by previous variational inference methods, and that the GMM approximation can be used for drawing samples that are on par with samples created by state-of-the-art MCMC samplers while requiring up to three orders of magnitude less computational resources. Oleg Arenz, Mingjun Zhong, Gerhard Neumann |
J. Mach. Learn. Res. | 1 |
| 2018 | Efficient Gradient-Free Variational Inference using Policy SearchabstractInference from complex distributions is a common problem in machine learning needed for many Bayesian methods. We propose an efficient, gradient-free method for learning general GMM approximations of multimodal distributions based on recent insights from stochastic search methods. Our method establishes information-geometric trust regions to ensure efficient exploration of the sampling space and stability of the GMM updates, allowing for efficient estimation of multi-variate Gaussian variational distributions. For GMMs, we apply a variational lower bound to decompose the learning objective into sub-problems given by learning the individual mixture components and the coefficients. The number of mixture components is adapted online in order to allow for arbitrary exact approximations. We demonstrate on several domains that we can learn significantly better approximations than competing variational inference methods and that the quality of samples drawn from our approximations is on par with samples created by state-of-the-art MCMC samplers that require significantly more computational resources. Oleg Arenz, Mingjun Zhong, Gerhard Neumann |
ICML | 1 |
| 2016 | Optimal control and inverse optimal control by distribution matchingabstractOptimal control is a powerful approach to achieve optimal behavior. However, it typically requires a manual specification of a cost function which often contains several objectives, such as reaching goal positions at different time steps or energy efficiency. Manually trading-off these objectives is often difficult and requires a high engineering effort. In this paper, we present a new approach to specify optimal behavior. We directly specify the desired behavior by a distribution over future states or features of the states. For example, the experimenter could choose to reach certain mean positions with given accuracy/variance at specified time steps. Our approach also unifies optimal control and inverse optimal control in one framework. Given a desired state distribution, we estimate a cost function such that the optimal controller matches the desired distribution. If the desired distribution is estimated from expert demonstrations, our approach performs inverse optimal control. We evaluate our approach on several optimal and inverse optimal control tasks on non-linear systems using incremental linearizations similar to differential dynamic programming approaches. Oleg Arenz, Hany Abdulsamad, Gerhard Neumann |
IROS | 1 |