EDBT 2026 Demo / reviewers in the wild / expert
Adrian Sosic
dblp:118/1267
· DBLP profile ↗
6ranked-venue papers
3as first author
0since 2021 · last 2019
0000-0003-2845-6635ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 35% Probabilistic and Bayesian machine learning · 23% Multi-agent systems · 8% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.7 | 2 | 2019 | Correlation Priors for Reinforcement Learning · NeurIPS 2019 Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling · J. Mach. Learn. Res. 2018 |
Machine learning › Reinforcement learning
bayesian reinforcement learning |
0.4 | 1 | 2019 | Correlation Priors for Reinforcement Learning · NeurIPS 2019 |
Machine learning › Kernel, tree and ensemble methods › kernel embedding
mean embedding |
0.4 | 1 | 2019 | Deep Reinforcement Learning for Swarm Systems · J. Mach. Learn. Res. 2019 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.4 | 1 | 2019 | Deep Reinforcement Learning for Swarm Systems · J. Mach. Learn. Res. 2019 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation |
0.4 | 1 | 2019 | Deep Reinforcement Learning for Swarm Systems · J. Mach. Learn. Res. 2019 |
Knowledge, reasoning and agents › Multi-agent systems
swarm systems |
0.4 | 1 | 2019 | Deep Reinforcement Learning for Swarm Systems · J. Mach. Learn. Res. 2019 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.3 | 1 | 2018 | Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling · J. Mach. Learn. Res. 2018 |
Robotics › Robot manipulation
learning from demonstration |
0.3 | 1 | 2018 | A Bayesian Approach to Policy Recognition and State Representation Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Computer vision › Video understanding and tracking
spatio-temporal modeling |
0.3 | 1 | 2018 | Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling · J. Mach. Learn. Res. 2018 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning |
0.3 | 1 | 2018 | A Bayesian Approach to Policy Recognition and State Representation Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Machine learning › Reinforcement learning
imitation learning |
0.1 | 1 | 2019 | Correlation Priors for Reinforcement Learning · NeurIPS 2019 |
Robotics › Motion planning and robot control
system identification |
0.1 | 1 | 2019 | Correlation Priors for Reinforcement Learning · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
radial basis functions · 0.4pólya-gamma augmentation · 0.4neural network features · 0.4deep reinforcement learning · 0.4bayesian modeling · 0.4nonparametric bayesian modeling · 0.3non-parametric model · 0.3bayesian inference · 0.3active learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Correlation Priors for Reinforcement LearningabstractMany decision-making problems naturally exhibit pronounced structures inherited from the characteristics of the underlying environment. In a Markov decision process model, for example, two distinct states can have inherently related semantics or encode resembling physical state configurations. This often implies locally correlated transition dynamics among the states. In order to complete a certain task in such environments, the operating agent usually needs to execute a series of temporally and spatially correlated actions. Though there exists a variety of approaches to capture these correlations in continuous state-action domains, a principled solution for discrete environments is missing. In this work, we present a Bayesian learning framework based on Pólya-Gamma augmentation that enables an analogous reasoning in such cases. We demonstrate the framework on a number of common decision-making related problems, such as imitation learning, subgoal extraction, system identification and Bayesian reinforcement learning. By explicitly modeling the underlying correlation structures of these problems, the proposed approach yields superior predictive performance compared to correlation-agnostic models, even when trained on data sets that are an order of magnitude smaller in size. Bastian Alt, Adrian Sosic, Heinz Koeppl |
NeurIPS | 2 |
| 2019 | Deep Reinforcement Learning for Swarm SystemsabstractRecently, deep reinforcement learning (RL) methods have been applied successfully to multi-agent scenarios. Typically, the observation vector for decentralized decision making is represented by a concatenation of the (local) information an agent gathers about other agents. However, concatenation scales poorly to swarm systems with a large number of homogeneous agents as it does not exploit the fundamental properties inherent to these systems: (i) the agents in the swarm are interchangeable and (ii) the exact number of agents in the swarm is irrelevant. Therefore, we propose a new state representation for deep multi-agent RL based on mean embeddings of distributions, where we treat the agents as samples and use the empirical mean embedding as input for a decentralized policy. We define different feature spaces of the mean embedding using histograms, radial basis functions and neural networks trained end-to-end. We evaluate the representation on two well-known problems from the swarm literature in a globally and locally observable setup. For the local setup we furthermore introduce simple communication protocols. Of all approaches, the mean embedding representation using neural network features enables the richest information exchange between neighboring agents, facilitating the development of complex collective strategies. Maximilian Hüttenrauch, Adrian Sosic, Gerhard Neumann |
J. Mach. Learn. Res. | 2 |
| 2018 | Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal ModelingabstractAdvances in the field of inverse reinforcement learning (IRL) have led to sophisticated inference frameworks that relax the original modeling assumption of observing an agent behavior that reflects only a single intention. Instead of learning a global behavioral model, recent IRL methods divide the demonstration data into parts, to account for the fact that different trajectories may correspond to different intentions, e.g., because they were generated by different domain experts. In this work, we go one step further: using the intuitive concept of subgoals, we build upon the premise that even a single trajectory can be explained more efficiently locally within a certain context than globally, enabling a more compact representation of the observed behavior. Based on this assumption, we build an implicit intentional model of the agent's goals to forecast its behavior in unobserved situations. The result is an integrated Bayesian prediction framework that significantly outperforms existing IRL solutions and provides smooth policy estimates consistent with the expert's plan. Most notably, our framework naturally handles situations where the intentions of the agent change over time and classical IRL algorithms fail. In addition, due to its probabilistic nature, the model can be straightforwardly applied in active learning scenarios to guide the demonstration process of the expert. Adrian Sosic, Elmar Rueckert, Jan Peters 0001, Abdelhak M. Zoubir, Heinz Koeppl |
J. Mach. Learn. Res. | 1 |
| 2018 | A Bayesian Approach to Policy Recognition and State Representation LearningabstractLearning from demonstration (LfD) is the process of building behavioral models of a task from demonstrations provided by an expert. These models can be used, e.g., for system control by generalizing the expert demonstrations to previously unencountered situations. Most LfD methods, however, make strong assumptions about the expert behavior, e.g., they assume the existence of a deterministic optimal ground truth policy or require direct monitoring of the expert's controls, which limits their practical use as part of a general system identification framework. In this work, we consider the LfD problem in a more general setting where we allow for arbitrary stochastic expert policies, without reasoning about the optimality of the demonstrations. Following a Bayesian methodology, we model the full posterior distribution of possible expert controllers that explain the provided demonstration data. Moreover, we show that our methodology can be applied in a nonparametric context to infer the complexity of the state representation used by the expert, and to learn task-appropriate partitionings of the system state space. Adrian Sosic, Abdelhak M. Zoubir, Heinz Koeppl |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Policy recognition via expectation maximizationabstractLearning from Demonstrations (LfD) has proven to be a powerful concept for solving optimal control problems in high-dimensional state spaces where demonstrations can be used to facilitate the search for efficient control policies. However, many existing LfD approaches suffer from either theoretical, practical, or computational drawbacks such as the need to learn a latent reward model, to monitor the expert's controls, or to repeatedly solve potentially demanding planning problems. In this work, we consider the LfD objective from a system identification perspective and propose a probabilistic policy recognition framework based on expectation maximization that operates directly on the observed expert trajectories, avoiding the aforementioned problems. Using a spatial prior over policies, we are able to make accurate predictions in regions of the state space that are scarcely explored. Adrian Sosic, Abdelhak M. Zoubir, Heinz Koeppl |
ICASSP | 1 |
| 2014 | sNN-LDS: Spatio-temporal Non-negative Sparse Coding for Human Action Recognition
Thomas Guthier, Adrian Sosic, Volker Willert, Julian Eggert |
ICANN | 2 |