Adrian Sosic

dblp:118/1267 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2019
0000-0003-2845-6635ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 35% Probabilistic and Bayesian machine learning · 23% Multi-agent systems · 8%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.722019
Correlation Priors for Reinforcement Learning · NeurIPS 2019
Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling · J. Mach. Learn. Res. 2018
Machine learning › Reinforcement learning
bayesian reinforcement learning
0.412019
Correlation Priors for Reinforcement Learning · NeurIPS 2019
Machine learning › Kernel, tree and ensemble methods › kernel embedding
mean embedding
0.412019
Deep Reinforcement Learning for Swarm Systems · J. Mach. Learn. Res. 2019
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412019
Deep Reinforcement Learning for Swarm Systems · J. Mach. Learn. Res. 2019
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation
0.412019
Deep Reinforcement Learning for Swarm Systems · J. Mach. Learn. Res. 2019
Knowledge, reasoning and agents › Multi-agent systems
swarm systems
0.412019
Deep Reinforcement Learning for Swarm Systems · J. Mach. Learn. Res. 2019
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.312018
Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling · J. Mach. Learn. Res. 2018
Robotics › Robot manipulation
learning from demonstration
0.312018
A Bayesian Approach to Policy Recognition and State Representation Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Video understanding and tracking
spatio-temporal modeling
0.312018
Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling · J. Mach. Learn. Res. 2018
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.312018
A Bayesian Approach to Policy Recognition and State Representation Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Machine learning › Reinforcement learning
imitation learning
0.112019
Correlation Priors for Reinforcement Learning · NeurIPS 2019
Robotics › Motion planning and robot control
system identification
0.112019
Correlation Priors for Reinforcement Learning · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

radial basis functions · 0.4pólya-gamma augmentation · 0.4neural network features · 0.4deep reinforcement learning · 0.4bayesian modeling · 0.4nonparametric bayesian modeling · 0.3non-parametric model · 0.3bayesian inference · 0.3active learning · 0.3
YearPublicationVenuePosition
2019 Correlation Priors for Reinforcement Learning
abstract
Many decision-making problems naturally exhibit pronounced structures inherited from the characteristics of the underlying environment. In a Markov decision process model, for example, two distinct states can have inherently related semantics or encode resembling physical state configurations. This often implies locally correlated transition dynamics among the states. In order to complete a certain task in such environments, the operating agent usually needs to execute a series of temporally and spatially correlated actions. Though there exists a variety of approaches to capture these correlations in continuous state-action domains, a principled solution for discrete environments is missing. In this work, we present a Bayesian learning framework based on Pólya-Gamma augmentation that enables an analogous reasoning in such cases. We demonstrate the framework on a number of common decision-making related problems, such as imitation learning, subgoal extraction, system identification and Bayesian reinforcement learning. By explicitly modeling the underlying correlation structures of these problems, the proposed approach yields superior predictive performance compared to correlation-agnostic models, even when trained on data sets that are an order of magnitude smaller in size.
Bastian Alt, Adrian Sosic, Heinz Koeppl
NeurIPS2
2019 Deep Reinforcement Learning for Swarm Systems
abstract
Recently, deep reinforcement learning (RL) methods have been applied successfully to multi-agent scenarios. Typically, the observation vector for decentralized decision making is represented by a concatenation of the (local) information an agent gathers about other agents. However, concatenation scales poorly to swarm systems with a large number of homogeneous agents as it does not exploit the fundamental properties inherent to these systems: (i) the agents in the swarm are interchangeable and (ii) the exact number of agents in the swarm is irrelevant. Therefore, we propose a new state representation for deep multi-agent RL based on mean embeddings of distributions, where we treat the agents as samples and use the empirical mean embedding as input for a decentralized policy. We define different feature spaces of the mean embedding using histograms, radial basis functions and neural networks trained end-to-end. We evaluate the representation on two well-known problems from the swarm literature in a globally and locally observable setup. For the local setup we furthermore introduce simple communication protocols. Of all approaches, the mean embedding representation using neural network features enables the richest information exchange between neighboring agents, facilitating the development of complex collective strategies.
Maximilian Hüttenrauch, Adrian Sosic, Gerhard Neumann
J. Mach. Learn. Res.2
2018 Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling
abstract
Advances in the field of inverse reinforcement learning (IRL) have led to sophisticated inference frameworks that relax the original modeling assumption of observing an agent behavior that reflects only a single intention. Instead of learning a global behavioral model, recent IRL methods divide the demonstration data into parts, to account for the fact that different trajectories may correspond to different intentions, e.g., because they were generated by different domain experts. In this work, we go one step further: using the intuitive concept of subgoals, we build upon the premise that even a single trajectory can be explained more efficiently locally within a certain context than globally, enabling a more compact representation of the observed behavior. Based on this assumption, we build an implicit intentional model of the agent's goals to forecast its behavior in unobserved situations. The result is an integrated Bayesian prediction framework that significantly outperforms existing IRL solutions and provides smooth policy estimates consistent with the expert's plan. Most notably, our framework naturally handles situations where the intentions of the agent change over time and classical IRL algorithms fail. In addition, due to its probabilistic nature, the model can be straightforwardly applied in active learning scenarios to guide the demonstration process of the expert.
Adrian Sosic, Elmar Rueckert, Jan Peters 0001, Abdelhak M. Zoubir, Heinz Koeppl
J. Mach. Learn. Res.1
2018 A Bayesian Approach to Policy Recognition and State Representation Learning
abstract
Learning from demonstration (LfD) is the process of building behavioral models of a task from demonstrations provided by an expert. These models can be used, e.g., for system control by generalizing the expert demonstrations to previously unencountered situations. Most LfD methods, however, make strong assumptions about the expert behavior, e.g., they assume the existence of a deterministic optimal ground truth policy or require direct monitoring of the expert's controls, which limits their practical use as part of a general system identification framework. In this work, we consider the LfD problem in a more general setting where we allow for arbitrary stochastic expert policies, without reasoning about the optimality of the demonstrations. Following a Bayesian methodology, we model the full posterior distribution of possible expert controllers that explain the provided demonstration data. Moreover, we show that our methodology can be applied in a nonparametric context to infer the complexity of the state representation used by the expert, and to learn task-appropriate partitionings of the system state space.
Adrian Sosic, Abdelhak M. Zoubir, Heinz Koeppl
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Policy recognition via expectation maximization
abstract
Learning from Demonstrations (LfD) has proven to be a powerful concept for solving optimal control problems in high-dimensional state spaces where demonstrations can be used to facilitate the search for efficient control policies. However, many existing LfD approaches suffer from either theoretical, practical, or computational drawbacks such as the need to learn a latent reward model, to monitor the expert's controls, or to repeatedly solve potentially demanding planning problems. In this work, we consider the LfD objective from a system identification perspective and propose a probabilistic policy recognition framework based on expectation maximization that operates directly on the observed expert trajectories, avoiding the aforementioned problems. Using a spatial prior over policies, we are able to make accurate predictions in regions of the state space that are scarcely explored.
Adrian Sosic, Abdelhak M. Zoubir, Heinz Koeppl
ICASSP1
2014 sNN-LDS: Spatio-temporal Non-negative Sparse Coding for Human Action Recognition
Thomas Guthier, Adrian Sosic, Volker Willert, Julian Eggert
ICANN2