EDBT 2026 Demo / reviewers in the wild / expert
Benjamin Rosman
dblp:45/4591 · also Benjamin Saul Rosman
· DBLP profile ↗
40ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0002-0284-4114ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 3 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 3 since 2021Systems, architecture and hardware · 6 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource LanguagesabstractJenalea Rajab, Anuoluwapo Aremu, Everlyn Asiko Chimoto, Dale Dunbar, Graham Morrissey, Fadel Thior, Luandrie Potgieter, Jessica Ojo, Atnafu Lambebo Tonja, Wilhelmina NdapewaOnyothi Nekoto, Pelonomi Moiloa, Jade Abbott, Vukosi Marivate, Benjamin Rosman. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jenalea Rajab, Aremu Anuoluwapo, Everlyn Chimoto, Dale Dunbar, Graham Morrissey, Fadel Thior, Luandrie Potgieter, Jessica Ojo, Atnafu Lambebo Tonja, Wilhelmina Nekoto, Pelonomi Moiloa, Jade Z. Abbott, Vukosi Marivate, Benjamin Rosman |
ACL (1) | 14 |
| 2025 | Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU NetworksabstractIn spite of finite dimension ReLU neural networks being a consistent factor behind recent deep learning successes, a theory of feature learning in these models remains elusive. Currently, insightful theories still rely on assumptions including the linearity of the network computations, unstructured input data and architectural constraints such as infinite width or a single hidden layer. To begin to address this gap we establish an equivalence between ReLU networks and Gated Deep Linear Networks, and use their greater tractability to derive dynamics of learning. We then consider multiple variants of a core task reminiscent of multi-task learning or contextual control which requires both feature learning and nonlinearity. We make explicit that, for these tasks, the ReLU networks possess an inductive bias towards latent representations which are *not* strictly modular or disentangled but are still highly structured and reusable between contexts. This effect is amplified with the addition of more contexts and hidden layers. Thus, we take a step towards a theory of feature learning in finite ReLU networks and shed light on how structured mixed-selective latent representations can emerge due to a bias for node-reuse and learning speed. Devon Jarvis, Richard Klein 0002, Benjamin Rosman, Andrew M. Saxe |
ICLR | 3 |
| 2025 | M-SAT: Multi-State-Action Tokenisation in Decision Transformers for Multi-Discrete ActionsabstractEffective decision-making in complex environments with multi-discrete action spaces poses significant challenges for agent architectures, particularly in image-based settings. While Decision Transformers have shown promise in various domains, their performance often suffers in environments where agents must handle multi-discrete actions. Existing enhancements to Decision Transformer architectures have yet to address this critical issue, limiting their ability to support agents in learning robust policies in these environments. To address this gap, we propose Multi-State Action Tokenisation (M-SAT), a novel approach designed to improve agent decision-making by tokenising actions at the individual action level and incorporating auxiliary state information. This disentanglement of actions improves both the performance of agents and the interpretability of individual actions within attention layers, fostering better visibility into agent decision processes. Importantly, M-SAT facilitates the development of more interpretable and transparent agents capable of making complex decisions in dynamic environments involving multi-discrete action spaces. We evaluate M-SAT on the challenging ViZDoom environments, focusing on scenarios with multi-discrete action spaces and image-based observations, such as Deadly Corridor, My Way Home and Death Match. Our approach demonstrates superior performance compared to baseline Decision Transformers, with no additional data or significant computational overheads. Furthermore, we observe that M-SAT does not require positional encoding to achieve high performance, with its removal occasionally leading to further improvements. These findings suggest that M-SAT enables more efficient and interpretable agent-based decision-making in multi-discrete action spaces. Perusha Moodley, Dhillu Thambi, Mark Trovinger, Pramod Kaushik, Praveen Paruchuri, Xia Hong 0001, Benjamin Rosman |
IJCNN | 7 |
| 2025 | Composition and Zero-Shot Transfer with Lattice Structures in Reinforcement LearningabstractAn important property of long-lived agents is the ability to reuse existing knowledge to solve new tasks. An appealing approach towards obtaining such agents is by leveraging logical composition over tasks, where new tasks are defined by applying logic operators to previously-solved ones. This composition is particularly powerful since it provides a human-understandable mechanism for task specification. However, no unifying formalism for applying logic operators to tasks and generalising combinatorially over them has yet been developed. We address the problem by formally defining logical composition as operators acting on a set of tasks in a lattice structure—the algebraic structure that generalises the study of Boolean logic. This provides a theoretically rigorous method for composing tasks, allowing us to formulate new tasks in terms of the negation, disjunction, and conjunction of a set of base tasks. We prove that by learning a new type of goal-oriented value function model free, called the world value function, an agent can solve composite tasks involving arbitrary logical operators with no further learning. We verify our approach in high-dimensional domains—including a video game environment and continuous-control task—where an agent first learns to solve a set of base tasks, and then composes these solutions to solve a super-exponential number of new tasks. Geraud Nangue Tasse, Steven James 0001, Benjamin Rosman |
J. Artif. Intell. Res. | 3 |
| 2025 | Source-Free Elastic Model Adaptation for Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) requires an agent to follow given instructions to navigate. Despite the significant progress, the model trained on seen environments has a performance drop on unseen environments due to distribution shift. To improve the generalization, existing method attempts to apply test-time adaptation to VLN. However, it needs to access the training data and all testing data for updating the model before inference. The setting is not suitable for the real application because it is hard for the agent to access training data and all testing data when the agent is applied in a new environment. In this paper, we consider a more practical setting with source-free and online-inference test-time adaption. In other words, the model can only access one testing sample for test-time adaptation. In this setting, the model may suffer from catastrophic forgetting of the learned knowledge and unstable parameter update issues. To solve these challenges, we propose an elastic adaptation model (EAM) that consists of an auxiliary decision model and a sample replay mechanism. We use the online testing samples to adapt the auxiliary decision model to new environments, which cooperates with the frozen original model to make better action decisions. The sample replay mechanism stores the historical testing samples to make the adaptation process more stable. Our method is model-agnostic and is effortless to be applied to most existing methods. Experimental results show that our method achieves stable performance improvement based on three existing methods on three VLN benchmark datasets. Mingkui Tan, Peihao Chen, Hongyan Zhi, Jiajie Mai, Benjamin Rosman, Dongyu Ji, Runhao Zeng |
IEEE Trans. Multim. | 5 |
| 2024 | The Zeno's Paradox of 'Low-Resource' LanguagesabstractThe disparity in the languages commonly studied in Natural Language Processing (NLP) is typically reflected by referring to languages as low vs high-resourced.However, there is limited consensus on what exactly qualifies as a 'low-resource language.'To understand how NLP papers define and study 'low resource' languages, we qualitatively analyzed 150 papers from the ACL Anthology and popular speechprocessing conferences that mention the keyword 'low-resource.' Based on our analysis, we show how several interacting axes contribute to 'low-resourcedness' of a language and why that makes it difficult to track progress for each individual language.We hope our work (1) elicits explicit definitions of the terminology when it is used in papers and (2) provides grounding for the different axes to consider when connoting a language as low-resource. Hellina Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio, Monojit Choudhury |
EMNLP | 3 |
| 2024 | Skill Machines: Temporal Logic Skill Composition in Reinforcement LearningabstractIt is desirable for an agent to be able to solve a rich variety of problems that can be specified through language in the same environment. A popular approach towards obtaining such agents is to reuse skills learned in prior tasks to generalise compositionally to new ones. However, this is a challenging problem due to the curse of dimensionality induced by the combinatorially large number of ways high-level goals can be combined both logically and temporally in language. To address this problem, we propose a framework where an agent first learns a sufficient set of skill primitives to achieve all high-level goals in its environment. The agent can then flexibly compose them both logically and temporally to provably achieve temporal logic specifications in any regular language, such as regular fragments of linear temporal logic. This provides the agent with the ability to map from complex temporal logic task specifications to near-optimal behaviours zero-shot. We demonstrate this experimentally in a tabular setting, as well as in a high-dimensional video game and continuous control environment. Finally, we also demonstrate that the performance of skill machines can be improved with regular off-policy reinforcement learning algorithms when optimal behaviours are desired. Geraud Nangue Tasse, Devon Jarvis, Steven James 0001, Benjamin Rosman |
ICLR | 4 |
| 2024 | Transferable dynamics models for efficient object-oriented reinforcement learningabstractThe Reinforcement Learning (RL) framework offers a general paradigm for constructing autonomous agents that can make effective decisions when solving tasks. An important area of study within the field of RL is transfer learning, where an agent utilizes knowledge gained from solving previous tasks to solve a new task more efficiently. While the notion of transfer learning is conceptually appealing, in practice, not all RL representations are amenable to transfer learning. Moreover, much of the research on transfer learning in RL is purely empirical. Previous research has shown that object-oriented representations are suitable for the purposes of transfer learning with theoretical efficiency guarantees. Such representations leverage the notion of object classes to learn lifted rules that apply to grounded object instantiations. In this paper, we extend previous research on object-oriented representations and introduce two formalisms: the first is based on deictic predicates, and is used to learn a transferable transition dynamics model; the second is based on propositions, and is used to learn a transferable reward dynamics model. In addition, we extend previously introduced efficient learning algorithms for object-oriented representations to our proposed formalisms. Our frameworks are then combined into a single efficient algorithm that learns transferable transition and reward dynamics models across a domain of related tasks. We illustrate our proposed algorithm empirically on an extended version of the Taxi domain, as well as the more difficult Sokoban domain, showing the benefits of our approach with regards to efficient learning and transfer. Ofir Marom, Benjamin Rosman |
Artif. Intell. | 2 |
| 2024 | Hierarchically Composing Level Generators for the Creation of Complex StructuresabstractProcedural content generation (PCG) is a growing field, with numerous applications in the video game industry and great potential to help create better games at a fraction of the cost of manual creation. However, much of the work in PCG is focused on generating relatively straightforward levels in simple games, as it is challenging to design an optimizable objective function for complex settings. This limits the applicability of PCG to more complex and modern titles, hindering its adoption in the industry. Our work aims to address this limitation by introducing a compositional level generation method that recursively composes simple low-level generators to construct large and complex creations. This approach allows for easily-optimizable objectives and the ability to design a complex structure in an interpretable way by referencing lower-level components. We empirically demonstrate that our method outperforms a noncompositional baseline by more accurately satisfying a designer's functional requirements in several tasks. Finally, we provide a qualitative showcase (inMinecraft) illustrating the large and complex, but still coherent, structures that were generated using simple base generators. Michael Beukman, Manuel Fokam, Marcel Kruger, Guy Axelrod, Muhammad Umair Nasir, Branden Ingram, Benjamin Rosman, Steven James 0001 |
IEEE Trans. Games | 7 |
| 2023 | On The Specialization of Neural Modules
Devon Jarvis, Richard Klein 0002, Benjamin Rosman, Andrew M. Saxe |
ICLR | 3 |
| 2023 | Dynamics Generalisation in Reinforcement Learning via Adaptive Context-Aware PoliciesabstractWhile reinforcement learning has achieved remarkable successes in several domains, its real-world application is limited due to many methods failing to generalise to unfamiliar conditions. In this work, we consider the problem of generalising to new transition dynamics, corresponding to cases in which the environment's response to the agent's actions differs. For example, the gravitational force exerted on a robot depends on its mass and changes the robot's mobility. Consequently, in such cases, it is necessary to condition an agent's actions on extrinsic state information and pertinent contextual information reflecting how the environment responds. While the need for context-sensitive policies has been established, the manner in which context is incorporated architecturally has received less attention. Thus, in this work, we present an investigation into how context information should be incorporated into behaviour learning to improve generalisation. To this end, we introduce a neural network architecture, the Decision Adapter, which generates the weights of an adapter module and conditions the behaviour of an agent on the context information. We show that the Decision Adapter is a useful generalisation of a previously proposed architecture and empirically demonstrate that it results in superior generalisation performance compared to previous approaches in several environments. Beyond this, the Decision Adapter is more robust to irrelevant distractor variables than several alternative methods. Michael Beukman, Devon Jarvis, Richard Klein 0002, Steven James 0001, Benjamin Rosman |
NeurIPS | 5 |
| 2023 | Who should I trust? Cautiously learning with unreliable experts
Tamlin Love, Ritesh Ajoodha, Benjamin Rosman |
Neural Comput. Appl. | 3 |
| 2023 | Generating Interpretable Play-Style Descriptions Through Deep Unsupervised Clustering of TrajectoriesabstractIn any game, play style is a concept that describes the technique and strategy employed by a player to achieve a goal. Identifying a player's style is desirable as it can enlighten players on which approaches work better or worse in different scenarios and inform developers of the value of design decisions. In previous work, we demonstrated an unsupervised LSTM-autoencoder clustering approach for play-style identification capable of handling multidimensional variable length player trajectories. The efficacy of our model was demonstrated on both complete and partial trajectories in both a simulated and natural environment. Lastly, through state frequency analysis, the properties of each of the play styles were identified and compared. This work expands on this approach by demonstrating a process by which we utilize temporal information to identify the decision boundaries related to particular clusters. Additionally, we demonstrate further robustness by applying the same techniques toMiniDungeons, another popular domain for player modeling research. Finally, we also propose approaches for determining mean play-style examples suitable for describing general play-style behaviors and for determining the correct number of represented play-styles. Branden Ingram, Clint J. van Alten, Richard Klein 0002, Benjamin Rosman |
IEEE Trans. Games | 4 |
| 2023 | FABRIC: A Framework for the Design and Evaluation of Collaborative Robots with Extended Human AdaptationabstractA limitation for collaborative robots (cobots) is their lack of ability to adapt to human partners, who typically exhibit an immense diversity of behaviors. We present an autonomous framework as a cobot’s real-time decision-making mechanism to anticipate a variety of human characteristics and behaviors, including human errors, toward a personalized collaboration. Our framework handles such behaviors in two levels: (1) short-term human behaviors are adapted through our novel Anticipatory Partially Observable Markov Decision Process (A-POMDP) models, covering a human’s changing intent (motivation), availability, and capability; (2) long-term changing human characteristics are adapted by our novel Adaptive Bayesian Policy Selection (ABPS) mechanism that selects a short-term decision model, e.g., an A-POMDP, according to an estimate of a human’s workplace characteristics, such as her expertise and collaboration preferences. To design and evaluate our framework over a diversity of human behaviors, we propose a pipeline where we first train and rigorously test the framework in simulation over novel human models. Then, we deploy and evaluate it on our novel physical experiment setup that induces cognitive load on humans to observe their dynamic behaviors, including their mistakes, and their changing characteristics such as their expertise. We conduct user studies and show that our framework effectively collaborates non-stop for hours and adapts to various changing human behaviors and characteristics in real-time. That increases the efficiency and naturalness of the collaboration with a higher perceived collaboration, positive teammate traits, and human trust. We believe that such an extended human-adaptation is a key to the long-term use of cobots. O. Can Görür, Benjamin Rosman, Fikret Sivrikaya, Sahin Albayrak |
ACM Trans. Hum. Robot Interact. | 2 |
| 2022 | Play-style Identification through Deep Unsupervised Clustering of TrajectoriesabstractIn any game, play-style is a concept that describes the technique and strategy employed by a player to achieve a goal. Being able to identify the play-style of a player is desirable as it can enlighten players on which approaches work better or worse in different scenarios, as well as inform developers of the value of design decisions. In this paper, we propose a novel approach to play-style identification based on an unsupervised LSTM-autoencoder clustering approach for multi-dimensional trajectory-based data of variable length. We evaluate our approach on two domains and show that not only is our model capable of identifying these play-styles from entire trajectories but it is also capable of this during gameplay from partial trajectories. Additionally, it is demonstrated through state frequency analysis that the properties of each of the play-styles can be identified and compared. Through these processes, we can extract useful information which describes the different behaviours or play-styles present within a domain useful to both players and developers. Branden Ingram, Benjamin Rosman, Clint J. van Alten, Richard Klein 0002 |
CoG | 2 |
| 2022 | Improved Action Prediction through Multiple Model Processing of Player TrajectoriesabstractAction prediction in video games is the process of extracting useful information in order to predict the future actions of a player. Long-range dependencies and the dynamic nature of video games make it difficult for most algorithms to accurately predict the future actions of players. We propose a novel machine learning approach to improving future action prediction from video game trajectories. This method requires having first clustered player trajectories based on behaviour similarities. Our model consists of a set of LSTM based prediction modules each trained on a subset of data based upon a respective cluster. The effectiveness of our model is analysed on both a synthetic and natural dataset. We find that our future action prediction approach of leveraging multiple models trained on individual data subsets results in greater accuracy over a single model on a complete dataset. Branden Ingram, Benjamin Rosman, Clint J. van Alten, Richard Klein 0002 |
CoG | 2 |
| 2022 | Autonomous Learning of Object-Centric Abstractions for High-Level Planning
Steven James 0001, Benjamin Rosman, George Dimitri Konidaris |
ICLR | 2 |
| 2022 | Generalisation in Lifelong Reinforcement Learning through Logical Composition
Geraud Nangue Tasse, Steven James 0001, Benjamin Rosman |
ICLR | 3 |
| 2022 | Reducing the Planning Horizon Through Reinforcement Learning
Logan Dunbar, Benjamin Rosman, Anthony G. Cohn 0001, Matteo Leonetti |
ECML/PKDD (4) | 2 |
| 2020 | Learning Portable Representations for High-Level PlanningabstractWe present a framework for autonomously learning a portable representation that describes a collection of low-level continuous environments. We show that these abstract representations can be learned in a task-independent egocentric space specific to the agent that, when grounded with problem-specific information, are provably sufficient for planning. We demonstrate transfer in two different domains, where an agent learns a portable, task-independent symbolic vocabulary, as well as operators expressed in that vocabulary, and then learns to instantiate those operators on a per-task basis. This reduces the number of samples required to learn a representation of a new task. Steven James 0001, Benjamin Rosman, George Dimitri Konidaris |
ICML | 2 |
| 2020 | A Boolean Task Algebra for Reinforcement LearningabstractThe ability to compose learned skills to solve new tasks is an important property for lifelong-learning agents. In this work we formalise the logical composition of tasks as a Boolean algebra. This allows us to formulate new tasks in terms of the negation, disjunction and conjunction of a set of base tasks. We then show that by learning goal-oriented value functions and restricting the transition dynamics of the tasks, an agent can solve these new tasks with no further learning. We prove that by composing these value functions in specific ways, we immediately recover the optimal policies for all tasks expressible under the Boolean algebra. We verify our approach in two domains---including a high-dimensional video game environment requiring function approximation---where an agent first learns a set of base skills, and then composes them to solve a super-exponential number of new tasks. Geraud Nangue Tasse, Steven James 0001, Benjamin Rosman |
NeurIPS | 3 |
| 2020 | If dropout limits trainable depth, does critical initialisation still matter? A large-scale statistical analysis on ReLU networks
Arnu Pretorius, Elan Van Biljon, Benjamin van Niekerk, Ryan Eloff, Matthew Reynard, Steven James 0001, Benjamin Rosman, Herman Kamper, Steve Kroon |
Pattern Recognit. Lett. | 7 |
| 2019 | Composing Value Functions in Reinforcement LearningabstractAn important property for lifelong-learning agents is the ability to combine existing skills to solve new unseen tasks. In general, however, it is unclear how to compose existing skills in a principled manner. Under the assumption of deterministic dynamics, we prove that optimal value function composition can be achieved in entropy-regularised reinforcement learning (RL), and extend this result to the standard RL setting. Composition is demonstrated in a high-dimensional video game, where an agent with an existing library of skills is immediately able to solve new tasks without the need for further learning. Benjamin van Niekerk, Steven James 0001, Adam Christopher Earle, Benjamin Rosman |
ICML | 4 |
| 2018 | Belief Reward Shaping in Reinforcement LearningabstractA key challenge in many reinforcement learning problems is delayed rewards, which can significantly slow down learning. Although reward shaping has previously been introduced to accelerate learning by bootstrapping an agent with additional information, this can lead to problems with convergence. We present a novel Bayesian reward shaping framework that augments the reward distribution with prior beliefs that decay with experience. Formally, we prove that under suitable conditions a Markov decision process augmented with our framework is consistent with the optimal policy of the original MDP when using the Q-learning algorithm. However, in general our method integrates seamlessly with any reinforcement learning algorithm that learns a value or action-value function through experience. Experiments are run on a gridworld and a more complex backgammon domain that show that we can learn tasks significantly faster when we specify intuitive priors on the reward distribution. Ofir Marom, Benjamin Rosman |
AAAI | 2 |
| 2018 | Social Cobots: Anticipatory Decision-Making for Collaborative Robots Incorporating Unexpected Human BehaviorsabstractWe propose an architecture as a robot»s decision-making mechanism to anticipate a human»s state of mind, and so plan accordingly during a human-robot collaboration task. At the core of the architecture lies a novel stochastic decision-making mechanism that implements a partially observable Markov decision process anticipating a human»s state of mind in two-stages. In the first stage it anticipates the human»s task related availability, intent (motivation), and capability during the collaboration. In the second, it further reasons about these states to anticipate the human»s true need for help. Our contribution lies in the ability of our model to handle these unexpected conditions: 1) when the human»s intention is estimated to be irrelevant to the assigned task and may be unknown to the robot, e.g., motivation is lost, another assignment is received, onset of tiredness, and 2) when the human»s intention is relevant but the human doesn»t want the robot»s assistance in the given context, e.g., because of the human»s changing emotional states or the human»s task-relevant distrust for the robot. Our results show that integrating this model into a robot»s decision-making process increases the efficiency and naturalness of the collaboration. O. Can Görür, Benjamin Rosman, Fikret Sivrikaya, Sahin Albayrak |
HRI | 2 |
| 2018 | Hierarchical Subtask Discovery with Non-Negative Matrix Factorization
Adam Christopher Earle, Andrew M. Saxe, Benjamin Rosman |
ICLR (Poster) | 3 |
| 2018 | Accelerating Model Learning with Inter-Robot Knowledge TransferabstractOnline learning of a robot's inverse dynamics model for trajectory tracking necessitates an interaction between the robot and its environment to collect training data. This is challenging for physical robots in the real world, especially for humanoids and manipulators due to their large and high dimensional state and action spaces, as a large amount of data must be collected over time. This can put the robot in danger when learning tabula rasa and can also be a time-intensive process especially in a multi-robot setting, where each robot is learning its model from scratch. We propose accelerating learning of the inverse dynamics model for trajectory tracking tasks in this multi-robot setting using knowledge transfer, where robots share and re-use data collected by preexisting robots, in order to speed up learning for new robots. We propose a scheme for collecting a sample of correspondences from the robots for training transfer models, and demonstrate, in simulations, the benefit of knowledge transfer in accelerating online learning of the inverse dynamics model between several robots, including between a low-cost Interbotix PhantomX Pincher arm, and a more expensive and relatively heavier Kuka youBot arm. We show that knowledge transfer can save up to 63% of training time of the youBot arm compared to learning from scratch, and about 58% for the lighter Pincher arm. Ndivhuwo Makondo, Benjamin Rosman, Osamu Hasegawa |
ICRA | 2 |
| 2018 | Real-Time Motion Planning in Changing Environments Using Topology-Based Encoding of Past KnowledgeabstractTrajectory planning and replanning in complex environments often reuses very little information from the previous solutions. This is particularly evident when the motion is repeated multiple times with only a limited amount of variation between each run. To address this issue, we propose the DRM-connect algorithm, a combination of dynamic reachability maps (DRM) with lazy collision checking and a fallback strategy based on the RRT-connect algorithm which is used to repair the roadmap through further exploration. This fallback allows us to use much sparser roadmaps. Furthermore, we investigate using an approximate Reeb graph to capture the topology-persistent features of the past solutions of the problem utilising this sparsity. We evaluate DRM-connect with a Reeb graph on reaching tasks, and we compare it to state-of-the-art methods. We show that the proposed method outperforms both RRT-connect and BKPIECE algorithms in the number of collision checks required and we show that our method has the potential to scale to systems with higher number degrees of freedom. Richard Fisher, Benjamin Rosman, Vladimir Ivan |
IROS | 2 |
| 2018 | Zero-Shot Transfer with Deictic Object-Oriented Representation in Reinforcement LearningabstractObject-oriented representations in reinforcement learning have shown promise in transfer learning, with previous research introducing a propositional object-oriented framework that has provably efficient learning bounds with respect to sample complexity. However, this framework has limitations in terms of the classes of tasks it can efficiently learn. In this paper we introduce a novel deictic object-oriented framework that has provably efficient learning bounds and can solve a broader range of tasks. Additionally, we show that this framework is capable of zero-shot transfer of transition dynamics across tasks and demonstrate this empirically for the Taxi and Sokoban domains. Ofir Marom, Benjamin Rosman |
NeurIPS | 2 |
| 2017 | An Analysis of Monte Carlo Tree SearchabstractMonte Carlo Tree Search (MCTS) is a family of directed search algorithms that has gained widespread attention in recent years. Despite the vast amount of research into MCTS, the effect of modifications on the algorithm, as well as the manner in which it performs in various domains, is still not yet fully known. In particular, the effect of using knowledge-heavy rollouts in MCTS still remains poorly understood, with surprising results demonstrating that better-informed rollouts often result in worse-performing agents. We present experimental evidence suggesting that, under certain smoothness conditions, uniformly random simulation policies preserve the ordering over action preferences. This explains the success of MCTS despite its common use of these rollouts to evaluate states. We further analyse non-uniformly random rollout policies and describe conditions under which they offer improved performance. Steven James 0001, George Dimitri Konidaris, Benjamin Rosman |
AAAI | 3 |
| 2017 | Fingerprint minutiae extraction using deep learningabstractThe high variability of fingerprint data (owing to, e.g., differences in quality, moisture conditions, and scanners) makes the task of minutiae extraction challenging, particularly when approached from a stance that relies on tunable algorithmic components, such as image enhancement. We pose minutiae extraction as a machine learning problem and propose a deep neural network - MENet, for Minutiae Extraction Network - to learn a data-driven representation of minutiae points. By using the existing capabilities of several minutiae extraction algorithms, we establish a voting scheme to construct training data, and so train MENet in an automated fashion on a large dataset for robustness and portability, thus eliminating the need for tedious manual data labelling. We present a post-processing procedure that determines precise minutiae locations from the output of MENet. We show that MENet performs favourably in comparisons against existing minutiae extractors. Luke Nicholas Darlow, Benjamin Rosman |
IJCB | 2 |
| 2017 | Hierarchy Through Composition with Multitask LMDPsabstractHierarchical architectures are critical to the scalability of reinforcement learning methods. Most current hierarchical frameworks execute actions serially, with macro-actions comprising sequences of primitive actions. We propose a novel alternative to these control hierarchies based on concurrent execution of many actions in parallel. Our scheme exploits the guaranteed concurrent compositionality provided by the linearly solvable Markov decision process (LMDP) framework, which naturally enables a learning agent to draw on several macro-actions simultaneously to solve new tasks. We introduce the Multitask LMDP module, which maintains a parallel distributed representation of tasks and may be stacked to form deep hierarchies abstracted in space and time. Andrew M. Saxe, Adam Christopher Earle, Benjamin Rosman |
ICML | 3 |
| 2017 | Online Constrained Model-based Reinforcement Learning
Benjamin van Niekerk, Andreas Damianou, Benjamin Rosman |
UAI | 3 |
| 2016 | Trajectory learning from human demonstrations via manifold mappingabstractThis work proposes a framework that enables arbitrary robots with unknown kinematics models to imitate human demonstrations to acquire a skill, and reproduce it in real-time. The diversity of robots active in non-laboratory environments is growing constantly, and to this end we present an approach for users to be able to easily teach a skill to a robot with any body configuration. Our proposed method requires a motion trajectory obtained from human demonstrations via a Kinect sensor, which is then projected onto a corresponding human skeleton model. The kinematics mapping between the robot and the human model is learned by employing Local Procrustes Analysis, which enables the transfer of the demonstrated trajectory from the human model to the robot. Finally, the transferred trajectory is modeled using Dynamic Movement Primitives, allowing it to be reproduced in real time. Experiments in simulation on a 4 degree of freedom robot show that our method is able to correctly imitate various skills demonstrated by a human. Michihisa Hiratsuka, Ndivhuwo Makondo, Benjamin Rosman, Osamu Hasegawa |
IROS | 3 |
| 2016 | Bayesian policy reuse
Benjamin Rosman, Majd Hawasly, Subramanian Ramamoorthy |
Mach. Learn. | 1 |
| 2015 | Nonparametric Bayesian reward segmentation for skill discovery using inverse reinforcement learningabstractWe present a method for segmenting a set of unstructured demonstration trajectories to discover reusable skills using inverse reinforcement learning (IRL). Each skill is characterised by a latent reward function which the demonstrator is assumed to be optimizing. The skill boundaries and the number of skills making up each demonstration are unknown. We use a Bayesian nonparametric approach to propose skill segmentations and maximum entropy inverse reinforcement learning to infer reward functions from the segments. This method produces a set of Markov Decision Processes (MDPs) that best describe the input trajectories. We evaluate this approach in a car driving domain and a simulated quadcopter obstacle course, showing that it is able to recover demonstrated skills more effectively than existing methods. Pravesh Ranchod, Benjamin Rosman, George Dimitri Konidaris |
IROS | 2 |
| 2014 | Giving advice to agents with hidden goalsabstractThis paper considers the problem of providing advice to an autonomous agent when neither the behavioural policy nor the goals of that agent are known to the advisor. We present an approach based on building a model of “common sense” behaviour in the domain, from an aggregation of different users performing various tasks, modelled as MDPs, in the same domain. From this model, we estimate the normalcy of the trajectory given by a new agent in the domain, and provide behavioural advice based on an approximation of the trade-off in utility between potential benefits to the exploring agent and the costs incurred in giving this advice. This model is evaluated on a maze world domain by providing advice to different types of agents, and we show that this leads to a considerable and unanimous improvement in the completion rate of their tasks. Benjamin Rosman, Subramanian Ramamoorthy |
ICRA | 1 |
| 2014 | On user behaviour adaptation under interface changeabstractDifferent interfaces allow a user to achieve the same end goal through different action sequences, e.g., command lines vs. drop down menus. Interface efficiency can be described in terms of a cost incurred, e.g., time taken, by the user in typical tasks. Realistic users arrive at evaluations of efficiency, hence making choices about which interface to use, over time, based on trial and error experience. Their choices are also determined by prior experience, which determines how much learning time is required. These factors have substantial effect on the adoption of new interfaces. In this paper, we aim at understanding how users adapt under interface change, how much time it takes them to learn to interact optimally with an interface, and how this learning could be expedited through intermediate interfaces. We present results from a series of experiments that make four main points: (a) different interfaces for accomplishing the same task can elicit significant variability in performance, (b) switching interfaces can result in adverse sharp shifts in performance, (c) subject to some variability, there are individual thresholds on tolerance to this kind of performance degradation with an interface, causing users to potentially abandon what may be a pretty good interface, and (d) our main result -- shaping user learning through the presentation of intermediate interfaces can mitigate the adverse shifts in performance while still enabling the eventual improved performance with the complex interface upon the user becoming suitably accustomed. In our experiments, human users use keyboard based interfaces to navigate a simulated ball through a maze. Our results are a first step towards interface adaptation algorithms that architect choice to accommodate personality traits of realistic users. Benjamin Rosman, Subramanian Ramamoorthy, M. M. Hassan Mahmud, Pushmeet Kohli |
IUI | 1 |
| 2010 | A game-theoretic procedure for learning hierarchically structured strategiesabstractThis paper addresses the problem of acquiring a hierarchically structured robotic skill in a nonstationary environment. This is achieved through a combination of learning primitive strategies from observation of an expert, and autonomously synthesising composite strategies from that basis. Both aspects of this problem are approached from a game theoretic viewpoint, building on prior work in the area of multiplicative weights learning algorithms. The utility of this procedure is demonstrated through simulation experiments motivated by the problem of autonomous driving. We show that this procedure allows the agent to come to terms with two forms of uncertainty in the world - continually varying goals (due to oncoming traffic) and nonstationarity of optimisation criteria (e.g., driven by changing navigability of the road). We argue that this type of factored task specification and learning is a necessary ingredient for robust autonomous behaviour in a “large-world” setting. Benjamin Rosman, Subramanian Ramamoorthy |
ICRA | 1 |
| 2006 | Language performance at high school and success in first year computer scienceabstractWe describe the first part of a study investigating the usefulness of high school language results as a predictor of success in first year computer science courses at a university where students have widely varying English language skills. Our results indicate that contrary to the generally accepted view that achievement in high school mathematics courses is the best individual predictor of success in undergraduate computer science, success in English at the first-language level in high school correlates better with actual performance. We discuss the implications of this for universities whose medium of teaching is English, operating in social contexts where many students are not native English speakers. Sarah Rauchas, Benjamin Rosman, George Dimitri Konidaris, Ian Douglas Sanders |
SIGCSE | 2 |