EDBT 2026 Demo / reviewers in the wild / expert
Thommen George Karimpanal
dblp:133/3358 · also Thommen Karimpanal George
· DBLP profile ↗
19ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-8918-3314ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Human Feedback for Semantically-Relevant Skill Discovery
Maxence Hussonnois, Thommen George Karimpanal, Santu Rana |
ICPR (12) | 2 |
| 2026 | Improving Multilingual Language Models by Aligning Representations through SteeringabstractImproving Multilingual Language Models by Aligning Representations through Steering Omar Mahmoud 0001, Buddhika Laknath Semage, Thommen George Karimpanal, Santu Rana |
LREC | 3 |
| 2025 | MAGIK: Mapping to Analogous Goals via Imagination-Enabled Knowledge TransferabstractHumans excel at analogical reasoning - applying knowledge from one task to a related one with minimal relearning. In contrast, reinforcement learning (RL) agents typically require extensive retraining even when new tasks share structural similarities with previously learned ones. In this work, we propose MAGIK, a novel framework that enables RL agents to transfer knowledge to analogous tasks without interacting with the target environment. Our approach leverages an imagination mechanism to map entities in the target task to their analogues in the source domain, allowing the agent to reuse its original policy. Experiments on custom MiniGrid and MuJoCo tasks show that MAGIK achieves effective zero-shot transfer using only a small number of human-labelled examples. We compare our approach to related baselines and highlight how it offers a novel and effective mechanism for knowledge transfer via imagination-based analogy mapping. Ajsal Shereef Palattuparambil, Thommen George Karimpanal, Santu Rana |
ECAI | 2 |
| 2025 | Human-Aligned Skill Discovery: Balancing Behaviour Exploration and Alignment
Maxence Hussonnois, Thommen George Karimpanal, Santu Rana |
AAMAS | 2 |
| 2025 | Beyond the Known: Decision Making with Counterfactual Reasoning Decision TransformerabstractDecision Transformers (DT) play a crucial role in modern reinforcement learning, leveraging offline datasets to achieve impressive results across various domains. However, DT requires high-quality, comprehensive data to perform optimally. In real-world applications, the lack of training data and the scarcity of optimal behaviours make training on offline datasets challenging, as suboptimal data can hinder performance. To address this, we propose the Counterfactual Reasoning Decision Transformer (CRDT), a novel framework inspired by counterfactual reasoning. CRDT enhances DT’s ability to reason beyond known data by generating and utilizing counterfactual experiences, enabling improved decision-making in unseen scenarios. Experiments across Atari and D4RL benchmarks, including scenarios with limited data and altered dynamics, demonstrate that CRDT outperforms conventional DT approaches. Additionally, reasoning counterfactually allows the DT agent to obtain stitching abilities, combining suboptimal trajectories, without architectural modifications. These results highlight the potential of counterfactual reasoning to enhance reinforcement learning agents' performance and generalization capabilities. Linh Le Pham Van, Thommen George Karimpanal, Sunil Gupta 0001, Hung Le 0002 |
IJCAI | 3 |
| 2025 | Human-informed skill discovery: Controlled diversity with preference in reinforcement learningabstractAutonomously learning diverse behaviours without an extrinsic reward signal has been a problem of interest in reinforcement learning. However, the nature of learning in such mechanisms is unconstrained, often resulting in the accumulation of several unusable, unsafe or misaligned skills. In order to avoid such issues and to ensure the discovery of safe and human-aligned skills, it is necessary to incorporate humans into the unsupervised training process, which remains a largely unexplored topic. In this work, we propose Controlled Diversity with Preference (CDP), a novel, collaborative human-guided mechanism for an agent to learn a set of skills that is diverse as well as desirable. The key principle is to restrict the discovery of skills to regions that are deemed to be desirable as per a preference model trained using human preference labels on trajectory pairs. We evaluate our approach on 2D navigation and Mujoco environments and demonstrate the ability to discover diverse, yet desirable skills. We also provide principled guidelines for selecting suitable hyperparameter values along with comprehensive sensitivity analyses of the various factors influencing the performance of our approach. Maxence Hussonnois, Thommen George Karimpanal, Mayank Shekhar Jha, Santu Rana |
Expert Syst. Appl. | 2 |
| 2025 | LaGR-SEQ: Language-guided reinforcement learning with sample-efficient queryingabstractAbstract Large language models (LLMs) have recently demonstrated their impressive ability to provide context-aware responses via text. This ability could potentially be used to predict plausible solutions in sequential decision making tasks pertaining to pattern completion. For example, by observing a partial stack of cubes, LLMs can predict the correct sequence in which the remaining cubes should be stacked by extrapolating the observed patterns (e.g., cube sizes, colors or other attributes) in the partial stack. In this work, we introduce LaGR (language-guided reinforcement learning), which uses this predictive ability of LLMs to propose solutions to tasks that have been partially completed by a primary reinforcement learning (RL) agent, in order to subsequently guide the latter’s training. However, as RL training is generally not sample-efficient, deploying this approach would inherently imply that the LLM be repeatedly queried for solutions; a process that can be expensive and infeasible. To address this issue, we introduce SEQ (sample-efficient querying), where we simultaneously train a secondary RL agent to decide when the LLM should be queried for solutions. Specifically, we use the quality of the solutions emanating from the LLM as the reward to train this agent. We show that our proposed framework LaGR-SEQ enables more efficient primary RL training, while simultaneously minimizing the number of queries to the LLM. We demonstrate our approach on a series of tasks and highlight the advantages of our approach, along with its limitations and potential future research directions. Thommen George Karimpanal, Buddhika Laknath Semage, Santu Rana, Hung Le 0002, Truyen Tran 0001, Sunil Gupta 0001, Svetha Venkatesh |
Neural Comput. Appl. | 1 |
| 2024 | Personalisation via Dynamic Policy FusionabstractReward-optimal policies obtained by training deep reinforcement learning agents may not be aligned with one’s personal preferences. Rectifying this by retraining the agent with a user-specific reward function is impractical, as such functions are not readily available. In addition, retraining is associated with high costs. Instead, we propose to adapt via policy fusion, the already trained policy with the user’s intent, which we in turn infer from trajectory-level feedback. We design the policy fusion process to be dynamic, such that the resulting policy is neither dominated by the task goals nor the user needs. We empirically demonstrate that our method consistently balances these objectives across various environments. Ajsal Shereef Palattuparambil, Thommen George Karimpanal, Santu Rana |
HAI | 2 |
| 2024 | EMOTE: An Explainable Architecture for Modelling the Other through Empathy
Manisha Senadeera, Thommen George Karimpanal, Stephan Jacobs, Sunil Gupta 0001, Santu Rana |
IJCAI | 2 |
| 2023 | Balanced Q-learning: Combining the influence of optimistic and pessimistic targetsabstractThe optimistic nature of the Q−learning target leads to an overestimation bias, which is an inherent problem associated with standard Q−learning. Such a bias fails to account for the possibility of low returns, particularly in risky scenarios. However, the existence of biases, whether overestimation or underestimation, need not necessarily be undesirable. In this paper, we analytically examine the utility of biased learning, and show that specific types of biases may be preferable, depending on the scenario. Based on this finding, we design a novel reinforcement learning algorithm, Balanced Q-learning, in which the target is modified to be a convex combination of a pessimistic and an optimistic term, whose associated weights are determined online, analytically. Such a balanced target inherently promotes risk-averse behavior, which we examine through the lens of the agent's exploration. We prove the convergence of this algorithm in a tabular setting, and empirically demonstrate its consistently good learning performance in various environments. Thommen George Karimpanal, Hung Le 0002, Majid Abdolshah, Santu Rana, Sunil Gupta 0001, Truyen Tran 0001, Svetha Venkatesh |
Artif. Intell. | 1 |
| 2023 | Human-aligned reinforcement learning for autonomous agents and robots
Francisco Cruz 0002, Thommen George Karimpanal, Miguel A. Solis, Pablo V. A. Barros, Richard Dazeley |
Neural Comput. Appl. | 2 |
| 2022 | Episodic Policy Gradient TrainingabstractWe introduce a novel training procedure for policy gradient methods wherein episodic memory is used to optimize the hyperparameters of reinforcement learning algorithms on-the-fly. Unlike other hyperparameter searches, we formulate hyperparameter scheduling as a standard Markov Decision Process and use episodic memory to store the outcome of used hyperparameters and their training contexts. At any policy update step, the policy learner refers to the stored experiences, and adaptively reconfigures its learning algorithm with the new hyperparameters determined by the memory. This mechanism, dubbed as Episodic Policy Gradient Training (EPGT), enables an episodic learning process, and jointly learns the policy and the learning algorithm's hyperparameters within a single run. Experimental results on both continuous and discrete environments demonstrate the advantage of using the proposed method in boosting the performance of various policy gradient algorithms. Hung Le 0002, Majid Abdolshah, Thommen George Karimpanal, Kien Do, Dung Nguyen 0001, Svetha Venkatesh |
AAAI | 3 |
| 2022 | Fast Model-based Policy Search for Universal Policy NetworksabstractAdapting an agent’s behaviour to new environments has been one of the primary focus areas of physics based reinforcement learning. Although recent approaches such as universal policy networks partially address this issue by enabling the storage of multiple policies trained in simulation on a wide range of dynamic/latent factors, efficiently identifying the most appropriate policy for a given environment remains a challenge. In this work, we propose a Gaussian Process-based prior learned in simulation, that captures the likely performance of a policy when transferred to a previously unseen environment. We integrate this prior with a Bayesian Optimisation-based policy search process to improve the efficiency of identifying the most appropriate policy from the universal policy network. We empirically evaluate our approach in a range of continuous and discrete control environments, and show that it outperforms other competing baselines. Buddhika Laknath Semage, Thommen George Karimpanal, Santu Rana, Svetha Venkatesh |
ICPR | 2 |
| 2022 | Uncertainty Aware System Identification with Universal PoliciesabstractSim2real transfer is primarily concerned with transferring policies trained in simulation to potentially noisy real world environments. A common problem associated with sim2real transfer is estimating the real-world environmental parameters to ground the simulated environment to. Although existing methods such as Domain Randomisation (DR) can produce robust policies by sampling from a distribution of parameters during training, there is no established method for identifying the parameters of the corresponding distribution for a given real-world setting. In this work, we propose Uncertainty-aware policy search (UncAPS), where we use Universal Policy Network (UPN) to store simulation-trained task-specific policies across the full range of environmental parameters and then subsequently employ robust Bayesian optimisation to craft robust policies for the given environment by combining relevant UPN policies in a DR like fashion. Such policy-driven grounding is expected to be more efficient as it estimates only task-relevant sets of parameters. Further, we also account for the estimation uncertainties in the search process to produce policies that are robust against both aleatoric and epistemic uncertainties. We empirically evaluate our approach in a range of noisy, continuous control environments, and show its improved performance compared to competing baselines. Buddhika Laknath Semage, Thommen George Karimpanal, Santu Rana, Svetha Venkatesh |
ICPR | 2 |
| 2022 | Learning to Constrain Policy Optimization with Virtual Trust RegionabstractWe introduce a constrained optimization method for policy gradient reinforcement learning, which uses two trust regions to regulate each policy update. In addition to using the proximity of one single old policy as the first trust region as done by prior works, we propose forming a second trust region by constructing another virtual policy that represents a wide range of past policies. We then enforce the new policy to stay closer to the virtual policy, which is beneficial if the old policy performs poorly. We propose a mechanism to automatically build the virtual policy from a memory buffer of past policies, providing a new capability for dynamically selecting appropriate trust regions during the optimization process. Our proposed method, dubbed Memory-Constrained Policy Optimization (MCPO), is examined in diverse environments, including robotic locomotion control, navigation with sparse rewards and Atari games, consistently demonstrating competitive performance against recent on-policy constrained policy gradient methods. Hung Le 0002, Thommen George Karimpanal, Majid Abdolshah, Dung Nguyen 0001, Kien Do, Sunil Gupta 0001, Svetha Venkatesh |
NeurIPS | 2 |
| 2021 | A New Representation of Successor Features for Transfer across Dissimilar EnvironmentsabstractTransfer in reinforcement learning is usually achieved through generalisation across tasks. Whilst many studies have investigated transferring knowledge when the reward function changes, they have assumed that the dynamics of the environments remain consistent. Many real-world RL problems require transfer among environments with different dynamics. To address this problem, we propose an approach based on successor features in which we model successor feature functions with Gaussian Processes permitting the source successor features to be treated as noisy measurements of the target successor feature function. Our theoretical analysis proves the convergence of this approach as well as the bounded error on modelling successor feature functions with Gaussian Processes in environments with both different dynamics and rewards. We demonstrate our method on benchmark datasets and show that it outperforms current baselines. Majid Abdolshah, Hung Le 0002, Thommen George Karimpanal, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
ICML | 3 |
| 2021 | Model-Based Episodic Memory Induces Dynamic Hybrid ControlsabstractEpisodic control enables sample efficiency in reinforcement learning by recalling past experiences from an episodic memory. We propose a new model-based episodic memory of trajectories addressing current limitations of episodic control. Our memory estimates trajectory values, guiding the agent towards good policies. Built upon the memory, we construct a complementary learning model via a dynamic hybrid control unifying model-based, episodic and habitual learning into a single architecture. Experiments demonstrate that our model allows significantly faster and better learning than other strong reinforcement learning agents across a variety of environments including stochastic and non-Markovian settings. Hung Le 0002, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran 0001, Svetha Venkatesh |
NeurIPS | 2 |
| 2020 | Learning Transferable Domain Priors for Safe Exploration in Reinforcement LearningabstractPrior access to domain knowledge could significantly improve the performance of a reinforcement learning agent. In particular, it could help agents avoid potentially catastrophic exploratory actions, which would otherwise have to be experienced during learning. In this work, we identify consistently undesirable actions in a set of previously learned tasks, and use pseudo-rewards associated with them to learn a prior policy. In addition to enabling safer exploratory behaviors in subsequent tasks in the domain, we show that these priors are transferable to similar environments, and can be learned off-policy and in parallel with the learning of other tasks in the domain. We compare our approach to established, state-of-the-art algorithms in both discrete as well as continuous environments, and demonstrate that it exhibits a safer exploratory behavior while learning to perform arbitrary tasks in the domain. We also present a theoretical analysis to support these results, and briefly discuss the implications and some alternative formulations of this approach, which could also be useful in certain scenarios. Thommen George Karimpanal, Santu Rana, Sunil Gupta 0001, Truyen Tran 0001, Svetha Venkatesh |
IJCNN | 1 |
| 2017 | Identification and off-policy learning of multiple objectives using adaptive clustering
Thommen George Karimpanal, Erik Wilhelm |
Neurocomputing | 1 |