EDBT 2026 Demo / reviewers in the wild / expert
Trevor McInroe
dblp:304/2817 · also Trevor A. McInroe
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 39% Robot manipulation · 33% Representation and self-supervised learning · 19% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
1.3 | 2 | 2023 | Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023 Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning · ICLR 2023 |
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 1 | 2025 | Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning · ICLR 2025 |
Robotics › Robot manipulation
contact task |
0.9 | 1 | 2025 | Enhancing Tactile-based Reinforcement Learning for Robotic Control · NeurIPS 2025 |
Robotics › Robot manipulation
dexterous manipulation |
0.9 | 1 | 2025 | Enhancing Tactile-based Reinforcement Learning for Robotic Control · NeurIPS 2025 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation |
0.9 | 1 | 2025 | Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning · ICLR 2025 |
Robotics › Robot manipulation
tactile sensing |
0.9 | 1 | 2025 | Enhancing Tactile-based Reinforcement Learning for Robotic Control · NeurIPS 2025 |
Machine learning › Learning theory
generalization |
0.7 | 1 | 2023 | Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning
generalization in reinforcement learning |
0.7 | 1 | 2023 | Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning · ICLR 2023 |
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
0.7 | 1 | 2023 | Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
mutual information minimization |
0.2 | 1 | 2023 | Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.5self-supervised learning · 0.9on-policy algorithms · 0.9empirical study · 0.9conditional mutual information · 0.7auxiliary task learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLM-Personalize: Aligning LLM Planners with Human Preferences via Reinforced Self-Training for Housekeeping RobotsabstractLarge language models (LLMs) have shown significant potential for robotics applications, particularly task planning, by harnessing their language comprehension and text generation capabilities. However, in applications such as household robotics, a critical gap remains in the personalization of these models to household preferences. For example, an LLM planner may find it challenging to perform tasks that require personalization, such as deciding where to place mugs in a kitchen based on specific household preferences. We introduce LLM-Personalize, a novel framework designed to personalize LLM planners for household robotics. LLM-Personalize uses an LLM planner to perform iterative planning in multi-room, partially-observable household environments, utilizing a scene graph built dynamically from local observations. To personalize the LLM planner towards user preferences, our optimization pipeline integrates imitation learning and reinforced Self-Training. We evaluate LLM-Personalize on Housekeep, a challenging simulated real-world 3D benchmark for household rearrangements, demonstrating a more than 30 percent increase in success rate over existing LLM planners, showcasing significantly improved alignment with human preferences. Dongge Han, Trevor McInroe, Adam Jelley, Stefano V. Albrecht, Peter Bell 0001, Amos J. Storkey |
COLING | 2 |
| 2025 | Studying the Interplay Between the Actor and Critic Representations in Reinforcement LearningabstractExtracting relevant information from a stream of high-dimensional observations is a central challenge for deep reinforcement learning agents. Actor-critic algorithms add further complexity to this challenge, as it is often unclear whether the same information will be relevant to both the actor and the critic. To this end, we here explore the principles that underlie effective representations for the actor and for the critic in on-policy algorithms. We focus our study on understanding whether the actor and critic will benefit from separate, rather than shared, representations. Our primary finding is that when separated, the representations for the actor and critic systematically specialise in extracting different types of information from the environment---the actor's representation tends to focus on action-relevant information, while the critic's representation specialises in encoding value and dynamics information. We conduct a rigourous empirical study to understand how different representation learning approaches affect the actor and critic's specialisations and their downstream performance, in terms of sample efficiency and generation capabilities. Finally, we discover that a separated critic plays an important role in exploration and data collection during training. Our code, trained models and data are accessible at https://github.com/francelico/deac-rep. Samuel Garcin, Trevor McInroe, Pablo Samuel Castro, Christopher G. Lucas, David Abel, Prakash Panangaden, Stefano V. Albrecht |
ICLR | 2 |
| 2025 | Enhancing Tactile-based Reinforcement Learning for Robotic ControlabstractAchieving safe, reliable real-world robotic manipulation requires agents to evolve beyond vision and incorporate tactile sensing to overcome sensory deficits and reliance on idealised state information. Despite its potential, the efficacy of tactile sensing in reinforcement learning (RL) remains inconsistent. We address this by developing self-supervised learning (SSL) methodologies to more effectively harness tactile observations, focusing on a scalable setup of proprioception and sparse binary contacts. We empirically demonstrate that sparse binary tactile signals are critical for dexterity, particularly for interactions that proprioceptive control errors do not register, such as decoupled robot-object motions. Our agents achieve superhuman dexterity in complex contact tasks (ball bouncing and Baoding ball rotation). Furthermore, we find that decoupling the SSL memory from the on-policy memory can improve performance. We release the Robot Tactile Olympiad ($\texttt{RoTO}$) benchmark to standardise and promote future research in tactile-based manipulation. Project page: https://elle-miller.github.io/tactile_rl. Elle Miller, Trevor McInroe, David Abel, Oisin Mac Aodha, Sethu Vijayakumar |
NeurIPS | 2 |
| 2023 | Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning
Mhairi Dunion, Trevor McInroe, Kevin S. Luck, Josiah Hanna, Stefano V. Albrecht |
ICLR | 2 |
| 2023 | Conditional Mutual Information for Disentangled Representations in Reinforcement LearningabstractReinforcement Learning (RL) environments can produce training data with spurious correlations between features due to the amount of training data or its limited feature coverage. This can lead to RL agents encoding these misleading correlations in their latent representation, preventing the agent from generalising if the correlation changes within the environment or when deployed in the real world. Disentangled representations can improve robustness, but existing disentanglement techniques that minimise mutual information between features require independent features, thus they cannot disentangle correlated features. We propose an auxiliary task for RL algorithms that learns a disentangled representation of high-dimensional observations with correlated features by minimising the conditional mutual information between features in the representation. We demonstrate experimentally, using continuous control tasks, that our approach improves generalisation under correlation shifts, as well as improving the training performance of RL algorithms in the presence of correlated features. Mhairi Dunion, Trevor McInroe, Kevin S. Luck, Josiah Hanna, Stefano V. Albrecht |
NeurIPS | 2 |