Trevor McInroe

dblp:304/2817 · also Trevor A. McInroe · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 39% Robot manipulation · 33% Representation and self-supervised learning · 19%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
1.322023
Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023
Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning · ICLR 2023
Machine learning › Reinforcement learning
actor-critic methods
0.912025
Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning · ICLR 2025
Robotics › Robot manipulation
contact task
0.912025
Enhancing Tactile-based Reinforcement Learning for Robotic Control · NeurIPS 2025
Robotics › Robot manipulation
dexterous manipulation
0.912025
Enhancing Tactile-based Reinforcement Learning for Robotic Control · NeurIPS 2025
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation
0.912025
Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning · ICLR 2025
Robotics › Robot manipulation
tactile sensing
0.912025
Enhancing Tactile-based Reinforcement Learning for Robotic Control · NeurIPS 2025
Machine learning › Learning theory
generalization
0.712023
Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
generalization in reinforcement learning
0.712023
Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning · ICLR 2023
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.712023
Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023
Machine learning › Representation and self-supervised learning
mutual information minimization
0.212023
Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.5self-supervised learning · 0.9on-policy algorithms · 0.9empirical study · 0.9conditional mutual information · 0.7auxiliary task learning · 0.7
YearPublicationVenuePosition
2025 LLM-Personalize: Aligning LLM Planners with Human Preferences via Reinforced Self-Training for Housekeeping Robots
abstract
Large language models (LLMs) have shown significant potential for robotics applications, particularly task planning, by harnessing their language comprehension and text generation capabilities. However, in applications such as household robotics, a critical gap remains in the personalization of these models to household preferences. For example, an LLM planner may find it challenging to perform tasks that require personalization, such as deciding where to place mugs in a kitchen based on specific household preferences. We introduce LLM-Personalize, a novel framework designed to personalize LLM planners for household robotics. LLM-Personalize uses an LLM planner to perform iterative planning in multi-room, partially-observable household environments, utilizing a scene graph built dynamically from local observations. To personalize the LLM planner towards user preferences, our optimization pipeline integrates imitation learning and reinforced Self-Training. We evaluate LLM-Personalize on Housekeep, a challenging simulated real-world 3D benchmark for household rearrangements, demonstrating a more than 30 percent increase in success rate over existing LLM planners, showcasing significantly improved alignment with human preferences.
Dongge Han, Trevor McInroe, Adam Jelley, Stefano V. Albrecht, Peter Bell 0001, Amos J. Storkey
COLING2
2025 Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
abstract
Extracting relevant information from a stream of high-dimensional observations is a central challenge for deep reinforcement learning agents. Actor-critic algorithms add further complexity to this challenge, as it is often unclear whether the same information will be relevant to both the actor and the critic. To this end, we here explore the principles that underlie effective representations for the actor and for the critic in on-policy algorithms. We focus our study on understanding whether the actor and critic will benefit from separate, rather than shared, representations. Our primary finding is that when separated, the representations for the actor and critic systematically specialise in extracting different types of information from the environment---the actor's representation tends to focus on action-relevant information, while the critic's representation specialises in encoding value and dynamics information. We conduct a rigourous empirical study to understand how different representation learning approaches affect the actor and critic's specialisations and their downstream performance, in terms of sample efficiency and generation capabilities. Finally, we discover that a separated critic plays an important role in exploration and data collection during training. Our code, trained models and data are accessible at https://github.com/francelico/deac-rep.
Samuel Garcin, Trevor McInroe, Pablo Samuel Castro, Christopher G. Lucas, David Abel, Prakash Panangaden, Stefano V. Albrecht
ICLR2
2025 Enhancing Tactile-based Reinforcement Learning for Robotic Control
abstract
Achieving safe, reliable real-world robotic manipulation requires agents to evolve beyond vision and incorporate tactile sensing to overcome sensory deficits and reliance on idealised state information. Despite its potential, the efficacy of tactile sensing in reinforcement learning (RL) remains inconsistent. We address this by developing self-supervised learning (SSL) methodologies to more effectively harness tactile observations, focusing on a scalable setup of proprioception and sparse binary contacts. We empirically demonstrate that sparse binary tactile signals are critical for dexterity, particularly for interactions that proprioceptive control errors do not register, such as decoupled robot-object motions. Our agents achieve superhuman dexterity in complex contact tasks (ball bouncing and Baoding ball rotation). Furthermore, we find that decoupling the SSL memory from the on-policy memory can improve performance. We release the Robot Tactile Olympiad ($\texttt{RoTO}$) benchmark to standardise and promote future research in tactile-based manipulation. Project page: https://elle-miller.github.io/tactile_rl.
Elle Miller, Trevor McInroe, David Abel, Oisin Mac Aodha, Sethu Vijayakumar
NeurIPS2
2023 Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning
Mhairi Dunion, Trevor McInroe, Kevin S. Luck, Josiah Hanna, Stefano V. Albrecht
ICLR2
2023 Conditional Mutual Information for Disentangled Representations in Reinforcement Learning
abstract
Reinforcement Learning (RL) environments can produce training data with spurious correlations between features due to the amount of training data or its limited feature coverage. This can lead to RL agents encoding these misleading correlations in their latent representation, preventing the agent from generalising if the correlation changes within the environment or when deployed in the real world. Disentangled representations can improve robustness, but existing disentanglement techniques that minimise mutual information between features require independent features, thus they cannot disentangle correlated features. We propose an auxiliary task for RL algorithms that learns a disentangled representation of high-dimensional observations with correlated features by minimising the conditional mutual information between features in the representation. We demonstrate experimentally, using continuous control tasks, that our approach improves generalisation under correlation shifts, as well as improving the training performance of RL algorithms in the presence of correlated features.
Mhairi Dunion, Trevor McInroe, Kevin S. Luck, Josiah Hanna, Stefano V. Albrecht
NeurIPS2