Peter Vamplew 0001

dblp:v/PeterVamplew · DBLP profile ↗
← Back
43ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-8687-4424ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 8 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 4Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 ES-C51: Expected sarsa based C51 distributional reinforcement learning algorithm
Rijul Tandon, Peter Vamplew 0001, Cameron Foale
Neural Networks2
2025 On Generalization Across Environments In Multi-Objective Reinforcement Learning
abstract
Real-world sequential decision-making tasks often require balancing trade-offs between multiple conflicting objectives, making Multi-Objective Reinforcement Learning (MORL) an increasingly prominent field of research. Despite recent advances, existing MORL literature has narrowly focused on performance within static environments, neglecting the importance of generalizing across diverse settings. Conversely, existing research on generalization in RL has always assumed scalar rewards, overlooking the inherent multi-objectivity of real-world problems. Generalization in the multi-objective context is fundamentally more challenging, as it requires learning a Pareto set of policies addressing varying preferences across multiple objectives. In this paper, we formalize the concept of generalization in MORL and how it can be evaluated. We then contribute a novel benchmark featuring diverse multi-objective domains with parameterized environment configurations to facilitate future studies in this area. Our baseline evaluations of state-of-the-art MORL algorithms on this benchmark reveals limited generalization capabilities, suggesting significant room for improvement. Our empirical findings also expose limitations in the expressivity of scalar rewards, emphasizing the need for multi-objective specifications to achieve effective generalization. We further analyzed the algorithmic complexities within current MORL approaches that could impede the transfer in performance from the single- to multiple-environment settings. This work fills a critical gap and lays the groundwork for future research that brings together two key areas in reinforcement learning: solving multi-objective decision-making problems and generalizing across diverse environments. We make our code available at [https://github.com/JaydenTeoh/MORL-Generalization](https://github.com/JaydenTeoh/MORL-Generalization).
Jayden Teoh, Pradeep Varakantham, Peter Vamplew 0001
ICLR3
2024 Position: Intent-aligned AI Systems Must Optimize for Agency Preservation
abstract
A central approach to AI-safety research has been to generate aligned AI systems: i.e. systems that do not deceive users and yield actions or recommendations that humans might judge as consistent with their intentions and goals. Here we argue that truthful AIs aligned solely to human intent are insufficient and that preservation of long-term agency of humans may be a more robust standard that may need to be separated and explicitly optimized for. We discuss the science of intent and control and how human intent can be manipulated and we provide a formal definition of agency-preserving AI-human interactions focusing on forward-looking explicit agency evaluations. Our work points to a novel pathway for human harm in AI-human interactions and proposes solutions to this challenge.
Catalin Mitelut, Benjamin J. Smith, Peter Vamplew 0001
ICML3
2024 Elastic step DQN: A novel multi-step algorithm to alleviate overestimation in Deep Q-Networks
abstract
Deep Q-Networks algorithm (DQN) was the first reinforcement learning algorithm using deep neural network to successfully surpass human level performance in a number of Atari learning environments. However, divergent and unstable behaviour have been long standing issues in DQNs. The unstable behaviour is often characterised by overestimation in the Q-values, commonly referred to as the overestimation bias. To address the overestimation bias and the divergent behaviour, a number of heuristic extensions have been proposed. Notably, multi-step updates have been shown to drastically reduce unstable behaviour while improving agent’s training performance. However, agents are often highly sensitive to the selection of the multi-step update horizon (n), and our empirical experiments show that a poorly chosen static value for n can in many cases lead to worse performance than single-step DQN. Inspired by the success of n-step DQN and the effects that multi-step updates have on overestimation bias, this paper proposes a new algorithm that we call ‘Elastic Step DQN’ (ES-DQN) to alleviate overestimation bias in DQNs. ES-DQN dynamically varies the step size horizon in multi-step updates based on the similarity between states visited. Our empirical evaluation shows that ES-DQN out-performs n-step with fixed n updates, Double DQN and Average DQN in several OpenAI Gym environments while at the same time alleviating the overestimation bias.
Adrian Ly, Richard Dazeley, Peter Vamplew 0001, Francisco Cruz 0002, Sunil Aryal
Neurocomputing3
2023 Elastic step DDPG: Multi-step reinforcement learning for improved sample efficiency
abstract
A major challenge in deep reinforcement learning is that it requires more data to converge to an policy for complex problems. One way to improve sample efficiency is to use n-step updates to reduce the number of samples required to converge to a good policy. However n-step updates are known to be brittle and difficult to tune. Elastic Step DQN has shown that it is possible to automate the value of$n$in DQN to solve problems involving discrete action spaces, however the efficacy of the technique when applied on more complex problems and against problems with continuous action spaces is yet to be shown. In this paper we adapt the innovations proposed by Elastic Step DQN onto the DDPG algorithm and show empirically that Elastic Step DDPG is able to achieve a much stronger final training policy and is more sample efficient than DDPG.
Adrian Ly, Richard Dazeley, Peter Vamplew 0001, Francisco Cruz 0002, Sunil Aryal
IJCNN3
2023 Human engagement providing evaluative and informative advice for interactive reinforcement learning
abstract
Abstract Interactive reinforcement learning proposes the use of externally sourced information in order to speed up the learning process. When interacting with a learner agent, humans may provide either evaluative or informative advice. Prior research has focused on the effect of human-sourced advice by including real-time feedback on the interactive reinforcement learning process, specifically aiming to improve the learning speed of the agent, while minimising the time demands on the human. This work focuses on answering which of two approaches, evaluative or informative, is the preferred instructional approach for humans. Moreover, this work presents an experimental setup for a human trial designed to compare the methods people use to deliver advice in terms of human engagement. The results obtained show that users giving informative advice to the learner agents provide more accurate advice, are willing to assist the learner agent for a longer time, and provide more advice per episode. Additionally, self-evaluation from participants using the informative approach has indicated that the agent’s ability to follow the advice is higher, and therefore, they feel their own advice to be of higher accuracy when compared to people providing evaluative advice.
Adam Bignold, Francisco Cruz 0002, Richard Dazeley, Peter Vamplew 0001, Cameron Foale
Neural Comput. Appl.4
2023 Persistent rule-based interactive reinforcement learning
Adam Bignold, Francisco Cruz 0002, Richard Dazeley, Peter Vamplew 0001, Cameron Foale
Neural Comput. Appl.4
2023 Explainable robotic systems: understanding goal-driven actions in a reinforcement learning scenario
Francisco Cruz 0002, Richard Dazeley, Peter Vamplew 0001, Ithan Moreira
Neural Comput. Appl.3
2023 Explainable reinforcement learning for broad-XAI: a conceptual framework and survey
abstract
Abstract Broad-XAI moves away from interpreting individual decisions based on a single datum and aims to provide integrated explanations from multiple machine learning algorithms into a coherent explanation of an agent’s behaviour that is aligned to the communication needs of the explainee. Reinforcement Learning (RL) methods, we propose, provide a potential backbone for the cognitive model required for the development of Broad-XAI. RL represents a suite of approaches that have had increasing success in solving a range of sequential decision-making problems. However, these algorithms operate as black-box problem solvers, where they obfuscate their decision-making policy through a complex array of values and functions. EXplainable RL (XRL) aims to develop techniques to extract concepts from the agent’s: perception of the environment; intrinsic/extrinsic motivations/beliefs; Q-values, goals and objectives. This paper aims to introduce the Causal XRL Framework (CXF), that unifies the current XRL research and uses RL as a backbone to the development of Broad-XAI. CXF is designed to incorporate many standard RL extensions and integrated with external ontologies and communication facilities so that the agent can answer questions that explain outcomes its decisions. This paper aims to: establish XRL as a distinct branch of XAI; introduce a conceptual framework for XRL; review existing approaches explaining agent behaviour; and identify opportunities for future research. Finally, this paper discusses how additional information can be extracted and ultimately integrated into models of communication, facilitating the development of Broad-XAI.
Richard Dazeley, Peter Vamplew 0001, Francisco Cruz 0002
Neural Comput. Appl.2
2023 AI apology: interactive multi-objective reinforcement learning for human-aligned AI
abstract
Abstract For an Artificially Intelligent (AI) system to maintain alignment between human desires and its behaviour, it is important that the AI account for human preferences. This paper proposes and empirically evaluates the first approach to aligning agent behaviour to human preference via an apologetic framework. In practice, an apology may consist of an acknowledgement, an explanation and an intention for the improvement of future behaviour. We propose that such an apology, provided in response to recognition of undesirable behaviour, is one way in which an AI agent may both be transparent and trustworthy to a human user. Furthermore, that behavioural adaptation as part of apology is a viable approach to correct against undesirable behaviours. The Act-Assess-Apologise framework potentially could address both the practical and social needs of a human user, to recognise and make reparations against prior undesirable behaviour and adjust for the future. Applied to a dual-auxiliary impact minimisation problem, the apologetic agent had a near perfect determination and apology provision accuracy in several non-trivial configurations. The agent subsequently demonstrated behaviour alignment with success that included up to complete avoidance of the impacts described by these objectives in some scenarios.
Hadassah Harland, Richard Dazeley, Bahareh Nakisa, Francisco Cruz 0002, Peter Vamplew 0001
Neural Comput. Appl.5
2022 Evaluating Human-like Explanations for Robot Actions in Reinforcement Learning Scenarios
abstract
Explainable artificial intelligence is a research field that tries to provide more transparency for autonomous intelligent systems. Explainability has been used, particularly in reinforcement learning and robotic scenarios, to better understand the robot decision-making process. Previous work, however, has been widely focused on providing technical explanations that can be better understood by AI practitioners than non-expert end-users. In this work, we make use of human-like explanations built from the probability of success to complete the goal that an autonomous robot shows after performing an action. These explanations are intended to be understood by people who have no or very little experience with artificial intelligence methods. This paper presents a user trial to study whether these explanations that focus on the probability an action has of succeeding in its goal constitute a suitable explanation for non-expert end-users. The results obtained show that non-expert participants rate robot explanations that focus on the probability of success higher and with less variance than technical explanations generated from Q-values, and also favor counterfactual explanations over standalone explanations.
Francisco Cruz 0002, Charlotte Young, Richard Dazeley, Peter Vamplew 0001
IROS4
2022 A practical guide to multi-objective reinforcement learning and planning
abstract
Abstract Real-world sequential decision-making tasks are generally complex, requiring trade-offs between multiple, often conflicting, objectives. Despite this, the majority of research in reinforcement learning and decision-theoretic planning either assumes only a single objective, or that multiple objectives can be adequately handled via a simple linear combination. Such approaches may oversimplify the underlying problem and hence produce suboptimal results. This paper serves as a guide to the application of multi-objective methods to difficult problems, and is aimed at researchers who are already familiar with single-objective reinforcement learning and planning methods who wish to adopt a multi-objective perspective on their research, as well as practitioners who encounter multi-objective decision problems in practice. It identifies the factors that may influence the nature of the desired solution, and illustrates by example how these influence the design of multi-objective decision-making systems for complex problems.
Conor F. Hayes, Roxana Radulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M. Zintgraf, Richard Dazeley, Fredrik Heintz, Enda Howley, Athirai Aravazhi Irissappane, Patrick Mannion, Ann Nowé, Gabriel de Oliveira Ramos, Marcello Restelli, Peter Vamplew 0001, Diederik M. Roijers
Auton. Agents Multi Agent Syst.17
2022 Scalar reward is not enough: a response to Silver, Singh, Precup and Sutton (2021)
abstract
Abstract The recent paper “Reward is Enough” by Silver, Singh, Precup and Sutton posits that the concept of reward maximisation is sufficient to underpin all intelligence, both natural and artificial, and provides a suitable basis for the creation of artificial general intelligence. We contest the underlying assumption of Silver et al. that such reward can be scalar-valued. In this paper we explain why scalar rewards are insufficient to account for some aspects of both biological and computational intelligence, and argue in favour of explicitly multi-objective models of reward maximisation. Furthermore, we contend that even if scalar reward functions can trigger intelligent behaviour in specific cases, this type of reward is insufficient for the development of human-aligned artificial general intelligence due to unacceptable risks of unsafe or unethical behaviour.
Peter Vamplew 0001, Benjamin J. Smith, Johan Källström, Gabriel de Oliveira Ramos, Roxana Radulescu, Diederik M. Roijers, Conor F. Hayes, Fredrik Heintz, Patrick Mannion, Pieter Libin, Richard Dazeley, Cameron Foale
Auton. Agents Multi Agent Syst.1
2022 Discrete-to-deep reinforcement learning methods
Budi Kurniawan, Peter Vamplew 0001, Michael Papasimeon, Richard Dazeley, Cameron Foale
Neural Comput. Appl.2
2022 The impact of environmental stochasticity on value-based multiobjective reinforcement learning
Peter Vamplew 0001, Cameron Foale, Richard Dazeley
Neural Comput. Appl.1
2021 Language Representations for Generalization in Reinforcement Learning
Goodger Nikolaj, Peter Vamplew 0001, Cameron Foale, Richard Dazeley
ACML2
2021 Levels of explainable artificial intelligence for human-aligned conversational explanations
Richard Dazeley, Peter Vamplew 0001, Cameron Foale, Charlotte Young, Sunil Aryal, Francisco Cruz 0002
Artif. Intell.2
2021 Potential-based multiobjective reinforcement learning approaches to low-impact agents for AI safety
Peter Vamplew 0001, Cameron Foale, Richard Dazeley, Adam Bignold
Eng. Appl. Artif. Intell.1
2021 A Prioritized objective actor-critic method for deep reinforcement learning
Ngoc Duy Nguyen, Thanh Thi Nguyen 0001, Peter Vamplew 0001, Richard Dazeley, Saeid Nahavandi
Neural Comput. Appl.3
2020 API Based Discrimination of Ransomware and Benign Cryptographic Programs
Paul Black, Ammar Sohail, Iqbal Gondal, Joarder Kamruzzaman, Peter Vamplew 0001, Paul A. Watters
ICONIP (2)5
2020 Identifying Cross-Version Function Similarity Using Contextual Features
abstract
The identification of similar functions in malware assists analysis by supporting the exclusion of functions that have been previously analysed, allows the identification of new variants, supports authorship attribution, and the analysis of malware phylogeny. A function's context is a set comprising the function itself and all the program functions that may be executed when this function is called. Contextual features consist of data that is extracted from the functions contained in the function context. This paper presents a novel technique called Cross Version Contextual Function Similarity (CVCFS) to identify function pairs in two programs using features based on both individual functions and function context. The CVCFS technique uses Support Vector Machine (SVM) machine learning of function similarity features to pre-filter function pairs and then applies an edit distance technique using function semantics to reduce false positives. A case study is provided where individual and contextual features are extracted from three versions of Zeus malware. The SVM pre-filtering, followed by the use of an edit distance technique to filter false positives, gives a function pair identification accuracy of 85 percent.
Paul Black, Iqbal Gondal, Peter Vamplew 0001, Arun Lakhotia
TrustCom3
2020 A multi-objective deep reinforcement learning framework
Thanh Thi Nguyen 0001, Ngoc Duy Nguyen, Peter Vamplew 0001, Saeid Nahavandi, Richard Dazeley, Chee Peng Lim
Eng. Appl. Artif. Intell.3
2019 Enhancing Model Performance for Fraud Detection by Feature Engineering and Compact Unified Expressions
Ikram Ul Haq, Iqbal Gondal, Peter Vamplew 0001
ICA3PP (2)3
2019 Survey of intrusion detection systems: techniques, datasets and challenges
abstract
Cyber-attacks are becoming more sophisticated and thereby presenting increasing challenges in accurately detecting intrusions. Failure to prevent the intrusions could degrade the credibility of security services, e.g. data confidentiality, integrity, and availability. Numerous intrusion detection methods have been proposed in the literature to tackle computer security threats, which can be broadly classified into Signature-based Intrusion Detection Systems (SIDS) and Anomaly-based Intrusion Detection Systems (AIDS). This survey paper presents a taxonomy of contemporary IDS, a comprehensive review of notable recent works, and an overview of the datasets commonly used for evaluation purposes. It also presents evasion techniques used by attackers to avoid detection and discusses future research challenges to counter such techniques so as to make computer systems more secure.
Ansam Khraisat, Iqbal Gondal, Peter Vamplew 0001, Joarder Kamruzzaman
Cybersecur.3
2018 Non-functional regression: A new challenge for neural networks
Peter Vamplew 0001, Richard Dazeley, Cameron Foale, Tanveer A. Choudhury
Neurocomputing1
2017 Evaluating Accuracy in Prudence Analysis for Cyber Security
Omaru Maruatona, Peter Vamplew 0001, Richard Dazeley, Paul A. Watters
ICONIP (5)2
2017 A taxonomy of griefer type by motivation in massively multiplayer online role-playing games
abstract
There is an anti-social phenomenon known as griefing that occurs in online games. Griefing refers to the act of one player intentionally disrupting another player’s game experience for personal pleasure and possibly potential gain. Achterbosch [2015. “Causes, Magnitude and Implications of Griefing in Massively Multiplayer Online Role-Playing Games.” PhD thesis, Faculty of Science and Technology, Federation University Australia] carried out a substantial two-phase mixed method investigation into the behaviour and experiences of both griefers and griefed players in massively multiplayer online role-playing games. The first phase consisted of a survey that attracted 1188 participants of a representative player population. The second phase consisted of interviews with 15 participants to expand the findings with more personalised data. The data were analysed from the perspectives of different demographics and different associations to griefing. One of the most unique findings is the factors that motivated a player to cause grief to another player. This paper analyses these factors to propose a taxonomy of ‘Griefer’ types (griefer being the individual who imposes upon others). The taxonomy consisted of eight types of griefers, based on their motivation for griefing. Some types related to previous studies, although new types of griefers were discovered such as the retaliator and elitist and these are discussed in detail in the article.
Leigh Achterbosch, Charlynn Miller, Peter Vamplew 0001
Behav. Inf. Technol.3
2017 Special issue on multi-objective reinforcement learning
Madalina M. Drugan, Marco A. Wiering, Peter Vamplew 0001, Madhu Chetty
Neurocomputing3
2017 Softmax exploration strategies for multiobjective reinforcement learning
Peter Vamplew 0001, Richard Dazeley, Cameron Foale
Neurocomputing1
2017 Steering approaches to Pareto-optimal multiobjective reinforcement learning
Peter Vamplew 0001, Rustam Issabekov, Richard Dazeley, Cameron Foale, Adam Berry, Tim Moore, Douglas C. Creighton
Neurocomputing1
2015 Patient admission prediction using a pruned fuzzy min-max neural network with rule extraction
Jin Wang 0002, Chee Peng Lim, Douglas C. Creighton, Abbas Khosravi, Saeid Nahavandi, Julien Ugon, Peter Vamplew 0001, Andrew Stranieri, Anton Freischmidt
Neural Comput. Appl.7
2013 A Survey of Multi-Objective Sequential Decision-Making
abstract
Sequential decision-making problems with multiple objectives arise naturally in practice and pose unique challenges for research in decision-theoretic planning and learning, which has largely focused on single-objective settings. This article surveys algorithms designed for sequential decision-making problems with multiple objectives. Though there is a growing body of literature on this subject, little of it makes explicit under what circumstances special methods are needed to solve multi-objective problems. Therefore, we identify three distinct scenarios in which converting such a problem to a single-objective one is impossible, infeasible, or undesirable. Furthermore, we propose a taxonomy that classifies multi-objective methods according to the applicable scenario, the nature of the scalarization function (which projects multi-objective values to scalar ones), and the type of policies considered. We show how these factors determine the nature of an optimal solution, which can be a single policy, a convex hull, or a Pareto front. Using this taxonomy, we survey the literature on multi-objective methods for planning and learning. Finally, we discuss key applications of such methods and outline opportunities for future work.
Diederik M. Roijers, Peter Vamplew 0001, Shimon Whiteson, Richard Dazeley
J. Artif. Intell. Res.2
2012 RM and RDM, a Preliminary Evaluation of Two Prudent RDR Techniques
Omaru Maruatona, Peter Vamplew 0001, Richard Dazeley
PKAW2
2012 Using psycholinguistic features for profiling first language of authors
abstract
This study empirically evaluates the effectiveness of different feature types for the classification of the first language of an author. In particular, it examines the utility of psycholinguistic features, extracted by the Linguistic Inquiry and Word Count (LIWC) tool, that have not previously been applied to the task of author profiling. As LIWC is a tool that has been developed in the psycholinguistic field rather than the computational linguistics field, it was hypothesized that it would be effective, both as a single type feature set because of its psycholinguistic basis, and in combination with other feature sets, because it should be sufficiently different to add insight rather than redundancy. It was found that LIWC features were competitive with previously used feature types in identifying the first language of an author, and that combined feature sets including LIWC features consistently showed better accuracy rates and average F measures than were achieved by the same feature sets without the LIWC features. As a secondary issue, this study also examined how effectively first language classification scaled up to a larger number of possible languages. It was found that the classification scheme scaled up effectively to the entire 16 language collection from the International Corpus of Learner English, when compared with results achieved on just 5 languages in previous research.
Rosemary Torney, Peter Vamplew 0001, John Yearwood
J. Assoc. Inf. Sci. Technol.2
2011 Reinforcement Learning Approach to AIBO Robot's Decision Making Process in Robosoccer's Goal Keeper Problem
abstract
Robocup is a popular test bed for AI programs around the world. Robosoccer is one of the two major parts of Robocup, in which AIBO entertainment robots take part in the middle sized soccer event. The three key challenges that robots need to face in this event are manoeuvrability, image recognition and decision making skills. This paper focuses on the decision making problem in Robosoccer -- The goal keeper problem. We investigate whether reinforcement learning (RL) as a form of semi-supervised learning can effectively contribute to the goal keeper's decision making process when penalty shot and two attacker problem are considered. Currently, the decision making process in Robosoccer is carried out using rule-base system. RL also is used for quadruped locomotion and navigation purpose in Robosoccer using AIBO. In this paper, we propose a reinforcement learning based approach that uses a dynamic state-action mapping using back propagation of reward and space quantized Q-learning (SQQL) for the choice of high level functions in order to save the goal. The novelty of our approach is that the agent learns while playing and can take independent decision which overcomes the limitations of rule-base system due to fixed and limited predefined decision rules. Performance of the proposed method has been verified against the bench mark data set made with Upenn'03 code logic. It was found that the efficiency of our SQQL approach in goalkeeping was better than the rule based approach. The SQQL develops a semi-supervised learning process over the rule-base system's input-output mapping process, given in the Upenn'03 code.
Subhasis Mukherjee, John Yearwood, Peter Vamplew 0001, Md. Shamsul Huda
SNPD3
2011 Empirical evaluation methods for multiobjective reinforcement learning algorithms
Peter Vamplew 0001, Richard Dazeley, Adam Berry, Rustam Issabekov, Evan Dekker
Mach. Learn.1
2010 The Ballarat Incremental Knowledge Engine
Richard Dazeley, Philip Warner, Peter Vamplew 0001
PKAW4
2010 Automated opinion detection: Implications of the level of agreement between human raters
Deanna J. Osman, John Yearwood, Peter Vamplew 0001
Inf. Process. Manag.3
2006 An efficient approach to unbounded bi-objective archives -: introducing the mak_tree algorithm
abstract
Given the prominence of elite archiving in contemporary multiobjective optimisation research and the limitations inherent in bounded population sizes, it is unusual that the vast majority of popular techniques aggressively truncate the capacity of archives and are based upon inefficient list representations. By forming better data structures and algorithms for the storage of archival members, the need for truncation is reduced and unbounded elite sets become viable. While work does exist in this vein, it is always of a general nature and significant improvements can be made in the bi-objective case. As such, this paper elucidates the unique properties of two-dimensional non-dominated sets and capitalises on these notions to develop the highly efficient and specialised bi-objective Mak_Tree algorithm. Theoretical results indicate that the specialised approach is preferable to pre-existing general techniques, while empirical analysis illustrates improved performance over both unbounded and bounded list techniques.
Adam Berry, Peter Vamplew 0001
GECCO2
2005 The Combative Accretion Model - Multiobjective Optimisation Without Explicit Pareto Ranking
Adam Berry, Peter Vamplew 0001
EMO2
2005 On-Line Reinforcement Learning Using Cascade Constructive Neural Networks
Peter Vamplew 0001, Robert Ollington
KES (3)1
2005 Concurrent Q-learning: Reinforcement learning for dynamic goals and environments
abstract
This article presents a powerful new algorithm for reinforcement learning in problems where the goals and also the environment may change. The algorithm is completely goal independent, allowing the mechanics of the environment to be learned independently of the task that is being undertaken. Conventional reinforcement learning techniques, such as Q-learning, are goal dependent. When the goal or reward conditions change, previous learning interferes with the new task that is being learned, resulting in very poor performance. Previously, the Concurrent Q-Learning algorithm was developed, based on Watkin's Q-learning, which learns the relative proximity of all states simultaneously. This learning is completely independent of the reward experienced at those states and, through a simple action selection strategy, may be applied to any given reward structure. Here it is shown that the extra information obtained may be used to replace the eligibility traces of Watkin's Q-learning, allowing many more value updates to be made at each time step. The new algorithm is compared to the previous version and also to DG-learning in tasks involving changing goals and environments. The new algorithm is shown to perform significantly better than these alternatives, especially in situations involving novel obstructions. The algorithm adapts quickly and intelligently to changes in both the environment and reward structure, and does not suffer interference from training undertaken prior to those changes. © 2005 Wiley Periodicals, Inc. Int J Int Syst 20: 1037–1052, 2005.
Robert Ollington, Peter Vamplew 0001
Int. J. Intell. Syst.2
2003 A simplified artificial life model for multiobjective optimisation: a preliminary report
abstract
Recent research in the field of multiobjective optimisation (MOO) has been focused on achieving the Pareto optimal front by explicitly analysing the dominance level of individual solutions. While such approaches have produced good results for a variety of problems, they are computationally expensive due to the complexities of deriving the dominance level for each solution against the entire population. TB/spl I.bar/MOO (threshold based multiobjective optimisation) is a new artificial life approach to MOO problems that does not analyse dominance, nor perform any agent-agent comparisons. This reduction in complexity results in a significant decrease in processing overhead. Results show that TB/spl I.bar/MOO performs comparably, and often better, than its more complicated counter-parts with respect to distance from the Pareto optimal front, but is slightly weaker in terms of distribution and extent.
Adam Berry, Peter Vamplew 0001
IEEE Congress on Evolutionary Computation2