EDBT 2026 Demo / reviewers in the wild / expert
Daniel Kudenko
dblp:46/2964
· DBLP profile ↗
62ranked-venue papers
0as first author
20since 2021 · last 2026
0000-0003-3359-3255ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 since 2021Human-computer interaction and ubiquitous computing · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6Databases, data management, data science and information retrieval · 4 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards unbiased action value estimation in reinforcement learning
Yuan Xue 0005, Daniel Kudenko, Megha Khosla |
Neurocomputing | 2 |
| 2025 | Improving the Effectiveness of Potential-based Reward Shaping in Reinforcement Learning
Henrik Müller, Daniel Kudenko |
AAMAS | 2 |
| 2025 | Streamlined Integration of GR(1) Synthesis and Reinforcement Learning for Optimizing Critical Cyber-Physical Systems
Eric Wete, Joel Greenyer, Tom Yaacov, Daniel Kudenko, Wolfgang Nejdl |
MODELS | 4 |
| 2025 | Hybrid pathfinding optimization for the Lightning Network with Reinforcement Learning
Danila Valko, Daniel Kudenko |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | TCR: topologically consistent reweighting for XGBoost in regression tasksabstractAbstract Gradient boosted tree ensembles (GBTEs) such as XGBoost continue to outperform other machine learning models on tabular data. However, the plethora of adjustable hyperparameters can exacerbate optimisation, especially in regression tasks with no intuitive performance measures such as accuracy and confidence. Automated machine learning frameworks alleviate the hyperparameter search for users, but if the optimisation procedure ends prematurely due to resource constraints, it is questionable whether users receive good models. To tackle this problem, we introduce a cost-efficient method to retrofit previously optimised XGBoost models by retraining them with a new weight distribution over the training instances. We base our approach on topological results, which allows us to infer model-agnostic weights for specific regions of the data distribution where the targets are more susceptible to input perturbations. By linking our theory to the training procedure of XGBoost regressors, we then establish a topologically consistent reweighting scheme, which is independent of the specific model instance. Empirically, we verify that our approach improves prediction performance, outperforms other reweighting methods and is much faster than a hyperparameter search. To enable users to find the optimal weights for their data, we provide guides based on our findings on 20 datasets. Our code is available at: https://github.com/montymaxzuehlke/tcr . Monty-Maximilian Zühlke, Daniel Kudenko |
Mach. Learn. | 2 |
| 2025 | Using incomplete and incorrect plans to shape reinforcement learning in long-sequence sparse-reward tasksabstractAbstract Reinforcement learning (RL) agents naturally struggle with long-sequence sparse-reward tasks due to the lack of reward feedback during exploration and the problem of identifying the necessary action sequences required to reach the goal. Previous works have used abstract symbolic task knowledge models to speed up RL agents in these tasks by either splitting the task into easier to solve sub-tasks or by creating an artificial dense reward function. These approaches are often limited by their requirement of perfect symbolic knowledge models, which cannot be guaranteed when the abstract symbolic models are provided by humans and in real-world tasks. We introduce exponential plan-based reward shaping, which is able to leverage the ability to learn from experience of RL to compensate deficiencies in incomplete and incorrect abstract symbolic plans and use them to solve difficult tasks faster, while guaranteeing convergence to the optimal policy. Our approach is able to work with plans that miss important steps, include unnecessary extra steps, contain steps that refer ambiguously to both important and useless states, or encode an incorrect order of steps. We use action representations designed by human experts to automatically compute plans to capture the high-level task structure. The abstract symbolic subgoals defined by the plan are used to create dense reward feedback, which signals important states to the RL agent that should be achieved and explored to reach the goal. We show the theoretical advantages of our approach for plans with many steps and show its effectiveness empirically on multiple tasks with different kinds of incomplete or incorrect knowledge. Henrik Müller, Lukas Berg, Daniel Kudenko |
Neural Comput. Appl. | 3 |
| 2025 | Increasing energy efficiency of bitcoin infrastructure with reinforcement learning and one-shot path planning for the lightning network
Danila Valko, Daniel Kudenko |
Neural Comput. Appl. | 2 |
| 2025 | Graph learning-based generation of abstractions for reinforcement learningabstractAbstract The application of reinforcement learning (RL) algorithms is often hindered by the combinatorial explosion of the state space. Previous works have leveraged abstractions which condense large state spaces to find tractable solutions. However, they assumed that the abstractions are provided by a domain expert. In this work, we propose a new approach to automatically construct abstract Markov decision processes (AMDPs) for potential-based reward shaping to improve the sample efficiency of RL algorithms. Our approach to constructing abstract states is inspired by graph representation learning methods, it effectively encodes the topological and reward structure of the ground-level MDP. We perform large-scale quantitative experiments on a range of navigation and gathering tasks under both stationary and stochastic settings. Our approach shows improvements of up to 8.5 times in sample efficiency and up to 3 times in run time over the baseline approach. Besides, with our qualitative analyses of the generated AMDPs, we are able to visually demonstrate the capability of our approach to preserve the topological and reward structure of the ground-level MDP. Yuan Xue 0005, Daniel Kudenko, Megha Khosla |
Neural Comput. Appl. | 2 |
| 2024 | Effectively Capturing Label Correlation for Tabular Multi-Label ClassificationabstractMulti-label data is prevalent across various applications, where instances can be annotated with a set of classes. Although multi-label data can take various forms, such as images and text, tabular multi-label data stands out as the predominant data type in many real-world scenarios. Over the past decades, numerous methods have been proposed for tabular multi-label classification. Effectively addressing challenges like class imbalance, correlation among labels and features, and scalability is crucial for a high-performance multi-label classifier. However, many existing methods fall short of fully considering the correlation between labels and features. In cases where attempts are made, they often encounter high computational costs, rendering them impractical for large datasets. This paper in- troduces an innovative classification method for tabular multi-label data, utilizing a fusion of transformers and graph convolutional networks (GCN). The central concept of the proposed approach involves transforming tabular data into images, leveraging state-of-the-art methods in image processing, including image-based transformers and pre-trained models to capture correlation among labels effectively. Our approach jointly learns the representation of feature space and the correlation among labels within a unified network. To substantiate the performance of our proposed method, we conducted a rigorous series of experiments across diverse multi-label datasets1. The results underscore the superior performance and scalability of our approach compared to other existing state-of-the-art methods. This work not only contributes a novel perspective to the field of tabular multi-label classification but also showcases advancements in both accuracy and scalability. Sajjad Kamali Siahroudi, Zahra Ahmadi, Daniel Kudenko |
CIKM | 3 |
| 2024 | Entity Matching Across Small Networks Using Node AttributesabstractEntity matching, also known as user identity linkage, is a critical task in data integration. While established techniques primarily focus on large-scale networks, there are several applications where small networks pose challenges due to limited training data and sparsity. This study addresses entity matching in the field of criminology, where small networks are common and the number of known matching nodes is restricted. To support this research, we exploit a multimodal dataset, collected as part of a security-related project, consisting of an intercepted telephone calls network (i.e., ROXSD data) and a network of social forum interactions (i.e., ROXHOOD data) collected in a simulated environment, although following real investigation scenario. To improve accuracy and efficiency, we propose a novel approach for entity matching across these two small networks using node attributes. Existing techniques often merely focus on topology consistency between two networks and overlook valuable information, such as network node attributes, making them vulnerable to structural changes. Inspired by the remarkable success of deep learning, we present UGC-DeepLink, an end-to-end semi-supervised learning framework that leverages user-generated content. UGC-DeepLink encodes network nodes into vector representations, capturing both local and global network structures to align anchor nodes using deep neural networks. A dual learning paradigm and the policy gradient method transfer knowledge and update the linkage. Additionally, node attributes, such as call contents and forum exchanged texts, enhance the ranking of matching nodes. Experimental results on ROXSD and ROXHOOD demonstrate that UGC-DeepLink surpasses baselines and state-of-the-art methods in terms of identity-match ranking. The code and dataset are available at https://github.com/erichoang/UGC-DeepLink. Zahra Ahmadi, Sergio Burdisso, Srikanth R. Madikeri, Petr Motlícek, Erinç Dikici, Gerhard Backfried, Marek Kovác, Kvetoslav Malý, Daniel Kudenko |
ECAI | 11 |
| 2024 | Reducing CO2 emissions in a peer-to-peer distributed payment network: Does geography matter in the lightning network?
Danila Valko, Daniel Kudenko |
Comput. Networks | 2 |
| 2023 | MDE and Learning for flexible Planning and optimized Execution of Multi-Robot ChoreographiesabstractMulti-Robot systems in automotive are safety-critical systems that consist of collaborating-aware robots and components that interact with external components, the environment, or humans at run-time. This implies a significant complexity for the system engineer to design, model, validate the system, and optimize the cycle time, including considering unexpected events at run-time. This paper addresses this challenge by describing a model-driven engineering approach that formally designs the system under the consideration of uncertainties and at run-time optimizes the system actions using learning-based approaches. We implemented this approach in an industrial-inspired case study of a spot-welding multi-robot cell. Based on the system requirements, we generate valid system strategies that consider unexpected events such as robot interruptions and failures. Considering movement and interruption time models, we implemented a reinforcement learning method to optimize system actions at run-time. We show that via simulations and learning, our approach can be used to synthesize time-efficient schedules for robot task assignments that improve the overall cycle time. Eric Wete, Joel Greenyer, Andreas Wortmann 0001, Daniel Kudenko, Wolfgang Nejdl |
ETFA | 4 |
| 2023 | Partial Multi-label Learning via Constraint Clustering
Sajjad Kamali Siahroudi, Daniel Kudenko |
ICONIP (11) | 2 |
| 2023 | An effective single-model learning for multi-label data
Sajjad Kamali Siahroudi, Daniel Kudenko |
Expert Syst. Appl. | 2 |
| 2022 | Imitating Playstyle with Dynamic Time Warping ImitationabstractImitation learning has been demonstrated as a useful technique in automatic game testing and the development of believable Non-Player Characters (NPCs). However, imitation learning methods typically focus on learning a policy to complete a task without consideration about the playstyle used. In this work we consider the case where the task is to imitate a given playstyle. We defined a player’s playstyle based on the strategies they use in order to complete the overall task. This has been achieved by rewarding a learning agent based on the similarity of the agent and demonstration trajectories, within a learnt representation space. This allows the playstyle to be learnt in levels that differ to the one the demonstrations were collected in. Mark Ferguson, Sam Devlin, Daniel Kudenko, James Alfred Walker |
FDG | 3 |
| 2022 | Assured Multi-agent Reinforcement Learning with Robust Agent-Interaction Adaptability
Joshua Riley, Radu Calinescu, Colin Paterson, Daniel Kudenko, Alec Banks |
KES-IDT | 4 |
| 2021 | Reinforcement Learning with Quantitative Verification for Assured Multi-Agent PoliciesabstractIn multi-agent reinforcement learning, several agents converge together towards optimal policies that solve complex decision-making problems.This convergence process is inherently stochastic, meaning that its use in safety-critical domains can be problematic.To address this issue, we introduce a new approach that combines multi-agent reinforcement learning with a formal verification technique termed quantitative verification.Our assured multi-agent reinforcement learning approach constrains agent behaviours in ways that ensure the satisfaction of requirements associated with the safety, reliability, and other non-functional aspects of the decision-making problem being solved.The approach comprises three stages.First, it models the problem as an abstract Markov decision process, allowing quantitative verification to be applied.Next, this abstract model is used to synthesise a policy which satisfies safety, reliability, and performance constraints.Finally, the synthesised policy is used to constrain agent behaviour within the low-level problem with a greatly lowered risk of constraint violations.We demonstrate our approach using a safety-critical multi-agent patrolling problem. Joshua Riley, Radu Calinescu, Colin Paterson, Daniel Kudenko, Alec Banks |
ICAART (2) | 4 |
| 2021 | Utilising Assured Multi-Agent Reinforcement Learning within Safety-Critical ScenariosabstractMulti-agent reinforcement learning allows a team of agents to learn how to work together to solve complex decision-making problems in a shared environment. However, this learning process utilises stochastic mechanisms, meaning that its use in safety-critical domains can be problematic. To overcome this issue, we propose an Assured Multi-Agent Reinforcement Learning (AMARL) approach that uses a model checking technique called quantitative verification to provide formal guarantees of agent compliance with safety, performance, and other non-functional requirements during and after the reinforcement learning process. We demonstrate the applicability of our AMARL approach in three different patrolling navigation domains in which multi-agent systems must learn to visit key areas by using different types of reinforcement learning algorithms (temporal difference learning, game theory, and direct policy search). Furthermore, we compare the effectiveness of these algorithms when used in combination with and without our approach. Our extensive experiments with both homogeneous and heterogeneous multi-agent systems of different sizes show that the use of AMARL leads to safety requirements being consistently satisfied and to better overall results than standard reinforcement learning. Joshua Riley, Radu Calinescu, Colin Paterson, Daniel Kudenko, Alec Banks |
KES | 4 |
| 2021 | An Online Learning Algorithm for Non-stationary Imbalanced Data by Extra-Charging Minority Class
Sajjad Kamali Siahroudi, Daniel Kudenko |
PAKDD (1) | 2 |
| 2021 | Improving cold-start recommendations using item-based stereotypesabstractAbstract Recommender systems (RSs) have become key components driving the success of e-commerce and other platforms where revenue and customer satisfaction is dependent on the user’s ability to discover desirable items in large catalogues. As the number of users and items on a platform grows, the computational complexity and the sparsity problem constitute important challenges for any recommendation algorithm. In addition, the most widely studied filtering-based RSs, while effective in providing suggestions for established users and items, are known for their poor performance for the new user and new item (cold-start) problems. Stereotypical modelling of users and items is a promising approach to solving these problems. A stereotype represents an aggregation of the characteristics of the items or users which can be used to create general user or item classes. We propose a set of methodologies for the automatic generation of stereotypes to address the cold-start problem. The novelty of the proposed approach rests on the findings that stereotypes built independently of the user-to-item ratings improve both recommendation metrics and computational performance during cold-start phases. The resulting RS can be used with any machine learning algorithm as a solver, and the improved performance gains due to rate-agnostic stereotypes are orthogonal to the gains obtained using more sophisticated solvers. The paper describes how such item-based stereotypes can be evaluated via a series of statistical tests prior to being used for recommendation. The proposed approach improves recommendation quality under a variety of metrics and significantly reduces the dimension of the recommendation model. Nourah A. ALRossais, Daniel Kudenko, Tommy Yuan |
User Model. User Adapt. Interact. | 2 |
| 2020 | Uniform State Abstraction for Reinforcement LearningabstractPotential Based Reward Shaping combined with a potential function based on appropriately defined abstract knowledge has been shown to significantly improve learning speed in Reinforcement Learning. MultiGrid Reinforcement Learning (MRL) has further shown that such abstract knowledge in the form of a potential function can be learned almost solely from agent interaction with the environment. However, we show that MRL faces the problem of not extending well to work with Deep Learning. In this paper we extend and improve MRL to take advantage of modern Deep Learning algorithms such as Deep Q-Networks (DQN). We show that DQN augmented with our approach perform significantly better on continuous control tasks than its Vanilla counterpart and DQN augmented with MRL. John Burden, Daniel Kudenko |
ECAI | 2 |
| 2020 | Player Style Clustering without Game VariablesabstractPlayer clustering when applied to the field of video games has several potential applications. For example, the evaluation of the composition of a player base or the generation of AI agents with identified playing styles. These agents can then be used for either the testing of new game content or used directly to enhance a player’s gaming experience. Most current player clustering techniques focus on the use of internal game variables. This raises two main issues: (1) the availability of game variables, as source code access is required to log them and hence limits the data sources that can be used, and (2) the choice of game variables can introduce unintended bias in the types of play style extracted. In this work, a hybrid unsupervised frame encoder and a ‘reference-based’ clustering algorithm are both proposed and combined to allow clustering from raw game play videos. It is shown that the proposed methods are most beneficial when the types of play styles are unknown. Mark Ferguson, Sam Devlin, Daniel Kudenko, James Alfred Walker |
FDG | 3 |
| 2020 | Automatic Similarity Detection in LEGO Ducks
Mark Ferguson, Sebastian Deterding, Andreas Lieberoth, Marc Malmdorf Andersen, Sam Devlin, Daniel Kudenko, James Alfred Walker |
ICCC | 6 |
| 2019 | Guest Editorial: Special Issue on Intelligent Robotics and Multi-Agent SystemsabstractThis special issue includes extended versions of selected papers presented in the Intelligent Robotics and Multi-Agent Systems (IRMAS) track of the 34th ACM/SIGAPP Symposium on Applied Computing (S... Rui P. Rocha, Daniel Kudenko |
Cybern. Syst. | 2 |
| 2018 | Learning to Run with Potential-Based Reward Shaping and Demonstrations from Video DataabstractLearning to produce efficient movement behaviour for humanoid robots from scratch is a hard problem, as has been illustrated by the “Learning to run” competition at NIPS 2017. The goal of this competition was to train a two-legged model of a humanoid body to run in a simulated race course with maximum speed. All submissions took a tabula rasa approach to reinforcement learning (RL) and were able to produce relatively fast, but not optimal running behaviour. In this paper, we demonstrate how data from videos of human running (e.g. taken from YouTube) can be used to shape the reward of the humanoid learning agent to speed up the learning and produce a better result. Specifically, we are using the positions of key body parts at regular time intervals to define a potential function for potential-based reward shaping (PBRS). Since PBRS does not change the optimal policy, this approach allows the RL agent to overcome sub-optimalities in the human movements that are shown in the videos. We present experiments in which we combine selected techniques from the top ten approaches from the NIPS competition with further optimizations to create an high-performing agent as a baseline. We then demonstrate how video-based reward shaping improves the performance further, resulting in an RL agent that runs twice as fast as the baseline in 12 hours of training. We furthermore show that our approach can overcome sub-optimal running behaviour in videos, with the learned policy significantly outperforming that of the running agent from the video. Aleksandra Malysheva, Daniel Kudenko, Aleksei Shpilman |
ICARCV | 2 |
| 2018 | Continuous Gesture Recognition from sEMG Sensor Data with Recurrent Neural Networks and Adversarial Domain AdaptationabstractMovement control of artificial limbs has made big advances in recent years. New sensor and control technology enhanced the functionality and usefulness of artificial limbs to the point that complex movements, such as grasping, can be performed to a limited extent. To date, the most successful results were achieved by applying recurrent neural networks (RNNs), However, in the domain of artificial hands, experiments so far were limited to non-mobile wrists, which significantly reduces the functionality of such prostheses. In this paper, for the first time, we present empirical results on gesture recognition with both mobile and non-mobile wrists. Furthermore, we demonstrate that recurrent neural networks with simple recurrent units (SRU) outperform regular RNNs in both cases in terms of gesture recognition accuracy, on data acquired by an arm band sensing electromagnetic signals from arm muscles (via surface electromyography or sEMG). Finally, we show that adding domain adaptation techniques to continuous gesture recognition with RNN improves the transfer ability between subjects, where a limb controller trained on data from one person is used for another person. Ivan Sosin, Daniel Kudenko, Aleksei Shpilman |
ICARCV | 2 |
| 2018 | A Comparative Evaluation of Machine Learning Methods for Robot Navigation Through Human CrowdsabstractRobot navigation through crowds poses a difficult challenge to AI systems, since the methods should result in fast and efficient movement but at the same time are not allowed to compromise safety. Most approaches to date were focused on the combination of pathfinding algorithms with machine learning for pedestrian walking prediction. More recently, reinforcement learning techniques have been proposed in the research literature. In this paper, we perform a comparative evaluation of pathfinding/prediction and reinforcement learning approaches on a crowd movement dataset collected from surveillance videos taken at Grand Central Station in New York. The results demonstrate the strong superiority of state-of-the-art reinforcement learning approaches over pathfinding with state-of-the-art behavior prediction techniques. Anastasia Gaydashenko, Daniel Kudenko, Aleksei Shpilman |
ICMLA | 2 |
| 2018 | Applying Cartesian Genetic Programming to Evolve Rules for Intrusion Detection System
Hasanen Alyasiri, John A. Clark, Daniel Kudenko |
IJCCI | 3 |
| 2017 | A Reinforcement Learning Based Workflow Application Scheduling Approach in Dynamic Cloud Environment
Daniel Kudenko, Shijun Liu, Li Pan 0001, Lei Wu 0002, Xiangxu Meng |
CollaborateCom | 2 |
| 2017 | Distribution Data Across Multiple Cloud Storage using Reinforcement Learning Method
Abdullah Fayez H. Algarni, Daniel Kudenko |
ICAART (2) | 2 |
| 2017 | Assured Reinforcement Learning with Formally Verified Abstract PoliciesabstractWe present a new reinforcement learning (RL) approach that enables an autonomous agent to solve decision making problems under constraints.Our assured reinforcement learning approach models the uncertain environment as a high-level, abstract Markov decision process (AMDP), and uses probabilistic model checking to establish AMDP policies that satisfy a set of constraints defined in probabilistic temporal logic.These formally verified abstract policies are then used to restrict the RL agent's exploration of the solution space so as to avoid constraint violations.We validate our RL approach by using it to develop autonomous agents for a flag-collection navigation task and an assisted-living planning problem. George Mason, Radu Calinescu, Daniel Kudenko, Alec Banks |
ICAART (2) | 3 |
| 2017 | Deep Learning of Cell Classification Using Microscope Images of Intracellular Microtubule NetworksabstractMicrotubule networks (MTs) are a component of a cell that may indicate the presence of various chemical compounds and can be used to recognize properties such as treatment resistance. Therefore, the classification of MT images is of great relevance for cell diagnostics. Human experts find it particularly difficult to recognize the levels of chemical compound exposure of a cell. Improving the accuracy with automated techniques would have a significant impact on cell therapy. In this paper we present the application of Deep Learning to MT image classification and evaluate it on a large MT image dataset of animal cells with three degrees of exposure to a chemical agent. The results demonstrate that the learned deep network performs on par or better at the corresponding cell classification task than human experts. Specifically, we show that the task of recognizing different levels of chemical agent exposure can be handled significantly better by the neural network than by human experts. Aleksei Shpilman, Dmitry Boikiy, Marina Polyakova, Daniel Kudenko, Anton Burakov, Elena Nadezhdina |
ICMLA | 4 |
| 2016 | Potential-based reward shaping for finite horizon online POMDP planning
Adam Eck, Leen-Kiat Soh, Sam Devlin, Daniel Kudenko |
Auton. Agents Multi Agent Syst. | 4 |
| 2015 | Distributed reinforcement learning for adaptive and robust network intrusion responseabstractDistributed denial of service (DDoS) attacks constitute a rapidly evolving threat in the current Internet. Multiagent Router Throttling is a novel approach to defend against DDoS attacks where multiple reinforcement learning agents are installed on a set of routers and learn to rate-limit or throttle traffic towards a victim server. The focus of this paper is on online learning and scalability. We propose an approach that incorporates task decomposition, team rewards and a form of reward shaping called difference rewards. One of the novel characteristics of the proposed system is that it provides a decentralised coordinated response to the DDoS problem, thus being resilient to DDoS attacks themselves. The proposed system learns remarkably fast, thus being suitable for online learning. Furthermore, its scalability is successfully demonstrated in experiments involving 1000 learning agents. We compare our approach against a baseline and a popular state-of-the-art throttling technique from the network security literature and show that the proposed approach is more effective, adaptive to sophisticated attack rate dynamics and robust to agent failures. Kleanthis Malialis, Sam Devlin, Daniel Kudenko |
Connect. Sci. | 3 |
| 2015 | Distributed response to network intrusions using multiagent reinforcement learning
Kleanthis Malialis, Daniel Kudenko |
Eng. Appl. Artif. Intell. | 2 |
| 2014 | Combining Multiple Correlated Reward and Shaping Signals by Measuring ConfidenceabstractMulti-objective problems with correlated objectives are a class of problems that deserve specific attention. In contrast to typical multi-objective problems, they do not require the identification of trade-offs between the objectives, as (near-) optimal solutions for any objective are (near-) optimal for every objective. Intelligently combining the feedback from these objectives, instead of only looking at a single one, can improve optimization. This class of problems is very relevant in reinforcement learning, as any single-objective reinforcement learning problem can be framed as such a multi-objective problem using multiple reward shaping functions. After discussing this problem class, we propose a solution technique for such reinforcement learning problems, called adaptive objective selection. This technique makes a temporal difference learner estimate the Q-function for each objective in parallel, and introduces a way of measuring confidence in these estimates. This confidence metric is then used to choose which objective's estimates to use for action selection. We show significant improvements in performance over other plausible techniques on two problem domains. Finally, we provide an intuitive analysis of the technique's decisions, yielding insights into the nature of the problems being solved. Tim Brys, Ann Nowé, Daniel Kudenko, Matthew E. Taylor |
AAAI | 3 |
| 2014 | Coordinated Team Learning and Difference Rewards for Distributed Intrusion ResponseabstractDistributed denial of service attacks constitute a rapidly evolving threat in the current Internet. Multiagent Router Throttling is a novel approach to respond to such attacks. We demonstrate that our approach can significantly scale-up using hierarchical communication and coordinated team learning. Furthermore, we incorporate a form of reward shaping called difference rewards and show that the scalability of our system is significantly improved in experiments involving over 100 reinforcement learning agents. We also demonstrate that difference rewards constitute an ideal online learning mechanism for network intrusion response. We compare our proposed approach against a popular state-of-the-art router throttling technique from the network security literature, and we show that our proposed approach significantly outperforms it. We note that our approach can be useful in other related multiagent domains. Kleanthis Malialis, Sam Devlin, Daniel Kudenko |
ECAI | 3 |
| 2014 | Improving Robustness of Gaussian Process-Based Inferential Control System Using Kernel Principle Component AnalysisabstractThe plausibility and robustness of an inferential control system entirely depend on the prediction accuracy of the estimator used as the feedback element. This paper is based on a previously proposed Gaussian process inferential controller that employs Gaussian process soft sensor as an estimator. The paper enhances the robustness and the reliability of the control system, particularly, during sensor input failures. The contribution of the paper is i) alleviating the affect of the failure on the prediction accuracy of feedback element (soft sensor) and thus improving the robustness of the overall control system. ii) Hybridising Kernel Principal Component Analysis with Gaussian process Inferential Control System to achieve this robustness during all process operating conditions. The paper empirically shows the effectiveness and the plausibility of the processed hybrid system on a simulated chemical reactor process. Ali Abusnina, Daniel Kudenko, Rolf Roth |
ICMLA | 2 |
| 2014 | A Phylogenetic Classification of the Video-Game Industry's Business Model Ecosystem
Nikolaos Goumagias, Ignazio Cabras, Kiran Jude Fernandes 0001, Feng Li 0021, Alberto Nucciarelli, Peter I. Cowling, Sam Devlin, Daniel Kudenko |
PRO-VE | 8 |
| 2014 | Multi-objectivization of reinforcement learning problems by reward shapingabstractMulti-objectivization is the process of transforming a single objective problem into a multi-objective problem. Research in evolutionary optimization has demonstrated that the addition of objectives that are correlated with the original objective can make the resulting problem easier to solve compared to the original single-objective problem. In this paper we investigate the multi-objectivization of reinforcement learning problems. We propose a novel method for the multi-objectivization of Markov Decision problems through the use of multiple reward shaping functions. Reward shaping is a technique to speed up reinforcement learning by including additional heuristic knowledge in the reward signal. The resulting composite reward signal is expected to be more informative during learning, leading the learner to identify good actions more quickly. Good reward shaping functions are by definition correlated with the target value function for the base reward signal, and we show in this paper that adding several correlated signals can help to solve the basic single objective problem faster and better. We prove that the total ordering of solutions, and by consequence the optimality of solutions, is preserved in this process, and empirically demonstrate the usefulness of this approach on two reinforcement learning tasks: a pathfinding problem and the Mario domain. Tim Brys, Anna Harutyunyan, Peter Vrancx, Matthew E. Taylor, Daniel Kudenko, Ann Nowé |
IJCNN | 5 |
| 2014 | A comparison of plan-based and abstract MDP reward shapingabstractReward shaping has been shown to significantly improve an agent's performance in reinforcement learning. As attention is shifting away from tabula-rasa approaches many different reward shaping methods have been developed. In this paper, we compare two different methods for reward shaping; plan-based, in which an agent is provided with a plan and extra rewards are given according to the steps of the plan the agent satisfies, and reward shaping via abstract Markov decision process (MDPs), in which an abstract high-level MDP of the environment is solved and the resulting value function is used to shape the agent. The comparison is conducted in terms of total reward, convergence speed and scaling up to more complex environments. Empirical results demonstrate the need to correctly select and set up reward shaping methods according to the needs of the environment the agents are acting in. This leads to the more interesting question, is there a reward shaping method which is universally better than all other approaches regardless of the environment dynamics? Kyriakos Efthymiadis, Daniel Kudenko |
Connect. Sci. | 2 |
| 2013 | Multiagent Router Throttling: Decentralized Coordinated Response Against DDoS AttacksabstractDistributed denial of service (DDoS) attacks constitute a rapidly evolving threat in the current Internet. In this paper we introduce Multiagent Router Throttling, a decentralized DDoS response mechanism in which a set of upstream routers independently learn to throttle traffic towards a victim server. We compare our approach against a baseline and a popular throttling technique from the literature, and we show that our proposed approach is more secure, reliable and cost-effective. Furthermore, our approach outperforms the baseline technique and either outperforms or has the same performance as the popular one. Kleanthis Malialis, Daniel Kudenko |
IAAI | 2 |
| 2010 | Online learning of shaping rewards in reinforcement learning
Marek Grzes, Daniel Kudenko |
Neural Networks | 2 |
| 2009 | Educational Narrative and Student Modeling for Ill-Defined DomainsabstractThis paper introduces an adaptive narrative-based learning environment (AEINS) which supports teaching in the domain of ethics and citizenship [1]. AEINS is an inquiry-based teaching system that adopts educational theories and classroom strategies, such as the Socratic Method. AEINS targets students aged 8 to 11 years. The idea is centered around presenting and involving students in different moral dilemmas (teaching moments), where the Socratic Method is used to lead the student and allows him to examine the validity of an opinion or belief. AEINS monitors and analyzes students' actions in order to provide an individualized story-path and a personalized learning process. We present preliminary evaluation results. Rania Hodhod, Daniel Kudenko, Paul A. Cairns |
AIED | 2 |
| 2009 | Theoretical and Empirical Analysis of Reward Shaping in Reinforcement LearningabstractReinforcement learning suffers scalability problems due to the state space explosion and the temporal credit assignment problem. Knowledge-based approaches have received a significant attention in the area. Reward shaping is a particular approach to incorporate domain knowledge into reinforcement learning. Theoretical and empirical analysis of this paper reveals important properties of this principle, especially the influence of the reward type, MDP discount factor, and the way of evaluating the potential function on the performance. Marek Grzes, Daniel Kudenko |
ICMLA | 2 |
| 2009 | Duality of Actor and Character Goals in Virtual Drama
María Arinbjarnar, Daniel Kudenko |
IVA | 2 |
| 2009 | Finding Unexpected Navigation Behaviour in Clickstream Data for Website Design Improvement
I-Hsien Ting, Chris Kimble, Daniel Kudenko |
J. Web Eng. | 3 |
| 2009 | Generation of Adaptive Dilemma-Based Interactive NarrativesabstractThe generator of adaptive dilemma-based interactive narratives (GADIN) presented in this paper dynamically generates interactive narratives which are focused on dilemmas to create dramatic tension. The system is provided with knowledge of generic story actions and dilemmas based on those clichE¿s encountered in many storytelling domains. The domain designer is only required to provide domain-specific information, for example, regarding characters and their relationships, locations, and actions. A planner creates sequences of actions that all lead to a dilemma for a character (who can be the user). The user interacts with the storyworld by making decisions on relevant dilemmas and by freely choosing their own actions. Using this input, the system chooses and adapts future storylines according to the user's past behavior. Previous interactive narrative systems often have content creation and ordering requirements that restrict the possibility for sustaining the dramatic interest of the narrative over a long time period. In addition, many of these systems are not easily transferable between domains. In this paper, the GADIN system is demonstrated to both be able to maintain the dramatic interest of generated narratives over a long time period and to have a core architecture that is applicable to any domain. Heather Barber, Daniel Kudenko |
IEEE Trans. Comput. Intell. AI Games | 2 |
| 2008 | Multi-Agent Reinforcement Learning for Intrusion Detection: A case study and evaluationabstractIn this paper we propose a novel approach to train Multi-Agent Reinforcement Learning (MARL) agents to cooperate to detect intrusions in the form of normal and abnormal states in the network. We present an architecture of distributed sensor and decision agents that learn how to identify normal and abnormal states of the network using Reinforcement Learning (RL). Sensor agents extract network-state information using tile-coding as a function approximation technique and send communication signals in the form of actions to decision agents. By means of an on line process, sensor and decision agents learn the semantics of the communication actions. In this paper we detail the learning process and the operation of the agent architecture. We also present tests and results of our research work in an intrusion detection case study, using a realistic network simulation where sensor and decision agents learn to identify normal and abnormal states of the network. Arturo Servin, Daniel Kudenko |
ECAI | 2 |
| 2008 | Multigrid Reinforcement Learning with Reward Shaping
Marek Grzes, Daniel Kudenko |
ICANN (1) | 2 |
| 2008 | Schemas in Directed Emergent Drama
María Arinbjarnar, Daniel Kudenko |
ICIDS | 2 |
| 2008 | Generation of Dilemma-Based Narratives: Method and Turing Test Evaluation
Heather Barber, Daniel Kudenko |
ICIDS | 2 |
| 2007 | An Analysis of Problem Difficulty for a Class of Optimisation Heuristics
Enda Ridge, Daniel Kudenko |
EvoCOP | 2 |
| 2007 | Analyzing heuristic performance with response surface models: prediction, optimization and robustnessabstractThis research uses a Design of Experiments (DOE) approach to build a predictive model of the performance of a combinatorial optimization heuristic over a range of heuristic tuning parameter settings and problem instance characteristics. The heuristic is Ant Colony System (ACS) for the Travelling Salesperson Problem. 10 heurstic tuning parameters and 2 problem characteristics are considered. Response Surface Models (RSM) of the solution quality and solution time predicted ACS performance on both new instances from a publicly available problem generator and new real-world instances from the TSPLIB benchmark library. A numerical optimisation of the RSMs is used to find the tuning parameter settings that yield optimal performance in terms of solution quality and solution time. This paper is the first use of desirability functions, a well-established technique in DOE, to simultaneously optimise these conflicting goals. Finally, overlay plots are used to examine the robustness of the performance of the optimised heuristic across a range of problem instance characteristics. These plots give predictions on the range of problem instances for which a given solutionquality can be expected within a given solution time. Enda Ridge, Daniel Kudenko |
GECCO | 2 |
| 2007 | Screening the parameters affecting heuristic performanceabstractThis research screens the tuning parameters of a combinatorial optimization heuristic. Specifically, it presents a Design of Experiments (DOE) approach that uses a Fractional Factorial Design to screen the tuning parameters of Ant Colony System (ACS) for the Travelling Sales person problem. Screening is a preliminary step towards building a full Response Surface Model (RSM) [2]. It identifies parametersthat have little influence on performance and can be omittedfrom the RSM design. This reduces the complexity andexpense of the RSM design. 10 algorithm parameters and 2 problem characteristics are considered. Open questionson the effect of 3 parameters on performance are answered.A further parameter, sometimes assumed important, was shown to have no effect on performance. A new problem characteristic that effects performance was identified. A full version of this paper is available [3]. Enda Ridge, Daniel Kudenko |
GECCO | 2 |
| 2007 | Applying Web Usage Mining Techniques to Discover Potential Browsing Problems of UsersabstractIn this paper, a web usage mining based approach is proposed to discover potential browsing problems. Two web usage mining techniques in the approach are introduced, including Automatic Pattern Discovery (APD) and Co-occurrence Pattern Mining with Distance Measurement (CPMDM). A combination method is also discussed to show how potential browsing problems can be identified I-Hsien Ting, Chris Kimble, Daniel Kudenko |
ICALT | 3 |
| 2006 | A Study of Concurrency in the Ant Colony System AlgorithmabstractThis paper reports the results of a study of a specific type of concurrency in the Ant Colony System (ACS) algorithm. Studies of Cellular Automata (CA) have shown that the update mechanism used can have a dramatic influence on the dynamics of the CA. ACS is usually implemented with a sequential update mechanism. A new method for controlling the concurrency in a nature-inspired algorithm is introduced. Comprehensive tests on a wide range of problem instances are reported. The study found that concurrency levels had no statistically significant effect on ACS performance. This result is interesting because it contradicts what has been observed in another form of nature-inspired algorithm, namely CAs. Enda Ridge, Daniel Kudenko, Dimitar Kazakov |
IEEE Congress on Evolutionary Computation | 2 |
| 2005 | A Pattern Restore Method for Restoring Missing Patterns in Server Side Clickstream Data
I-Hsien Ting, Chris Kimble, Daniel Kudenko |
APWeb | 3 |
| 2005 | UBB Mining: Finding Unexpected Browsing Behaviour in Clickstream Data to Improve a Web Site's DesignabstractThis paper describes a novel Web usage mining approach to discover patterns in the navigation of Web sites known as unexpected browsing behaviours (UBBs). By reviewing these UBBs, a Web site designer can choose to modify the design of their Web site or redesign the site completely. UBB mining is based on the continuous common subsequence (CCS), a special instance of common subsequence (CS), which is used to define a set of expected routes. The predefined expected routes are then treated as rules and stored in a rule base. By using the predefined route and the UBB mining algorithm, interesting browsing behaviours can be discovered. This paper introduces the format of the expected route and describes the UBB algorithms. The paper also describes a series of experiments designed to evaluate how well UBB mining algorithms work. I-Hsien Ting, Chris Kimble, Daniel Kudenko |
Web Intelligence | 3 |
| 2004 | Algorithms for Distributed Exploration
Thomas Walker, Daniel Kudenko, Malcolm J. A. Strens |
ECAI | 2 |
| 1994 | An Empirical Analysis of Terminological Representation Systems
Jochen Heinsohn, Daniel Kudenko, Bernhard Nebel, Hans-Jürgen Profitlich |
Artif. Intell. | 2 |
| 1992 | An Empirical Analysis of Terminological Representation Systems
Jochen Heinsohn, Daniel Kudenko, Bernhard Nebel, Hans-Jürgen Profitlich |
AAAI | 2 |