VLDB 2026 Research / reviewers in the wild / expert
Ivana Dusparic
dblp:77/723
· DBLP profile ↗
33ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0003-0621-5400ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 10 since 2021Computer networks · 8 · 6 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geographically-aware Transformer-based Traffic Forecasting for Urban Motorway Digital Twins
Kresimir Kusic, Vinny Cahill, Ivana Dusparic |
IV | 3 |
| 2026 | Introduction to the Special Issue on Artificial Intelligence for Adaptive and Autonomous Cloud/Edge Computing Systems
Gabriele Russo Russo, Valeria Cardellini, Ivana Dusparic, Stefano Iannucci |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2025 | Optimistic Exploration for Risk-Averse Constrained Reinforcement LearningabstractRisk-averse Constrained Reinforcement Learning (RaCRL) aims to learn policies that minimise the likelihood of rare and catastrophic constraint violations caused by an environment’s inherent randomness. In general, risk-aversion leads to conservative exploration of the environment which typically results in converging to sub-optimal policies that fail to adequately maximise reward or, in some cases, fail to achieve the goal. In this paper, we propose an exploration-based approach for RaCRL called Optimistic Risk-averse Actor Critic (ORAC), which constructs an exploratory policy by maximising a local upper confidence bound of the state-action reward value function whilst minimising a local lower confidence bound of the risk-averse state-action cost value function. Specifically, at each step, the weighting assigned to the cost value is increased or decreased if it exceeds or falls below the safety constraint value. This way the policy is encouraged to explore uncertain regions of the environment to discover high reward states whilst still satisfying the safety constraints. Our experimental results demonstrate that the ORAC approach prevents convergence to sub-optimal policies and improves significantly the reward-cost trade-off in various continuous control tasks such as Safety-Gymnasium and a complex building energy management environment CityLearn. James McCarthy, Radu Marinescu 0002, Elizabeth Daly, Ivana Dusparic |
ECAI | 4 |
| 2025 | Drama: Mamba-Enabled Model-Based Reinforcement Learning Is Sample and Parameter EfficientabstractModel-based reinforcement learning (RL) offers a solution to the data inefficiency that plagues most model-free RL algorithms. However, learning a robust world model often requires complex and deep architectures, which are computationally expensive and challenging to train. Within the world model, sequence models play a critical role in accurate predictions, and various architectures have been explored, each with its own challenges. Currently, recurrent neural network (RNN)-based world models struggle with vanishing gradients and capturing long-term dependencies. Transformers, on the other hand, suffer from the quadratic memory and computational complexity of self-attention mechanisms, scaling as $O(n^2)$, where $n$ is the sequence length.
To address these challenges, we propose a state space model (SSM)-based world model, Drama, specifically leveraging Mamba, that achieves $O(n)$ memory and computational complexity while effectively capturing long-term dependencies and enabling efficient training with longer sequences. We also introduce a novel sampling method to mitigate the suboptimality caused by an incorrect world model in the early training stages. Combining these techniques, Drama achieves a normalised score on the Atari100k benchmark that is competitive with other state-of-the-art (SOTA) model-based RL algorithms, using only a 7 million-parameter world model. Drama is accessible and trainable on off-the-shelf hardware, such as a standard laptop. Our code is available at https://github.com/realwenlongwang/Drama.git. Ivana Dusparic, Vinny Cahill |
ICLR | 2 |
| 2024 | Semifactual Explanations for Reinforcement LearningabstractReinforcement Learning (RL) is a learning paradigm in which the agent learns from its environment through trial and error. Deep reinforcement learning (DRL) algorithms represent the agent’s policies using neural networks, making their decisions difficult to interpret. Explaining the behaviour of DRL agents is necessary to advance user trust, increase engagement, and facilitate integration with real-life tasks. Semifactual explanations aim to explain an outcome by providing “even if” scenarios, such as “even if the car were moving twice as slowly, it would still have to swerve to avoid crashing”. Semifactuals help users understand the effects of different factors on the outcome and support the optimisation of resources. While extensively studied in psychology and even utilised in supervised learning, semifactuals have not been used to explain the decisions of RL systems. In this work, we develop a first approach to generating semifactual explanations for RL agents. We start by defining five properties of desirable semifactual explanations in RL and then introducing SGRL-Rewind and SGRL-Advance, the first algorithms for generating semifactual explanations in RL. We evaluate the algorithms in two standard RL environments and find that they generate semifactuals that are easier to reach, represent the agent’s policy better, and are more diverse compared to baselines. Lastly, we conduct and analyse a user study to assess the participant’s perception of semifactual explanations of the agent’s actions. Jasmina Gajcin, Jovan Jeromela, Ivana Dusparic |
HAI | 3 |
| 2024 | Applying Neural Monte Carlo Tree Search to Unsignalized Multi-intersection Scheduling for Autonomous VehiclesabstractDynamic scheduling of access to shared resources by autonomous systems is a challenging problem, characterized as being NP-hard. The complexity of this task leads to a combinatorial explosion of possibilities in highly dynamic systems where arriving requests must be continuously scheduled subject to strong safety and time constraints. An example of such a system is an unsignalized intersection, where automated vehicles’ access to potential conflict zones must be dynamically scheduled. In this paper, we apply Neural Monte Carlo Tree Search (NMCTS) to the challenging task of scheduling platoons of vehicles crossing unsignalized intersections. Crucially, we introduce a transformation model that maps successive sequences of potentially conflicting road-space reservation requests from platoons of vehicles into a series of board-game-like problems and use NMCTS to search for solutions representing optimal road-space allocation schedules in the context of past allocations. To optimize search, we incorporate a prioritized re-sampling method with parallel NMCTS (PNMCTS) to improve the quality of training data. To optimize training, a curriculum learning strategy is used to train the agent to schedule progressively more complex boards culminating in overlapping boards that represent busy intersections. In a busy single four-way unsignalized intersection simulation, PNMCTS solved 95% of unseen scenarios, reducing crossing time by 43% in light and 52% in heavy traffic versus first-in, first-out control. In a 3x3 multi-intersection network, the proposed method maintained free-flow in light traffic when all intersections are under control of PNMCTS and outperformed state-of-the-art RL-based traffic-light controllers in average travel time by 74.5% and total throughput by 16% in heavy traffic. Xiaowen Tao, Ivana Dusparic, Vinny Cahill |
IROS | 4 |
| 2024 | FedSBS: Federated-Learning participant-selection method for Intrusion Detection Systems
Hélio N. Cunha Neto, Jernej Hribar, Ivana Dusparic, Natalia Castro Fernandes, Diogo M. F. Mattos |
Comput. Networks | 3 |
| 2023 | Prevalence of Code Smells in Reinforcement Learning ProjectsabstractReinforcement Learning (RL) is being increasingly used to learn and adapt application behavior in many domains, including large-scale and safety critical systems, as for example, autonomous driving. With the advent of plug-n-play RL libraries, its applicability has further increased, enabling integration of RL algorithms by users. We note, however, that the majority of such code is not developed by RL engineers, which as a consequence, may lead to poor program quality yielding bugs, suboptimal performance, maintainability, and evolution problems for RL-based projects. In this paper we begin the exploration of this hypothesis, specific to code utilizing RL, analyzing different projects found in the wild, to assess their quality from a software engineering perspective. Our study includes 24 popular RL-based Python projects, analyzed with standard software engineering metrics. Our results, aligned with similar analyses for ML code in general, show that popular and widely reused RL repositories contain many code smells (3.95% of the code base on average), significantly affecting the projects’ maintainability. The most common code smells detected are long method and long method chain, highlighting problems in the definition and interaction of agents. Detected code smells suggest problems in responsibility separation, and the appropriateness of current abstractions for the definition of RL algorithms. Nicolás Cardozo, Ivana Dusparic, Christian Cabrera 0001 |
CAIN | 2 |
| 2023 | Expert-Free Online Transfer Learning in Multi-Agent Reinforcement LearningabstractTransfer learning in Reinforcement Learning (RL) has been widely studied to overcome training challenges in Deep-RL, i.e., exploration cost, data availability and convergence time, by bootstrapping external knowledge to enhance learning phase. While this overcomes the training issues on a novice agent, a good understanding of the task by the expert agent is required for such a transfer to be effective. As an alternative, in this paper we propose Expert-Free Online Transfer Learning (EF-OnTL), an algorithm that enables expert-free real-time dynamic transfer learning in multi-agent system. No dedicated expert agent exists, and transfer source agent and knowledge to be transferred are dynamically selected at each transfer step based on agents’ performance and level of uncertainty. To improve uncertainty estimation, we also propose State Action Reward Next-State Random Network Distillation (sars-RND), an extension of RND that estimates uncertainty from RL agent-environment interaction. We demonstrate EF-OnTL effectiveness against a no-transfer scenario and state-of-the-art advice-based baselines, with and without expert agents, in three benchmark tasks: Cart-Pole, a grid-based Multi-Team Predator-Prey (MT-PP) and Half Field Offense (HFO). Our results show that EF-OnTL achieves overall comparable performance to that of advice-based approaches, while not requiring expert agents, external input, nor threshold tuning. EF-OnTL outperforms no-transfer with an improvement related to the complexity of the task addressed. Alberto Castagna, Ivana Dusparic |
ECAI | 2 |
| 2023 | Iterative Reward Shaping Using Human Feedback for Correcting Reward MisspecificationabstractA well-defined reward function is crucial for successful training of an reinforcement learning (RL) agent. However, defining a suitable reward function is a notoriously challenging task, especially in complex, multi-objective environments. Developers often have to resort to starting with an initial, potentially misspecified reward function, and iteratively adjusting its parameters, based on observed learned behavior. In this work, we aim to automate this process by proposing ITERS, an iterative reward shaping approach using human feedback for mitigating the effects of a misspecified reward function. Our approach allows the user to provide trajectory-level feedback on agent’s behavior during training, which can be integrated as a reward shaping signal in the following training iteration. We also allow the user to provide explanations of their feedback, which are used to augment the feedback and reduce user effort and feedback frequency. We evaluate ITERS in three environments and show that it can successfully correct misspecified reward functions. Jasmina Gajcin, James McCarthy, Rahul Nair 0004, Radu Marinescu 0002, Elizabeth Daly, Ivana Dusparic |
ECAI | 6 |
| 2023 | Deep W-Networks: Solving Multi-Objective Optimisation Problems with Deep Reinforcement LearningabstractIn this paper, we build on advances introduced by the Deep Q-Networks (DQN) approach to extend the multi-objective tabular Reinforcement Learning (RL) algorithm W-learning to large state spaces. W-learning algorithm can naturally solve the competition between multiple single policies in multi-objective environments. However, the tabular version does not scale well to environments with large state spaces. To address this issue, we replace underlying Q-tables with DQN, and propose an addition of W-Networks, as a replacement for tabular weights (W) representations. We evaluate the resulting Deep W-Networks (DWN) approach in two widely-accepted multi-objective RL benchmarks: deep sea treasure and multi-objective mountain car. We show that DWN solves the competition between multiple policies while outperforming the baseline in the form of a DQN solution. Additionally, we demonstrate that the proposed algorithm can find the Pareto front in both tested environments. Jernej Hribar, Luke Hackett, Ivana Dusparic |
ICAART (2) | 3 |
| 2023 | Reservation of Virtualized Resources with Optimistic Online LearningabstractThe virtualization of wireless networks enables new services to access network resources made available by the Network Operator (NO) through a Network Slicing market. The different service providers (SPs) have the opportunity to lease the network resources from the NO to constitute slices that address the demand of their specific network service. The goal of any SP is to maximize its service utility and minimize costs from leasing resources while facing uncertainties of the prices of the resources and the users' demand. In this paper, we propose a solution that allows the SP to decide its online reservation policy, which aims to maximize its service utility and minimize its cost of reservation simultaneously. We design the Optimistic Online Learning for Reservation (OOLR) solution, a decision algorithm built upon the Follow-the-Regularized Leader (FTRL), that incorporates key predictions to assist the decision-making process. Our solution achieves a$\mathcal{O}(\sqrt{T})$regret bound where$T$represents the horizon. We integrate a prediction model into the OOLR solution and we demonstrate through numerical results the efficacy of the combined models' solution against the FTRL baseline. Jean-Baptiste Monteil, George Iosifidis, Ivana Dusparic |
ICC | 3 |
| 2023 | Density-Aware Reinforcement Learning to Optimise Energy Efficiency in UAV-Assisted NetworksabstractUnmanned aerial vehicles (UAVs) serving as aerial base stations can be deployed to provide wireless connectivity to mobile users, such as vehicles. However, the density of vehicles on roads often varies spatially and temporally primarily due to mobility and traffic situations in a geographical area, making it difficult to provide ubiquitous service. Moreover, as energy-constrained UAVs hover in the sky while serving mobile users, they may be faced with interference from nearby UAV cells or other access points sharing the same frequency band, thereby impacting the system’s energy efficiency (EE). Recent multiagent reinforcement learning (MARL) approaches applied to optimise the users’ coverage worked well in reasonably even densities but might not perform as well in uneven users’ distribution, i.e., in urban road networks with uneven concentration of vehicles. In this work, we propose a density-aware communication-enabled multi-agent decentralised double deep Q-network (DACEMAD–DDQN) approach that maximises the total system’s EE by jointly optimising the trajectory of each UAV, the number of connected users, and the UAVs’ energy consumption while keeping track of dense and uneven users’ distribution. Our result outperforms state-of-the-art MARL approaches in terms of EE by as much as 65% – 85%. Omoniwa Babatunji, Boris Galkin, Ivana Dusparic |
WiMob | 3 |
| 2023 | Auto-COP: Adaptation generation in Context-oriented Programming using Reinforcement Learning optionsabstractSelf-adaptive software systems continuously adapt in response to internal and external changes in their execution environment, captured as contexts. The Context-oriented Programming (COP) paradigm posits a technique for the development of self-adaptive systems, capturing their main characteristics with specialized programming language constructs. In COP, adaptations are specified as independent modules that are composed in and out of the base system as contexts are activated and deactivated in response to sensed circumstances from the surrounding environment. However, the definition of adaptations, their contexts and associated specialized behavior, need to be specified at design time. In complex cyber physical systems this is intractable, if not impossible, due to new unpredicted operating conditions arising. In this paper, we propose Auto-COP, a new technique to enable generation of adaptations at run time. Auto-COP uses Reinforcement Learning (RL) options to build action sequences, based on the previous instances of the system execution (for example, atomic system actions enacted by human operators). Options are further explored in interaction with the environment, and the most suitable options for each context are used to generate the adaptations, exploiting COP abstractions. To validate Auto-COP, we present two case studies exhibiting different system characteristics and application domains: a driving assistant and a robot delivery system. We present examples of Auto-COP to illustrate the types of circumstances (contexts) requiring adaptation at run time, and the corresponding generated adaptations for each context. We confirm that the generated adaptations exhibit correct system behavior measured by domain-specific performance metrics (e.g., conformance to specified speed limit), while reducing the number of required execution/actuation steps by a factor of two showing that the adaptations are regularly selected by the running system as adaptive behavior is more appropriate than the execution of atomic actions. Therefore, we demonstrate that Auto-COP is able to increase system adaptivity by enabling run-time generation of new adaptations for conditions detected at run time, while retaining the modularity offered by COP languages, and reducing the upfront specification required by system developers. Nicolás Cardozo, Ivana Dusparic |
Inf. Softw. Technol. | 2 |
| 2022 | Prequential Model Selection for Time Series Forecasting based on Saliency MapsabstractOver the last few years, incremental machine learning for streaming data has gained significant attention due to the need to learn from a constantly evolving stream of data without the need to store it. The advent of big data has further fuelled research on developing systems that can cope with continuously changing data streams and tackle the challenges associated with historical data requirements. The problem of time series forecasting has been studied using varied approaches like neural networks, ensemble methods, decision trees and rules, support vector machines to name a few. However, neural network models have gained particular attention in dealing with changing data distribution due to their generalization abilities. In this paper, we propose a prequential framework named PS-PGSM which involves incrementally training the base models and online Regions of Competence (ROC) computation followed by selection of the best forecaster for the task of time series forecasting using saliency maps. We build upon the state-of-the-art approach named OS-PGSM (Online Model Selection using Performance Gradient based Saliency Maps) in which the model training and ROC computation is performed offline. Past research has demonstrated that a set of different models enables specialization for each model compared to a single forecasting model which is particularly useful when predicting for an evolving time series sequence. Our approach uses saliency maps for prequential calculation of ROC for each model to find the best forecaster based on the performance of each model. We evaluate the proposed approach against OS-PGSM, as well as against previous best performing model by first conducting preliminary experiments on 10 real-world time series datasets and then using 2 real-world big datasets to showcase its applicability to big data. Experimental results not only validate the effectiveness of our approach for big data but also demonstrate superior performance in terms of prediction accuracy and computational time efficiency while also handling concept drift. Shivani Tomar, Seshu Tirupathi, Dhaval Salwala, Ivana Dusparic, Elizabeth Daly |
IEEE Big Data | 4 |
| 2022 | Energy-aware optimization of UAV base stations placement via decentralized multi-agent Q-learningabstractUnmanned aerial vehicles serving as aerial base stations (UAV-BSs) can be deployed to provide wireless connectivity to ground devices in events of increased network demand, points-of-failure in existing infrastructure, or disasters. However, it is challenging to conserve the energy of UAVs during prolonged coverage tasks, considering their limited on-board battery capacity. Reinforcement learning-based (RL) approaches have been previously used to improve energy utilization of multiple UAVs, however, a central cloud controller is assumed to have complete knowledge of the end-devices’ locations, i.e., the controller periodically scans and sends updates for UAV decision-making. This assumption is impractical in dynamic network environments with UAVs serving mobile ground devices. To address this problem, we propose a decentralized Q-learning approach, where each UAVBS is equipped with an autonomous agent that maximizes the connectivity of mobile ground devices while improving its energy utilization. Experimental results show that the proposed design significantly outperforms the centralized approaches in jointly maximizing the number of connected ground devices and the energy utilization of the UAV-BSs. Omoniwa Babatunji, Boris Galkin, Ivana Dusparic |
CCNC | 3 |
| 2022 | Multi-agent Transfer Learning in Reinforcement Learning-based Ride-sharing SystemsabstractReinforcement learning (RL) has been used in a range of simulated real-world tasks, e.g., sensor coordination, traffic light control, and on-demand mobility services. However, real world deployments are rare, as RL struggles with dynamic nature of real world environments, requiring time for learning a task and adapting to changes in the environment. Transfer Learning (TL) can help lower these adaptation times. In particular, there is a significant potential of applying TL in multi-agent RL systems, where multiple agents can share knowledge with each other, as well as with new agents that join the system. To obtain the most from inter-agent transfer, transfer roles (i.e., determining which agents act as sources and which as targets), as well as relevant transfer content parameters (e.g., transfer size) should be selected dynamically in each particular situation. As a first step towards fully dynamic transfers, in this paper we investigate the impact of TL transfer parameters with fixed source and target roles. Specifically, we label every agent-environment interaction with agent's epistemic confidence, and we filter the shared examples using varying threshold levels and sample sizes. We investigate impact of these parameters in two scenarios, a standard predator-prey RL benchmark and a simulation of a ride-sharing system with 200 vehicle agents and 10,000 ride-requests. Alberto Castagna, Ivana Dusparic |
ICAART (2) | 2 |
| 2022 | Multi-Agent Deep Reinforcement Learning For Optimising Energy Efficiency of Fixed-Wing UAV Cellular Access PointsabstractUnmanned Aerial Vehicles (UAVs) promise to become an intrinsic part of next generation communications, as they can be deployed to provide wireless connectivity to ground users to supplement existing terrestrial networks. The majority of the existing research into the use of UAV access points for cellular coverage considers rotary-wing UAV designs (i.e. quadcopters). However, we expect fixed-wing UAVs to be more appropriate for connectivity purposes in scenarios where long flight times are necessary (such as for rural coverage), as fixed-wing UAVs rely on a more energy-efficient form of flight when compared to the rotary-wing design. As fixed-wing UAVs are typically incapable of hovering in place, their deployment optimisation involves optimising their individual flight trajectories in a way that allows them to deliver high quality service to the ground users in an energy-efficient manner. In this paper, we propose a multi-agent deep reinforcement learning approach to optimise the energy efficiency of fixed-wing UAV cellular access points while still allowing them to deliver high-quality service to users on the ground. In our decentralized approach, each UAV is equipped with a Dueling Deep Q-Network (DDQN) agent which can adjust the 3D trajectory of the UAV over a series of timesteps. By coordinating with their neighbours, the UAVs adjust their individual flight trajectories in a manner that optimises the total system energy efficiency. We benchmark the performance of our approach against a series of heuristic trajectory planning strategies, and demonstrate that our method can improve the system energy efficiency by as much as 70%. Boris Galkin, Omoniwa Babatunji, Ivana Dusparic |
ICC | 3 |
| 2022 | FedSA: Accelerating Intrusion Detection in Collaborative Environments with Federated Simulated AnnealingabstractFast identification of new network attack patterns is crucial for improving network security. Nevertheless, identifying an ongoing attack in a heterogeneous network is a non-trivial task. Federated learning emerges as a solution to collaborative training for an Intrusion Detection System (IDS). The federated learning-based IDS trains a global model using local machine learning models provided by federated participants without sharing local data. However, optimization challenges are intrinsic to federated learning. This paper proposes the Federated Simulated Annealing (FedSA) metaheuristic to select the hyperparameters and a subset of participants for each aggregation round in federated learning. FedSA optimizes hyperparameters linked to the global model convergence. The proposal reduces aggregation rounds and speeds up convergence. Thus, FedSA accelerates learning extraction from local models, requiring fewer IDS updates. The proposal assessment shows that the FedSA global model converges in less than ten communication rounds. The proposal requires up to 50% fewer aggregation rounds to achieve approximately 97% accuracy in attack detection than the conventional aggregation approach. Hélio N. Cunha Neto, Ivana Dusparic, Diogo M. F. Mattos, Natalia Castro Fernandes |
NetSoft | 2 |
| 2022 | Enabling Deep Reinforcement Learning on Energy Constrained Devices at the Edge of the NetworkabstractDeep Reinforcement Learning (DRL) solutions are becoming pervasive at the edge of the network as they enable autonomous decision-making in a dynamic environment. However, to be able to adapt to the ever-changing environment, the DRL solution implemented on an embedded device has to continue to occasionally take exploratory actions even after initial convergence. In other words, the device has to occasionally take random actions and update the value function, i.e., re-train the Artificial Neural Network (ANN), to ensure its performance remains optimal. Unfortunately, embedded devices often lack processing power and energy required to train the ANN. The energy aspect is particularly challenging when the edge device is powered only by a means of Energy Harvesting (EH). To overcome this problem, we propose a two-part algorithm in which the DRL process is trained at the sink. Then the weights of the fully trained underlying ANN are periodically transferred to the EH-powered embedded device taking actions. Using an EH-powered sensor, real-world measurements dataset, and optimizing for Age of Information (AoI) metric, we demonstrate that such a DRL solution can operate without any degradation in the performance, with only a few ANN updates per day. Jernej Hribar, Ivana Dusparic |
WCNC | 2 |
| 2022 | Timely and sustainable: Utilising correlation in status updates of battery-powered and energy-harvesting sensors using Deep Reinforcement LearningabstractIn a system with energy-constrained sensors, each transmitted observation comes at a price. The price is the energy the sensor expends to obtain and send a new measurement. The system has to ensure that sensors’ updates are timely, i.e., their updates represent the observed phenomenon accurately, enabling services to make informed decisions based on the information provided. If there are multiple sensors observing the same physical phenomenon, it is likely that their measurements are correlated in time and space. To take advantage of this correlation to reduce the energy use of sensors, in this paper we consider a system in which a gateway sets the intervals at which each sensor broadcasts its readings. We consider the presence of battery-powered sensors as well as sensors that rely on Energy Harvesting (EH) to replenish their energy. We propose a Deep Reinforcement Learning (DRL)-based scheduling mechanism that learns the appropriate update interval for each sensor, by considering the timeliness of the information collected measured through the Age of Information (AoI) metric, the spatial and temporal correlation between readings, and the energy capabilities of each sensor. We show that our proposed scheduler can achieve near-optimal performance in terms of the expected network lifetime. Jernej Hribar, Luiz A. DaSilva, Sheng Zhou 0001, Zhiyuan Jiang, Ivana Dusparic |
Comput. Commun. | 5 |
| 2021 | Analyse or Transmit: Utilising Correlation at the Edge with Deep Reinforcement LearningabstractMillions of sensors, cameras, meters, and other edge devices are deployed in networks to collect and analyse data. In many cases, such devices are powered only by Energy Harvesting (EH) and have limited energy available to analyse acquired data. When edge infrastructure is available, a device has a choice: to perform analysis locally or offload the task to other resource-rich devices such as cloudlet servers. However, such a choice carries a price in terms of consumed energy and accuracy. On the one hand, transmitting raw data can result in a higher energy cost in comparison to the required energy to process data locally. On the other hand, performing data analytics on servers can improve the task's accuracy. Additionally, due to the correlation between information sent by multiple devices, accuracy might not be affected if some edge devices decide to neither process nor send data and preserve energy instead. For such a scenario, we propose a Deep Reinforcement Learning (DRL) based solution capable of learning and adapting the policy to the time-varying energy arrival due to EH patterns. We leverage two datasets, one to model energy an EH device can collect and the other to model the correlation between cameras. Furthermore, we compare the proposed solution performance to three baseline policies. Our results show that we can increase accuracy by 15% in comparison to conventional approaches while preventing outages. Jernej Hribar, Ryoichi Shinkuma, George Iosifidis, Ivana Dusparic |
GLOBECOM | 4 |
| 2019 | Variational Policy Chaining for Lifelong Reinforcement LearningabstractWith increasing applications of reinforcement learning in real life problems, it is becoming essential that agents are able to update their knowledge continually. Lifelong learning approaches aim to enable agents to retain the knowledge they learn and to selectively transfer knowledge to new tasks. Recent techniques for lifelong reinforcement learning have shown great success in getting an agent to generalise over several tasks. However, scalability becomes an issue when agents learn numerous tasks, as each task's information must be remembered. To address this issue, this paper proposes the approach of Variational Policy Chaining (VPC) which enables a reinforcement learning agent to generalise effectively in a scalable manner when presented with continuous task updates, without storing multiple historic experiences. VPC uses Kullback-Leibler divergence based method to isolate the most common pieces of knowledge, and condenses the important knowledge into a single policy chain. We evaluate VPC in a GridWorld environment and compare it to vanilla policy gradient methods, showing that VPC's ability to reuse knowledge from previously encountered tasks reduces learning time in new tasks by up to 50%. Christopher Doyle, Maxime Guériau, Ivana Dusparic |
ICTAI | 3 |
| 2019 | Parallel Transfer Learning in Multi-Agent Systems: What, when and how to transfer?abstractMulti-agent Reinforcement Learning (RL) is frequently used in large-scale autonomous systems to learn the behaviours that best suit the system's operating environment. Learning can take a significant amount of time during which an RL system's performance is necessarily suboptimal. Transfer learning (TL), a method of reusing knowledge which has been gained in one task to improve the performance in another, has been used to speed up learning in single RL agent systems. TL requires learning on a source task to complete before transferring it to a target task, i.e., transfer is done offline. Parallel Transfer Learning (PTL), a technique which enables the source and target tasks to run concurrently, has been proposed to enable online transfers. However, the online selection of knowledge to be transferred, as well as online ways of integration of that knowledge on the receiving agents remain open issues. This paper proposes methods for selecting the knowledge to be transferred in PTL, frequency and size of transfers, and methods for knowledge integration into the target task. We evaluate the proposed approaches in two canonical RL examples: Mountain Car and Co-operative Predator Prey Pursuit. We show that PTL, similarly to RL, is highly sensitive to parameter selection and that suitable parameters differ per scenario. Adam Taylor, Ivana Dusparic, Maxime Guériau, Siobhán Clarke |
IJCNN | 2 |
| 2019 | An RL-based Approach to Improve Communication Performance and Energy Utilization in Fog-based IoTabstractRecent research has shown the potential of using available mobile fog devices (such as smartphones, drones, domestic and industrial robots) as relays to minimize communication outages between sensors and destination devices, where localized Internet-of-Things services (e.g., manufacturing process control, health and security monitoring) are delivered. However, these mobile relays deplete energy when they move and transmit to distant destinations. As such, power-control mechanisms and intelligent mobility of the relay devices are critical in improving communication performance and energy utilization. In this paper, we propose a Q-learning-based decentralized approach where each mobile fog relay agent (MFRA) is controlled by an autonomous agent which uses reinforcement learning to simultaneously improve communication performance and energy utilization. Each autonomous agent learns based on the feedback from the destination and its own energy levels whether to remain active and forward the message, or become passive for that transmission phase. We evaluate the approach by comparing with the centralized approach, and observe that with lesser number of MFRAs, our approach is able to ensure reliable delivery of data and reduce overall energy cost by 56.76% - 88.03%. Omoniwa Babatunji, Maxime Guériau, Ivana Dusparic |
WiMob | 3 |
| 2019 | Using Social Dependence to Enable Neighbourly Behaviour in Open Multi-Agent SystemsabstractAgents frequently collaborate to achieve a shared goal or to accomplish a task that they cannot do alone. However, collaboration is difficult in open multi-agent systems where agents share constrained resources to achieve both individual and shared goals. In current approaches to collaboration, agents are organised into disjoint groups and social reasoning is used to capture their capabilities when selecting a qualified set of collaborators. These approaches are not useful when agents are in multiple, overlapping groups; depend on each other when using shared resources; have multiple goals to achieve simultaneously; and have to share the overall costs and benefits. In this article, agents use social reasoning to enhance their understanding of other agents’ goals and their dependencies, and self-adaptive techniques to adapt their level of self-interest in a collaborative process, with a view to contributing to lowering shared costs or increasing shared benefits. This model aims at improving the extent to which agents’ goals are met while improving shared resource usage efficiency. For example, in a public transport system where each mode of transport has limited capacity, commuters will be enabled to make choices that avoid over-capacity in different modes, or in a smart energy grid with limited capacity, users can make choices as to when they increase their demand. The model simultaneously helps avoid overloading a shared resource while allowing users to achieve their own goals. The proposed model is evaluated in an open multi-agent system with 100 agents operating in multiple overlapping groups and sharing multiple constrained resources. The impact of agents’ varying levels of social dependencies, mobility, and their groups’ density on their individual and shared goal achievement is analysed. Fatemeh Golpayegani, Ivana Dusparic, Siobhán Clarke |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2018 | Adaptive Reward Allocation for Participatory SensingabstractParticipatory sensing is a paradigm through which mobile device users (or participants) collect and share data about their environments. The data captured by participants is typically submitted to an intermediary (the service provider) who will build a service based upon this data. For a participatory sensing system to attract the data submissions it requires, its users often need to be incentivized. However, as an environment is constantly changing (for example, an accident causing a buildup of traffic and elevated pollution levels), the value of a given data item to the service provider is likely to change significantly over time, and therefore an incentivization scheme must be able to adapt the rewards it offers in real‐time to match the environmental conditions and current participation rates, thereby optimizing the consumption of the service provider’s budget. This paper presents adaptive reward allocation (ARA), which uses the Lyapunov Optimization method to provide adaptive reward allocation that optimizes the consumption of the service provider’s budget. ARA is evaluated using a simulated participatory sensing environment with experimental results showing that the rewards offered to participants are adjusted so as to ensure that the data captured matches the dynamic changes occurring in the sensing environment and takes the response rate into account while also seeking to optimize budget consumption. Martin Connolly, Ivana Dusparic, George Iosifidis, Mélanie Bouroche |
Wirel. Commun. Mob. Comput. | 2 |
| 2017 | Prediction-Based Multi-Agent Reinforcement Learning in Inherently Non-Stationary EnvironmentsabstractMulti-agent reinforcement learning (MARL) is a widely researched technique for decentralised control in complex large-scale autonomous systems. Such systems often operate in environments that are continuously evolving and where agents’ actions are non-deterministic, so called inherently non-stationary environments. When there are inconsistent results for agents acting on such an environment, learning and adapting is challenging. In this article, we propose P-MARL, an approach that integrates prediction and pattern change detection abilities into MARL and thus minimises the effect of non-stationarity in the environment. The environment is modelled as a time-series, with future estimates provided using prediction techniques. Learning is based on the predicted environment behaviour, with agents employing this knowledge to improve their performance in realtime. We illustrate P-MARL’s performance in a real-world smart grid scenario, where the environment is heavily influenced by non-stationary power demand patterns from residential consumers. We evaluate P-MARL in three different situations, where agents’ action decisions are independent, simultaneous, and sequential. Results show that all methods outperform traditional MARL, with sequential P-MARL achieving best results. Andrei Marinescu, Ivana Dusparic, Siobhán Clarke |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2016 | Multi-agent Collaboration for Conflict Management in Residential Demand Response
Fatemeh Golpayegani, Ivana Dusparic, Adam Taylor, Siobhán Clarke |
Comput. Commun. | 2 |
| 2014 | A dynamic forecasting method for small scale residential electrical demandabstractSmall scale electrical demand forecasting is an emerging field motivated by the penetration of renewable energy sources and the growth of microgrids and virtual power plants. These advances pose more complex forecasting challenges compared to the already established large scale forecasting approaches. Current short term load forecasting methods deal with two types of day, normal and anomalous, which are predicted separately. Anomalous days are classified as such ahead of time, based on key calendar events such as public holidays. However, there are some anomalous days which are not always predictable on a day ahead basis. Due to unforeseen events, a seemingly normal day can progress towards an anomalous case causing high errors in prediction. We propose a new dynamic forecasting mechanism that actively monitors residential electrical demand along a forecasted day, and detects anomalous pattern changes from a previously predicted demand of the day. A self-organising map is employed to detect anomalous days as they progress. Once an anomaly is detected, a neural network based prediction system changes its input neurons according to a previously detected and recorded match found in a database of anomalous days, in order to accommodate the anomalous day prediction. Results are based on measured power demands recorded in Ireland from domestic smart-meters between 2009-2011, and focus on small scale residential electrical demands of up to 350 kWh. During anomalous days our dynamic prediction approach achieves forecasting results within 3.63% of the real load, down from the 7.37% obtained by the initial prediction algorithm and the 5.41% achieved by standalone re-prediction, without pattern matching. Andrei Marinescu, Ivana Dusparic, Colin Harris, Vinny Cahill, Siobhán Clarke |
IJCNN | 2 |
| 2014 | Accelerating Learning in multi-objective systems through Transfer LearningabstractLarge-scale, multi-agent systems are too complex for optimal control strategies to be known at design time and as a result good strategies must be learned at runtime. Learning in such systems, particularly those with multiple objectives, takes a considerable amount of time because of the size of the environment and dependencies between goals. Transfer Learning (TL) has been shown to reduce learning time in single-agent, single-objective applications. It is the process of sharing knowledge between two learning tasks called the source and target. The source is required to have been completed prior to the target task. This work proposes extending TL to multi-agent, multi-objective applications. To achieve this, an on-line version of TL called Parallel Transfer Learning (PTL) is presented. The issues involved in extending this algorithm to a multi-objective form are discussed. The effectiveness of this approach is evaluated in a smart grid scenario. When using PTL in this scenario learning is significantly accelerated. PTL achieves comparable performance to the base line in one third of the time. Adam Taylor, Ivana Dusparic, Edgar Galván López, Siobhán Clarke, Vinny Cahill |
IJCNN | 2 |
| 2012 | Autonomic multi-policy optimization in pervasive systems: Overview and evaluationabstractThis article describes Distributed W-Learning (DWL), a reinforcement learning-based algorithm for collaborative agent-based optimization of pervasive systems. DWL supports optimization towards multiple heterogeneous policies and addresses the challenges arising from the heterogeneity of the agents that are charged with implementing them. DWL learns and exploits the dependencies between agents and between policies to improve overall system performance. Instead of always executing the locally-best action, agents learn how their actions affect their immediate neighbors and execute actions suggested by neighboring agents if their importance exceeds the local action's importance when scaled using a predefined or learned collaboration coefficient. We have evaluated DWL in a simulation of an Urban Traffic Control (UTC) system, a canonical example of the large-scale pervasive systems that we are addressing. We show that DWL outperforms widely deployed fixed-time and simple adaptive UTC controllers under a variety of traffic loads and patterns. Our results also confirm that enabling collaboration between agents is beneficial as is the ability for agents to learn the degree to which it is appropriate for them to collaborate. These results suggest that DWL is a suitable basis for optimization in other large-scale systems with similar characteristics. Ivana Dusparic, Vinny Cahill |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2009 | Using Reinforcement Learning for Multi-policy Optimization in Decentralized Autonomic Systems - An Experimental Evaluation
Ivana Dusparic, Vinny Cahill |
ATC | 1 |