Frits de Nijs

dblp:116/2730 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-4466-2447ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021
YearPublicationVenuePosition
2026 CrewAId: Interactive Optimisation for Human-In-The-Loop Crew Rostering and Rerostering
abstract
Constraint programming technology allows optimisation experts to solve a broad category of personnel rostering problems, such as nurse rostering, airline crew rostering or retail worker scheduling. However, for problem domain experts to use this technology, the optimisation system must bridge the gap for users to easily explore solutions and influence constraints. Working with our energy industry partner for several years, we identified rostering problems involving multi-skilled shift workers present on site for extended periods. Their existing workflow for handling rostering (crew allocation), and rerostering (dealing with inevitable employee absences) and for time-limited formation of dedicated maintenance crews is labour intensive and complex, requiring in-depth knowledge of personnel files and skill competencies. To address this, we propose an interactive decision support system for crew rostering and rerostering, currently being deployed by our industry partner, that provides interactive tools for domain experts to perform exploration, validation, and conflict recovery.
Matthias Klapperstück, Frits de Nijs, Ilankaikone Senthooran, Matteo Miceli, Michael Wybrow
CP2
2026 Constraint-Aware Self-Supervised Learning for Edge Selection
abstract
Many edge-selection problems, such as the Traveling Salesman Problem and Orienteering Problem, are NP-hard, making them expensive to solve with exact methods and challenging to address with hand-crafted heuristics. Learning-based approaches provide an efficient alternative, while self-supervised methods avoid costly solution labels. However, existing approaches often still rely on heavy post-processing or narrow problem-specific designs. We propose a reusable self-supervised framework for edge-selection optimization that learns directly from unlabeled instances. The framework uses differentiable surrogate objectives and feasibility-driven penalties to encourage the model to learn feasibility-aware solution structure during training. To support efficient inference, we introduce a lightweight graph architecture centered on a cost-attention convolution, where edge costs and feasibility information directly shape message passing. Experiments on three problem families demonstrate strong solution quality and efficient inference across diverse edge-selection settings.
Xinda Zheng, Frits de Nijs, Edward Lam 0001
CP2
2025 Using Assistance Rewards Without Introducing Bias: Overcoming Sparse Rewards in Multi-Agent Reinforcement Learning
Bernd Meyer 0001, Frits de Nijs
AAMAS3
2025 Learning to cooperate against ensembles of diverse opponents
abstract
Abstract The emergence of cooperation in decentralized multi-agent systems is challenging; naive implementations of learning algorithms typically fail to converge or converge to equilibria without cooperation. Opponent modeling techniques, combined with reinforcement learning, have been successful in promoting cooperation, but face challenges when other agents are plentiful or anonymous. We envision environments in which agents face a sequence of interactions with different and heterogeneous agents. Inspired by models of evolutionary game theory, we introduce RL agents that forgo explicit modeling of others. Instead, they augment their reward signal by considering how to best respond to others assumed to be rational against their own strategy. This technique not only scales well in environments with many agents, but can also outperform opponent modeling techniques across a range of cooperation games. Agents that use the algorithm we propose can successfully maintain and establish cooperation when playing against an ensemble of diverse agents. This finding is robust across different kinds of games and can also be shown not to disadvantage agents in purely competitive interactions. While cooperation in pairwise settings is foundational, interactions across large groups of diverse agents are likely to be the norm in future applications where cooperation is an emergent property of agent design, rather than a design goal at the system level. The algorithm we propose here is a simple and scalable step in this direction.
Isuri Perera, Frits de Nijs, Julian García
Neural Comput. Appl.2
2023 Exploring Hydrogen Supply/Demand Networks: Modeller and Domain Expert Views
Matthias Klapperstück, Frits de Nijs, Ilankaikone Senthooran, Jack Lee-Kopij, Maria Garcia de la Banda, Michael Wybrow
CP2
2021 Evaluating Meta-Reinforcement Learning through a HVAC Control Benchmark (Student Abstract)
abstract
Meta-Reinforcement Learning (RL) algorithms promise to leverage prior task experience to quickly learn new unseen tasks. Unfortunately, evaluating meta-RL algorithms is complicated by a lack of suitable benchmarks. In this paper we propose adapting a challenging real-world heating, ventilation and air-conditioning (HVAC) control benchmark for meta-RL. Unlike existing benchmark problems, HVAC control has a broader task distribution, and sources of exogenous stochasticity from price and weather predictions which can be shared across task definitions. This can enable greater differentiation between the performance of current meta-RL approaches, and open the way for future research into algorithms that can adapt to entirely new tasks not sampled from the current task distribution.
Yashvir S. Grewal, Frits de Nijs, Sarah Goodwin
AAAI2
2021 Constrained Multiagent Markov Decision Processes: a Taxonomy of Problems and Algorithms
abstract
In domains such as electric vehicle charging, smart distribution grids and autonomous warehouses, multiple agents share the same resources. When planning the use of these resources, agents need to deal with the uncertainty in these domains. Although several models and algorithms for such constrained multiagent planning problems under uncertainty have been proposed in the literature, it remains unclear when which algorithm can be applied. In this survey we conceptualize these domains and establish a generic problem class based on Markov decision processes. We identify and compare the conditions under which algorithms from the planning literature for problems in this class can be applied: whether constraints are soft or hard, whether agents are continuously connected, whether the domain is fully observable, whether a constraint is momentarily (instantaneous) or on a budget, and whether the constraint is on a single resource or on multiple. Further we discuss the advantages and disadvantages of these algorithms. We conclude by identifying open problems that are directly related to the conceptualized domains, as well as in adjacent research areas.
Frits de Nijs, Erwin Walraven, Mathijs de Weerdt, Matthijs T. J. Spaan
J. Artif. Intell. Res.1
2020 Large Neighborhood Search for Temperature Control with Demand Response
Edward Lam 0001, Frits de Nijs, Peter J. Stuckey, Donald Azuatalam, Ariel Liebman
CP2
2018 Preallocation and Planning Under Stochastic Resource Constraints
abstract
Resource constraints frequently complicate multi-agent planning problems. Existing algorithms for resource-constrained, multi-agent planning problems rely on the assumption that the constraints are deterministic. However, frequently resource constraints are themselves subject to uncertainty from external influences. Uncertainty about constraints is especially challenging when agents must execute in an environment where communication is unreliable, making on-line coordination difficult. In those cases, it is a significant challenge to find coordinated allocations at plan time depending on availability at run time. To address these limitations, we propose to extend algorithms for constrained multi-agent planning problems to handle stochastic resource constraints. We show how to factorize resource limit uncertainty and use this to develop novel algorithms to plan policies for stochastic constraints. We evaluate the algorithms on a search-and-rescue problem and on a power-constrained planning domain where the resource constraints are decided by nature. We show that plans taking into account all potential realizations of the constraint obtain significantly better utility than planning for the expectation, while causing fewer constraint violations.
Frits de Nijs, Matthijs T. J. Spaan, Mathijs de Weerdt
AAAI1
2017 Bounding the Probability of Resource Constraint Violations in Multi-Agent MDPs
abstract
Multi-agent planning problems with constraints on global resource consumption occur in several domains. Existing algorithms for solving Multi-agent Markov Decision Processes can compute policies that meet a resource constraint in expectation, but these policies provide no guarantees on the probability that a resource constraint violation will occur. We derive a method to bound constraint violation probabilities using Hoeffding's inequality. This method is applied to two existing approaches for computing policies satisfying constraints: the Constrained MDP framework and a Column Generation approach. We also introduce an algorithm to adaptively relax the bound up to a given maximum violation tolerance. Experiments on a hard toy problem show that the resulting policies outperform static optimal resource allocations to an arbitrary level. By testing the algorithms on more realistic planning domains from the literature, we demonstrate that the adaptive bound is able to efficiently trade off violation probability with expected value, outperforming state-of-the-art planners.
Frits de Nijs, Erwin Walraven, Mathijs de Weerdt, Matthijs T. J. Spaan
AAAI1
2016 Decoupling a Resource Constraint Through Fictitious Play in Multi-Agent Sequential Decision Making
abstract
When multiple independent agents use a limited shared resource, they need to coordinate and thereby their planning problems become coupled. We present a resource assignment strategy that decouples agents using marginal utility cost, allowing them to plan individually. We show that agents converge to an expected cost curve by keeping a history of plans, inspired by fictitious play. This performs slightly better than a state-of-the-art best-response approach and is significantly more scalable than a preallocation Mixed-Integer Linear Programming formulation, providing a good trade-off between performance and quality.
Frits de Nijs, Matthijs T. J. Spaan, Mathijs de Weerdt
ECAI1
2015 Best-Response Planning of Thermostatically Controlled Loads under Power Constraints
abstract
Renewable power sources such as wind and solar are inflexible in their energy production, which requires demand to rapidly follow supply in order to maintain energy balance. Promising controllable demands are air-conditioners and heat pumps which use electric energy to maintain a temperature at a setpoint. Such Thermostatically Controlled Loads (TCLs) have been shown to be able to follow a power curve using reactive control. In this paper we investigate the use of planning under uncertainty to pro-actively control an aggregation of TCLs to overcome temporary grid imbalance. We present a formal definition of the planning problem under consideration, which we model using the Multi-Agent Markov Decision Process (MMDP) framework. Since we are dealing with hundreds of agents, solving the resulting MMDPs directly is intractable. Instead, we propose to decompose the problem by decoupling the interactions through arbitrage. Decomposition of the problem means relaxing the joint power consumption constraint, which means that joining the plans together can cause overconsumption. Arbitrage acts as a conflict resolution mechanism during policy execution, using the future expected value of policies to determine which TCLs should receive the available energy. We experimentally compare several methods to plan with arbitrage, and conclude that a best response-like mechanism is a scalable approach that returns near-optimal solutions.
Frits de Nijs, Matthijs T. J. Spaan, Mathijs de Weerdt
AAAI1