VLDB 2026 Research / reviewers in the wild / expert
Frits de Nijs
dblp:116/2730
· DBLP profile ↗
12ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-4466-2447ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CrewAId: Interactive Optimisation for Human-In-The-Loop Crew Rostering and RerosteringabstractConstraint programming technology allows optimisation experts to solve a broad category of personnel rostering problems, such as nurse rostering, airline crew rostering or retail worker scheduling. However, for problem domain experts to use this technology, the optimisation system must bridge the gap for users to easily explore solutions and influence constraints. Working with our energy industry partner for several years, we identified rostering problems involving multi-skilled shift workers present on site for extended periods. Their existing workflow for handling rostering (crew allocation), and rerostering (dealing with inevitable employee absences) and for time-limited formation of dedicated maintenance crews is labour intensive and complex, requiring in-depth knowledge of personnel files and skill competencies. To address this, we propose an interactive decision support system for crew rostering and rerostering, currently being deployed by our industry partner, that provides interactive tools for domain experts to perform exploration, validation, and conflict recovery. Matthias Klapperstück, Frits de Nijs, Ilankaikone Senthooran, Matteo Miceli, Michael Wybrow |
CP | 2 |
| 2026 | Constraint-Aware Self-Supervised Learning for Edge SelectionabstractMany edge-selection problems, such as the Traveling Salesman Problem and Orienteering Problem, are NP-hard, making them expensive to solve with exact methods and challenging to address with hand-crafted heuristics. Learning-based approaches provide an efficient alternative, while self-supervised methods avoid costly solution labels. However, existing approaches often still rely on heavy post-processing or narrow problem-specific designs. We propose a reusable self-supervised framework for edge-selection optimization that learns directly from unlabeled instances. The framework uses differentiable surrogate objectives and feasibility-driven penalties to encourage the model to learn feasibility-aware solution structure during training. To support efficient inference, we introduce a lightweight graph architecture centered on a cost-attention convolution, where edge costs and feasibility information directly shape message passing. Experiments on three problem families demonstrate strong solution quality and efficient inference across diverse edge-selection settings. Xinda Zheng, Frits de Nijs, Edward Lam 0001 |
CP | 2 |
| 2025 | Using Assistance Rewards Without Introducing Bias: Overcoming Sparse Rewards in Multi-Agent Reinforcement Learning
Bernd Meyer 0001, Frits de Nijs |
AAMAS | 3 |
| 2025 | Learning to cooperate against ensembles of diverse opponentsabstractAbstract The emergence of cooperation in decentralized multi-agent systems is challenging; naive implementations of learning algorithms typically fail to converge or converge to equilibria without cooperation. Opponent modeling techniques, combined with reinforcement learning, have been successful in promoting cooperation, but face challenges when other agents are plentiful or anonymous. We envision environments in which agents face a sequence of interactions with different and heterogeneous agents. Inspired by models of evolutionary game theory, we introduce RL agents that forgo explicit modeling of others. Instead, they augment their reward signal by considering how to best respond to others assumed to be rational against their own strategy. This technique not only scales well in environments with many agents, but can also outperform opponent modeling techniques across a range of cooperation games. Agents that use the algorithm we propose can successfully maintain and establish cooperation when playing against an ensemble of diverse agents. This finding is robust across different kinds of games and can also be shown not to disadvantage agents in purely competitive interactions. While cooperation in pairwise settings is foundational, interactions across large groups of diverse agents are likely to be the norm in future applications where cooperation is an emergent property of agent design, rather than a design goal at the system level. The algorithm we propose here is a simple and scalable step in this direction. Isuri Perera, Frits de Nijs, Julian García |
Neural Comput. Appl. | 2 |
| 2023 | Exploring Hydrogen Supply/Demand Networks: Modeller and Domain Expert Views
Matthias Klapperstück, Frits de Nijs, Ilankaikone Senthooran, Jack Lee-Kopij, Maria Garcia de la Banda, Michael Wybrow |
CP | 2 |
| 2021 | Evaluating Meta-Reinforcement Learning through a HVAC Control Benchmark (Student Abstract)abstractMeta-Reinforcement Learning (RL) algorithms promise to leverage prior task experience to quickly learn new unseen tasks. Unfortunately, evaluating meta-RL algorithms is complicated by a lack of suitable benchmarks. In this paper we propose adapting a challenging real-world heating, ventilation and air-conditioning (HVAC) control benchmark for meta-RL. Unlike existing benchmark problems, HVAC control has a broader task distribution, and sources of exogenous stochasticity from price and weather predictions which can be shared across task definitions. This can enable greater differentiation between the performance of current meta-RL approaches, and open the way for future research into algorithms that can adapt to entirely new tasks not sampled from the current task distribution. Yashvir S. Grewal, Frits de Nijs, Sarah Goodwin |
AAAI | 2 |
| 2021 | Constrained Multiagent Markov Decision Processes: a Taxonomy of Problems and AlgorithmsabstractIn domains such as electric vehicle charging, smart distribution grids and autonomous warehouses, multiple agents share the same resources. When planning the use of these resources, agents need to deal with the uncertainty in these domains. Although several models and algorithms for such constrained multiagent planning problems under uncertainty have been proposed in the literature, it remains unclear when which algorithm can be applied. In this survey we conceptualize these domains and establish a generic problem class based on Markov decision processes. We identify and compare the conditions under which algorithms from the planning literature for problems in this class can be applied: whether constraints are soft or hard, whether agents are continuously connected, whether the domain is fully observable, whether a constraint is momentarily (instantaneous) or on a budget, and whether the constraint is on a single resource or on multiple. Further we discuss the advantages and disadvantages of these algorithms. We conclude by identifying open problems that are directly related to the conceptualized domains, as well as in adjacent research areas. Frits de Nijs, Erwin Walraven, Mathijs de Weerdt, Matthijs T. J. Spaan |
J. Artif. Intell. Res. | 1 |
| 2020 | Large Neighborhood Search for Temperature Control with Demand Response
Edward Lam 0001, Frits de Nijs, Peter J. Stuckey, Donald Azuatalam, Ariel Liebman |
CP | 2 |
| 2018 | Preallocation and Planning Under Stochastic Resource ConstraintsabstractResource constraints frequently complicate multi-agent planning problems. Existing algorithms for resource-constrained, multi-agent planning problems rely on the assumption that the constraints are deterministic. However, frequently resource constraints are themselves subject to uncertainty from external influences. Uncertainty about constraints is especially challenging when agents must execute in an environment where communication is unreliable, making on-line coordination difficult. In those cases, it is a significant challenge to find coordinated allocations at plan time depending on availability at run time. To address these limitations, we propose to extend algorithms for constrained multi-agent planning problems to handle stochastic resource constraints. We show how to factorize resource limit uncertainty and use this to develop novel algorithms to plan policies for stochastic constraints. We evaluate the algorithms on a search-and-rescue problem and on a power-constrained planning domain where the resource constraints are decided by nature. We show that plans taking into account all potential realizations of the constraint obtain significantly better utility than planning for the expectation, while causing fewer constraint violations. Frits de Nijs, Matthijs T. J. Spaan, Mathijs de Weerdt |
AAAI | 1 |
| 2017 | Bounding the Probability of Resource Constraint Violations in Multi-Agent MDPsabstractMulti-agent planning problems with constraints on global resource consumption occur in several domains. Existing algorithms for solving Multi-agent Markov Decision Processes can compute policies that meet a resource constraint in expectation, but these policies provide no guarantees on the probability that a resource constraint violation will occur. We derive a method to bound constraint violation probabilities using Hoeffding's inequality. This method is applied to two existing approaches for computing policies satisfying constraints: the Constrained MDP framework and a Column Generation approach. We also introduce an algorithm to adaptively relax the bound up to a given maximum violation tolerance. Experiments on a hard toy problem show that the resulting policies outperform static optimal resource allocations to an arbitrary level. By testing the algorithms on more realistic planning domains from the literature, we demonstrate that the adaptive bound is able to efficiently trade off violation probability with expected value, outperforming state-of-the-art planners. Frits de Nijs, Erwin Walraven, Mathijs de Weerdt, Matthijs T. J. Spaan |
AAAI | 1 |
| 2016 | Decoupling a Resource Constraint Through Fictitious Play in Multi-Agent Sequential Decision MakingabstractWhen multiple independent agents use a limited shared resource, they need to coordinate and thereby their planning problems become coupled. We present a resource assignment strategy that decouples agents using marginal utility cost, allowing them to plan individually. We show that agents converge to an expected cost curve by keeping a history of plans, inspired by fictitious play. This performs slightly better than a state-of-the-art best-response approach and is significantly more scalable than a preallocation Mixed-Integer Linear Programming formulation, providing a good trade-off between performance and quality. Frits de Nijs, Matthijs T. J. Spaan, Mathijs de Weerdt |
ECAI | 1 |
| 2015 | Best-Response Planning of Thermostatically Controlled Loads under Power ConstraintsabstractRenewable power sources such as wind and solar are inflexible in their energy production, which requires demand to rapidly follow supply in order to maintain energy balance. Promising controllable demands are air-conditioners and heat pumps which use electric energy to maintain a temperature at a setpoint. Such Thermostatically Controlled Loads (TCLs) have been shown to be able to follow a power curve using reactive control. In this paper we investigate the use of planning under uncertainty to pro-actively control an aggregation of TCLs to overcome temporary grid imbalance. We present a formal definition of the planning problem under consideration, which we model using the Multi-Agent Markov Decision Process (MMDP) framework. Since we are dealing with hundreds of agents, solving the resulting MMDPs directly is intractable. Instead, we propose to decompose the problem by decoupling the interactions through arbitrage. Decomposition of the problem means relaxing the joint power consumption constraint, which means that joining the plans together can cause overconsumption. Arbitrage acts as a conflict resolution mechanism during policy execution, using the future expected value of policies to determine which TCLs should receive the available energy. We experimentally compare several methods to plan with arbitrage, and conclude that a best response-like mechanism is a scalable approach that returns near-optimal solutions. Frits de Nijs, Matthijs T. J. Spaan, Mathijs de Weerdt |
AAAI | 1 |