Aske Plaat

dblp:53/5607 · DBLP profile ↗
← Back
56ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-7202-3322ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 3 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 13 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Systems, architecture and hardware · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 6Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Software engineering, systems software and programming languages · 3Security and privacy · 1Theory of computation · 1
YearPublicationVenuePosition
2026 How Does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
abstract
Chain‑of‑thought (CoT) prompting boosts Large Language Models accuracy on multi‑step tasks, yet whether the generated ``thoughts'' reflect the true internal reasoning process is unresolved. We present the first feature‑level causal study of CoT faithfulness. Combining sparse autoencoders with activation patching, we extract monosemantic features from Pythia‑70M and Pythia‑2.8B while they tackle GSM8K math problems under CoT and plain (noCoT) prompting. Swapping a small set of CoT‑reasoning features into a noCoT run raises answer log‑probabilities significantly in the 2.8B model, but has no reliable effect in 70M, revealing a clear contrast for these two scales. CoT also leads to significantly higher activation sparsity and feature interpretability scores in the larger model, signalling more modular internal computation. For example, the model's confidence in generating correct answers improves from 1.2 to 4.3. We introduce patch‑curves and random‑feature patching baselines, showing that useful CoT information is not only present in the top-K patches but widely distributed. Overall, our results indicate that CoT can induce more interpretable internal structures in high-capacity LLMs, validating its role as a structured prompting method.
Aske Plaat, Niki van Stein
AAAI2
2026 Assessing Reproducibility in Evolutionary Computation: A Case Study using Human- and LLM-based Assessment
abstract
Reproducibility is an important requirement in evolutionary computation, where results largely depend on computational experiments. In practice, reproducibility relies on how algorithms, experimental protocols, and artifacts are documented and shared. Despite growing awareness, there is still limited empirical evidence on the actual reproducibility levels of published work in the field. In this paper, we study the reproducibility practices in papers published in the Evolutionary Combinatorial Optimization and Metaheuristics track of the Genetic and Evolutionary Computation Conference over a ten-year period. We introduce a structured reproducibility checklist and apply it through a systematic manual assessment of the selected corpus. In addition, we propose RECAP (REproducibility Checklist Automation Pipeline), an LLM-based system that automatically evaluates reproducibility signals from paper text and associated code repositories. Our analysis shows that papers achieve an average completeness score of 0.62, and that 36.90% of them provide additional material beyond the manuscript itself. We demonstrate that automated assessment is feasible: RECAP achieves substantial agreement with human evaluators (Cohen's κ of 0.67). Together, these results highlight persistent gaps in reproducibility reporting and suggest that automated tools can effectively support large-scale, systematic monitoring of reproducibility practices.
Francesca Da Ros, Tarik Zaciragic, Aske Plaat, Thomas Bäck, Niki van Stein
GECCO3
2025 EconoJax: A Fast & Scalable Economic Simulation in JAX
Koen Ponse, Aske Plaat, Niki van Stein, Thomas M. Moerland
AAMAS2
2025 Baba Is LLM: Reasoning in a Game with Dynamic Rules
Fien van Wetten, Aske Plaat, Max J. van Duijn
IJCCI (1)2
2025 LLM-enhanced Interactions in Human-Robot Collaborative Drawing with Older Adults
abstract
The goal of this study is to identify factors that support and enhance older adults’ creative experiences in human-robot co-creativity. Because the research into the use of robots for creativity support with older adults remains underexplored, we carried out an exploratory case study. We took a participatory approach and collaborated with professional art educators to design a course "Drawing with Robots" for adults aged 65 and over. The course featured human-human and human-robot drawing activities with various types of robots. We observed collaborative drawing interactions, interviewed participants on their experiences, and analyzed collected data. Findings show that participants preferred acting as curators, evaluating creative suggestions from the robot in a teacher or coach role. When we enhanced a robot with a multimodal Large Language Model (LLM), participants appreciated its spoken dialogue capabilities. They reported however, that the robot’s feedback sometimes lacked an understanding of the context, and sensitivity to their artistic goals and preferences. Our findings highlight the potential of LLM-enhanced robots to support creativity and offer future directions for advancing human-robot co-creativity with older adults.
Marianne Bossema, Somaya Ben Allouch, Aske Plaat, Rob Saunders
RO-MAN3
2025 Agentic Large Language Models, a Survey
abstract
Background: There is great interest in agentic LLMs, large language models that act as agents. Objectives: We review the growing body of work in this area and provide a research agenda. Methods: Agentic LLMs are LLMs that (1) reason, (2) act, and (3) interact. We organize the literature according to these three categories. Results: The research in the first category focuses on reasoning, reflection, and retrieval, aiming to improve decision making; the second category focuses on action models, robots, and tools, aiming for agents that act as useful assistants; the third category focuses on multi-agent systems, aiming for collaborative task solving and simulating interaction to study emergent social behavior. We find that works mutually benefit from results in other categories: retrieval enables tool use, reflection improves multi-agent collaboration, and reasoning benefits all categories. Conclusions: We discuss applications of agentic LLMs and provide an agenda for further research. Important applications are in medical diagnosis, logistics and financial market analysis. Meanwhile, self-reflective agents playing roles and interacting with one another augment the process of scientific research itself. Further, agentic LLMs provide a solution for the problem of LLMs running out of training data: inference-time behavior generates new training states, such that LLMs can keep learning without needing ever larger datasets. We note that there is risk associated with LLM assistants taking action in the real world—safety, liability and security are open problems—while agentic LLMs are also likely to benefit society.
Aske Plaat, Max J. van Duijn, Niki van Stein, Mike Preuss, Peter van der Putten, Kees Joost Batenburg
J. Artif. Intell. Res.1
2024 A Hybrid Intelligence Method for Argument Mining
abstract
Large-scale survey tools enable the collection of citizen feedback in opinion corpora. Extracting the key arguments from a large and noisy set of opinions helps in understanding the opinions quickly and accurately. Fully automated methods can extract arguments but (1) require large labeled datasets that induce large annotation costs and (2) work well for known viewpoints, but not for novel points of view. We propose HyEnA, a hybrid (human + AI) method for extracting arguments from opinionated texts, combining the speed of automated processing with the understanding and reasoning capabilities of humans. We evaluate HyEnA on three citizen feedback corpora. We find that, on the one hand, HyEnA achieves higher coverage and precision than a state-of-the-art automated method when compared to a common set of diverse opinions, justifying the need for human insight. On the other hand, HyEnA requires less human effort and does not compromise quality compared to (fully manual) expert analysis, demonstrating the benefit of combining human and artificial intelligence.
Michiel van der Meer, Enrico Liscio, Catholijn M. Jonker, Aske Plaat, Piek Vossen, Pradeep K. Murukannaiah
J. Artif. Intell. Res.4
2024 Subspace Adaptation Prior for Few-Shot Learning
abstract
Abstract Gradient-based meta-learning techniques aim to distill useful prior knowledge from a set of training tasks such that new tasks can be learned more efficiently with gradient descent. While these methods have achieved successes in various scenarios, they commonly adapt all parameters of trainable layers when learning new tasks. This neglects potentially more efficient learning strategies for a given task distribution and may be susceptible to overfitting, especially in few-shot learning where tasks must be learned from a limited number of examples. To address these issues, we propose Subspace Adaptation Prior (SAP), a novel gradient-based meta-learning algorithm that jointly learns good initialization parameters (prior knowledge) and layer-wise parameter subspaces in the form of operation subsets that should be adaptable. In this way, SAP can learn which operation subsets to adjust with gradient descent based on the underlying task distribution, simultaneously decreasing the risk of overfitting when learning new tasks. We demonstrate that this ability is helpful as SAP yields superior or competitive performance in few-shot image classification settings (gains between 0.1% and 3.9% in accuracy). Analysis of the learned subspaces demonstrates that low-dimensional operations often yield high activation strengths, indicating that they may be important for achieving good few-shot learning performance. For reproducibility purposes, we publish all our research code publicly.
Mike Huisman, Aske Plaat, Jan N. van Rijn
Mach. Learn.2
2024 Understanding transfer learning and gradient-based meta-learning techniques
abstract
Abstract Deep neural networks can yield good performance on various tasks but often require large amounts of data to train them. Meta-learning received considerable attention as one approach to improve the generalization of these networks from a limited amount of data. Whilst meta-learning techniques have been observed to be successful at this in various scenarios, recent results suggest that when evaluated on tasks from a different data distribution than the one used for training, a baseline that simply finetunes a pre-trained network may be more effective than more complicated meta-learning techniques such as MAML, which is one of the most popular meta-learning techniques. This is surprising as the learning behaviour of MAML mimics that of finetuning: both rely on re-using learned features. We investigate the observed performance differences between finetuning, MAML, and another meta-learning technique called Reptile, and show that MAML and Reptile specialize for fast adaptation in low-data regimes of similar data distribution as the one used for training. Our findings show that both the output layer and the noisy training conditions induced by data scarcity play important roles in facilitating this specialization for MAML. Lastly, we show that the pre-trained features as obtained by the finetuning baseline are more diverse and discriminative than those learned by MAML and Reptile. Due to this lack of diversity and distribution specialization, MAML and Reptile may fail to generalize to out-of-distribution tasks whereas finetuning can fall back on the diversity of the learned features.
Mike Huisman, Aske Plaat, Jan N. van Rijn
Mach. Learn.2
2023 Fine-grained Affective Processing Capabilities Emerging from Large Language Models
abstract
Large language models, in particular generative pre-trained transformers (GPTs), show impressive results on a wide variety of language-related tasks. In this paper, we explore ChatGPT’s zero-shot ability to perform affective computing tasks using prompting alone. We show that ChatGPT a) performs meaningful sentiment analysis in the Valence, Arousal and Dominance dimensions, b) has meaningful emotion representations in terms of emotion categories and these affective dimensions, and c) can perform basic appraisal-based emotion elicitation of situations based on a prompt-based computational implementation of the OCC appraisal model. These findings are highly relevant: First, they show that the ability to solve complex affect processing tasks emerges from language-based token prediction trained on extensive data sets. Second, they show the potential of large language models for simulating, processing and analyzing human emotions, which has important implications for various applications such as sentiment analysis, socially interactive agents, and social robotics.
Joost Broekens, Bernhard Hilpert, Suzan Verberne, Kim Baraka, Patrick Gebhard, Aske Plaat
ACII6
2023 Continuous Episodic Control
abstract
Non-parametric episodic memory can be used to quickly latch onto high-rewarded experience in reinforcement learning tasks. In contrast to parametric deep reinforcement learning approaches in which reward signals need to be back-propagated slowly, these methods only need to discover the solution once, and may then repeatedly solve the task. However, episodic control solutions are stored in discrete tables, and this approach has so far only been applied to discrete action space problems. Therefore, this paper introduces Continuous Episodic Control (CEC), a novel non-parametric episodic memory algorithm for sequential decision making in problems with a continuous action space. Results on several sparse-reward continuous control environments show that our proposed method learns faster than state-of-the-art model-free RL and memory-augmented RL algorithms, while maintaining good long-run performance as well. In short, CEC can be a fast approach for learning in continuous control tasks.1
Zhao Yang 0003, Thomas M. Moerland, Mike Preuss, Aske Plaat
CoG4
2023 Two-Memory Reinforcement Learning
abstract
While deep reinforcement learning has shown important empirical success, it tends to learn relatively slow due to slow propagation of rewards information and slow update of parametric neural networks. Non-parametric episodic memory, on the other hand, provides a faster learning alternative that does not require representation learning and uses maximum episodic return as state-action values for action selection. Episodic memory and reinforcement learning both have their own strengths and weaknesses. Notably, humans can leverage multiple memory systems concurrently during learning and benefit from all of them. In this work, we propose a method called Two-Memory reinforcement learning agent (2M) that combines episodic memory and reinforcement learning to distill both of their strengths. The 2M agent exploits the speed of the episodic memory part and the optimality and the generalization capacity of the reinforcement learning part to complement each other. Our experiments demonstrate that the 2M agent is more data efficient and outperforms both pure episodic memory and pure reinforcement learning, as well as a state-of-the-art memory-augmented RL agent. Moreover, the proposed approach provides a general framework that can be used to combine any episodic memory agent with other off-policy reinforcement learning algorithms.1
Zhao Yang 0003, Thomas M. Moerland, Mike Preuss, Aske Plaat
CoG4
2023 First Go, then Post-Explore: The Benefits of Post-Exploration in Intrinsic Motivation
abstract
Computer Systems, Imagery and Media
Zhao Yang 0003, Thomas M. Moerland, Mike Preuss, Aske Plaat
ICAART (2)4
2023 Human-Robot Co-creativity: A Scoping Review : Informing a Research Agenda for Human-Robot Co-Creativity with Older Adults
abstract
This review is the first step in a long-term research project exploring how social robotics and AI-generated content can contribute to the creative experiences of older adults, with a focus on collaborative drawing and painting. We systematically searched and selected literature on human-robot co-creativity, and analyzed articles to identify methods and strategies for researching co-creative robotics. We found that none of the studies involved older adults, which shows the gap in the literature for this often involved participant group in robotics research. The analyzed literature provides valuable insights into the design of human-robot co-creativity and informs a research agenda to further investigate the topic with older adults. We argue that future research should focus on ecological and developmental perspectives on creativity, on how system behavior can be aligned with the values of older adults, and on the system structures that support this best.
Marianne Bossema, Somaya Ben Allouch, Aske Plaat, Rob Saunders
RO-MAN3
2023 Are LSTMs good few-shot learners?
abstract
Abstract Deep learning requires large amounts of data to learn new tasks well, limiting its applicability to domains where such data is available. Meta-learning overcomes this limitation by learning how to learn. Hochreiter et al. (International conference on artificial neural networks, Springer, 2001) showed that an LSTM trained with backpropagation across different tasks is capable of meta-learning. Despite promising results of this approach on small problems, and more recently, also on reinforcement learning problems, the approach has received little attention in the supervised few-shot learning setting. We revisit this approach and test it on modern few-shot learning benchmarks. We find that LSTM, surprisingly, outperform the popular meta-learning technique MAML on a simple few-shot sine wave regression benchmark, but that LSTM, expectedly, fall short on more complex few-shot image classification benchmarks. We identify two potential causes and propose a new method called Outer Product LSTM (OP-LSTM) that resolves these issues and displays substantial performance gains over the plain LSTM. Compared to popular meta-learning baselines, OP-LSTM yields competitive performance on within-domain few-shot image classification, and performs better in cross-domain settings by 0.5–1.9% in accuracy score. While these results alone do not set a new state-of-the-art, the advances of OP-LSTM are orthogonal to other advances in the field of meta-learning, yield new insights in how LSTM work in image classification, allowing for a whole range of new research directions. For reproducibility purposes, we publish all our research code publicly.
Mike Huisman, Thomas M. Moerland, Aske Plaat, Jan N. van Rijn
Mach. Learn.3
2022 Towards verifiable Benchmarks for Reinforcement Learning
abstract
Reinforcement Learning (RL) is one of the most dynamic research areas in Game AI and AI as a whole, and a wide variety of games are used as its prominent test problems. However, it is subject to the replicability crisis that currently affects most algorithmic AI research. Benchmarking in Reinforcement Learning could be improved through verifiable results. There are numerous benchmark environments whose scores are used to compare different algorithms, such as Atari. Nevertheless, reviewers must trust that figures represent truthful values, as it is difficult to reproduce an exact training curve. We propose improving this situation by providing access to the original evaluation data to validate study results. To that end, we rely on the concept of replay traces. These allow re-simulation of action sequences in deterministic RL environments and, in turn, enable reviewers to verify, re-use, and manually inspect evaluation results without needing large compute clusters. It also permits validation of presented reward graphs, an inspection of individual episodes, and re-use of result data (baselines) for proper comparison in follow-up papers. We offer plug-and-play code that works with Gym so that our measures fit well in the existing RL and reproducibility eco-system. Our approach is freely available, easy to use, and adds minimal overhead, as replay traces allow a data compression ratio of up to $\approx 10^{4}$: 1 (94 GB to 8 MB for Atari Pong) compared to a regular MDP trace used in offline RL datasets. The paper presents proof-of-concept results for a variety of games.
Matthias Müller-Brockhausen, Aske Plaat, Mike Preuss
CoG2
2022 Academic Games - Mapping the Use of Video Games in Research Contexts
abstract
Video games have been used as tools for non-entertainment purposes, including research contexts. This paper defines ‘academic games’ as games that are used and developed within academic institutions for the generation, evaluation, or dissemination of knowledge. Broad intentions related to this unique use of games are rarely explicitly discussed. When they are mentioned, they tend to be specific to an individual game’s implementation, or the field of study in which it is situated. This article maps the different fundamental purposes that motivate the use of games in research contexts: involvement as stimulus, intervention, incentive, or as modeling platform. A compact review of existing literature is provided, complemented by a discussion of different facets shaping the use of games in research contexts: the flow of information, the dependency between academic effort and game artifact, and the specificity that is required. This discussion is informed by the analysis of various example games from previous work. A research agenda for the future professionalization of academic game development and its discourse concludes the article.
Marcello A. Gómez Maureira, Max J. van Duijn, Carolien Rieffe, Aske Plaat
FDG4
2022 Stateless neural meta-learning using second-order gradients
abstract
Abstract Meta-learning can be used to learn a good prior that facilitates quick learning; two popular approaches are MAML and the meta-learner LSTM. These two methods represent important and different approaches in meta-learning. In this work, we study the two and formally show that the meta-learner LSTM subsumes MAML, although MAML, which is in this sense less general, outperforms the other. We suggest the reason for this surprising performance gap is related to second-order gradients. We construct a new algorithm (named TURTLE) to gain more insight into the importance of second-order gradients. TURTLE is simpler than the meta-learner LSTM yet more expressive than MAML and outperforms both techniques at few-shot sine wave regression and 50% of the tested image classification settings (without any additional hyperparameter tuning) and is competitive otherwise, at a computational cost that is comparable to second-order MAML. We find that second-order gradients also significantly increase the accuracy of the meta-learner LSTM. When MAML was introduced, one of its remarkable features was the use of second-order gradients. Subsequent work focused on cheaper first-order approximations. On the basis of our findings, we argue for more attention for second-order gradients.
Mike Huisman, Aske Plaat, Jan N. van Rijn
Mach. Learn.2
2021 A New Challenge: Approaching Tetris Link with AI
abstract
Decades of research have been invested in making computer programs for playing games such as Chess and Go. This paper introduces a board game, Tetris Link, that is yet unexplored and appears to be highly challenging. Tetris Link has a large branching factor and lines of play that can be very deceptive, that search has a hard time uncovering. Finding good moves is very difficult for a computer player, our experiments show. We explore heuristic planning and two other approaches: Reinforcement Learning and Monte Carlo tree search. Curiously, a naive heuristic approach that is fueled by expert knowledge is still stronger than the planning and learning approaches. We, therefore, presume that Tetris Link is more difficult than expected. We offer our findings to the community as a challenge to improve upon.
Matthias Müller-Brockhausen, Mike Preuss, Aske Plaat
CoG3
2021 Procedural Content Generation: Better Benchmarks for Transfer Reinforcement Learning
abstract
The idea of transfer in reinforcement learning (TRL) is intriguing: being able to transfer knowledge from one problem to another problem without learning everything from scratch. This promises quicker learning and learning more complex methods. To gain an insight into the field and to detect emerging trends, we performed a database search. We note a surprisingly late adoption of deep learning that starts in 2018. The introduction of deep learning has not yet solved the greatest challenge of TRL: generalization. Transfer between different domains works well when domains have strong similarities (e.g. MountainCar to Cartpole), and most TRL publications focus on different tasks within the same domain that have few differences. Most TRL applications we encountered compare their improvements against self-defined baselines, and the field is still missing unified benchmarks. We consider this to be a disappointing situation. For the future, we note that: (1) A clear measure of task similarity is needed. (2) Generalization needs to improve. Promising approaches merge deep learning with planning via MCTS or introduce memory through LSTMs. (3) The lack of benchmarking tools will be remedied to enable meaningful comparison and measure progress. Already Alchemy and Meta-World are emerging as interesting benchmark suites. We note that another development, the increase in procedural content generation (PCG), can improve both benchmarking and generalization in TRL.
Matthias Müller-Brockhausen, Mike Preuss, Aske Plaat
CoG3
2021 Adaptive Warm-Start MCTS in AlphaZero-Like Deep Reinforcement Learning
Hui Wang 0053, Mike Preuss, Aske Plaat
PRICAI (3)3
2021 Level Design Patterns That Invoke Curiosity-Driven Exploration: An Empirical Study Across Multiple Conditions
abstract
Video games frequently feature 'open world' environments, designed to motivate exploration. Level design patterns are implemented to invoke curiosity and to guide player behavior. However, evidence of the efficacy of such patterns has remained theoretical. This study presents an empirical study of how level design patterns impact curiosity-driven exploration in a 3D open-world video game. 254 participants played a game in an empirical study using a between-subjects factorial design, testing 4 variables: presence or absence of patterns, goal or open-ended, nature and alien aesthetic, and assured or unassured compensation. Data collection consisted of in-game metrics and emotion word prompts as well as post-game questionnaires. Results show that design patterns invoke heightened exploration, but this effect is influenced by the presence of an explicit goal or monetary compensation. There appear to be many motivations behind exploratory behavior in games, with patterns raising expectations in players. A disposition for curiosity (i.e. 'trait curiosity') was not found to influence exploration. We interpret and discuss the impact of the conditions, individual patterns, and player motivations.
Marcello A. Gómez Maureira, Isabelle Kniestedt, Max J. van Duijn, Carolien Rieffe, Aske Plaat
Proc. ACM Hum. Comput. Interact.5
2020 Warm-Start AlphaZero Self-play Search Enhancements
Hui Wang 0053, Mike Preuss, Aske Plaat
PPSN (2)3
2019 Automated Semantic Annotation of Species Names in Handwritten Texts
Lise Stork, Andreas Weber 0008, H. Jaap van den Herik, Aske Plaat, Fons J. Verbeek, Katy Wolstencroft
ECIR (1)4
2019 Semantic annotation of natural history collections
abstract
Large collections of historical biodiversity expeditions are housed in natural history museums throughout the world. Potentially they can serve as rich sources of data for cultural historical and biodiversity research. However, they exist as only partially catalogued specimen repositories and images of unstructured, non-standardised, hand-written text and drawings. Although many archival collections have been digitised, disclosing their content is challenging. They refer to historical place names and outdated taxonomic classifications and are written in multiple languages. Efforts to transcribe the hand-written text can make the content accessible, but semantically describing and interlinking the content would further facilitate research. We propose a semantic model that serves to structure the named entities in natural history archival collections. In addition, we present an approach for the semantic annotation of these collections whilst documenting their provenance. This approach serves as an initial step for an adaptive learning approach for semi-automated extraction of named entities from natural history archival collections. The applicability of the semantic model and the annotation approach is demonstrated using image scans from a collection of 8, 000 field book pages gathered by the Committee for Natural History of the Netherlands Indies between 1820 and 1850, and evaluated together with domain experts from the field of natural and cultural history.
Lise Stork, Andreas Weber 0008, Eulàlia Gassó Miracle, Fons J. Verbeek, Aske Plaat, H. Jaap van den Herik, Katy Wolstencroft
J. Web Semant.5
2018 Towards Affordable Fault-Tolerant Nanosatellite Computing with Commodity Hardware
abstract
Modern embedded and mobile-market processor technology is a cornerstone of miniaturized satellite design. This type of lighter, cheaper, and rapidly developed spacecraft has enabled a variety of new commercial and scientific missions. However micro-and nanosatellites (<100kg) currently are not considered suitable for critical, high-priority, and complex multi-phased missions, due to their low reliability. The hardware fault tolerance (FT) concepts used aboard larger spacecraft can usually not be used, due to tight energy and mass constraints, as well as disproportional costs. Thus, we developed a hardware-software hybrid FT-approach, which enables FT through software-side coarse-grain lockstep, FPGA reconfiguration, and thread-level mixed criticality. This allows our FPGA-based proof-of-concept implementation to deliver strong fault coverage even for missions with a long duration, but also to adapt to varying performance requirements during the mission. In this paper, we present the implementation results on a tiled multiprocessor system-on-a-chip (MPSoC) design we developed as an ideal platform for our approach. We provide details on the validation of our approach through fault injection, which show that our lockstep implementation is effective and efficient for providing FDIR within our system, and show in direct comparison that our results are consistent with related work. These results show that our architecture is effective, overhead efficient, and remains within the tight energy, complexity, and cost limitations of even very small spacecraft such as CubeSats. To our knowledge, this is the first fault mitigation approach offering strong fault tolerance, which can uphold computational correctness viable for miniaturized spacecraft and is not dependent on proprietary processor cores.
Christian M. Fuchs, Nadia Murillo, Aske Plaat, Erik van der Kouwe
ATS3
2018 First Results Solving Arbitrarily Structured Maximum Independent Set Problems Using Quantum Annealing
abstract
Commercial quantum processing units (QPUs) such as those made by D-Wave Systems are being increasingly used for solving complex combinatorial optimization problems. In this paper, we review a canonical NP-hard problem, the Maximum Independent Set (MIS) problem. We show how to map MIS problems to quadratic unconstrained binary optimization (QUBO) problems, and use a D-Wave 2000Q QPU to solve them. We compare the results from the D-Wave system to classical algorithms such as simulated thermal annealing and the graphical networks package NetworkX. To our knowledge, these are the first results of experiments involving arbitrarily-structured MIS inputs using a D-Wave QPU. We find that the QPU can be used as a heuristic optimizer for randomly generated inputs, but due to physical control errors, can be outperformed by simulated thermal annealing.
Sheir Yarkoni, Aske Plaat, Thomas Bäck
CEC2
2018 Fast and Reproducible LOFAR Workflows with AGLOW
Alexandar P. Mechev, Raymond Oonk, Timothy W. Shimwell, Aske Plaat, Huib Intema, Huub Rottgerin
eScience4
2018 From Handwritten Manuscripts to Linked Data
Lise Stork, Andreas Weber 0008, H. Jaap van den Herik, Aske Plaat, Fons J. Verbeek, Katy Wolstencroft
TPDL4
2018 A Lock-free Algorithm for Parallel MCTS
Sayyed Ali Mirsoleimani, H. Jaap van den Herik, Aske Plaat, J. A. M. Vermaseren
ICAART (2)3
2018 Pipeline Pattern for Parallel MCTS
Sayyed Ali Mirsoleimani, H. Jaap van den Herik, Aske Plaat, J. A. M. Vermaseren
ICAART (2)3
2018 Real-Time Excavation Detection at Construction Sites using Deep Learning
Bas van Boven, Peter van der Putten, Anders Åström, Hakim Khalafi, Aske Plaat
IDA5
2018 Software-Defined Dependable Computing for Spacecraft
abstract
In this contribution, we provide insights on the practical feasibility, effectiveness, and validation of a software-based fault-tolerance architecture we developed for use aboard small satellites. We exploit thread-level coarse-grain lockstep to facilitate forward-error-correction and assures computational correctness on an FPGA-based MPSoC. It can be implemented using standard open-source and FPGA design tools, requires only standard COTS components, and is processor architecture and operating system agnostic.
Christian M. Fuchs, Nadia Murillo, Aske Plaat, Erik van der Kouwe, Daniel Harsono
PRDC3
2017 Bringing Fault-Tolerant GigaHertz-Computing to Space: A Multi-stage Software-Side Fault-Tolerance Approach for Miniaturized Spacecraft
abstract
Modern embedded technology is a driving factor in satellite miniaturization, contributing to a massive boom in satellite launches and a rapidly evolving new space industry. Miniaturized satellites, however, suffer from low reliability, as traditional hardware-based fault-tolerance (FT) concepts are ineffective for on-board computers (OBCs) utilizing modern systems-on-a-chip (SoC). Therefore, larger satellites continue to rely on proven processors with large feature sizes. Software-based concepts have largely been ignored by the space industry as they were researched only in theory, and have not yet reached the level of maturity necessary for implementation. We present the first integral, real-world solution to enable fault-tolerant general-purpose computing with modern multiprocessor-SoCs (MPSoCs) for spaceflight, thereby enabling their use in future high-priority space missions. The presented multi-stage approach consists of three FT stages, combining coarse-grained thread-level distributed self-validation, FPGA reconfiguration, and mixed criticality to assure long-term FT and excellent scalability for both resource constrained and critical high-priority space missions. Early benchmark results indicate a drastic performance increase over state-of-the-art radiation-hard OBC designs and considerably lower software- and hardware development costs. This approach was developed for a 4-year European Space Agency (ESA) project, and we are implementing a tiled MPSoC prototype jointly with two industrial partners.
Christian M. Fuchs, Todor P. Stefanov, Nadia Murillo, Aske Plaat
ATS4
2017 An Analysis of Virtual Loss in Parallel MCTS
Sayyed Ali Mirsoleimani, Aske Plaat, H. Jaap van den Herik, J. A. M. Vermaseren
ICAART (2)2
2016 Rapid Adaptation of Air Combat Behaviour
abstract
Adaptive behaviour for computer generated forces enriches training simulations with appropriate challenge levels. For adequate insight into the range of possible behaviour, the adaptation has to take place in a rapid fashion. Ideally, each new behaviour model should remain readable by (and thereby under the control of) human experts. Although various attempts have been made at creating adaptive behaviour, current solutions require large numbers of simulations. Moreover, usability by end users has been of subordinate interest, as is compliance with doctrine and ethics. In this work, we present a machine learning method that enables fast behaviour adaptation, while keeping the behaviour models in a human-readable format. We demonstrate the effectiveness of the proposed method in beyond-visual-range air combat simulations.
Armon Toubman, Jan Joris Roessingh, Pieter Spronck, Aske Plaat, H. Jaap van den Herik
ECAI4
2016 Ensemble UCT Needs High Exploitation
abstract
Recent results have shown that the MCTS algorithm (a new, adaptive, randomized optimization algorithm) is effective in a remarkably diverse set of applications in Artificial Intelligence, Operations Research, and High Energy Physics. MCTS can find good solutions without domain dependent heuristics, using the UCT formula to balance exploitation and exploration. It has been suggested that the optimum in the exploitation-exploration balance differs for different search tree sizes: small search trees needs more exploitation; large search trees need more exploration. Small search trees occur in variations of MCTS, such as parallel and ensemble approaches. This paper investigates the possibility of improving the performance of Ensemble UCT by increasing the level of exploitation. As the search trees become smaller we achieve an improved performance. The results are important for improving the performance of large scale parallelism of MCTS.
Sayyed Ali Mirsoleimani, Aske Plaat, H. Jaap van den Herik, J. A. M. Vermaseren
ICAART (2)2
2016 On the Impact of Data Set Size in Transfer Learning Using Deep Neural Networks
Deepak Soekhoe, Peter van der Putten, Aske Plaat
IDA3
2016 Recent Advances in Computer Games
H. Jaap van den Herik, Walter A. Kosters, Aske Plaat
Theor. Comput. Sci.3
2015 Transfer Learning of Air Combat Behavior
abstract
Machine learning techniques can help to automatically generate behavior for computer generated forces inhabiting air combat training simulations. However, as the complexity of scenarios increases, so does the time to learn optimal behavior. Transfer learning has the potential to significantly shorten the learning time between domains that are sufficiently similar. In this paper, we transfer air combat agents with experience fighting in 2-versus-1 scenarios to various 2-versus-2 scenarios. The performance of the transferred agents is compared to that of agents that learn from scratch in the 2v2 scenarios. The experiments show that the experience gained in the 2v1 scenarios is very beneficial in the plain 2v2 scenarios, where further learning is minimal. In difficult 2v2 scenarios transfer also occurs, and further learning ensues. The results pave the way for fast generation of behavior rules for air combat agents for new, complex scenarios using existing behavior models.
Armon Toubman, Jan Joris Roessingh, Pieter Spronck, Aske Plaat, H. Jaap van den Herik
ICMLA4
2015 Scaling Monte Carlo Tree Search on Intel Xeon Phi
abstract
Many algorithms have been parallelized successfully on the Intel Xeon Phi coprocessor, especially those with regular, balanced, and predictable data access patterns and instruction flows. Irregular and unbalanced algorithms are harder to parallelize efficiently. They are, for instance, present in artificial intelligence search algorithms such as Monte Carlo Tree Search (MCTS). In this paper we study the scaling behavior of MCTS, on a highly optimized real-world application, on real hardware. The Intel Xeon Phi allows shared memory scaling studies up to 61 cores and 244 hardware threads. We compare work-stealing (Cilk Plus and TBB) and work-sharing (FIFO scheduling) approaches. Interestingly, we find that a straightforward thread pool with a work-sharing FIFO queue shows the best performance. A crucial element for this high performance is the controlling of the grain size, an approach that we call Grain Size Controlled Parallel MCTS. Our subsequent comparing with the Xeon CPUs shows an even more comprehensible distinction in performance between different threading libraries. We achieve, to the best of our knowledge, the fastest implementation of a parallel MCTS on the 61 core (= 244 hardware threads) Intel Xeon Phi using a real application (47 times faster than a sequential run).
Sayyed Ali Mirsoleimani, Aske Plaat, H. Jaap van den Herik, J. A. M. Vermaseren
ICPADS2
2015 Rewarding Air Combat Behavior in Training Simulations
abstract
Computer generated forces (CGFs) inhabiting air combat training simulations must show realistic and adaptive behavior to effectively perform their roles as allies and adversaries. In earlier work, behavior for these CGFs was successfully generated using reinforcement learning. However, due to missile hits being subject to chance (a.k.a. The probability of-kill), the CGFs have in certain cases been improperly rewarded and punished. We surmise that taking this probability of-kill into account in the reward function will improve performance. To remedy the false rewards and punishments, a new reward function is proposed that rewards agents based on the expected outcome of their actions. Tests show that the use of this function significantly increases the performance of the CGFs in various scenarios, compared to the previous reward function and a naïve baseline. Based on the results, the new reward function allows the CGFs to generate more intelligent behavior, which enables better training simulations.
Armon Toubman, Jan Joris Roessingh, Pieter Spronck, Aske Plaat, H. Jaap van den Herik
SMC4
2015 Past Our Prime: A Study of Age and Play Style Development in Battlefield 3
abstract
In recent decades, video games have come to appeal to people of all ages. The effect of age on how people play games is not fully understood. In this paper, we delve into the question how age relates to an individual's play style. “Play style” is defined as any (set of) patterns in game actions performed by a player. Based on data from 10 416 Battlefield 3 players, we found that age strongly correlates to how people start out playing a game (initial play style), and to how they change their play style over time (play style development). Our data shows three major trends: 1) correlations between age and initial play style peak around the age of 20; 2) performance decreases with age; and 3) speed of play decreases with age. The relationship between age and play style may be explained by the neurocognitive effects of aging: as people grow older, their cognitive performance decays, their personalities shift to a more conscientious style, and their gaming motivations become less achievement-oriented.
Shoshannah Tekofsky, Pieter Spronck, Martijn Goudbeek, Aske Plaat, H. Jaap van den Herik
IEEE Trans. Comput. Intell. AI Games4
2014 Combining Simulated Annealing and Monte Carlo Tree Search for Expression Simplification
abstract
Abstract: In many applications of computer algebra large expressions must be simplified to make repeated numerical evaluations tractable. Previous works presented heuristically guided improvements, e.g., for Horner schemes. The remaining expression is then further reduced by common subexpression elimination. A recent approach successfully applied a relatively new algorithm, Monte Carlo Tree Search (MCTS) with UCT as the selection criterion, to find better variable orderings. Yet, this approach is fit for further improvements since it is sensi-tive to the so-called “exploration-exploitation ” constant Cp and the number of tree updates N. In this paper we propose a new selection criterion called Simulated Annealing UCT (SA-UCT) that has a dynamic exploration-exploitation parameter, which decreases with the iteration number i and thus reduces the importance of explo-ration over time. First, we provide an intuitive explanation in terms of the exploration-exploitation behavior of the algorithm. Then, we test our algorithm on three large expressions of different origins. We observe that SA-UCT widens the interval of good initial values Cp where best results are achieved. The improvement is large (more than a tenfold) and facilitates the selection of an appropriate Cp. 1
Ben Ruijl, J. A. M. Vermaseren, Aske Plaat, H. Jaap van den Herik
ICAART (1)3
2014 Dynamic Scripting with Team Coordination in Air Combat Simulation
Armon Toubman, Jan Joris Roessingh, Pieter Spronck, Aske Plaat, H. Jaap van den Herik
IEA/AIE (1)4
2014 Virtual Reflexes
Catholijn M. Jonker, Joost Broekens, Aske Plaat
IVA3
2013 Towards high performance software teamwork
abstract
Context: Research indicates that software quality, to a large extent, depends on cooperation within software teams [1] Since software development is a creative process that involves human interaction in the context of a team, it is important to understand the teamwork factors that influence performance. Objective: We present a study design in which we aim to examine the factors within software development teams that have significant influence on the performance of the team. We propose to consider factors such as communication, coordination of expertise, cohesion, trust, cooperation, and value diversity. The study investigates whether and to which extent these factors correlate with a performance of the team. In order to capture a variety of relevant teamwork factors, we created a new model extending the work of Hoegl and Gemuenden [2] and Liang et al. [3] Method: The study is based on quantitative research by means of an online questionnaire. We invited more than 20 software development teams in the Netherlands to participate in our team performance assessment, evaluating the teamwork and performance of the team. Based on an average team size of five people, one would therefore expect at least 100 participants in total. Also, product stakeholders will be asked to give their independent assessments of the performance of the team. Expected result: By analyzing the correlation between teamwork factors and team performance, we expect to gain a deeper understanding of how teamwork factors influence team performance. We also expect to validate the implemented extensions of teamwork model with respect to earlier work. Conclusion: Software teamwork factors are important to understand. In order to get a better understanding of the role of teamwork factors, this study should be conducted.
Emily Weimar, Ariadi Nugroho, Joost Visser 0001, Aske Plaat
EASE4
2013 PsyOps: Personality assessment through gaming behavior
Shoshannah Tekofsky, Pieter Spronck, Aske Plaat, H. Jaap van den Herik, Jan M. Broersen
FDG3
2002 Analysis of Transposition-Table-Driven Work Scheduling in Distributed Search
abstract
This paper discusses a new work-scheduling algorithm for parallel search of single-agent state spaces, called transposition-table-driven work scheduling, that places the transposition table at the heart of the parallel work scheduling. The scheme results in less synchronization overhead, less processor idle time, and less redundant search effort. Measurements on a 128-processor parallel machine show that the scheme achieves close-to-linear speedups; for large problems the speedups are even superlinear due to better memory usage. On the same machine, the algorithm is 1.6 to 12.9 times faster than traditional work-stealing-based schemes.
John W. Romein, Henri E. Bal, Jonathan Schaeffer 0001, Aske Plaat
IEEE Trans. Parallel Distributed Syst.4
2001 Sensitivity of parallel applications to large differences in bandwidth and latency in two-layer interconnects
Aske Plaat, Henri E. Bal, Rutger F. H. Hofman, Thilo Kielmann
Future Gener. Comput. Syst.1
2001 Unifying single-agent and two-player search
Jonathan Schaeffer 0001, Aske Plaat, Andreas Junghanns
Inf. Sci.2
1999 Sensitivity of Parallel Applications to Large Differences in Bandwidth and Latency in Two-Layer Interconnects
abstract
This paper studies application performance on systems with strongly non-uniform remote memory access. In current generation NUMAs the speed difference between the slowest and fastest link in an interconnect-the "NUMA gap"-is typically less than an order of magnitude, and many conventional parallel programs achieve good performance. We study how different NUMA gaps influence application performance, up to and including typical wide-area latencies and bandwidths. We find that for gaps larger than those of current generation NUMAs, performance suffers considerably (for applications that were designed for a uniform access interconnect). For many applications, however, performance can be greatly improved with comparatively simple changes: traffic over slow links can be reduced by making communication patterns hierarchical-like the interconnect. We find that in four out of our six applications the size of the gap can be increased by an order of magnitude or more without severely impacting speedup. We analyze why the improvements are needed, why they work so well, and how much non-uniformity they can mask.
Aske Plaat, Henri E. Bal, Rutger F. H. Hofman
HPCA1
1999 MagPIe: MPI's Collective Communication Operations for Clustered Wide Area Systems
abstract
Writing parallel applications for computational grids is a challenging task. To achieve good performance, algorithms designed for local area networks must be adapted to the differences in link speeds. An important class of algorithms are collective operations, such as broadcast and reduce. We have developed MAGPIE, a library of collective communication operations optimized for wide area systems. MAGPIE's algorithms send the minimal amount of data over the slow wide area links, and only incur a single wide area latency. Using our system, existing MPI applications can be run unmodified on geographically distributed systems. On moderate cluster sizes, using a wide area latency of 10 milliseconds and a bandwidth of 1 MByte/s, MAGPIE executes operations up to 10 times faster than MPICH, a widely used MPI implementation; application kernels improve by up to a factor of 4. Due to the structure of our algorithms, MAGPIE's advantage increases for higher wide area latencies.
Thilo Kielmann, Rutger F. H. Hofman, Henri E. Bal, Aske Plaat, Raoul Bhoedjang
PPoPP4
1999 An Efficient Implementation of Java's Remote Method Invocation
abstract
Java offers interesting opportunities for parallel computing. In particular, Java Remote Method Invocation provides an unusually flexible kind of Remote Procedure Call. Unlike RPC, RMI supports polymorphism, which requires the system to be able to download remote classes into a running application. Sun's RMI implementation achieves this kind of flexibility by passing around object type information and processing it at run time, which causes a major run time overhead. Using Sun's JDK 1.1.4 on a Pentium Pro/Myri.net cluster, for example, the latency for a null RMI (without parameters or a return value) is 1228 μsec, which is about a factor of 40 higher than that of a user-level RPC. In this paper, we study an alternative approach for implementing RMI, based on native compilation. This approach allows for better optimization, eliminates the need for processing of type information at run time, and makes a light weight communication protocol possible. We have built a Java system based on a native compiler, which supports both compile time and run time generation of marshallers. We find that almost all of the run time overhead of RMI can be pushed to compile time. With this approach, the latency of a null RMI is reduced to 34 μsec, while still supporting polymorphic RMIs (and allowing interoperability with other JVMs).
Jason Maassen, Rob van Nieuwpoort, Ronald Veldema, Henri E. Bal, Aske Plaat
PPoPP5
1996 Best-First Fixed-Depth Minimax Algorithms
Aske Plaat, Jonathan Schaeffer 0001, Wim Pijls, Arie de Bruin
Artif. Intell.1
1995 Best-First Fixed-Depth Game-Tree Search in Practice
Aske Plaat, Jonathan Schaeffer 0001, Wim Pijls, Arie de Bruin
IJCAI1