VLDB 2026 Research / reviewers in the wild / expert
Mustafa Mert Çelikok
dblp:227/2902
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-2331-4697ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 60% Transfer learning and domain adaptation · 18% Learning theory · 14% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-AI interaction · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › computational learning theory
machine teaching |
1.0 | 2 | 2023 | Teaching to Learn: Sequential Teaching of Learners with Internal States · AAAI 2023 Machine Teaching of Active Sequential Learners · NeurIPS 2019 |
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 1 | 2025 | Value Improved Actor Critic Algorithms · NeurIPS 2025 |
Machine learning › Reinforcement learning › policy optimization
policy improvement |
0.9 | 1 | 2025 | Value Improved Actor Critic Algorithms · NeurIPS 2025 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.9 | 1 | 2025 | Value Improved Actor Critic Algorithms · NeurIPS 2025 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.9 | 1 | 2025 | Value Improved Actor Critic Algorithms · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.7 | 1 | 2023 | Teaching to Learn: Sequential Teaching of Learners with Internal States · AAAI 2023 |
Machine learning › Transfer learning and domain adaptation › meta-learning
online meta-learning |
0.7 | 1 | 2023 | Teaching to Learn: Sequential Teaching of Learners with Internal States · AAAI 2023 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.6 | 1 | 2022 | Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › data-efficient learning
sample-efficient training |
0.6 | 1 | 2022 | Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems · NeurIPS 2022 |
Machine learning › Reinforcement learning
multi-armed bandit |
0.4 | 1 | 2019 | Machine Teaching of Active Sequential Learners · NeurIPS 2019 |
Human-AI interaction
interactive machine learning |
0.4 | 1 | 2019 | Machine Teaching of Active Sequential Learners · NeurIPS 2019 |
Human-AI interaction › interactive machine learning
machine teaching |
0.4 | 1 | 2019 | Machine Teaching of Active Sequential Learners · NeurIPS 2019 |
Performance modeling and evaluation › simulation › parallel and distributed simulation
distributed simulation |
0.2 | 1 | 2022 | Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
multi-objective control · 1.3inductive bias modeling · 1.3parallel simulation · 1.1influence-augmented local simulators · 1.1greedification operator · 0.9generalized policy iteration · 0.9planning · 0.8markov decision process · 0.8learned dynamics models · 0.6learned dynamics model · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Complexity of Learning to Cooperate in Populations of Socially Rational Agents
Saptarashmi Bandyopadhyay, Mustafa Mert Çelikok, Robert Loftin |
AAMAS | 2 |
| 2025 | Value Improved Actor Critic AlgorithmsabstractTo learn approximately optimal acting policies for decision problems, modern Actor Critic algorithms rely on deep Neural Networks (DNNs) to parameterize the acting policy and greedification operators to iteratively improve it.
The reliance on DNNs suggests an improvement that is gradient based, which is per step much less greedy than the improvement possible by greedier operators such as the greedy update used by Q-learning algorithms.
On the other hand, slow and steady changes to the policy can also be beneficial for the stability of the learning process, resulting in a tradeoff between greedification and stability.
To address this tradeoff, we propose to extend the standard framework of actor critic algorithms with value-improvement: a second greedification operator applied only when updating the policy's value estimate.
In this framework the agent can evaluate non-parameterized policies and perform much greedier updates while maintaining the steady gradient-based improvement to the parameterized acting policy.
We prove that this approach converges in the popular analysis scheme of generalized Policy Iteration in the finite-horizon domain.
Empirically, incorporating value-improvement into the popular off-policy actor-critic algorithms TD3 and SAC significantly improves or matches performance over their respective baselines, across different environments from the DeepMind continuous control domain, with negligible compute and implementation cost. Yaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart, Mustafa Mert Çelikok, Wendelin Böhmer, Matthijs T. J. Spaan |
NeurIPS | 4 |
| 2023 | Teaching to Learn: Sequential Teaching of Learners with Internal StatesabstractIn sequential machine teaching, a teacher’s objective is to provide the optimal sequence of inputs to sequential learners in order to guide them towards the best model. However, this teaching objective considers a restricted class of learners with fixed inductive biases. In this paper, we extend the machine teaching framework to learners that can improve their inductive biases, represented as latent internal states, in order to generalize to new datasets. We introduce a novel framework in which learners’ inductive biases may change with the teaching interaction, which affects the learning performance in future tasks. In order to teach such learners, we propose a multi-objective control approach that takes the future performance of the learner after teaching into account. This framework provides tools for modelling learners with internal states, humans and meta-learning algorithms alike. Furthermore, we distinguish manipulative teaching, which can be done by effectively hiding data and also used for indoctrination, from teaching to learn which aims to help the learner become better at learning from new datasets in the absence of a teacher. Our empirical results demonstrate that our framework is able to reduce the number of required tasks for online meta-learning, and increases independent learning performance of simulated human users in future tasks. Mustafa Mert Çelikok, Pierre-Alexandre Murena, Samuel Kaski |
AAAI | 1 |
| 2023 | Differentiable user modelsabstractProbabilistic user modeling is essential for building machine learning systems in the ubiquitous cases with humans in the loop. However, modern advanced user models, often designed as cognitive behavior simulators, are incompatible with modern machine learning pipelines and computationally prohibitive for most practical applications. We address this problem by introducing widely-applicable differentiable surrogates for bypassing this computational bottleneck; the surrogates enable computationally efficient inference with modern cognitive models. We show experimentally that modeling capabilities comparable to the only available solution, existing likelihood-free inference methods, are achievable with a computational cost suitable for online applications. Finally, we demonstrate how AI-assistants can now use cognitive models for online interaction in a menu-search task, which has so far required hours of computation during interaction. Alex Hämäläinen, Mustafa Mert Çelikok, Samuel Kaski |
UAI | 2 |
| 2022 | Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked SystemsabstractDue to its high sample complexity, simulation is, as of today, critical for the successful application of reinforcement learning. Many real-world problems, however, exhibit overly complex dynamics, making their full-scale simulation computationally slow. In this paper, we show how to factorize large networked systems of many agents into multiple local regions such that we can build separate simulators that run independently and in parallel. To monitor the influence that the different local regions exert on one another, each of these simulators is equipped with a learned model that is periodically trained on real trajectories. Our empirical results reveal that distributing the simulation among different processes not only makes it possible to train large multi-agent systems in just a few hours but also helps mitigate the negative effects of simultaneous learning. Miguel Suau, Jinke He, Mustafa Mert Çelikok, Matthijs T. J. Spaan, Frans A. Oliehoek |
NeurIPS | 3 |
| 2019 | Machine Teaching of Active Sequential LearnersabstractMachine teaching addresses the problem of finding the best training data that can guide a learning algorithm to a target model with minimal effort. In conventional settings, a teacher provides data that are consistent with the true data distribution. However, for sequential learners which actively choose their queries, such as multi-armed bandits and active learners, the teacher can only provide responses to the learner’s queries, not design the full data. In this setting, consistent teachers can be sub-optimal for finite horizons. We formulate this sequential teaching problem, which current techniques in machine teaching do not address, as a Markov decision process, with the dynamics nesting a model of the learner and the actions being the teacher's responses. Furthermore, we address the complementary problem of learning from a teacher that plans: to recognise the teaching intent of the responses, the learner is endowed with a model of the teacher. We test the formulation with multi-armed bandit learners in simulated experiments and a user study. The results show that learning is improved by (i) planning teaching and (ii) the learner having a model of the teacher. The approach gives tools to taking into account strategic (planning) behaviour of users of interactive intelligent systems, such as recommendation engines, by considering them as boundedly optimal teachers. Tomi Peltola, Mustafa Mert Çelikok, Pedram Daee, Samuel Kaski |
NeurIPS | 2 |