Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mustafa Mert Çelikok

dblp:227/2902 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-2331-4697ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 60% Transfer learning and domain adaptation · 18% Learning theory · 14%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › computational learning theory
machine teaching
1.022023
Teaching to Learn: Sequential Teaching of Learners with Internal States · AAAI 2023
Machine Teaching of Active Sequential Learners · NeurIPS 2019
Machine learning › Reinforcement learning
actor-critic methods
0.912025
Value Improved Actor Critic Algorithms · NeurIPS 2025
Machine learning › Reinforcement learning › policy optimization
policy improvement
0.912025
Value Improved Actor Critic Algorithms · NeurIPS 2025
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.912025
Value Improved Actor Critic Algorithms · NeurIPS 2025
Machine learning › Reinforcement learning
value-based reinforcement learning
0.912025
Value Improved Actor Critic Algorithms · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
meta-learning
0.712023
Teaching to Learn: Sequential Teaching of Learners with Internal States · AAAI 2023
Machine learning › Transfer learning and domain adaptation › meta-learning
online meta-learning
0.712023
Teaching to Learn: Sequential Teaching of Learners with Internal States · AAAI 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.612022
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems · NeurIPS 2022
Machine learning › Efficient and distributed learning › data-efficient learning
sample-efficient training
0.612022
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems · NeurIPS 2022
Machine learning › Reinforcement learning
multi-armed bandit
0.412019
Machine Teaching of Active Sequential Learners · NeurIPS 2019
Human-AI interaction
interactive machine learning
0.412019
Machine Teaching of Active Sequential Learners · NeurIPS 2019
Human-AI interaction › interactive machine learning
machine teaching
0.412019
Machine Teaching of Active Sequential Learners · NeurIPS 2019
Performance modeling and evaluation › simulation › parallel and distributed simulation
distributed simulation
0.212022
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

multi-objective control · 1.3inductive bias modeling · 1.3parallel simulation · 1.1influence-augmented local simulators · 1.1greedification operator · 0.9generalized policy iteration · 0.9planning · 0.8markov decision process · 0.8learned dynamics models · 0.6learned dynamics model · 0.6
YearPublicationVenuePosition
2025 On the Complexity of Learning to Cooperate in Populations of Socially Rational Agents
Saptarashmi Bandyopadhyay, Mustafa Mert Çelikok, Robert Loftin
AAMAS2
2025 Value Improved Actor Critic Algorithms
abstract
To learn approximately optimal acting policies for decision problems, modern Actor Critic algorithms rely on deep Neural Networks (DNNs) to parameterize the acting policy and greedification operators to iteratively improve it. The reliance on DNNs suggests an improvement that is gradient based, which is per step much less greedy than the improvement possible by greedier operators such as the greedy update used by Q-learning algorithms. On the other hand, slow and steady changes to the policy can also be beneficial for the stability of the learning process, resulting in a tradeoff between greedification and stability. To address this tradeoff, we propose to extend the standard framework of actor critic algorithms with value-improvement: a second greedification operator applied only when updating the policy's value estimate. In this framework the agent can evaluate non-parameterized policies and perform much greedier updates while maintaining the steady gradient-based improvement to the parameterized acting policy. We prove that this approach converges in the popular analysis scheme of generalized Policy Iteration in the finite-horizon domain. Empirically, incorporating value-improvement into the popular off-policy actor-critic algorithms TD3 and SAC significantly improves or matches performance over their respective baselines, across different environments from the DeepMind continuous control domain, with negligible compute and implementation cost.
Yaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart, Mustafa Mert Çelikok, Wendelin Böhmer, Matthijs T. J. Spaan
NeurIPS4
2023 Teaching to Learn: Sequential Teaching of Learners with Internal States
abstract
In sequential machine teaching, a teacher’s objective is to provide the optimal sequence of inputs to sequential learners in order to guide them towards the best model. However, this teaching objective considers a restricted class of learners with fixed inductive biases. In this paper, we extend the machine teaching framework to learners that can improve their inductive biases, represented as latent internal states, in order to generalize to new datasets. We introduce a novel framework in which learners’ inductive biases may change with the teaching interaction, which affects the learning performance in future tasks. In order to teach such learners, we propose a multi-objective control approach that takes the future performance of the learner after teaching into account. This framework provides tools for modelling learners with internal states, humans and meta-learning algorithms alike. Furthermore, we distinguish manipulative teaching, which can be done by effectively hiding data and also used for indoctrination, from teaching to learn which aims to help the learner become better at learning from new datasets in the absence of a teacher. Our empirical results demonstrate that our framework is able to reduce the number of required tasks for online meta-learning, and increases independent learning performance of simulated human users in future tasks.
Mustafa Mert Çelikok, Pierre-Alexandre Murena, Samuel Kaski
AAAI1
2023 Differentiable user models
abstract
Probabilistic user modeling is essential for building machine learning systems in the ubiquitous cases with humans in the loop. However, modern advanced user models, often designed as cognitive behavior simulators, are incompatible with modern machine learning pipelines and computationally prohibitive for most practical applications. We address this problem by introducing widely-applicable differentiable surrogates for bypassing this computational bottleneck; the surrogates enable computationally efficient inference with modern cognitive models. We show experimentally that modeling capabilities comparable to the only available solution, existing likelihood-free inference methods, are achievable with a computational cost suitable for online applications. Finally, we demonstrate how AI-assistants can now use cognitive models for online interaction in a menu-search task, which has so far required hours of computation during interaction.
Alex Hämäläinen, Mustafa Mert Çelikok, Samuel Kaski
UAI2
2022 Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
abstract
Due to its high sample complexity, simulation is, as of today, critical for the successful application of reinforcement learning. Many real-world problems, however, exhibit overly complex dynamics, making their full-scale simulation computationally slow. In this paper, we show how to factorize large networked systems of many agents into multiple local regions such that we can build separate simulators that run independently and in parallel. To monitor the influence that the different local regions exert on one another, each of these simulators is equipped with a learned model that is periodically trained on real trajectories. Our empirical results reveal that distributing the simulation among different processes not only makes it possible to train large multi-agent systems in just a few hours but also helps mitigate the negative effects of simultaneous learning.
Miguel Suau, Jinke He, Mustafa Mert Çelikok, Matthijs T. J. Spaan, Frans A. Oliehoek
NeurIPS3
2019 Machine Teaching of Active Sequential Learners
abstract
Machine teaching addresses the problem of finding the best training data that can guide a learning algorithm to a target model with minimal effort. In conventional settings, a teacher provides data that are consistent with the true data distribution. However, for sequential learners which actively choose their queries, such as multi-armed bandits and active learners, the teacher can only provide responses to the learner’s queries, not design the full data. In this setting, consistent teachers can be sub-optimal for finite horizons. We formulate this sequential teaching problem, which current techniques in machine teaching do not address, as a Markov decision process, with the dynamics nesting a model of the learner and the actions being the teacher's responses. Furthermore, we address the complementary problem of learning from a teacher that plans: to recognise the teaching intent of the responses, the learner is endowed with a model of the teacher. We test the formulation with multi-armed bandit learners in simulated experiments and a user study. The results show that learning is improved by (i) planning teaching and (ii) the learner having a model of the teacher. The approach gives tools to taking into account strategic (planning) behaviour of users of interactive intelligent systems, such as recommendation engines, by considering them as boundedly optimal teachers.
Tomi Peltola, Mustafa Mert Çelikok, Pedram Daee, Samuel Kaski
NeurIPS2