VLDB 2026 Research / reviewers in the wild / expert
Michael K. Cohen
dblp:237/9460
· DBLP profile ↗
7ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0003-1749-875XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 61% Probabilistic and Bayesian machine learning · 22% Planning, search and constraint satisfaction · 8% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.6 | 1 | 2022 | Log-Linear-Time Gaussian Processes Using Binary Tree Kernels · NeurIPS 2022 |
Machine learning › Reinforcement learning
imitation learning |
0.6 | 1 | 2022 | Fully General Online Imitation Learning · J. Mach. Learn. Res. 2022 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
kernel design |
0.6 | 1 | 2022 | Log-Linear-Time Gaussian Processes Using Binary Tree Kernels · NeurIPS 2022 |
Machine learning › Reinforcement learning › imitation learning
online imitation learning |
0.6 | 1 | 2022 | Fully General Online Imitation Learning · J. Mach. Learn. Res. 2022 |
Algorithms and data structures
kernel methods |
0.6 | 1 | 2022 | Log-Linear-Time Gaussian Processes Using Binary Tree Kernels · NeurIPS 2022 |
Machine learning › Reinforcement learning
bayesian reinforcement learning |
0.4 | 1 | 2020 | Pessimism About Unknown Unknowns Inspires Conservatism · COLT 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
goal alignment |
0.4 | 1 | 2020 | Asymptotically Unambitious Artificial General Intelligence · AAAI 2020 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.4 | 1 | 2020 | Pessimism About Unknown Unknowns Inspires Conservatism · COLT 2020 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.4 | 1 | 2020 | Pessimism About Unknown Unknowns Inspires Conservatism · COLT 2020 |
Machine learning › Reinforcement learning
exploration |
0.4 | 1 | 2019 | A Strongly Asymptotically Optimal Agent in General Environments · IJCAI 2019 |
Methods — techniques the papers use, named apart from their topics
kernel approximation · 1.1gaussian process regression · 1.1conservative learning · 0.6bayesian inference · 0.6worst-case expected reward optimization · 0.4decision theory · 0.4bayesian model class · 0.4reinforcement learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can a Bayesian Oracle Prevent Harm from an Agent?abstractIs there a way to design powerful AI systems based on machine learning methods that would satisfy probabilistic safety guarantees? With the long-term goal of obtaining a probabilistic guarantee that would apply in every context, we consider estimating a context-dependent bound on the probability of violating a given safety specification. Such a risk evaluation would need to be performed at run-time to provide a guardrail against dangerous actions of an AI. Noting that different plausible hypotheses about the world could produce very different outcomes, and because we do not know which one is right, we derive bounds on the safety violation probability predicted under the true but unknown hypothesis. Such bounds could be used to reject potentially dangerous actions. Our main results involve searching for cautious but plausible hypotheses, obtained by a maximization that involves Bayesian posteriors over hypotheses. We consider two forms of this result, in the i.i.d. case and in the non-i.i.d. case, and conclude with open problems towards turning such theoretical results into practical AI guardrails. Yoshua Bengio, Michael K. Cohen, Nikolay Malkin, Matt MacDermott, Damiano Fornasiere, Pietro Greiner, Younesse Kaddar |
UAI | 2 |
| 2025 | RL, but don't do anything I wouldn't doabstractIn reinforcement learning (RL), if the agent’s reward differs from the designers’ true utility, even only rarely, the state distribution resulting from the agent’s policy can be very bad, in theory and in practice. When RL policies would devolve into undesired behavior, a common countermeasure is KL regularization to a trusted policy ("Don’t do anything I wouldn’t do"). All current cutting-edge language models are RL agents that are KL-regularized to a "base policy" that is purely predictive. Unfortunately, we demonstrate that when this base policy is a Bayesian predictive model of a trusted policy, the KL constraint is no longer reliable for controlling the behavior of an advanced RL agent. We demonstrate this theoretically using algorithmic information theory, and while systems today are too weak to exhibit this theorized failure precisely, we RL-finetune a language model and find evidence that our formal results are plausibly relevant in practice. We also propose a theoretical alternative that avoids this problem by replacing the "Don’t do anything I wouldn’t do" principle with "Don’t do anything I mightn’t do". Michael K. Cohen, Marcus Hutter, Yoshua Bengio, Stuart Russell 0001 |
UAI | 1 |
| 2022 | Log-Linear-Time Gaussian Processes Using Binary Tree KernelsabstractGaussian processes (GPs) produce good probabilistic models of functions, but most GP kernels require $O((n+m)n^2)$ time, where $n$ is the number of data points and $m$ the number of predictive locations. We present a new kernel that allows for Gaussian process regression in $O((n+m)\log(n+m))$ time. Our "binary tree" kernel places all data points on the leaves of a binary tree, with the kernel depending only on the depth of the deepest common ancestor. We can store the resulting kernel matrix in $O(n)$ space in $O(n \log n)$ time, as a sum of sparse rank-one matrices, and approximately invert the kernel matrix in $O(n)$ time. Sparse GP methods also offer linear run time, but they predict less well than higher dimensional kernels. On a classic suite of regression tasks, we compare our kernel against Mat\'ern, sparse, and sparse variational kernels. The binary tree GP assigns the highest likelihood to the test data on a plurality of datasets, usually achieves lower mean squared error than the sparse methods, and often ties or beats the Mat\'ern GP. On large datasets, the binary tree GP is fastest, and much faster than a Mat\'ern GP. Michael K. Cohen, Samuel Daulton, Michael A. Osborne |
NeurIPS | 1 |
| 2022 | Fully General Online Imitation LearningabstractIn imitation learning, imitators and demonstrators are policies for picking actions given past interactions with the environment. If we run an imitator, we probably want events to unfold similarly to the way they would have if the demonstrator had been acting the whole time. In general, one mistake during learning can lead to completely different events. In the special setting of environments that restart, existing work provides formal guidance in how to imitate so that events unfold similarly, but outside that setting, no formal guidance exists. We address a fully general setting, in which the (stochastic) environment and demonstrator never reset, not even for training purposes, and we allow our imitator to learn online from the demonstrator. Our new conservative Bayesian imitation learner underestimates the probabilities of each available action, and queries for more data with the remaining probability. Our main result: if an event would have been unlikely had the demonstrator acted the whole time, that event's likelihood can be bounded above when running the (initially totally ignorant) imitator instead. Meanwhile, queries to the demonstrator rapidly diminish in frequency. If any such event qualifies as "dangerous", our imitator would have the notable distinction of being relatively "safe". Michael K. Cohen, Marcus Hutter, Neel Nanda |
J. Mach. Learn. Res. | 1 |
| 2020 | Asymptotically Unambitious Artificial General IntelligenceabstractGeneral intelligence, the ability to solve arbitrary solvable problems, is supposed by many to be artificially constructible. Narrow intelligence, the ability to solve a given particularly difficult problem, has seen impressive recent development. Notable examples include self-driving cars, Go engines, image classifiers, and translators. Artificial General Intelligence (AGI) presents dangers that narrow intelligence does not: if something smarter than us across every domain were indifferent to our concerns, it would be an existential threat to humanity, just as we threaten many species despite no ill will. Even the theory of how to maintain the alignment of an AGI's goals with our own has proven highly elusive. We present the first algorithm we are aware of for asymptotically unambitious AGI, where “unambitiousness” includes not seeking arbitrary power. Thus, we identify an exception to the Instrumental Convergence Thesis, which is roughly that by default, an AGI would seek power, including over us. Michael K. Cohen, Badri N. Vellambi, Marcus Hutter |
AAAI | 1 |
| 2020 | Pessimism About Unknown Unknowns Inspires ConservatismabstractIf we could define the set of all bad outcomes, we could hard-code an agent which avoids them; however, in sufficiently complex environments, this is infeasible. We do not know of any general-purpose approaches in the literature to avoiding novel failure modes. Motivated by this, we define an idealized Bayesian reinforcement learner which follows a policy that maximizes the worst-case expected reward over a set of world-models. We call this agent pessimistic, since it optimizes assuming the worst case. A scalar parameter tunes the agent’s pessimism by changing the size of the set of world-models taken into account. Our first main contribution is: given an assumption about the agent’s model class, a sufficiently pessimistic agent does not cause “unprecedented events” with probability $1-\delta$, whether or not designers know how to precisely specify those precedents they are concerned with. Since pessimism discourages exploration, at each timestep, the agent may defer to a mentor, who may be a human or some known-safe policy we would like to improve. Our other main contribution is that the agent’s policy’s value approaches at least that of the mentor, while the probability of deferring to the mentor goes to 0. In high-stakes environments, we might like advanced artificial agents to pursue goals cautiously, which is a non-trivial problem even if the agent were allowed arbitrary computing power; we present a formal solution. Michael K. Cohen, Marcus Hutter |
COLT | 1 |
| 2019 | A Strongly Asymptotically Optimal Agent in General EnvironmentsabstractReinforcement Learning agents are expected to eventually perform well. Typically, this takes the form of a guarantee about the asymptotic behavior of an algorithm given some assumptions about the environment. We present an algorithm for a policy whose value approaches the optimal value with probability 1 in all computable probabilistic environments, provided the agent has a bounded horizon. This is known as strong asymptotic optimality, and it was previously unknown whether it was possible for a policy to be strongly asymptotically optimal in the class of all computable probabilistic environments. Our agent, Inquisitive Reinforcement Learner (Inq), is more likely to explore the more it expects an exploratory action to reduce its uncertainty about which environment it is in, hence the term inquisitive. Exploring inquisitively is a strategy that can be applied generally; for more manageable environment classes, inquisitiveness is tractable. We conducted experiments in "grid-worlds" to compare the Inquisitive Reinforcement Learner to other weakly asymptotically optimal agents. Michael K. Cohen, Elliot Catt, Marcus Hutter |
IJCAI | 1 |