EDBT 2026 Demo / reviewers in the wild / expert
Michael O. Duff
dblp:23/3703
· DBLP profile ↗
6ranked-venue papers
4as first author
0since 2021 · last 2003
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 56% Probabilistic and Bayesian machine learning · 22% Learning theory · 22% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
diffusion approximation |
0.0 | 1 | 2003 | Diffusion Approximation for Bayesian Markov Chains · ICML 2003 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
markov chain |
0.0 | 1 | 2003 | Diffusion Approximation for Bayesian Markov Chains · ICML 2003 |
Machine learning › Reinforcement learning
bandit |
0.0 | 2 | 1996 | Local Bandit Approximation for Optimal Learning Problems · NIPS 1996 Q-Learning for Bandit Problems · ICML 1995 |
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff |
0.0 | 1 | 1996 | Local Bandit Approximation for Optimal Learning Problems · NIPS 1996 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.0 | 1 | 1995 | Q-Learning for Bandit Problems · ICML 1995 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.0 | 1 | 1995 | Q-Learning for Bandit Problems · ICML 1995 |
Machine learning › Reinforcement learning › markov decision process
semi-markov decision process |
0.0 | 1 | 1994 | Reinforcement Learning Methods for Continuous-Time Markov Decision Problems · NIPS 1994 |
Machine learning › Reinforcement learning › dynamic programming
value iteration |
0.0 | 1 | 1994 | Reinforcement Learning Methods for Continuous-Time Markov Decision Problems · NIPS 1994 |
Machine learning › Reinforcement learning › model-free reinforcement learning
monte carlo reinforcement learning |
0.0 | 1 | 1993 | Monte Carlo Matrix Inversion and Reinforcement Learning · NIPS 1993 |
Algorithms and data structures › numerical linear algebra › generalized inverse
matrix inversion |
0.0 | 1 | 1993 | Monte Carlo Matrix Inversion and Reinforcement Learning · NIPS 1993 |
Methods — techniques the papers use, named apart from their topics
diffusion approximation · 0.0q-learning · 0.0monte carlo matrix inversion · 0.0bandit approximation · 0.0stochastic approximation · 0.0TD-learning · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2003 | Design for an Optimal Probe
Michael O. Duff |
ICML | 1 |
| 2003 | Diffusion Approximation for Bayesian Markov Chains
Michael O. Duff |
ICML | 1 |
| 1996 | Local Bandit Approximation for Optimal Learning Problems
Michael O. Duff, Andrew G. Barto |
NIPS | 1 |
| 1995 | Q-Learning for Bandit Problems
Michael O. Duff |
ICML | 1 |
| 1994 | Reinforcement Learning Methods for Continuous-Time Markov Decision ProblemsabstractSemi-Markov Decision Problems are continuous time generaliza(cid:173) tions of discrete time Markov Decision Problems. A number of reinforcement learning algorithms have been developed recently for the solution of Markov Decision Problems, based on the ideas of asynchronous dynamic programming and stochastic approxima(cid:173) tion. Among these are TD(,x), Q-Iearning, and Real-time Dynamic Programming. After reviewing semi-Markov Decision Problems and Bellman's optimality equation in that context, we propose al(cid:173) gorithms similar to those named above, adapted to the solution of semi-Markov Decision Problems. We demonstrate these algorithms by applying them to the problem of determining the optimal con(cid:173) trol for a simple queueing system. We conclude with a discussion of circumstances under which these algorithms may be usefully ap(cid:173) plied. Steven J. Bradtke, Michael O. Duff |
NIPS | 2 |
| 1993 | Monte Carlo Matrix Inversion and Reinforcement Learning
Andrew G. Barto, Michael O. Duff |
NIPS | 2 |