Michael O. Duff

dblp:23/3703 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2003
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 56% Probabilistic and Bayesian machine learning · 22% Learning theory · 22%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
diffusion approximation
0.012003
Diffusion Approximation for Bayesian Markov Chains · ICML 2003
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
markov chain
0.012003
Diffusion Approximation for Bayesian Markov Chains · ICML 2003
Machine learning › Reinforcement learning
bandit
0.021996
Local Bandit Approximation for Optimal Learning Problems · NIPS 1996
Q-Learning for Bandit Problems · ICML 1995
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.011996
Local Bandit Approximation for Optimal Learning Problems · NIPS 1996
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.011995
Q-Learning for Bandit Problems · ICML 1995
Machine learning › Reinforcement learning
value-based reinforcement learning
0.011995
Q-Learning for Bandit Problems · ICML 1995
Machine learning › Reinforcement learning › markov decision process
semi-markov decision process
0.011994
Reinforcement Learning Methods for Continuous-Time Markov Decision Problems · NIPS 1994
Machine learning › Reinforcement learning › dynamic programming
value iteration
0.011994
Reinforcement Learning Methods for Continuous-Time Markov Decision Problems · NIPS 1994
Machine learning › Reinforcement learning › model-free reinforcement learning
monte carlo reinforcement learning
0.011993
Monte Carlo Matrix Inversion and Reinforcement Learning · NIPS 1993
Algorithms and data structures › numerical linear algebra › generalized inverse
matrix inversion
0.011993
Monte Carlo Matrix Inversion and Reinforcement Learning · NIPS 1993

Methods — techniques the papers use, named apart from their topics

diffusion approximation · 0.0q-learning · 0.0monte carlo matrix inversion · 0.0bandit approximation · 0.0stochastic approximation · 0.0TD-learning · 0.0
YearPublicationVenuePosition
2003 Design for an Optimal Probe
Michael O. Duff
ICML1
2003 Diffusion Approximation for Bayesian Markov Chains
Michael O. Duff
ICML1
1996 Local Bandit Approximation for Optimal Learning Problems
Michael O. Duff, Andrew G. Barto
NIPS1
1995 Q-Learning for Bandit Problems
Michael O. Duff
ICML1
1994 Reinforcement Learning Methods for Continuous-Time Markov Decision Problems
abstract
Semi-Markov Decision Problems are continuous time generaliza(cid:173) tions of discrete time Markov Decision Problems. A number of reinforcement learning algorithms have been developed recently for the solution of Markov Decision Problems, based on the ideas of asynchronous dynamic programming and stochastic approxima(cid:173) tion. Among these are TD(,x), Q-Iearning, and Real-time Dynamic Programming. After reviewing semi-Markov Decision Problems and Bellman's optimality equation in that context, we propose al(cid:173) gorithms similar to those named above, adapted to the solution of semi-Markov Decision Problems. We demonstrate these algorithms by applying them to the problem of determining the optimal con(cid:173) trol for a simple queueing system. We conclude with a discussion of circumstances under which these algorithms may be usefully ap(cid:173) plied.
Steven J. Bradtke, Michael O. Duff
NIPS2
1993 Monte Carlo Matrix Inversion and Reinforcement Learning
Andrew G. Barto, Michael O. Duff
NIPS2