Constantinos Maglaras

dblp:92/4070 · also Costis Maglaras · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2019
0000-0002-4283-2177ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 87% Probabilistic and Bayesian machine learning · 13%
Theoretical computer science
1 paper
Approximation and online algorithms · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-armed bandit
0.412019
Thompson Sampling with Information Relaxation Penalties · NeurIPS 2019
Machine learning › Reinforcement learning
thompson sampling
0.412019
Thompson Sampling with Information Relaxation Penalties · NeurIPS 2019
Approximation and online algorithms › online learning
exploration-exploitation tradeoff
0.412019
Thompson Sampling with Information Relaxation Penalties · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning
bayesian decision theory
0.112019
Thompson Sampling with Information Relaxation Penalties · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

thompson sampling · 0.8performance bounds · 0.8information relaxation · 0.8
YearPublicationVenuePosition
2019 Thompson Sampling with Information Relaxation Penalties
abstract
We consider a finite-horizon multi-armed bandit (MAB) problem in a Bayesian setting, for which we propose an information relaxation sampling framework. With this framework, we define an intuitive family of control policies that include Thompson sampling (TS) and the Bayesian optimal policy as endpoints. Analogous to TS, which, at each decision epoch pulls an arm that is best with respect to the randomly sampled parameters, our algorithms sample entire future reward realizations and take the corresponding best action. However, this is done in the presence of “penalties” that seek to compensate for the availability of future information. We develop several novel policies and performance bounds for MAB problems that vary in terms of improving performance and increasing computational complexity between the two endpoints. Our policies can be viewed as natural generalizations of TS that simultaneously incorporate knowledge of the time horizon and explicitly consider the exploration-exploitation trade-off. We prove associated structural results on performance bounds and suboptimality gaps. Numerical experiments suggest that this new class of policies perform well, in particular in settings where the finite time horizon introduces significant exploration-exploitation tension into the problem.
Seungki Min, Constantinos Maglaras, Ciamac C. Moallemi
NeurIPS2