Recep Yusuf Bekci

dblp:274/1638 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 67% Learning theory · 33%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
delayed feedback
0.812024
Online Learning of Delayed Choices · NeurIPS 2024
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.812024
Online Learning of Delayed Choices · NeurIPS 2024
Machine learning › Learning theory
online learning
0.812024
Online Learning of Delayed Choices · NeurIPS 2024
Algorithmic game theory and mechanism design › decision theory
choice models
0.812024
Online Learning of Delayed Choices · NeurIPS 2024
Algorithmic game theory and mechanism design › decision theory
multinomial logit model
0.812024
Online Learning of Delayed Choices · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

optimism in the face of uncertainty · 1.5confidence bounds · 1.5
YearPublicationVenuePosition
2024 Online Learning of Delayed Choices
abstract
Choice models are essential for understanding decision-making processes in domains like online advertising, product recommendations, and assortment optimization. The Multinomial Logit (MNL) model is particularly versatile in selecting products or advertisements for display. However, challenges arise with unknown MNL parameters and delayed feedback, requiring sellers to learn customers’ choice behavior and make dynamic decisions with biased knowledge due to delays. We address these challenges by developing an algorithm that handles delayed feedback, balancing exploration and exploitation using confidence bounds and optimism. We first consider a censored setting where a threshold for considering feedback is imposed by business requirements. Our algorithm demonstrates a $\tilde{O}(\sqrt{NT})$ regret, with a matching lower bound up to a logarithmic term. Furthermore, we extend our analysis to environments with non-thresholded delays, achieving a $\tilde{O}(\sqrt{NT})$ regret. To validate our approach, we conduct experiments that confirm the effectiveness of our algorithm.
Recep Yusuf Bekci
NeurIPS1