Yiliu Wang

dblp:348/6504 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 67% Deep learning architectures and training · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › multi-armed bandit
combinatorial bandits
0.812024
Combinatorial Bandits for Maximum Value Reward Function under Value-Index Feedback · ICLR 2024
Machine learning › Deep learning architectures and training
feedback loop
0.812024
Combinatorial Bandits for Maximum Value Reward Function under Value-Index Feedback · ICLR 2024
Machine learning › Reinforcement learning
multi-armed bandit
0.812024
Combinatorial Bandits for Maximum Value Reward Function under Value-Index Feedback · ICLR 2024

Methods — techniques the papers use, named apart from their topics

regret analysis · 0.8biased arm replacement · 0.8
YearPublicationVenuePosition
2024 Combinatorial Bandits for Maximum Value Reward Function under Value-Index Feedback
abstract
We investigate the combinatorial multi-armed bandit problem where an action is to select $k$ arms from a set of base arms, and its reward is the maximum of the sample values of these $k$ arms, under a weak feedback structure that only returns the value and index of the arm with the maximum value. This novel feedback structure is much weaker than the semi-bandit feedback previously studied and is only slightly stronger than the full-bandit feedback, and thus it presents a new challenge for the online learning task. We propose an algorithm and derive a regret bound for instances where arm outcomes follow distributions with finite supports. Our algorithm introduces a novel concept of biased arm replacement to address the weak feedback challenge, and it achieves a distribution-dependent regret bound of $O((k/\Delta)\log(T))$ and a distribution-independent regret bound of $\tilde{O}(\sqrt{T})$, where $\Delta$ is the reward gap and $T$ is the time horizon. Notably, our regret bound is comparable to the bounds obtained under the more informative semi-bandit feedback. We demonstrate the effectiveness of our algorithm through experimental results.
Yiliu Wang, Milan Vojnovic
ICLR1