VLDB 2026 Research / reviewers in the wild / expert
Yiliu Wang
dblp:348/6504
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 67% Deep learning architectures and training · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › multi-armed bandit
combinatorial bandits |
0.8 | 1 | 2024 | Combinatorial Bandits for Maximum Value Reward Function under Value-Index Feedback · ICLR 2024 |
Machine learning › Deep learning architectures and training
feedback loop |
0.8 | 1 | 2024 | Combinatorial Bandits for Maximum Value Reward Function under Value-Index Feedback · ICLR 2024 |
Machine learning › Reinforcement learning
multi-armed bandit |
0.8 | 1 | 2024 | Combinatorial Bandits for Maximum Value Reward Function under Value-Index Feedback · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
regret analysis · 0.8biased arm replacement · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Combinatorial Bandits for Maximum Value Reward Function under Value-Index FeedbackabstractWe investigate the combinatorial multi-armed bandit problem where an action is to select $k$ arms from a set of base arms, and its reward is the maximum of the sample values of these $k$ arms, under a weak feedback structure that only returns the value and index of the arm with the maximum value. This novel feedback structure is much weaker than the semi-bandit feedback previously studied and is only slightly stronger than the full-bandit feedback, and thus it presents a new challenge for the online learning task. We propose an algorithm and derive a regret bound for instances where arm outcomes follow distributions with finite supports. Our algorithm introduces a novel concept of biased arm replacement to address the weak feedback challenge, and it achieves a distribution-dependent regret bound of $O((k/\Delta)\log(T))$ and a distribution-independent regret bound of $\tilde{O}(\sqrt{T})$, where $\Delta$ is the reward gap and $T$ is the time horizon.
Notably, our regret bound is comparable to the bounds obtained under the more informative semi-bandit feedback.
We demonstrate the effectiveness of our algorithm through experimental results. Yiliu Wang, Milan Vojnovic |
ICLR | 1 |