Anna Winnicki

dblp:251/5594 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 64% Learning theory · 36%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › markov decision process
episodic MDP
0.912025
Reinforcement Learning with Segment Feedback · ICML 2025
Machine learning › Learning theory › online learning
regret bounds
0.912025
Reinforcement Learning with Segment Feedback · ICML 2025
Machine learning › Reinforcement learning
policy optimization
0.812024
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization · ICML 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization · ICML 2024
Machine learning › Learning theory
sample complexity
0.812024
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization · ICML 2024
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.312025
Reinforcement Learning with Segment Feedback · ICML 2025
Machine learning › Reinforcement learning
reward learning
0.212024
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization · ICML 2024

Methods — techniques the papers use, named apart from their topics

sum feedback · 0.9regret analysis · 0.9binary feedback · 0.9trajectory-level elliptical potential analysis · 0.8policy cover-policy gradient · 0.8neural function approximation · 0.8
YearPublicationVenuePosition
2025 Reinforcement Learning with Segment Feedback
abstract
Standard reinforcement learning (RL) assumes that an agent can observe a reward for each state-action pair. However, in practical applications, it is often difficult and costly to collect a reward for each state-action pair. While there have been several works considering RL with trajectory feedback, it is unclear if trajectory feedback is inefficient for learning when trajectories are long. In this work, we consider a model named RL with segment feedback, which offers a general paradigm filling the gap between per-state-action feedback and trajectory feedback. In this model, we consider an episodic Markov decision process (MDP), where each episode is divided into $m$ segments, and the agent observes reward feedback only at the end of each segment. Under this model, we study two popular feedback settings: binary feedback and sum feedback, where the agent observes a binary outcome and a reward sum according to the underlying reward function, respectively. To investigate the impact of the number of segments $m$ on learning performance, we design efficient algorithms and establish regret upper and lower bounds for both feedback settings. Our theoretical and experimental results show that: under binary feedback, increasing the number of segments $m$ decreases the regret at an exponential rate; in contrast, surprisingly, under sum feedback, increasing $m$ does not reduce the regret significantly.
Yihan Du, Anna Winnicki, Gal Dalal, Shie Mannor, R. Srikant 0001
ICML2
2024 Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
abstract
Reinforcement Learning from Human Feedback (RLHF) has achieved impressive empirical successes while relying on a small amount of human feedback. However, there is limited theoretical justification for this phenomenon. Additionally, most recent studies focus on value-based algorithms despite the recent empirical successes of policy-based algorithms. In this work, we consider an RLHF algorithm based on policy optimization (PO-RLHF). The algorithm is based on the popular Policy Cover-Policy Gradient (PC-PG) algorithm, which assumes knowledge of the reward function. In PO-RLHF, knowledge of the reward function is not assumed and the algorithm relies on trajectory-based comparison feedback to infer the reward function. We provide performance bounds for PO-RLHF with low query complexity, which provides insight into why a small amount of human feedback may be sufficient to get good performance with RLHF. A key novelty is our trajectory-level elliptical potential analysis technique used to infer reward function parameters when comparison queries rather than reward observations are used. We provide and analyze algorithms in two settings: linear and neural function approximation, PG-RLHF and NN-PG-RLHF, respectively.
Yihan Du, Anna Winnicki, Gal Dalal, Shie Mannor, R. Srikant 0001
ICML2
2023 On The Convergence Of Policy Iteration-Based Reinforcement Learning With Monte Carlo Policy Evaluation
abstract
A common technique in reinforcement learning is to evaluate the value function from Monte Carlo simulations of a given policy, and use the estimated value function to obtain a new policy which is greedy with respect to the estimated value function. A well-known longstanding open problem in this context is to prove the convergence of such a scheme when the value function of a policy is estimated from data collected from a single sample path obtained from implementing the policy (see page 99 of [Sutton and Barto, 2018], page 8 of [Tsitsiklis, 2002]). We present a solution to the open problem by showing that a first-visit version of such a policy iteration scheme indeed converges to the optimal policy provided that the policy improvement step uses lookahead [Silver et al., 2016, Mnih et al., 2016, Silver et al., 2017b] rather than a simple greedy policy improvement. We provide results both for the original open problem in the tabular setting and also present extensions to the function approximation setting, where we show that the policy resulting from the algorithm performs close to the optimal policy within a function approximation error.
Anna Winnicki, R. Srikant 0001
AISTATS1