Amin Rakhsha

dblp:261/9027 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 63% Language models and text generation · 24% Trustworthy machine learning · 10%
Network and information security
2 papers
Security and privacy of machine learning · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
model-based reinforcement learning
1.322024
Maximum Entropy Model Correction in Reinforcement Learning · ICLR 2024
Operator Splitting Value Iteration · NeurIPS 2022
Machine learning › Reinforcement learning › dynamic programming
value iteration
1.322024
Maximum Entropy Model Correction in Reinforcement Learning · ICLR 2024
Operator Splitting Value Iteration · NeurIPS 2022
Machine learning › Reinforcement learning › robust reinforcement learning
adversarial reinforcement learning
0.922021
Policy Teaching in Reinforcement Learning via Environment Poisoning Attacks · J. Mach. Learn. Res. 2021
Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning · ICML 2020
Natural language and speech › Language models and text generation › decoding
best-of-n selection
0.912025
Majority of the Bests: Improving Best-of-N via Bootstrapping · NeurIPS 2025
Natural language and speech › Language models and text generation
self-consistency
0.912025
Majority of the Bests: Improving Best-of-N via Bootstrapping · NeurIPS 2025
Machine learning › Trustworthy machine learning
model correction
0.812024
Maximum Entropy Model Correction in Reinforcement Learning · ICLR 2024
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
dyna
0.612022
Operator Splitting Value Iteration · NeurIPS 2022
Machine learning › Reinforcement learning › human-in-the-loop reinforcement learning › interactive reinforcement learning
policy teaching
0.512021
Policy Teaching in Reinforcement Learning via Environment Poisoning Attacks · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
maximum entropy density estimation
0.212024
Maximum Entropy Model Correction in Reinforcement Learning · ICLR 2024

Methods — techniques the papers use, named apart from their topics

optimization · 1.9markov decision process · 1.0reward model · 0.9bootstrapping · 0.9regret minimization · 0.9value iteration · 0.8maximum entropy density estimation · 0.8dyna · 0.8operator splitting · 0.6numerical linear algebra · 0.6
YearPublicationVenuePosition
2025 Majority of the Bests: Improving Best-of-N via Bootstrapping
abstract
Sampling multiple outputs from a Large Language Model (LLM) and selecting the most frequent (Self-consistency) or highest-scoring (Best-of-N) candidate is a popular approach to achieve higher accuracy in tasks with discrete final answers. Best-of-N (BoN) selects the output with the highest reward, and with perfect rewards, it often achieves near-perfect accuracy. With imperfect rewards from reward models, however, BoN fails to reliably find the correct answer and its performance degrades drastically. We consider the distribution of BoN’s outputs and highlight that, although the correct answer does not usually have a probability close to one under imperfect rewards, it is often the most likely outcome. This suggests that the mode of this distribution can be more reliably correct than a sample from it. Based on this idea, we propose Majority-of-the-Bests (MoB), a novel selection mechanism that estimates the output distribution of BoN via bootstrapping and selects its mode. Experimental results across five benchmarks, three different base LLMs, and two reward models demonstrate consistent improvements over BoN in 25 out of 30 setups. We also provide theoretical results for the consistency of the bootstrapping. MoB serves as a simple, yet strong alternative to BoN and self-consistency, and more broadly, motivates further research in more nuanced selection mechanisms.
Amin Rakhsha, Kanika Madan, Tianyu Zhang 0003, Amir-massoud Farahmand, Amir Khasahmadi
NeurIPS1
2024 Maximum Entropy Model Correction in Reinforcement Learning
abstract
We propose and theoretically analyze an approach for planning with an approximate model in reinforcement learning that can reduce the adverse impact of model error. If the model is accurate enough, it accelerates the convergence to the true value function too. One of its key components is the MaxEnt Model Correction (MoCo) procedure that corrects the model’s next-state distributions based on a Maximum Entropy density estimation formulation. Based on MoCo, we introduce the Model Correcting Value Iteration (MoCoVI) algorithm, and its sampled-based variant MoCoDyna. We show that MoCoVI and MoCoDyna’s convergence can be much faster than the conventional model-free algorithms. Unlike traditional model-based algorithms, MoCoVI and MoCoDyna effectively utilize an approximate model and still converge to the correct value function.
Amin Rakhsha, Mete Kemertas, Mohammad Ghavamzadeh, Amir-massoud Farahmand
ICLR1
2022 Operator Splitting Value Iteration
abstract
We introduce new planning and reinforcement learning algorithms for discounted MDPs that utilize an approximate model of the environment to accelerate the convergence of the value function. Inspired by the splitting approach in numerical linear algebra, we introduce \emph{Operator Splitting Value Iteration} (OS-VI) for both Policy Evaluation and Control problems. OS-VI achieves a much faster convergence rate when the model is accurate enough. We also introduce a sample-based version of the algorithm called OS-Dyna. Unlike the traditional Dyna architecture, OS-Dyna still converges to the correct value function in presence of model approximation error.
Amin Rakhsha, Mohammad Ghavamzadeh, Amir-massoud Farahmand
NeurIPS1
2021 Policy Teaching in Reinforcement Learning via Environment Poisoning Attacks
abstract
We study a security threat to reinforcement learning where an attacker poisons the learning environment to force the agent into executing a target policy chosen by the attacker. As a victim, we consider RL agents whose objective is to find a policy that maximizes reward in infinite-horizon problem settings. The attacker can manipulate the rewards and the transition dynamics in the learning environment at training-time, and is interested in doing so in a stealthy manner. We propose an optimization framework for finding an optimal stealthy attack for different measures of attack cost. We provide lower/upper bounds on the attack cost, and instantiate our attacks in two settings: (i) an offline setting where the agent is doing planning in the poisoned environment, and (ii) an online setting where the agent is learning a policy with poisoned feedback. Our results show that the attacker can easily succeed in teaching any target policy to the victim under mild conditions and highlight a significant security threat to reinforcement learning agents in practice.
Amin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu 0001, Adish Singla
J. Mach. Learn. Res.1
2020 Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning
abstract
We study a security threat to reinforcement learning where an attacker poisons the learning environment to force the agent into executing a target policy chosen by the attacker. As a victim, we consider RL agents whose objective is to find a policy that maximizes average reward in undiscounted infinite-horizon problem settings. The attacker can manipulate the rewards or the transition dynamics in the learning environment at training-time and is interested in doing so in a stealthy manner. We propose an optimization framework for finding an \emph{optimal stealthy attack} for different measures of attack cost. We provide sufficient technical conditions under which the attack is feasible and provide lower/upper bounds on the attack cost. We instantiate our attacks in two settings: (i) an \emph{offline} setting where the agent is doing planning in the poisoned environment, and (ii) an \emph{online} setting where the agent is learning a policy using a regret-minimization framework with poisoned feedback. Our results show that the attacker can easily succeed in teaching any target policy to the victim under mild conditions and highlight a significant security threat to reinforcement learning agents in practice.
Amin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu 0001, Adish Singla
ICML1