VLDB 2026 Research / reviewers in the wild / expert
Amin Rakhsha
dblp:261/9027
· DBLP profile ↗
5ranked-venue papers
5as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 63% Language models and text generation · 24% Trustworthy machine learning · 10% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
model-based reinforcement learning |
1.3 | 2 | 2024 | Maximum Entropy Model Correction in Reinforcement Learning · ICLR 2024 Operator Splitting Value Iteration · NeurIPS 2022 |
Machine learning › Reinforcement learning › dynamic programming
value iteration |
1.3 | 2 | 2024 | Maximum Entropy Model Correction in Reinforcement Learning · ICLR 2024 Operator Splitting Value Iteration · NeurIPS 2022 |
Machine learning › Reinforcement learning › robust reinforcement learning
adversarial reinforcement learning |
0.9 | 2 | 2021 | Policy Teaching in Reinforcement Learning via Environment Poisoning Attacks · J. Mach. Learn. Res. 2021 Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning · ICML 2020 |
Natural language and speech › Language models and text generation › decoding
best-of-n selection |
0.9 | 1 | 2025 | Majority of the Bests: Improving Best-of-N via Bootstrapping · NeurIPS 2025 |
Natural language and speech › Language models and text generation
self-consistency |
0.9 | 1 | 2025 | Majority of the Bests: Improving Best-of-N via Bootstrapping · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
model correction |
0.8 | 1 | 2024 | Maximum Entropy Model Correction in Reinforcement Learning · ICLR 2024 |
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
dyna |
0.6 | 1 | 2022 | Operator Splitting Value Iteration · NeurIPS 2022 |
Machine learning › Reinforcement learning › human-in-the-loop reinforcement learning › interactive reinforcement learning
policy teaching |
0.5 | 1 | 2021 | Policy Teaching in Reinforcement Learning via Environment Poisoning Attacks · J. Mach. Learn. Res. 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
maximum entropy density estimation |
0.2 | 1 | 2024 | Maximum Entropy Model Correction in Reinforcement Learning · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
optimization · 1.9markov decision process · 1.0reward model · 0.9bootstrapping · 0.9regret minimization · 0.9value iteration · 0.8maximum entropy density estimation · 0.8dyna · 0.8operator splitting · 0.6numerical linear algebra · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Majority of the Bests: Improving Best-of-N via BootstrappingabstractSampling multiple outputs from a Large Language Model (LLM) and selecting the most frequent (Self-consistency) or highest-scoring (Best-of-N) candidate is a popular approach to achieve higher accuracy in tasks with discrete final answers. Best-of-N (BoN) selects the output with the highest reward, and with perfect rewards, it often achieves near-perfect accuracy. With imperfect rewards from reward models, however, BoN fails to reliably find the correct answer and its performance degrades drastically. We consider the distribution of BoN’s outputs and highlight that, although the correct answer does not usually have a probability close to one under imperfect rewards, it is often the most likely outcome. This suggests that the mode of this distribution can be more reliably correct than a sample from it. Based on this idea, we propose Majority-of-the-Bests (MoB), a novel selection mechanism that estimates the output distribution of BoN via bootstrapping and selects its mode. Experimental results across five benchmarks, three different base LLMs, and two reward models demonstrate consistent improvements over BoN in 25 out of 30 setups. We also provide theoretical results for the consistency of the bootstrapping. MoB serves as a simple, yet strong alternative to BoN and self-consistency, and more broadly, motivates further research in more nuanced selection mechanisms. Amin Rakhsha, Kanika Madan, Tianyu Zhang 0003, Amir-massoud Farahmand, Amir Khasahmadi |
NeurIPS | 1 |
| 2024 | Maximum Entropy Model Correction in Reinforcement LearningabstractWe propose and theoretically analyze an approach for planning with an approximate model in reinforcement learning that can reduce the adverse impact of model error. If the model is accurate enough, it accelerates the convergence to the true value function too. One of its key components is the MaxEnt Model Correction (MoCo) procedure that corrects the model’s next-state distributions based on a Maximum Entropy density estimation formulation. Based on MoCo, we introduce the Model Correcting Value Iteration (MoCoVI) algorithm, and its sampled-based variant MoCoDyna. We show that MoCoVI and MoCoDyna’s convergence can be much faster than the conventional model-free algorithms. Unlike traditional model-based algorithms, MoCoVI and MoCoDyna effectively utilize an approximate model and still converge to the correct value function. Amin Rakhsha, Mete Kemertas, Mohammad Ghavamzadeh, Amir-massoud Farahmand |
ICLR | 1 |
| 2022 | Operator Splitting Value IterationabstractWe introduce new planning and reinforcement learning algorithms for discounted MDPs that utilize an approximate model of the environment to accelerate the convergence of the value function. Inspired by the splitting approach in numerical linear algebra, we introduce \emph{Operator Splitting Value Iteration} (OS-VI) for both Policy Evaluation and Control problems. OS-VI achieves a much faster convergence rate when the model is accurate enough. We also introduce a sample-based version of the algorithm called OS-Dyna. Unlike the traditional Dyna architecture, OS-Dyna still converges to the correct value function in presence of model approximation error. Amin Rakhsha, Mohammad Ghavamzadeh, Amir-massoud Farahmand |
NeurIPS | 1 |
| 2021 | Policy Teaching in Reinforcement Learning via Environment Poisoning AttacksabstractWe study a security threat to reinforcement learning where an attacker poisons the learning environment to force the agent into executing a target policy chosen by the attacker. As a victim, we consider RL agents whose objective is to find a policy that maximizes reward in infinite-horizon problem settings. The attacker can manipulate the rewards and the transition dynamics in the learning environment at training-time, and is interested in doing so in a stealthy manner. We propose an optimization framework for finding an optimal stealthy attack for different measures of attack cost. We provide lower/upper bounds on the attack cost, and instantiate our attacks in two settings: (i) an offline setting where the agent is doing planning in the poisoned environment, and (ii) an online setting where the agent is learning a policy with poisoned feedback. Our results show that the attacker can easily succeed in teaching any target policy to the victim under mild conditions and highlight a significant security threat to reinforcement learning agents in practice. Amin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu 0001, Adish Singla |
J. Mach. Learn. Res. | 1 |
| 2020 | Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement LearningabstractWe study a security threat to reinforcement learning where an attacker poisons the learning environment to force the agent into executing a target policy chosen by the attacker. As a victim, we consider RL agents whose objective is to find a policy that maximizes average reward in undiscounted infinite-horizon problem settings. The attacker can manipulate the rewards or the transition dynamics in the learning environment at training-time and is interested in doing so in a stealthy manner. We propose an optimization framework for finding an \emph{optimal stealthy attack} for different measures of attack cost. We provide sufficient technical conditions under which the attack is feasible and provide lower/upper bounds on the attack cost. We instantiate our attacks in two settings: (i) an \emph{offline} setting where the agent is doing planning in the poisoned environment, and (ii) an \emph{online} setting where the agent is learning a policy using a regret-minimization framework with poisoned feedback. Our results show that the attacker can easily succeed in teaching any target policy to the victim under mild conditions and highlight a significant security threat to reinforcement learning agents in practice. Amin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu 0001, Adish Singla |
ICML | 1 |