Lior Shani

dblp:232/2141 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-1504-0534ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Reinforcement learning · 65% Language models and text generation · 12% Learning theory · 8%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 25 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
policy optimization
2.852024
Multi-turn Reinforcement Learning with Preference Human Feedback · NeurIPS 2024
Mirror Descent Policy Optimization · ICLR 2022
Online Apprenticeship Learning · AAAI 2022
Machine learning › Reinforcement learning
reinforcement learning from human feedback
2.432025
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward · NeurIPS 2025
Embedding-Aligned Language Models · NeurIPS 2024
Multi-turn Reinforcement Learning with Preference Human Feedback · NeurIPS 2024
Machine learning › Optimization for machine learning
mirror descent
1.322024
Multi-turn Reinforcement Learning with Preference Human Feedback · NeurIPS 2024
Mirror Descent Policy Optimization · ICLR 2022
Machine learning › Learning theory › online learning
regret bounds
1.122023
Reinforcement Learning with History Dependent Dynamic Contexts · ICML 2023
Optimistic Policy Optimization with Bandit Feedback · ICML 2020
Machine learning › Reinforcement learning › exploration › intrinsic motivation
curiosity-driven exploration
0.912025
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward · NeurIPS 2025
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.912025
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward · NeurIPS 2025
Natural language and speech › Question answering and dialogue systems
personalized dialogue
0.912025
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward · NeurIPS 2025
Machine learning › Reinforcement learning
model-based reinforcement learning
0.822023
Reinforcement Learning with History Dependent Dynamic Contexts · ICML 2023
Optimistic Policy Optimization with Bandit Feedback · ICML 2020
Natural language and speech › Language models and text generation
controllable text generation
0.812024
Embedding-Aligned Language Models · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
interpretable embedding
0.812024
Demystifying Embedding Spaces using Large Language Models · ICLR 2024
Machine learning › Reinforcement learning › markov decision process
contextual markov decision process
0.712023
Reinforcement Learning with History Dependent Dynamic Contexts · ICML 2023
Natural language and speech › Language models and text generation
text summarization
0.712023
Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback · ACL (1) 2023
Machine learning › Reinforcement learning
exploration
0.622022
Optimistic Policy Optimization with Bandit Feedback · ICML 2020
Online Apprenticeship Learning · AAAI 2022
Machine learning › Reinforcement learning
imitation learning
0.612022
Online Apprenticeship Learning · AAAI 2022
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.612022
Online Apprenticeship Learning · AAAI 2022
Machine learning › Reinforcement learning
markov decision process
0.612022
Reinforcement Learning with a Terminator · NeurIPS 2022
Machine learning › Learning theory › online learning
no-regret algorithms
0.612022
Online Apprenticeship Learning · AAAI 2022
Machine learning › Reinforcement learning › exploration › efficient exploration
provably efficient exploration
0.612022
Reinforcement Learning with a Terminator · NeurIPS 2022
Machine learning › Reinforcement learning
regret minimization
0.612022
Reinforcement Learning with a Terminator · NeurIPS 2022
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization
0.412020
Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs · AAAI 2020
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.412019
Exploration Conscious Reinforcement Learning Revisited · ICML 2019
Natural language and speech › Language models and text generation
alignment
0.212024
Multi-turn Reinforcement Learning with Preference Human Feedback · NeurIPS 2024
Machine learning › Reinforcement learning › exploration
optimistic exploration
0.212022
Online Apprenticeship Learning · AAAI 2022
Mathematical optimization › continuous optimization
convex optimization
0.112020
Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs · AAAI 2020
Mathematical optimization › numerical computation › numerical optimization › second-order methods
trust region methods
0.112020
Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs · AAAI 2020

Methods — techniques the papers use, named apart from their topics

mirror descent · 2.0reinforcement learning · 1.4optimism · 1.1user modeling · 0.9reinforcement learning from human feedback · 0.9mirror-descent-based policy optimization · 0.8latent embedding space · 0.8large language model · 0.8deep RL · 0.8concept activation vectors · 0.8trust-region method · 0.4
YearPublicationVenuePosition
2025 Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
abstract
Effective conversational agents must personalize their interactions to adapt to user preferences, personalities, and attributes across diverse domains like education and healthcare. Current methods like Reinforcement Learning from Human Feedback (RLHF), often prioritize helpfulness and safety but fall short in fostering truly empathetic, adaptive, and personalized dialogues. Existing personalization approaches typically rely on extensive user history, limiting their effectiveness for new or context-limited users. To address these limitations, we propose leveraging a user model to incorporate a curiosity-based intrinsic reward into multi-turn RLHF. This novel reward mechanism encourages the agent to actively infer user traits by optimizing conversations to improve its user model's accuracy. Consequently, the agent delivers more personalized interactions by learning more about the user. We demonstrate our method's effectiveness in two distinct domains: significantly improving personalization performance in a conversational recommendation task, and personalizing conversations for different learning styles in an educational setting with improved generalization capabilities compared to traditional multi-turn RLHF, all while maintaining conversation quality. Our method offers a promising solution for creating more personalized, adaptive, and engaging conversational agents.
Yanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani, Natasha Jaques
NeurIPS4
2024 Demystifying Embedding Spaces using Large Language Models
abstract
Embeddings have become a pivotal means to represent complex, multi-faceted information about entities, concepts, and relationships in a condensed and useful format. Nevertheless, they often preclude direct interpretation. While downstream tasks make use of these compressed representations, meaningful interpretation usually requires visualization using dimensionality reduction or specialized machine learning interpretability methods. This paper addresses the challenge of making such embeddings more interpretable and broadly useful, by employing large language models (LLMs) to directly interact with embeddings -- transforming abstract vectors into understandable narratives. By injecting embeddings into LLMs, we enable querying and exploration of complex embedding data. We demonstrate our approach on a variety of diverse tasks, including: enhancing concept activation vectors (CAVs), communicating novel embedded entities, and decoding user preferences in recommender systems. Our work couples the immense information potential of embeddings with the interpretative power of LLMs.
Guy Tennenholtz, Yinlam Chow, Jihwan Jeong, Lior Shani, Aza Tulepbergenov, Deepak Ramachandran, Martin Mladenov, Craig Boutilier
ICLR5
2024 Multi-turn Reinforcement Learning with Preference Human Feedback
abstract
Reinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate remarkable abilities in various tasks. Existing methods work by emulating the human preference at the single decision (turn) level, limiting their capabilities in settings that require planning or multi-turn interactions to achieve a long-term goal. In this paper, we address this issue by developing novel methods for Reinforcement Learning (RL) from preference feedback between two full multi-turn conversations. In the tabular setting, we present a novel mirror-descent-based policy optimization algorithm for the general multi-turn preference-based RL problem, and prove its convergence to Nash equilibrium. To evaluate performance, we create a new environment, Education Dialogue, where a teacher agent guides a student in learning a random topic, and show that a deep RL variant of our algorithm outperforms RLHF baselines. Finally, we show that in an environment with explicit rewards, our algorithm recovers the same performance as a reward-based RL baseline, despite relying solely on a weaker preference signal.
Lior Shani, Aviv Rosenberg 0002, Asaf Cassel, Oran Lang, Daniele Calandriello, Avital Zipori, Hila Noga, Orgad Keller, Bilal Piot, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Rémi Munos
NeurIPS1
2024 Embedding-Aligned Language Models
abstract
We propose a novel approach for training large language models (LLMs) to adhere to objectives defined within a latent embedding space. Our method leverages reinforcement learning (RL), treating a pre-trained LLM as an environment. Our embedding-aligned guided language (EAGLE) agent is trained to iteratively steer the LLM's generation towards optimal regions of the latent embedding space, w.r.t. some predefined criterion. We demonstrate the effectiveness of the EAGLE agent using the MovieLens 25M and Amazon Review datasets to surface content gaps that satisfy latent user demand. We also demonstrate the benefit of using an optimal design of a state-dependent action set to improve EAGLE's efficiency. Our work paves the way for controlled and grounded text generation using LLMs, ensuring consistency with domain-specific knowledge and data representations.
Guy Tennenholtz, Yinlam Chow, Lior Shani, Craig Boutilier
NeurIPS4
2023 Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback
abstract
Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Leonard Hussenot, Orgad Keller, Nikola Momchev, Sabela Ramos Garea, Piotr Stanczyk, Nino Vieillard, Olivier Bachem, Gal Elidan, Avinatan Hassidim, Olivier Pietquin, Idan Szpektor. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Léonard Hussenot, Orgad Keller, Nikola Momchev, Sabela Ramos, Piotr Stanczyk, Nino Vieillard, Olivier Bachem, Gal Elidan, Avinatan Hassidim, Olivier Pietquin, Idan Szpektor
ACL (1)3
2023 Reinforcement Learning with History Dependent Dynamic Contexts
abstract
We introduce Dynamic Contextual Markov Decision Processes (DCMDPs), a novel reinforcement learning framework for history-dependent environments that generalizes the contextual MDP framework to handle non-Markov environments, where contexts change over time. We consider special cases of the model, with a focus on logistic DCMDPs, which break the exponential dependence on history length by leveraging aggregation functions to determine context transitions. This special structure allows us to derive an upper-confidence-bound style algorithm for which we establish regret bounds. Motivated by our theoretical results, we introduce a practical model-based algorithm for logistic DCMDPs that plans in a latent space and uses optimism over history-dependent features. We demonstrate the efficacy of our approach on a recommendation task (using MovieLens data) where user behavior dynamics evolve in response to recommendations.
Guy Tennenholtz, Nadav Merlis, Lior Shani, Martin Mladenov, Craig Boutilier
ICML3
2022 Online Apprenticeship Learning
abstract
In Apprenticeship Learning (AL), we are given a Markov Decision Process (MDP) without access to the cost function. Instead, we observe trajectories sampled by an expert that acts according to some policy. The goal is to find a policy that matches the expert's performance on some predefined set of cost functions. We introduce an online variant of AL (Online Apprenticeship Learning; OAL), where the agent is expected to perform comparably to the expert while interacting with the environment. We show that the OAL problem can be effectively solved by combining two mirror descent based no-regret algorithms: one for policy optimization and another for learning the worst case cost. By employing optimistic exploration, we derive a convergent algorithm with O(sqrt(K)) regret, where K is the number of interactions with the MDP, and an additional linear error term that depends on the amount of expert trajectories available. Importantly, our algorithm avoids the need to solve an MDP at each iteration, making it more practical compared to prior AL methods. Finally, we implement a deep variant of our algorithm which shares some similarities to GAIL, but where the discriminator is replaced with the costs learned by OAL. Our simulations suggest that OAL performs well in high dimensional control problems.
Lior Shani, Tom Zahavy, Shie Mannor
AAAI1
2022 Mirror Descent Policy Optimization
Manan Tomar, Lior Shani, Yonathan Efroni, Mohammad Ghavamzadeh
ICLR2
2022 Reinforcement Learning with a Terminator
abstract
We present the problem of reinforcement learning with exogenous termination. We define the Termination Markov Decision Process (TerMDP), an extension of the MDP framework, in which episodes may be interrupted by an external non-Markovian observer. This formulation accounts for numerous real-world situations, such as a human interrupting an autonomous driving agent for reasons of discomfort. We learn the parameters of the TerMDP and leverage the structure of the estimation problem to provide state-wise confidence bounds. We use these to construct a provably-efficient algorithm, which accounts for termination, and bound its regret. Motivated by our theoretical analysis, we design and implement a scalable approach, which combines optimism (w.r.t. termination) and a dynamic discount factor, incorporating the termination probability. We deploy our method on high-dimensional driving and MinAtar benchmarks. Additionally, we test our approach on human data in a driving setting. Our results demonstrate fast convergence and significant improvement over various baseline approaches.
Guy Tennenholtz, Nadav Merlis, Lior Shani, Shie Mannor, Uri Shalit, Gal Chechik, Assaf Hallak, Gal Dalal
NeurIPS3
2020 Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs
abstract
Trust region policy optimization (TRPO) is a popular and empirically successful policy search algorithm in Reinforcement Learning (RL) in which a surrogate problem, that restricts consecutive policies to be ‘close’ to one another, is iteratively solved. Nevertheless, TRPO has been considered a heuristic algorithm inspired by Conservative Policy Iteration (CPI). We show that the adaptive scaling mechanism used in TRPO is in fact the natural “RL version” of traditional trust-region methods from convex analysis. We first analyze TRPO in the planning setting, in which we have access to the model and the entire state space. Then, we consider sample-based TRPO and establish Õ(1/√N) convergence rate to the global optimum. Importantly, the adaptive scaling mechanism allows us to analyze TRPO in regularized MDPs for which we prove fast rates of Õ(1/N), much like results in convex optimization. This is the first result in RL of better rates when regularizing the instantaneous cost or reward.
Lior Shani, Yonathan Efroni, Shie Mannor
AAAI1
2020 Optimistic Policy Optimization with Bandit Feedback
abstract
Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. Yet, so far, such methods have been mostly analyzed from an optimization perspective, without addressing the problem of exploration, or by making strong assumptions on the interaction with the environment. In this paper we consider model-based RL in the tabular finite-horizon MDP setting with unknown transitions and bandit feedback. For this setting, we propose an optimistic trust region policy optimization (TRPO) algorithm for which we establish $\tilde O(\sqrt{S^2 A H^4 K})$ regret for stochastic rewards. Furthermore, we prove $\tilde O( \sqrt{ S^2 A H^4 } K^{2/3} ) $ regret for adversarial rewards. Interestingly, this result matches previous bounds derived for the bandit feedback case, yet with known transitions. To the best of our knowledge, the two results are the first sub-linear regret bounds obtained for policy optimization algorithms with unknown transitions and bandit feedback.
Lior Shani, Yonathan Efroni, Aviv Rosenberg 0002, Shie Mannor
ICML1
2019 Exploration Conscious Reinforcement Learning Revisited
abstract
The Exploration-Exploitation tradeoff arises in Reinforcement Learning when one cannot tell if a policy is optimal. Then, there is a constant need to explore new actions instead of exploiting past experience. In practice, it is common to resolve the tradeoff by using a fixed exploration mechanism, such as $\epsilon$-greedy exploration or by adding Gaussian noise, while still trying to learn an optimal policy. In this work, we take a different approach and study exploration-conscious criteria, that result in optimal policies with respect to the exploration mechanism. Solving these criteria, as we establish, amounts to solving a surrogate Markov Decision Process. We continue and analyze properties of exploration-conscious optimal policies and characterize two general approaches to solve such criteria. Building on the approaches, we apply simple changes in existing tabular and deep Reinforcement Learning algorithms and empirically demonstrate superior performance relatively to their non-exploration-conscious counterparts, both for discrete and continuous action spaces.
Lior Shani, Yonathan Efroni, Shie Mannor
ICML1