Yannis Flet-Berliac

dblp:239/5247 · also Yannis Paul Raymond Flet-Berliac · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 78% Language models and text generation · 10% Planning, search and constraint satisfaction · 4%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
offline reinforcement learning
2.742024
OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators · NeurIPS 2024
Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets · NeurIPS 2023
Model-Based Offline Reinforcement Learning with Local Misspecification · AAAI 2023
Machine learning › Reinforcement learning › policy optimization
policy gradient
1.732024
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion · EMNLP 2024
Learning Value Functions in Deep Policy Gradients using Residual Variance · ICLR 2021
Only Relevant Information Matters: Filtering Out Noisy Samples To Boost RL · IJCAI 2020
Machine learning › Reinforcement learning
value-based reinforcement learning
0.912025
ShiQ: Bringing back Bellman to LLMs · NeurIPS 2025
Natural language and speech › Language models and text generation
alignment
0.812024
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion · EMNLP 2024
Machine learning › Reinforcement learning › off-policy evaluation
estimator selection
0.812024
OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators · NeurIPS 2024
Machine learning › Reinforcement learning
off-policy evaluation
0.812024
OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators · NeurIPS 2024
Machine learning › Reinforcement learning
policy evaluation
0.812024
OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators · NeurIPS 2024
Machine learning › Reinforcement learning › offline reinforcement learning
decision transformer
0.712023
Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets · NeurIPS 2023
Machine learning › Reinforcement learning › offline reinforcement learning
model-based offline reinforcement learning
0.712023
Model-Based Offline Reinforcement Learning with Local Misspecification · AAAI 2023
Machine learning › Reinforcement learning
policy selection
0.712023
Model-Based Offline Reinforcement Learning with Local Misspecification · AAAI 2023
Machine learning › Reinforcement learning › safe reinforcement learning
safe policy improvement
0.712023
Model-Based Offline Reinforcement Learning with Local Misspecification · AAAI 2023
Machine learning › Reinforcement learning
supervised learning for RL
0.712023
Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets · NeurIPS 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
algorithm selection
0.612022
Data-Efficient Pipeline for Offline Reinforcement Learning with Limited Data · NeurIPS 2022
Machine learning › Optimization for machine learning
hyperparameter optimization
0.612022
Data-Efficient Pipeline for Offline Reinforcement Learning with Limited Data · NeurIPS 2022
Machine learning › Reinforcement learning
actor-critic methods
0.512021
Adversarially Guided Actor-Critic · ICLR 2021
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.512021
Adversarially Guided Actor-Critic · ICLR 2021
Machine learning › Reinforcement learning
value function estimation
0.512021
Learning Value Functions in Deep Policy Gradients using Residual Variance · ICLR 2021
Machine learning › Reinforcement learning
sample efficiency
0.412020
Only Relevant Information Matters: Filtering Out Noisy Samples To Boost RL · IJCAI 2020
Machine learning › Deep learning architectures and training
transformer
0.212023
Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

token-wise learning · 0.9off-policy learning · 0.9bellman equation · 0.9statistical estimation · 0.8reweighting · 0.8contrastive policy gradient · 0.8waypoint conditioning · 0.7sequence modeling · 0.7pessimism approximation · 0.7lower bound analysis · 0.7
YearPublicationVenuePosition
2025 ShiQ: Bringing back Bellman to LLMs
abstract
The fine-tuning of pre-trained large language models (LLMs) using reinforcement learning (RL) is generally formulated as direct policy optimization. This approach was naturally favored as it efficiently improves a pretrained LLM with simple gradient updates. Another RL paradigm, Q-learning methods, has received far less attention in the LLM community while demonstrating major success in various non-LLM RL tasks. In particular, Q-learning effectiveness stems from its sample efficiency and ability to learn offline, which is particularly valuable given the high computational cost of sampling with LLM. However, naively applying a Q-learning–style update to the model’s logits is ineffective due to the specificity of LLMs. Our contribution is to derive theoretically grounded loss functions from Bellman equations to adapt Q-learning methods to LLMs. To do so, we interpret LLM logits as Q-values and carefully adapt insights from the RL literature to account for LLM-specific characteristics. It thereby ensures that the logits become reliable Q-value estimates. We then use this loss to build a practical algorithm, ShiQ for Shifted-Q, that supports off-policy, token-wise learning while remaining simple to implement. Finally, ShiQ is evaluated on both synthetic data and real-world benchmarks, e.g., UltraFeedback, BFCL-V3, demonstrating its effectiveness in both single-turn and multi-turn LLM settings.
Pierre Clavier, Nathan Grinsztajn, Raphaël Avalos, Yannis Flet-Berliac, Irem Ergün, Omar Darwiche Domingues, Olivier Pietquin, Pierre H. Richemond, Florian Strub, Matthieu Geist
NeurIPS4
2024 Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
abstract
Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi, Bill Wu, Chris Cremer, Arash Ahmadian, Yash Chandak, Mohammad Gheshlaghi Azar, Olivier Pietquin, Matthieu Geist. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi, Bill Wu, Chris Cremer, Arash Ahmadian, Yash Chandak, Mohammad Gheshlaghi Azar, Olivier Pietquin, Matthieu Geist
EMNLP1
2024 OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators
abstract
Offline policy evaluation (OPE) allows us to evaluate and estimate a new sequential decision-making policy's performance by leveraging historical interaction data collected from other policies. Evaluating a new policy online without a confident estimate of its performance can lead to costly, unsafe, or hazardous outcomes, especially in education and healthcare. Several OPE estimators have been proposed in the last decade, many of which have hyperparameters and require training. Unfortunately, choosing the best OPE algorithm for each task and domain is still unclear. In this paper, we propose a new algorithm that adaptively blends a set of OPE estimators given a dataset without relying on an explicit selection using a statistical procedure. We prove that our estimator is consistent and satisfies several desirable properties for policy evaluation. Additionally, we demonstrate that when compared to alternative approaches, our estimator can be used to select higher-performing policies in healthcare and robotics. Our work contributes to improving ease of use for a general-purpose, estimator-agnostic, off-policy evaluation framework for offline RL.
Allen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath, Yannis Flet-Berliac, Emma Brunskill
NeurIPS5
2023 Model-Based Offline Reinforcement Learning with Local Misspecification
abstract
We present a model-based offline reinforcement learning policy performance lower bound that explicitly captures dynamics model misspecification and distribution mismatch and we propose an empirical algorithm for optimal offline policy selection. Theoretically, we prove a novel safe policy improvement theorem by establishing pessimism approximations to the value function. Our key insight is to jointly consider selecting over dynamics models and policies: as long as a dynamics model can accurately represent the dynamics of the state-action pairs visited by a given policy, it is possible to approximate the value of that particular policy. We analyze our lower bound in the LQR setting and also show competitive performance to previous lower bounds on policy selection across a set of D4RL tasks.
Kefan Dong, Yannis Flet-Berliac, Allen Nie, Emma Brunskill
AAAI2
2023 Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets
abstract
Despite the recent advancements in offline reinforcement learning via supervised learning (RvS) and the success of the decision transformer (DT) architecture in various domains, DTs have fallen short in several challenging benchmarks. The root cause of this underperformance lies in their inability to seamlessly connect segments of suboptimal trajectories. To overcome this limitation, we present a novel approach to enhance RvS methods by integrating intermediate targets. We introduce the Waypoint Transformer (WT), using an architecture that builds upon the DT framework and conditioned on automatically-generated waypoints. The results show a significant increase in the final return compared to existing RvS methods, with performance on par or greater than existing state-of-the-art temporal difference learning-based methods. Additionally, the performance and stability improvements are largest in the most challenging environments and data configurations, including AntMaze Large Play/Diverse and Kitchen Mixed/Partial.
Anirudhan Badrinath, Yannis Flet-Berliac, Allen Nie, Emma Brunskill
NeurIPS2
2022 Data-Efficient Pipeline for Offline Reinforcement Learning with Limited Data
abstract
Offline reinforcement learning (RL) can be used to improve future performance by leveraging historical data. There exist many different algorithms for offline RL, and it is well recognized that these algorithms, and their hyperparameter settings, can lead to decision policies with substantially differing performance. This prompts the need for pipelines that allow practitioners to systematically perform algorithm-hyperparameter selection for their setting. Critically, in most real-world settings, this pipeline must only involve the use of historical data. Inspired by statistical model selection methods for supervised learning, we introduce a task- and method-agnostic pipeline for automatically training, comparing, selecting, and deploying the best policy when the provided dataset is limited in size. In particular, our work highlights the importance of performing multiple data splits to produce more reliable algorithm-hyperparameter selection. While this is a common approach in supervised learning, to our knowledge, this has not been discussed in detail in the offline RL setting. We show it can have substantial impacts when the dataset is small. Compared to alternate approaches, our proposed pipeline outputs higher-performing deployed policies from a broad range of offline policy learning algorithms and across various simulation domains in healthcare, education, and robotics. This work contributes toward the development of a general-purpose meta-algorithm for automatic algorithm-hyperparameter selection for offline RL.
Allen Nie, Yannis Flet-Berliac, Deon R. Jordan, William Steenbergen, Emma Brunskill
NeurIPS2
2022 Offline policy optimization with eligible actions
abstract
Offline policy optimization could have a large impact on many real-world decision-making problems, as online learning may be infeasible in many applications. Importance sampling and its variants are a common used type of estimator in offline policy evaluation, and such estimators typically do not require assumptions on the properties and representational capabilities of value function or decision process model function classes. In this paper, we identify an important overfitting phenomenon in optimizing the importance weighted return, in which it may be possible for the learned policy to essentially avoid making aligned decisions for part of the initial state space. We propose an algorithm to avoid this overfitting through a new per-state-neighborhood normalization constraint, and provide a theoretical justification of the proposed algorithm. We also show the limitations of previous attempts to this approach. We test our algorithm in a healthcare-inspired simulator, a logged dataset collected from real hospitals and continuous control tasks. These experiments show the proposed method yields less overfitting and better test performance compared to state-of-the-art batch reinforcement learning algorithms.
Yao Liu 0009, Yannis Flet-Berliac, Emma Brunskill
UAI2
2021 Adversarially Guided Actor-Critic
Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux, Matthieu Geist
ICLR1
2021 Learning Value Functions in Deep Policy Gradients using Residual Variance
Yannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard, Philippe Preux
ICLR1
2020 Only Relevant Information Matters: Filtering Out Noisy Samples To Boost RL
abstract
In reinforcement learning, policy gradient algorithms optimize the policy directly and rely on sampling efficiently an environment. Nevertheless, while most sampling procedures are based on direct policy sampling, self-performance measures could be used to improve such sampling prior to each policy update. Following this line of thought, we introduce SAUNA, a method where non-informative transitions are rejected from the gradient update. The level of information is estimated according to the fraction of variance explained by the value function: a measure of the discrepancy between V and the empirical returns. In this work, we use this criterion to select samples that are useful to learn from, and we demonstrate that this selection can significantly improve the performance of policy gradient methods. In this paper: (a) We introduce the SAUNA method to filter transitions. (b) We conduct experiments on a set of benchmark continuous control problems. SAUNA significantly improves performance. (c) We investigate how SAUNA reliably selects samples with the most positive impact on learning and study its improvement on both performance and sample efficiency.
Yannis Flet-Berliac, Philippe Preux
IJCAI1