Charline Le Lan

dblp:234/9001 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 53% Deep learning architectures and training · 18% Language models and text generation · 14%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
1.322023
Understanding Self-Predictive Learning for Reinforcement Learning · ICML 2023
Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks · ICLR 2023
Natural language and speech › Language models and text generation › alignment
human alignment
0.812024
Human Alignment of Large Language Models through Online Preference Optimisation · ICML 2024
Natural language and speech › Language models and text generation
preference optimization
0.812024
Human Alignment of Large Language Models through Online Preference Optimisation · ICML 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Human Alignment of Large Language Models through Online Preference Optimisation · ICML 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
auxiliary tasks
0.712023
Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks · ICLR 2023
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation
0.712023
Bootstrapped Representations in Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning
temporal difference learning
0.712023
Bootstrapped Representations in Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning
value-based reinforcement learning
0.712023
Understanding Self-Predictive Learning for Reinforcement Learning · ICML 2023
Machine learning › Deep learning architectures and training
equivariant neural network
0.512021
LieTransformer: Equivariant Self-Attention for Lie Groups · ICML 2021
Machine learning › Deep learning architectures and training › equivariant neural network
equivariant self-attention
0.512021
LieTransformer: Equivariant Self-Attention for Lie Groups · ICML 2021
Machine learning › Deep learning architectures and training › equivariant neural network
lie group equivariance
0.512021
LieTransformer: Equivariant Self-Attention for Lie Groups · ICML 2021
Machine learning › Reinforcement learning
markov decision process
0.512021
Metrics and Continuity in Reinforcement Learning · AAAI 2021
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.512021
LieTransformer: Equivariant Self-Attention for Lie Groups · ICML 2021
Machine learning › Representation and self-supervised learning
hierarchical representation
0.412019
Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning
0.412019
Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders · NeurIPS 2019
Machine learning › Generative modeling
variational autoencoder
0.412019
Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning
0.212023
Understanding Self-Predictive Learning for Reinforcement Learning · ICML 2023
Machine learning › Generative modeling › generative adversarial network
mode collapse mitigation
0.212023
Understanding Self-Predictive Learning for Reinforcement Learning · ICML 2023
Computer vision › 3D vision
point cloud processing
0.112021
LieTransformer: Equivariant Self-Attention for Lie Groups · ICML 2021
Machine learning › Reinforcement learning
reinforcement learning theory
0.112021
Metrics and Continuity in Reinforcement Learning · AAAI 2021

Methods — techniques the papers use, named apart from their topics

self-play · 0.8nash mirror descent · 0.8direct policy optimization · 0.8value function · 0.7temporal difference learning · 0.7spectral decomposition · 0.7semi-gradient updates · 0.7residual gradient · 0.7monte carlo · 0.7auxiliary tasks · 0.7
YearPublicationVenuePosition
2024 Human Alignment of Large Language Models through Online Preference Optimisation
abstract
Ensuring alignment of language model’s outputs with human preferences is critical to guarantee a useful, safe, and pleasant user experience. Thus, human alignment has been extensively studied recently and several methods such as Reinforcement Learning from Human Feedback (RLHF), Direct Policy Optimisation (DPO) and Sequence Likelihood Calibration (SLiC) have emerged. In this paper, our contribution is two-fold. First, we show the equivalence between two recent alignment methods, namely Identity Policy Optimisation (IPO) and Nash Mirror Descent (Nash-MD). Second, we introduce a generalisation of IPO, named IPO-MD, that leverages the regularised sampling approach proposed by Nash-MD. This equivalence may seem surprising at first sight, since IPO is an offline method whereas Nash-MD is an online method using a preference model. However, this equivalence can be proven when we consider the online version of IPO, that is when both generations are sampled by the online policy and annotated by a trained preference model. Optimising the IPO loss with such a stream of data becomes then equivalent to finding the Nash equilibrium of the preference model through self-play. Building on this equivalence, we introduce the IPO-MD algorithm that generates data with a mixture policy (between the online and reference policy) similarly as the general Nash-MD algorithm. We compare online-IPO and IPO-MD to different online versions of existing losses on preference data such as DPO and SLiC on a summarisation task.
Daniele Calandriello, Zhaohan Guo, Rémi Munos, Mark Rowland 0001, Yunhao Tang, Bernardo Ávila Pires, Pierre H. Richemond, Charline Le Lan, Michal Valko, Tianqi Liu 0002, Rishabh Joshi, Bilal Piot
ICML8
2023 A Novel Stochastic Gradient Descent Algorithm for Learning Principal Subspaces
abstract
Many machine learning problems encode their data as a matrix with a possibly very large number of rows and columns. In several applications like neuroscience, image compression or deep reinforcement learning, the principal subspace of such a matrix provides a useful, low-dimensional representation of individual data. Here, we are interested in determining the $d$-dimensional principal subspace of a given matrix from sample entries, i.e. from small random submatrices. Although a number of sample-based methods exist for this problem (e.g. Oja’s rule (Oja, 1982)), these assume access to full columns of the matrix or particular matrix structure such as symmetry and cannot be combined as-is with neural networks (Baldi et al., 1989). In this paper, we derive an algorithm that learns a principal subspace from sample entries, can be applied when the approximate subspace is represented by a neural network, and hence can be scaled to datasets with an effectively infinite number of rows and columns. Our method consists in defining a loss function whose minimizer is the desired principal subspace, and constructing a gradient estimate of this loss whose bias can be controlled. We complement our theoretical analysis with a series of experiments on synthetic matrices, the MNIST dataset (LeCun et al. 2010) and the reinforcement learning domain PuddleWorld (Sutton, 1995) demonstrating the usefulness of our approach.
Charline Le Lan, Joshua Greaves, Jesse Farebrother, Mark Rowland 0001, Fabian Pedregosa, Rishabh Agarwal, Marc G. Bellemare
AISTATS1
2023 Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks
Jesse Farebrother, Joshua Greaves, Rishabh Agarwal, Charline Le Lan, Ross Goroshin, Pablo Samuel Castro, Marc G. Bellemare
ICLR4
2023 Bootstrapped Representations in Reinforcement Learning
abstract
In reinforcement learning (RL), state representations are key to dealing with large or continuous state spaces. While one of the promises of deep learning algorithms is to automatically construct features well-tuned for the task they try to solve, such a representation might not emerge from end-to-end training of deep RL agents. To mitigate this issue, auxiliary objectives are often incorporated into the learning process and help shape the learnt state representation. Bootstrapping methods are today's method of choice to make these additional predictions. Yet, it is unclear which features these algorithms capture and how they relate to those from other auxiliary-task-based approaches. In this paper, we address this gap and provide a theoretical characterization of the state representation learnt by temporal difference learning (Sutton, 1988). Surprisingly, we find that this representation differs from the features learned by Monte Carlo and residual gradient algorithms for most transition structures of the environment in the policy evaluation setting. We describe the efficacy of these representations for policy evaluation, and use our theoretical analysis to design new auxiliary learning rules. We complement our theoretical results with an empirical comparison of these learning rules for different cumulant functions on classic domains such as the four-room domain (Sutton et al, 1999) and Mountain Car (Moore, 1990).
Charline Le Lan, Stephen Tu, Mark Rowland 0001, Anna Harutyunyan, Rishabh Agarwal, Marc G. Bellemare, Will Dabney
ICML1
2023 Understanding Self-Predictive Learning for Reinforcement Learning
abstract
We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their own future latent representations. Despite its recent empirical success, such algorithms have an apparent defect: trivial representations (such as constants) minimize the prediction error, yet it is obviously undesirable to converge to such solutions. Our central insight is that careful designs of the optimization dynamics are critical to learning meaningful representations. We identify that a faster paced optimization of the predictor and semi-gradient updates on the representation, are crucial to preventing the representation collapse. Then in an idealized setup, we show self-predictive learning dynamics carries out spectral decomposition on the state transition matrix, effectively capturing information of the transition dynamics. Building on the theoretical insights, we propose bidirectional self-predictive learning, a novel self-predictive algorithm that learns two representations simultaneously. We examine the robustness of our theoretical insights with a number of small-scale experiments and showcase the promise of the novel representation learning algorithm with large-scale experiments.
Yunhao Tang, Zhaohan Guo, Pierre H. Richemond, Bernardo Ávila Pires, Yash Chandak, Rémi Munos, Mark Rowland 0001, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, András György 0001, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko
ICML9
2022 On the Generalization of Representations in Reinforcement Learning
abstract
In reinforcement learning, state representations are used to tractably deal with large problem spaces. State representations serve both to approximate the value function with few parameters, but also to generalize to newly encountered states. Their features may be learned implicitly (as part of a neural network) or explicitly (for example, the successor representation of Dayan(1993). While the approximation properties of representations are reasonably well-understood, a precise characterization of how and when these representations generalize is lacking. In this work, we address this gap and provide an informative bound on the generalization error arising from a specific state representation. This bound is based on the notion of effective dimension which measures the degree to which knowing the value at one state informs the value at other states. Our bound applies to any state representation and quantifies the natural tension between representations that generalize well and those that approximate well. We complement our theoretical results with an empirical survey of classic representation learning methods from the literature and results on the Arcade Learning Environment, and find that the generalization behaviour of learned representations is well-explained by their effective dimension.
Charline Le Lan, Stephen Tu, Adam M. Oberman, Rishabh Agarwal, Marc G. Bellemare
AISTATS1
2021 Metrics and Continuity in Reinforcement Learning
abstract
In most practical applications of reinforcement learning, it is untenable to maintain direct estimates for individual states; in continuous-state systems, it is impossible. Instead, researchers often leverage {\em state similarity} (whether explicitly or implicitly) to build models that can generalize well from a limited set of samples. The notion of state similarity used, and the neighbourhoods and topologies they induce, is thus of crucial importance, as it will directly affect the performance of the algorithms. Indeed, a number of recent works introduce algorithms assuming the existence of "well-behaved" neighbourhoods, but leave the full specification of such topologies for future work. In this paper we introduce a unified formalism for defining these topologies through the lens of metrics. We establish a hierarchy amongst these metrics and demonstrate their theoretical implications on the Markov Decision Process specifying the reinforcement learning problem. We complement our theoretical results with empirical evaluations showcasing the differences between the metrics considered.
Charline Le Lan, Marc G. Bellemare, Pablo Samuel Castro
AAAI1
2021 LieTransformer: Equivariant Self-Attention for Lie Groups
abstract
Group equivariant neural networks are used as building blocks of group invariant neural networks, which have been shown to improve generalisation performance and data efficiency through principled parameter sharing. Such works have mostly focused on group equivariant convolutions, building on the result that group equivariant linear maps are necessarily convolutions. In this work, we extend the scope of the literature to self-attention, that is emerging as a prominent building block of deep learning models. We propose the LieTransformer, an architecture composed of LieSelfAttention layers that are equivariant to arbitrary Lie groups and their discrete subgroups. We demonstrate the generality of our approach by showing experimental results that are competitive to baseline methods on a wide range of tasks: shape counting on point clouds, molecular property regression and modelling particle trajectories under Hamiltonian dynamics.
Michael J. Hutchinson, Charline Le Lan, Sheheryar Zaidi, Emilien Dupont, Yee Whye Teh, Hyunjik Kim
ICML2
2019 Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders
abstract
The Variational Auto-Encoder (VAE) is a popular method for learning a generative model and embeddings of the data. Many real datasets are hierarchically structured. However, traditional VAEs map data in a Euclidean latent space which cannot efficiently embed tree-like structures. Hyperbolic spaces with negative curvature can. We therefore endow VAEs with a Poincaré ball model of hyperbolic geometry as a latent space and rigorously derive the necessary methods to work with two main Gaussian generalisations on that space. We empirically show better generalisation to unseen data than the Euclidean counterpart, and can qualitatively and quantitatively better recover hierarchical structures.
Emile Mathieu, Charline Le Lan, Chris J. Maddison, Ryota Tomioka, Yee Whye Teh
NeurIPS2