Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Olivier Delalleau

dblp:68/2192 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Reinforcement learning · 55% Language models and text generation · 30% Trustworthy machine learning · 12%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reward learning
reward modeling
2.842025
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages · NeurIPS 2025
HelpSteer2-Preference: Complementing Ratings with Preferences · ICLR 2025
HelpSteer 2: Open-source dataset for training top-performing reward models · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.722025
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages · NeurIPS 2025
HelpSteer2-Preference: Complementing Ratings with Preferences · ICLR 2025
Natural language and speech › Language models and text generation
alignment
1.622025
Diverging Preferences: When do Annotators Disagree and do Models Know? · ICML 2025
HelpSteer 2: Open-source dataset for training top-performing reward models · NeurIPS 2024
Machine learning › Trustworthy machine learning
annotator disagreement
0.912025
Diverging Preferences: When do Annotators Disagree and do Models Know? · ICML 2025
Machine learning › Trustworthy machine learning
fairness
0.912025
Diverging Preferences: When do Annotators Disagree and do Models Know? · ICML 2025
Machine learning › Reinforcement learning
human feedback
0.912025
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks · ACL (1) 2025
Natural language and speech › Language models and text generation
instruction tuning
0.912025
HelpSteer2-Preference: Complementing Ratings with Preferences · ICLR 2025
Natural language and speech › Language models and text generation › alignment
pluralistic alignment
0.912025
Diverging Preferences: When do Annotators Disagree and do Models Know? · ICML 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks · ACL (1) 2025
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.812024
IQL-TD-MPC: Implicit Q-Learning for Hierarchical Model Predictive Control · ICRA 2024
Machine learning › Reinforcement learning › offline reinforcement learning
model-based offline reinforcement learning
0.812024
IQL-TD-MPC: Implicit Q-Learning for Hierarchical Model Predictive Control · ICRA 2024
Machine learning › Reinforcement learning
model-based reinforcement learning
0.812024
IQL-TD-MPC: Implicit Q-Learning for Hierarchical Model Predictive Control · ICRA 2024
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
offline goal-conditioned reinforcement learning
0.212024
IQL-TD-MPC: Implicit Q-Learning for Hierarchical Model Predictive Control · ICRA 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.212024
IQL-TD-MPC: Implicit Q-Learning for Hierarchical Model Predictive Control · ICRA 2024
Natural language and speech › Language models and text generation › alignment
preference alignment
0.212024
HelpSteer 2: Open-source dataset for training top-performing reward models · NeurIPS 2024
Machine learning › Deep learning architectures and training
neural network expressivity
0.112011
Shallow vs. Deep Sum-Product Networks · NIPS 2011
Machine learning › Probabilistic and Bayesian machine learning › tractable probabilistic model
sum-product networks
0.112011
Shallow vs. Deep Sum-Product Networks · NIPS 2011
Machine learning › Deep learning architectures and training › feedforward neural network
convex neural network
0.112005
Convex Neural Networks · NIPS 2005
Machine learning › Learning theory
curse of dimensionality
0.112005
The Curse of Highly Variable Functions for Local Kernel Machines · NIPS 2005
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
out-of-sample extension
0.012003
Out-of-Sample Extensions for LLE, Isomap, MDS, Eigenmaps, and Spectral Clustering · NIPS 2003
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel machines
0.012005
The Curse of Highly Variable Functions for Local Kernel Machines · NIPS 2005
Mathematical optimization › continuous optimization
convex optimization
0.012005
Convex Neural Networks · NIPS 2005
Machine learning › Graph learning › graph clustering
spectral clustering
0.012003
Out-of-Sample Extensions for LLE, Isomap, MDS, Eigenmaps, and Spectral Clustering · NIPS 2003

Methods — techniques the papers use, named apart from their topics

reinforcement learning from human feedback · 0.9regression-based reward modeling · 0.9human annotation · 0.9bradley-terry model · 0.9bradley-terry · 0.9REINFORCE · 0.9LLM-as-judge · 0.9temporal difference learning · 0.8model predictive control · 0.8implicit q-learning · 0.8convex optimization · 0.1
YearPublicationVenuePosition
2025 HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
abstract
Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Daniel Egert, Ellie Evans, Hoo-Chang Shin, Felipe Soares, Yi Dong, Oleksii Kuchaiev. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Daniel Egert, Ellie Evans, Hoo-Chang Shin, Felipe Soares, Yi Dong 0003, Oleksii Kuchaiev
ACL (1)3
2025 HelpSteer2-Preference: Complementing Ratings with Preferences
abstract
Reward models are critical for aligning models to follow instructions, and are typically trained following one of two popular paradigms: Bradley-Terry style or Regression style. However, there is a lack of evidence that either approach is better than the other, when adequately matched for data. This is primarily because these approaches require data collected in different (but incompatible) formats, meaning that adequately matched data is not available in existing public datasets. To tackle this problem, we release preference annotations (designed for Bradley-Terry training) to complement existing ratings (designed for Regression style training) in the HelpSteer2 dataset. To improve data interpretability, preference annotations are accompanied with human-written justifications. Using this data, we conduct the first head-to-head comparison of Bradley-Terry and Regression models when adequately matched for data. Based on insights derived from such a comparison, we propose a novel approach to combine Bradley-Terry and Regression reward modeling. A Llama-3.1-70B-Instruct model tuned with this approach scores 94.1 on RewardBench, emerging top of more than 140 reward models as of 1 Oct 2024. This reward model can then be used with REINFORCE algorithm (RLHF) to align an Instruct model to reach 85.0 on Arena Hard, which is No. 1 as of 1 Oct 2024. We open-source this dataset (CC-BY-4.0 license) at https://huggingface.co/datasets/nvidia/HelpSteer2#preferences-new---1-oct-2024 and openly release the trained Reward and Instruct models at https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Reward and https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct .
Zhilin Wang, Alexander Bukharin, Olivier Delalleau, Daniel Egert, Gerald Shen, Jiaqi Zeng, Oleksii Kuchaiev, Yi Dong 0003
ICLR3
2025 Diverging Preferences: When do Annotators Disagree and do Models Know?
abstract
We examine diverging preferences in human-labeled preference datasets. We develop a taxonomy of disagreement sources spanning ten categories across four high-level classes and find that the majority of disagreements are due to factors such as task underspecification or response style. Our findings challenge a standard assumption in reward modeling methods that annotator disagreements can be attributed to simple noise. We then explore how these findings impact two areas of LLM development: reward modeling training and evaluation. In our experiments, we demonstrate how standard reward modeling (e.g., Bradley-Terry) and LLM-as-Judge evaluation methods fail to account for divergence between annotators. These findings highlight challenges in LLM evaluations, which are greatly influenced by divisive features like response style, and in developing pluralistically aligned LLMs. To address these issues, we develop methods for identifying diverging preferences to mitigate their influence in evaluations and during LLM training.
Michael J. Q. Zhang, Zhilin Wang, Jena D. Hwang, Yi Dong 0003, Olivier Delalleau, Yejin Choi 0001, Eunsol Choi, Xiang Ren 0001, Valentina Pyatkin
ICML5
2025 HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
abstract
Preference datasets are essential for training general-domain, instruction-following language models with Reinforcement Learning from Human Feedback (RLHF). Each subsequent data release raises expectations for future data collection, meaning there is a constant need to advance the quality and diversity of openly available preference data. To address this need, we introduce HelpSteer3-Preference, a permissively licensed (CC-BY-4.0), high-quality, human-annotated preference dataset comprising of over 40,000 samples. These samples span diverse real-world applications of large language models (LLMs), including tasks relating to STEM, coding and multilingual scenarios. Using HelpSteer3-Preference, we train Reward Models (RMs) that achieve top performance on RM-Bench (82.4%) and JudgeBench (73.7%). This represents a substantial improvement (~10% absolute) over the previously best-reported results from existing RMs. We demonstrate HelpSteer3-Preference can also be applied to train Generative RMs and how policy models can be aligned with RLHF using our RMs.
Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Hoo-Chang Shin, Felipe Soares, Alexander Bukharin, Ellie Evans, Yi Dong 0003, Oleksii Kuchaiev
NeurIPS3
2024 IQL-TD-MPC: Implicit Q-Learning for Hierarchical Model Predictive Control
abstract
Model-based reinforcement learning (RL) has shown great promise due to its sample efficiency, but still struggles with long-horizon sparse-reward tasks, especially in offline settings where the agent learns from a fixed dataset. We hypothesize that model-based RL agents struggle in these environments due to a lack of long-term planning capabilities, and that planning in a temporally abstract model of the environment can alleviate this issue. In this paper, we make two key contributions: 1) we introduce an offline model-based RL algorithm, IQL-TD-MPC, that extends the state- of-the-art Temporal Difference Learning for Model Predictive Control (TD-MPC) with Implicit Q-Learning (IQL); and 2) we propose to use IQL-TD-MPC as a Manager in a hierarchical setting with any off-the-shelf offline RL algorithm as a Worker. More specifically, we pre-train a temporally abstract IQL-TD-MPC Manager to predict "intent embeddings", which roughly correspond to subgoals, via planning. We show that augmenting state representations with intent embeddings generated by an IQL-TD-MPC manager significantly improves off-the-shelf offline RL agents' performance on some of the most challenging D4RL benchmark tasks. For instance, the offline RL algorithms AWAC, TD3-BC, DT, and CQL all get zero or near-zero normalized evaluation scores on the medium and large antmaze tasks, while our modification gives an average score over 40.
Rohan Chitnis, Yingchen Xu, Bobak Hashemi, Lucas Lehnert, Ürün Dogan, Zheqing Zhu, Olivier Delalleau
ICRA7
2024 HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM
abstract
Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Scowcroft, Neel Kant, Aidan Swope, Oleksii Kuchaiev. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhilin Wang, Yi Dong 0003, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Polak Scowcroft, Neel Kant, Aidan Swope, Oleksii Kuchaiev
NAACL-HLT7
2024 HelpSteer 2: Open-source dataset for training top-performing reward models
abstract
High-quality preference datasets are essential for training reward models that can effectively guide large language models (LLMs) in generating high-quality responses aligned with human preferences.As LLMs become stronger and better aligned, permissively licensed preference datasets, such as Open Assistant, HH-RLHF, and HelpSteer need to be updated to remain effective for reward modeling.Methods that distil preference data from proprietary LLMs such as GPT-4 have restrictions on commercial usage imposed by model providers.To improve upon both generated responses and attribute labeling quality, we release HelpSteer2, a permissively licensed preference dataset (CC-BY-4.0). Using a powerful Nemotron-4-340B base model trained on HelpSteer2, we are able to achieve the SOTA score (92.0%) on Reward-Bench's primary dataset, outperforming currently listed open and proprietary models, as of June 12th, 2024.Notably, HelpSteer2 consists of only ten thousand response pairs, an order of magnitude fewer than existing preference datasets (e.g., HH-RLHF), which makes it highly efficient for training reward models. Our extensive experiments demonstrate that reward models trained with HelpSteer2 are effective in aligning LLMs. Additionally, we propose SteerLM 2.0, a model alignment approach that can effectively make use of the rich multi-attribute score predicted by our reward models. HelpSteer2 is available at https://huggingface.co/datasets/nvidia/HelpSteer2 and code is available at https://github.com/NVIDIA/NeMo-Aligner
Zhilin Wang, Yi Dong 0003, Olivier Delalleau, Jiaqi Zeng, Gerald Shen, Daniel Egert, Jimmy Zhang, Makesh Narsimhan Sreedhar, Oleksii Kuchaiev
NeurIPS3
2012 Detonation Classification from acoustic Signature with the Restricted Boltzmann Machine
abstract
We compare the recently proposed Discriminative Restricted Boltzmann Machine (DRBM) to the classical Support Vector Machine (SVM) on a challenging classification task consisting in identifying weapon classes from audio signals. The three weapon classes considered in this work (mortar, rocket, and rocket‐propelled grenade), are difficult to reliably classify with standard techniques because they tend to have similar acoustic signatures. In addition, specificities of the data available in this study make it challenging to rigorously compare classifiers, and we address methodological issues arising from this situation. Experiments show good classification accuracy that could make these techniques suitable for fielding on autonomous devices. DRBMs appear to yield better accuracy than SVMs, and are less sensitive to the choice of signal preprocessing and model hyperparameters. This last property is especially appealing in such a task where the lack of data makes model validation difficult.
Yoshua Bengio, Nicolas Chapados, Olivier Delalleau, Hugo Larochelle, Xavier Saint-Mleux, Christian Hudon, Jérôme Louradour
Comput. Intell.3
2012 Beyond Skill Rating: Advanced Matchmaking in Ghost Recon Online
abstract
Player satisfaction is particularly difficult to ensure in online games, due to interactions with other players. In adversarial multiplayer games, matchmaking typically consists in trying to match together players of similar skill level. However, this is usually based on a single-skill value, and assumes the only factor of “fun” is the game balance. We present a more advanced matchmaking strategy developed for Ghost Recon Online, an upcoming team-focused first-person shooter (FPS) from Ubisoft (Montreal, QC, Canada). We first show how incorporating more information about players than their raw skill can lead to more balanced matches. We also argue that balance is not the only factor that matters, and present a strategy to explicitly maximize the players' fun, taking advantage of a rich player profile that includes information about player behavior and personal preferences. Ultimately, our goal is to ask players to provide direct feedback on match quality through an in-game survey. However, because such data were not available for this study, we rely here on heuristics tailored to this specific game. Experiments on data collected during Ghost Recon Online's beta tests show that neural networks can effectively be used to predict both balance and player enjoyment.
Olivier Delalleau, Emile Contal, Eric Thibodeau-Laufer, Raul Chandias Ferrari, Yoshua Bengio
IEEE Trans. Comput. Intell. AI Games1
2011 On the Expressive Power of Deep Architectures
Yoshua Bengio, Olivier Delalleau
ALT2
2011 On the Expressive Power of Deep Architectures
Yoshua Bengio, Olivier Delalleau
Discovery Science2
2011 Shallow vs. Deep Sum-Product Networks
abstract
We investigate the representational power of sum-product networks (computation networks analogous to neural networks, but whose individual units compute either products or weighted sums), through a theoretical analysis that compares deep (multiple hidden layers) vs. shallow (one hidden layer) architectures. We prove there exist families of functions that can be represented much more efficiently with a deep network than with a shallow one, i.e. with substantially fewer hidden units. Such results were not available until now, and contribute to motivate recent research involving learning of deep sum-product networks, and more generally motivate research in Deep Learning.
Olivier Delalleau, Yoshua Bengio
NIPS1
2010 Decision trees do not generalize to new variations
abstract
The family of decision tree learning algorithms is among the most widespread and studied. Motivated by the desire to develop learning algorithms that can generalize when learning highly varying functions such as those presumably needed to achieve artificial intelligence, we study some theoretical limitations of decision trees. We demonstrate formally that they can be seriously hurt by the curse of dimensionality in a sense that is a bit different from other nonparametric statistical methods, but most importantly, that they cannot generalize to variations not seen in the training set. This is because a decision tree creates a partition of the input space and needs at least one example in each of the regions associated with a leaf to make a sensible prediction in that region. A better understanding of the fundamental reasons for this limitation suggests that one should use forests or even deeper architectures instead of trees, which provide a form of distributed representation and can generalize to variations not encountered in the training data.
Yoshua Bengio, Olivier Delalleau, Clarence Simard
Comput. Intell.2
2009 Justifying and Generalizing Contrastive Divergence
abstract
We study an expansion of the log likelihood in undirected graphical models such as the restricted Boltzmann machine (RBM), where each term in the expansion is associated with a sample in a Gibbs chain alternating between two random variables (the visible vector and the hidden vector in RBMs). We are particularly interested in estimators of the gradient of the log likelihood obtained through this expansion. We show that its residual term converges to zero, justifying the use of a truncation--running only a short Gibbs chain, which is the main idea behind the contrastive divergence (CD) estimator of the log-likelihood gradient. By truncating even more, we obtain a stochastic reconstruction error, related through a mean-field approximation to the reconstruction error often used to train autoassociators and stacked autoassociators. The derivation is not specific to the particular parametric forms used in RBMs and requires only convergence of the Gibbs chain. We present theoretical and empirical evidence linking the number of Gibbs steps k and the magnitude of the RBM parameters to the bias in the CD estimator. These experiments also suggest that the sign of the CD estimator is correct most of the time, even when the bias is large, so that CD-k is a good descent direction even for small k.
Yoshua Bengio, Olivier Delalleau
Neural Comput.2
2005 The Curse of Highly Variable Functions for Local Kernel Machines
abstract
We present a series of theoretical arguments supporting the claim that a large class of modern learning algorithms that rely solely on the smoothness prior with similarity between examples expressed with a local kernel are sensitive to the curse of dimensionality, or more precisely to the variability of the target. Our discussion covers supervised, semisupervised and unsupervised learning algorithms. These algorithms are found to be local in the sense that crucial properties of the learned function at x depend mostly on the neighbors of x in the training set. This makes them sensitive to the curse of dimensionality, well studied for classical non-parametric statistical learning. We show in the case of the Gaussian kernel that when the function to be learned has many variations, these algorithms require a number of training examples proportional to the number of variations, which could be large even though there may exist short descriptions of the target function, i.e. their Kolmogorov complexity may be low. This suggests that there exist non-local learning algorithms that at least have the potential to learn about such structured but apparently complex functions (because locally they have many variations), while not using very specific prior domain knowledge.
Yoshua Bengio, Olivier Delalleau, Nicolas Le Roux
NIPS2
2005 Convex Neural Networks
abstract
Convexity has recently received a lot of attention in the machine learning community, and the lack of convexity has been seen as a major disadvantage of many learning algorithms, such as multi-layer artificial neural networks. We show that training multi-layer neural networks in which the number of hidden units is learned can be viewed as a convex optimization problem. This problem involves an infinite number of variables, but can be solved by incrementally inserting a hidden unit at a time, each time finding a linear classifier that minimizes a weighted sum of errors.
Yoshua Bengio, Nicolas Le Roux, Pascal Vincent, Olivier Delalleau, Patrice Marcotte
NIPS4
2004 Learning Eigenfunctions Links Spectral Embedding and Kernel PCA
abstract
In this letter, we show a direct relation between spectral embedding methods and kernel principal components analysis and how both are special cases of a more general learning problem: learning the principal eigenfunctions of an operator defined from a kernel and the unknown data-generating density. Whereas spectral embedding methods provided only coordinates for the training points, the analysis justifies a simple extension to out-of-sample examples (the Nyström formula) for multidimensional scaling (MDS), spectral clustering, Laplacian eigenmaps, locally linear embedding (LLE), and Isomap. The analysis provides, for all such spectral embedding methods, the definition of a loss function, whose empirical average is minimized by the traditional algorithms. The asymptotic expected value of that loss defines a generalization performance and clarifies what these algorithms are trying to learn. Experiments with LLE, Isomap, spectral clustering, and MDS show that this out-of-sample embedding formula generalizes well, with a level of error comparable to the effect of small perturbations of the training set on the embedding.
Yoshua Bengio, Olivier Delalleau, Nicolas Le Roux, Jean-François Paiement, Pascal Vincent, Marie Ouimet
Neural Comput.2
2003 Out-of-Sample Extensions for LLE, Isomap, MDS, Eigenmaps, and Spectral Clustering
abstract
Several unsupervised learning algorithms based on an eigendecompo- sition provide either an embedding or a clustering only for given train- ing points, with no straightforward extension for out-of-sample examples short of recomputing eigenvectors. This paper provides a unified frame- work for extending Local Linear Embedding (LLE), Isomap, Laplacian Eigenmaps, Multi-Dimensional Scaling (for dimensionality reduction) as well as for Spectral Clustering. This framework is based on seeing these algorithms as learning eigenfunctions of a data-dependent kernel. Numerical experiments show that the generalizations performed have a level of error comparable to the variability of the embedding algorithms due to the choice of training data.
Yoshua Bengio, Jean-François Paiement, Pascal Vincent, Olivier Delalleau, Nicolas Le Roux, Marie Ouimet
NIPS4