Nathan Grinsztajn

dblp:278/3008 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-6817-5972ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 58% Language models and text generation · 25% Multi-agent systems · 15%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
combinatorial optimization
1.322023
Winner Takes It All: Training Performant RL Populations for Combinatorial Optimization · NeurIPS 2023
Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023
Machine learning › Reinforcement learning
value-based reinforcement learning
0.912025
ShiQ: Bringing back Bellman to LLMs · NeurIPS 2025
Natural language and speech › Language models and text generation
alignment
0.812024
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion · EMNLP 2024
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments
0.812024
Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX · ICLR 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs · ICML 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination
0.812024
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs · ICML 2024
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent debate
0.812024
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs · ICML 2024
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.812024
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion · EMNLP 2024
Machine learning › Reinforcement learning
reinforcement learning environment
0.812024
Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX · ICLR 2024
Machine learning › Reinforcement learning
policy learning
0.712023
Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023
Machine learning › Reinforcement learning
exploration
0.512021
There Is No Turning Back: A Self-Supervised Approach for Reversibility-Aware Reinforcement Learning · NeurIPS 2021
Machine learning › Reinforcement learning › exploration › intrinsic motivation
self-supervised exploration
0.512021
There Is No Turning Back: A Self-Supervised Approach for Reversibility-Aware Reinforcement Learning · NeurIPS 2021
Natural language and speech › Language models and text generation › test-time scaling
inference-time search
0.212023
Winner Takes It All: Training Performant RL Populations for Combinatorial Optimization · NeurIPS 2023
Mathematical optimization › combinatorial optimization › network optimization
routing and scheduling
0.212023
Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.6token-wise learning · 0.9off-policy learning · 0.9bellman equation · 0.9self-consistency · 0.8prompting · 0.8ensembling · 0.8contrastive policy gradient · 0.8actor-critic · 0.8JAX · 0.8stochastic sampling · 0.7latent space search · 0.7beam search · 0.7
YearPublicationVenuePosition
2025 Memory-Enhanced Neural Solvers for Routing Problems
abstract
Routing Problems are central to many real-world applications, yet remain challenging due to their (NP-)hard nature. Amongst existing approaches, heuristics often offer the best trade-off between quality and scalability, making them suitable for industrial use. While Reinforcement Learning (RL) offers a flexible framework for designing heuristics, its adoption over handcrafted heuristics remains incomplete. Existing learned methods still lack the ability to adapt to specific instances and fully leverage the available computational budget. Current best methods either rely on a collection of pre-trained policies, or on RL fine-tuning; hence failing to fully utilize newly available information within the constraints of the budget. In response, we present MEMENTO, an approach that leverages memory to improve the search of neural solvers at inference. MEMENTO updates the action distribution dynamically based on the outcome of previous decisions. We validate its effectiveness on Traveling Salesman and Capacitated Vehicle Routing problems, demonstrating its superiority over tree-search and policy-gradient fine-tuning; and showing that it can be zero-shot combined with diversity-based solvers. We successfully train all RL auto-regressive solvers on large instances, and verify MEMENTO's scalability and data-efficiency: pushing the state-of-the-art on 11 out of 12 evaluated tasks.
Félix Chalumeau, Refiloe Shabe, Noah de Nicola, Arnu Pretorius, Tom Barrett, Nathan Grinsztajn
NeurIPS6
2025 ShiQ: Bringing back Bellman to LLMs
abstract
The fine-tuning of pre-trained large language models (LLMs) using reinforcement learning (RL) is generally formulated as direct policy optimization. This approach was naturally favored as it efficiently improves a pretrained LLM with simple gradient updates. Another RL paradigm, Q-learning methods, has received far less attention in the LLM community while demonstrating major success in various non-LLM RL tasks. In particular, Q-learning effectiveness stems from its sample efficiency and ability to learn offline, which is particularly valuable given the high computational cost of sampling with LLM. However, naively applying a Q-learning–style update to the model’s logits is ineffective due to the specificity of LLMs. Our contribution is to derive theoretically grounded loss functions from Bellman equations to adapt Q-learning methods to LLMs. To do so, we interpret LLM logits as Q-values and carefully adapt insights from the RL literature to account for LLM-specific characteristics. It thereby ensures that the logits become reliable Q-value estimates. We then use this loss to build a practical algorithm, ShiQ for Shifted-Q, that supports off-policy, token-wise learning while remaining simple to implement. Finally, ShiQ is evaluated on both synthetic data and real-world benchmarks, e.g., UltraFeedback, BFCL-V3, demonstrating its effectiveness in both single-turn and multi-turn LLM settings.
Pierre Clavier, Nathan Grinsztajn, Raphaël Avalos, Yannis Flet-Berliac, Irem Ergün, Omar Darwiche Domingues, Olivier Pietquin, Pierre H. Richemond, Florian Strub, Matthieu Geist
NeurIPS2
2024 Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
abstract
Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi, Bill Wu, Chris Cremer, Arash Ahmadian, Yash Chandak, Mohammad Gheshlaghi Azar, Olivier Pietquin, Matthieu Geist. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi, Bill Wu, Chris Cremer, Arash Ahmadian, Yash Chandak, Mohammad Gheshlaghi Azar, Olivier Pietquin, Matthieu Geist
EMNLP2
2024 Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX
abstract
Open-source reinforcement learning (RL) environments have played a crucial role in driving progress in the development of AI algorithms. In modern RL research, there is a need for simulated environments that are performant, scalable, and modular to enable their utilization in a wider range of potential real-world applications. Therefore, we present Jumanji, a suite of diverse RL environments specifically designed to be fast, flexible, and scalable. Jumanji provides a suite of environments focusing on combinatorial problems frequently encountered in industry, as well as challenging general decision-making tasks. By leveraging the efficiency of JAX and hardware accelerators like GPUs and TPUs, Jumanji enables rapid iteration of research ideas and large-scale experimentation, ultimately empowering more capable agents. Unlike existing RL environment suites, Jumanji is highly customizable, allowing users to tailor the initial state distribution and problem complexity to their needs. Furthermore, we provide actor-critic baselines for each environment, accompanied by preliminary findings on scaling and generalization scenarios. Jumanji aims to set a new standard for speed, adaptability, and scalability of RL environments.
Clément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana, Sasha Abramowitz, Paul Duckworth, Vincent Coyette, Laurence Illing Midgley, Elshadai Tegegn, Tristan Kalloniatis, Omayma Mahjoub, Matthew Macfarlane, Andries P. Smit, Nathan Grinsztajn, Raphaël Boige, Cemlyn N. Waters, Mohamed A. Mimouni, Ulrich A. Mbou Sob, Ruan de Kock, Siddarth Singh, Daniel Furelos-Blanco, Victor Le, Arnu Pretorius, Alexandre Laterre
ICLR14
2024 Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
abstract
Recent advancements in large language models (LLMs) underscore their potential for responding to inquiries in various domains. However, ensuring that generative agents provide accurate and reliable answers remains an ongoing challenge. In this context, multi-agent debate (MAD) has emerged as a promising strategy for enhancing the truthfulness of LLMs. We benchmark a range of debating and prompting strategies to explore the trade-offs between cost, time, and accuracy. Importantly, we find that multi-agent debating systems, in their current form, do not reliably outperform other proposed prompting strategies, such as self-consistency and ensembling using multiple reasoning paths. However, when performing hyperparameter tuning, several MAD systems, such as Multi-Persona, perform better. This suggests that MAD protocols might not be inherently worse than other approaches, but that they are more sensitive to different hyperparameter settings and difficult to optimize. We build on these results to offer insights into improving debating strategies, such as adjusting agent agreement levels, which can significantly enhance performance and even surpass all other non-debate protocols we evaluated. We provide an open-source repository to the community with several state-of-the-art protocols together with evaluation scripts to benchmark across popular research datasets.
Andries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett, Arnu Pretorius
ICML2
2023 Combinatorial Optimization with Policy Adaptation using Latent Space Search
abstract
Combinatorial Optimization underpins many real-world applications and yet, designing performant algorithms to solve these complex, typically NP-hard, problems remains a significant research challenge. Reinforcement Learning (RL) provides a versatile framework for designing heuristics across a broad spectrum of problem domains. However, despite notable progress, RL has not yet supplanted industrial solvers as the go-to solution. Current approaches emphasize pre-training heuristics that construct solutions, but often rely on search procedures with limited variance, such as stochastically sampling numerous solutions from a single policy, or employing computationally expensive fine-tuning of the policy on individual problem instances. Building on the intuition that performant search at inference time should be anticipated during pre-training, we propose COMPASS, a novel RL approach that parameterizes a distribution of diverse and specialized policies conditioned on a continuous latent space. We evaluate COMPASS across three canonical problems - Travelling Salesman, Capacitated Vehicle Routing, and Job-Shop Scheduling - and demonstrate that our search strategy (i) outperforms state-of-the-art approaches in 9 out of 11 standard benchmarking tasks and (ii) generalizes better, surpassing all other approaches on a set of 18 procedurally transformed instance distributions.
Félix Chalumeau, Shikha Surana, Clément Bonnet, Nathan Grinsztajn, Arnu Pretorius, Alexandre Laterre, Tom Barrett
NeurIPS4
2023 Winner Takes It All: Training Performant RL Populations for Combinatorial Optimization
abstract
Applying reinforcement learning (RL) to combinatorial optimization problems is attractive as it removes the need for expert knowledge or pre-solved instances. However, it is unrealistic to expect an agent to solve these (often NP-)hard problems in a single shot at inference due to their inherent complexity. Thus, leading approaches often implement additional search strategies, from stochastic sampling and beam-search to explicit fine-tuning. In this paper, we argue for the benefits of learning a population of complementary policies, which can be simultaneously rolled out at inference. To this end, we introduce Poppy, a simple training procedure for populations. Instead of relying on a predefined or hand-crafted notion of diversity, Poppy induces an unsupervised specialization targeted solely at maximizing the performance of the population. We show that Poppy produces a set of complementary policies, and obtains state-of-the-art RL results on three popular NP-hard problems: traveling salesman, capacitated vehicle routing, and job-shop scheduling.
Nathan Grinsztajn, Daniel Furelos-Blanco, Shikha Surana, Clément Bonnet, Tom Barrett
NeurIPS1
2022 Meta-learning from Learning Curves: Challenge Design and Baseline Results
abstract
Meta-Iearning has been widely studied and implemented in many Automated Machine Learning systems to improve the process of selecting and training Machine Learning models for new tasks, by leveraging expertise acquired on previously observed tasks. We design a novel meta-learning challenge aiming at learning-to-learn from one of the most essential model evaluation data, the learning curve. It consists of multiple model evaluations collected during the process of training. A meta-learner is expected to apply a learned policy to learning curves of partially trained models on the task at hand, to rapidly find the best task solution, without training all potential models to convergence. This implies learning the exploration-exploitation trade-off. Our challenge is split into two phases: a development phase and a final test phase. In each phase, a meta-learner is meta-trained and meta-tested on validation learning curves (development phase) or test learning curves (final test phase). During meta-training, the meta-learner is allowed to learn from the provided learning curves in any possible way. In meta-testing, we borrowed the common Reinforcement Learning setting in which an agent (a meta-learner) learns by interacting with an environment storing pre-computed learning curves. A meta-learner must pay a cost (corresponding to the actual training and testing time) to reveal learning curve information progressively. The meta-learner is evaluated and ranked based on the average area under its learning curves. This challenge was accepted as part of the official selection of WCCI 2022 competitions.
Lisheng Sun-Hosoya, Nathan Grinsztajn, Isabelle Guyon
IJCNN3
2021 READYS: A Reinforcement Learning Based Strategy for Heterogeneous Dynamic Scheduling
abstract
In this paper, we propose READYS, a reinforcement learning algorithm for the dynamic scheduling of computations modeled as a Directed Acyclic Graph (DAGs). Our goal is to develop a scheduling algorithm in which allocation and scheduling decisions are made at runtime, based on the state of the system, as performed in runtime systems such as StarPU or ParSEC. Reinforcement Learning is a natural candidate to achieve this task, since its general principle is to build step by step a strategy that, given the state of the system (the state of the resources and a view of the ready tasks and their successors in our case), makes a decision to optimize a global criterion. Moreover, the use of Reinforcement Learning is natural in a context where the duration of tasks (and communications) is stochastic. We propose READYS that combines Graph Convolutional Networks (GCN) with an Actor-Critic Algorithm (A2C): it builds an adaptive representation of the scheduling problem on the fly and learns a scheduling strategy, aiming at minimizing the makespan. A crucial point is that READYS builds a general scheduling strategy which is neither limited to only one specific application or task graph nor one particular problem size, and that can be used to schedule any DAG. We focus on different types of task graphs originating from linear algebra factorization kernels (CHOLESKY, LU, QR) and we consider heterogeneous platforms made of a few CPUs and GPUs. We first propose to analyze the performance of READYS when learning is performed on a given (platform, kernel, problem size) combination. Using simulations, we show that the scheduling agent obtains performances very similar or even superior to algorithms from the literature, and that it is especially powerful when the scheduling environment contains a lot of uncertainty. We additionally demonstrate that our agent exhibits very promising generalization capabilities. To the best of our knowledge, this is the first paper which shows that reinforcement learning can really be used for dynamic DAG scheduling on heterogeneous resources.
Nathan Grinsztajn, Olivier Beaumont, Emmanuel Jeannot, Philippe Preux
CLUSTER1
2021 There Is No Turning Back: A Self-Supervised Approach for Reversibility-Aware Reinforcement Learning
abstract
We propose to learn to distinguish reversible from irreversible actions for better informed decision-making in Reinforcement Learning (RL). From theoretical considerations, we show that approximate reversibility can be learned through a simple surrogate task: ranking randomly sampled trajectory events in chronological order. Intuitively, pairs of events that are always observed in the same order are likely to be separated by an irreversible sequence of actions. Conveniently, learning the temporal order of events can be done in a fully self-supervised way, which we use to estimate the reversibility of actions from experience, without any priors.We propose two different strategies that incorporate reversibility in RL agents, one strategy for exploration (RAE) and one strategy for control (RAC). We demonstrate the potential of reversibility-aware agents in several environments, including the challenging Sokoban game. In synthetic tasks, we show that we can learn control policies that never fail and reduce to zero the side-effects of interactions, even without access to the reward function.
Nathan Grinsztajn, Johan Ferret, Olivier Pietquin, Philippe Preux, Matthieu Geist
NeurIPS1