Bogdan Mazoure

dblp:209/7660 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-5117-6864ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 43% Vision and language · 11% Robot manipulation · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.622025
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons · CVPR 2025
Grounding Multimodal Large Language Models in Actions · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
1.122022
Improving Zero-Shot Generalization in Offline Reinforcement Learning using Generalized Similarity Functions · NeurIPS 2022
Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL · ICLR 2022
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
1.022022
Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL · ICLR 2022
Deep Reinforcement and InfoMax Learning · NeurIPS 2020
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent
0.912025
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons · CVPR 2025
Machine learning › Reinforcement learning
embodied agent training
0.912025
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons · CVPR 2025
Machine learning › Reinforcement learning › online decision making
online reinforcement learning
0.912025
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons · CVPR 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
On the Modeling Capabilities of Large Language Models for Sequential Decision Making · ICLR 2025
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
language-conditioned reinforcement learning
0.812024
Large Language Models as Generalizable Policies for Embodied Tasks · ICLR 2024
Robotics › Robot manipulation
learning from demonstration
0.712023
Learning About Progress From Experts · ICLR 2023
Machine learning › Reinforcement learning
reward learning
0.712023
Learning About Progress From Experts · ICLR 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Improving Zero-Shot Generalization in Offline Reinforcement Learning using Generalized Similarity Functions · NeurIPS 2022
Machine learning › Reinforcement learning
offline reinforcement learning
0.612022
Improving Zero-Shot Generalization in Offline Reinforcement Learning using Generalized Similarity Functions · NeurIPS 2022
Machine learning › Representation and self-supervised learning
mutual information maximization
0.412020
Deep Reinforcement and InfoMax Learning · NeurIPS 2020
Machine learning › Reinforcement learning › multi-armed bandit
adversarial bandit
0.412019
On-Line Adaptative Curriculum Learning for GANs · AAAI 2019
Machine learning › Generative modeling › generative adversarial network
GAN training
0.412019
On-Line Adaptative Curriculum Learning for GANs · AAAI 2019
Machine learning › Generative modeling
generative adversarial network
0.412019
On-Line Adaptative Curriculum Learning for GANs · AAAI 2019
Machine learning › Reinforcement learning
multi-armed bandit
0.412019
On-Line Adaptative Curriculum Learning for GANs · AAAI 2019
Bioinformatics and computational biology › drug discovery
high-throughput screening
0.312017
Detecting and removing multiplicative spatial bias in high-throughput screening technologies · Bioinform. 2017

Methods — techniques the papers use, named apart from their topics

large language model · 1.6supervised learning · 0.9online reinforcement learning · 0.9fine-tuning · 0.9action tokenizer · 0.9semantic action alignment · 0.8reinforcement learning · 0.8action tokenization · 0.8progress estimation · 0.7expert demonstration · 0.7statistical bias correction · 0.3multiplicative bias model · 0.3
YearPublicationVenuePosition
2025 From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
abstract
We examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typically trained on. Specifically, our focus lies in areas such as Embodied AI, Games, UI Control, and Planning. To this end, we introduce a process of adapting an MLLM to a Generalist Embodied Agent (GEA). GEA is a single unified model capable of grounding itself across these varied domains through a multi-embodiment action tokenizer. GEA is trained with supervised learning on a large dataset of embodied experiences and with online RL in interactive simulators. We explore the data and algorithmic choices necessary to develop such a model. Our findings reveal the importance of training with cross-domain data and online RL for building generalist agents. The final GEA model achieves strong generalization performance to unseen tasks across diverse benchmarks compared to other generalist models and benchmark-specific approaches.
Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, R. Devon Hjelm, Zhe Gan, Zsolt Kira, Alexander Toshev
CVPR2
2025 On the Modeling Capabilities of Large Language Models for Sequential Decision Making
abstract
Large pretrained models are showing increasingly better performance in reasoning and planning tasks across different modalities, opening the possibility to leverage them for complex sequential decision making problems. In this paper, we investigate the capabilities of Large Language Models (LLMs) for reinforcement learning (RL) across a diversity of interactive domains. We evaluate their ability to produce decision-making policies, either directly, by generating actions, or indirectly, by first generating reward models to train an agent with RL. Our results show that, even without task-specific fine-tuning, LLMs excel at reward modeling. In particular, crafting rewards through artificial intelligence (AI) feedback yields the most generally applicable approach and can enhance performance by improving credit assignment and exploration. Finally, in environments with unfamiliar dynamics, we explore how fine-tuning LLMs with synthetic data can significantly improve their reward modeling capabilities while mitigating catastrophic forgetting, further broadening their utility in sequential decision-making tasks.
Martin Klissarov, R. Devon Hjelm, Alexander Toshev, Bogdan Mazoure
ICLR4
2024 Large Language Models as Generalizable Policies for Embodied Tasks
abstract
We show that large language models (LLMs) can be adapted to be generalizable policies for embodied visual tasks. Our approach, called Large LAnguage model Reinforcement Learning Policy (LLaRP), adapts a pre-trained frozen LLM to take as input text instructions and visual egocentric observations and output actions directly in the environment. Using reinforcement learning, we train LLaRP to see and act solely through environmental interactions. We show that LLaRP is robust to complex paraphrasings of task instructions and can generalize to new tasks that require novel optimal behavior. In particular, on 1,000 unseen tasks it achieves 42% success rate, 1.7x the success rate of other common learned baselines or zero-shot applications of LLMs. Finally, to aid the community in studying language conditioned, massively multi-task, embodied AI problems we release a novel benchmark, Language Rearrangement, consisting of 150,000 training and 1,000 testing tasks for language-conditioned rearrangement.
Andrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure, Rin Metcalf Susa, Walter Talbott, Natalie Mackraz, R. Devon Hjelm, Alexander Toshev
ICLR4
2024 Grounding Multimodal Large Language Models in Actions
abstract
Multimodal Large Language Models (MLLMs) have demonstrated a wide range of capabilities across many domains including Embodied AI. In this work, we study how to best ground a MLLM into different embodiments and their associated action spaces, including both continuous and discrete actions. For continuous actions, a set of learned tokenizations that capture an action at various resolutions allows for sufficient modeling precision, yielding the best performance on downstream tasks. For discrete actions, semantically aligning these actions with the native output token space of the MLLM leads to the strongest performance. We arrive at these lessons via a thorough study of seven action grounding approaches on five different environments, encompassing over 114 embodied tasks.
Andrew Szot, Bogdan Mazoure, Harsh Agrawal, R. Devon Hjelm, Zsolt Kira, Alexander Toshev
NeurIPS2
2023 Learning About Progress From Experts
Jake Bruce, Ankit Anand, Bogdan Mazoure, Rob Fergus
ICLR3
2022 Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL
Bogdan Mazoure, Ahmed M. Ahmed 0004, R. Devon Hjelm, Andrey Kolobov, Patrick MacAlpine
ICLR1
2022 Improving Zero-Shot Generalization in Offline Reinforcement Learning using Generalized Similarity Functions
abstract
Reinforcement learning (RL) agents are widely used for solving complex sequential decision-making tasks, but still exhibit difficulty generalizing to scenarios not seen during training. While prior online approaches demonstrated that using additional signals beyond the reward function can lead to better generalization capabilities in RL agents, i.e. using self-supervised learning (SSL), they struggle in the offline RL setting, i.e. learning from a static dataset. We show that the performance of online algorithms for generalization in RL can be hindered in the offline setting due to poor estimation of similarity between observations. We propose a new theoretically-motivated framework called Generalized Similarity Functions (GSF), which uses contrastive learning to train an offline RL agent to aggregate observations based on the similarity of their expected future behavior, where we quantify this similarity using generalized value functions. We show that GSF is general enough to recover existing SSL objectives while improving zero-shot generalization performance on two complex pixel-based offline RL benchmarks.
Bogdan Mazoure, Ilya Kostrikov, Ofir Nachum, Jonathan Tompson
NeurIPS1
2022 Low-Rank Representation of Reinforcement Learning Policies
abstract
We propose a general framework for policy representation for reinforcement learning tasks. This framework involves finding a low-dimensional embedding of the policy on a reproducing kernel Hilbert space (RKHS). The usage of RKHS based methods allows us to derive strong theoretical guarantees on the expected return of the reconstructed policy. Such guarantees are typically lacking in black-box models, but are very desirable in tasks requiring stability and convergence guarantees. We conduct several experiments on classic RL domains. The results confirm that the policies can be robustly represented in a low-dimensional space while the embedded policy incurs almost no decrease in returns.
Bogdan Mazoure, Thang Doan, Tianyu Li 0008, Vladimir Makarenkov, Joelle Pineau, Doina Precup, Guillaume Rabusseau
J. Artif. Intell. Res.1
2021 A Theoretical Analysis of Catastrophic Forgetting through the NTK Overlap Matrix
abstract
Continual learning (CL) is a setting in which an agent has to learn from an incoming stream of data during its entire lifetime. Although major advances have been made in the field, one recurring problem which remains unsolved is that of Catastrophic Forgetting (CF). While the issue has been extensively studied empirically, little attention has been paid from a theoretical angle. In this paper, we show that the impact of CF increases as two tasks increasingly align. We introduce a measure of task similarity called the NTK overlap matrix which is at the core of CF. We analyze common projected gradient algorithms and demonstrate how they mitigate forgetting. Then, we propose a variant of Orthogonal Gradient Descent (OGD) which leverages structure of the data through Principal Component Analysis (PCA). Experiments support our theoretical findings and show how our method can help reduce CF on classical CL datasets.
Thang Doan, Mehdi Abbana Bennani, Bogdan Mazoure, Guillaume Rabusseau, Pierre Alquier
AISTATS3
2020 Efficient Planning under Partial Observability with Unnormalized Q Functions and Spectral Learning
abstract
Learning and planning in partially-observable domains is one of the most difficult problems in reinforcement learning. Traditional methods consider these two problems as independent, resulting in a classic two-stage paradigm: first learn the environment dynamics and then compute the optimal policy accordingly. This approach, however, disconnects the reward information from the learning of the environment model and can consequently lead to representations that are sample inefficient and time consuming for planning purpose. In this paper, we propose a novel algorithm that incorporate reward information into the representations of the environment to unify these two stages. Our algorithm is closely related to the spectral learning algorithm for predicitive state representations and offers appealing theoretical guarantees and time complexity. We empirically show on two domains that our approach is more sample and time efficient compared to classical methods.
Tianyu Li 0008, Bogdan Mazoure, Doina Precup, Guillaume Rabusseau
AISTATS2
2020 Deep Reinforcement and InfoMax Learning
abstract
We posit that a reinforcement learning (RL) agent will perform better when it uses representations that are better at predicting the future, particularly in terms of few-shot learning and domain adaptation. To test that hypothesis, we introduce an objective based on Deep InfoMax (DIM) which trains the agent to predict the future by maximizing the mutual information between its internal representation of successive timesteps. We provide an intuitive analysis of the convergence properties of our approach from the perspective of Markov chain mixing times, and argue that convergence of the lower bound on mutual information is related to the inverse absolute spectral gap of the transition model. We test our approach in several synthetic settings, where it successfully learns representations that are predictive of the future. Finally, we augment C51, a strong distributional RL agent, with our temporal DIM objective and demonstrate on a continual learning task (inspired by Ms.~PacMan) and on the recently introduced Procgen environment that our approach improves performance, which supports our core hypothesis.
Bogdan Mazoure, Remi Tachet des Combes, Thang Doan, Philip Bachman, R. Devon Hjelm
NeurIPS1
2019 On-Line Adaptative Curriculum Learning for GANs
abstract
Generative Adversarial Networks (GANs) can successfully approximate a probability distribution and produce realistic samples. However, open questions such as sufficient convergence conditions and mode collapse still persist. In this paper, we build on existing work in the area by proposing a novel framework for training the generator against an ensemble of discriminator networks, which can be seen as a one-student/multiple-teachers setting. We formalize this problem within the full-information adversarial bandit framework, where we evaluate the capability of an algorithm to select mixtures of discriminators for providing the generator with feedback during learning. To this end, we propose a reward function which reflects the progress made by the generator and dynamically update the mixture weights allocated to each discriminator. We also draw connections between our algorithm and stochastic optimization methods and then show that existing approaches using multiple discriminators in literature can be recovered from our framework. We argue that less expressive discriminators are smoother and have a general coarse grained view of the modes map, which enforces the generator to cover a wide portion of the data distribution support. On the other hand, highly expressive discriminators ensure samples quality. Finally, experimental results show that our approach improves samples quality and diversity over existing baselines by effectively learning a curriculum. These results also support the claim that weaker discriminators have higher entropy improving modes coverage.
Thang Doan, João Monteiro 0002, Isabela Albuquerque, Bogdan Mazoure, Audrey Durand, Joelle Pineau, R. Devon Hjelm
AAAI4
2019 Exploring Attention Mechanism for Acoustic-based Classification of Speech Utterances into System-directed and Non-system-directed
abstract
Voice controlled virtual assistants (VAs) are now available in smartphones, cars, and standalone devices in homes. In most cases, the user needs to first "wake-up" the VA by saying a particular word/phrase every time he/she wants the VA to do something. Eliminating the need for saying the wake-up word for every interaction could improve the user experience. This would require the VA to have the capability of understanding whether the user is talking to it or not. In other words, the challenge is to distinguish between system-directed and non-system-directed speech utterances. In this paper, we present a number of neural network architectures for tackling this classification problem based on using only the acoustic signal. It is shown that a model comprised of convolutional, recurrent, and feed-forward layers can achieve an equal error rate (EER) of below 20% for this task. In addition, we investigate the use of an attention mechanism for helping the model to focus on the more important parts of the signal and to improve handling of variable length inputs sequences. The results show that the proposed attention mechanism significantly improves the model accuracy achieving an EER of 16.25% and 15.62% on two distinct realistic datasets.
Atta Norouzian, Bogdan Mazoure, Dermot Connolly, Daniel Willett
ICASSP2
2017 Detecting and removing multiplicative spatial bias in high-throughput screening technologies
abstract
MOTIVATION: Considerable attention has been paid recently to improve data quality in high-throughput screening (HTS) and high-content screening (HCS) technologies widely used in drug development and chemical toxicity research. However, several environmentally- and procedurally-induced spatial biases in experimental HTS and HCS screens decrease measurement accuracy, leading to increased numbers of false positives and false negatives in hit selection. Although effective bias correction methods and software have been developed over the past decades, almost all of these tools have been designed to reduce the effect of additive bias only. Here, we address the case of multiplicative spatial bias. RESULTS: We introduce three new statistical methods meant to reduce multiplicative spatial bias in screening technologies. We assess the performance of the methods with synthetic and real data affected by multiplicative spatial bias, including comparisons with current bias correction methods. We also describe a wider data correction protocol that integrates methods for removing both assay and plate-specific spatial biases, which can be either additive or multiplicative. CONCLUSIONS: The methods for removing multiplicative spatial bias and the data correction protocol are effective in detecting and cleaning experimental data generated by screening technologies. As our protocol is of a general nature, it can be used by researchers analyzing current or next-generation high-throughput screens. AVAILABILITY AND IMPLEMENTATION: The AssayCorrector program, implemented in R, is available on CRAN. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Iurie Caraus, Bogdan Mazoure, Robert Nadon, Vladimir Makarenkov
Bioinform.2