EDBT 2026 Demo / reviewers in the wild / expert
Bogdan Mazoure
dblp:209/7660
· DBLP profile ↗
14ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-5117-6864ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Reinforcement learning · 43% Vision and language · 11% Robot manipulation · 10% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.6 | 2 | 2025 | From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons · CVPR 2025 Grounding Multimodal Large Language Models in Actions · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation
zero-shot transfer |
1.1 | 2 | 2022 | Improving Zero-Shot Generalization in Offline Reinforcement Learning using Generalized Similarity Functions · NeurIPS 2022 Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL · ICLR 2022 |
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
1.0 | 2 | 2022 | Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL · ICLR 2022 Deep Reinforcement and InfoMax Learning · NeurIPS 2020 |
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent |
0.9 | 1 | 2025 | From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons · CVPR 2025 |
Machine learning › Reinforcement learning
embodied agent training |
0.9 | 1 | 2025 | From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons · CVPR 2025 |
Machine learning › Reinforcement learning › online decision making
online reinforcement learning |
0.9 | 1 | 2025 | From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons · CVPR 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | On the Modeling Capabilities of Large Language Models for Sequential Decision Making · ICLR 2025 |
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
language-conditioned reinforcement learning |
0.8 | 1 | 2024 | Large Language Models as Generalizable Policies for Embodied Tasks · ICLR 2024 |
Robotics › Robot manipulation
learning from demonstration |
0.7 | 1 | 2023 | Learning About Progress From Experts · ICLR 2023 |
Machine learning › Reinforcement learning
reward learning |
0.7 | 1 | 2023 | Learning About Progress From Experts · ICLR 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | Improving Zero-Shot Generalization in Offline Reinforcement Learning using Generalized Similarity Functions · NeurIPS 2022 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.6 | 1 | 2022 | Improving Zero-Shot Generalization in Offline Reinforcement Learning using Generalized Similarity Functions · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning
mutual information maximization |
0.4 | 1 | 2020 | Deep Reinforcement and InfoMax Learning · NeurIPS 2020 |
Machine learning › Reinforcement learning › multi-armed bandit
adversarial bandit |
0.4 | 1 | 2019 | On-Line Adaptative Curriculum Learning for GANs · AAAI 2019 |
Machine learning › Generative modeling › generative adversarial network
GAN training |
0.4 | 1 | 2019 | On-Line Adaptative Curriculum Learning for GANs · AAAI 2019 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | On-Line Adaptative Curriculum Learning for GANs · AAAI 2019 |
Machine learning › Reinforcement learning
multi-armed bandit |
0.4 | 1 | 2019 | On-Line Adaptative Curriculum Learning for GANs · AAAI 2019 |
Bioinformatics and computational biology › drug discovery
high-throughput screening |
0.3 | 1 | 2017 | Detecting and removing multiplicative spatial bias in high-throughput screening technologies · Bioinform. 2017 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.6supervised learning · 0.9online reinforcement learning · 0.9fine-tuning · 0.9action tokenizer · 0.9semantic action alignment · 0.8reinforcement learning · 0.8action tokenization · 0.8progress estimation · 0.7expert demonstration · 0.7statistical bias correction · 0.3multiplicative bias model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Multimodal LLMs to Generalist Embodied Agents: Methods and LessonsabstractWe examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typically trained on. Specifically, our focus lies in areas such as Embodied AI, Games, UI Control, and Planning. To this end, we introduce a process of adapting an MLLM to a Generalist Embodied Agent (GEA). GEA is a single unified model capable of grounding itself across these varied domains through a multi-embodiment action tokenizer. GEA is trained with supervised learning on a large dataset of embodied experiences and with online RL in interactive simulators. We explore the data and algorithmic choices necessary to develop such a model. Our findings reveal the importance of training with cross-domain data and online RL for building generalist agents. The final GEA model achieves strong generalization performance to unseen tasks across diverse benchmarks compared to other generalist models and benchmark-specific approaches. Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, R. Devon Hjelm, Zhe Gan, Zsolt Kira, Alexander Toshev |
CVPR | 2 |
| 2025 | On the Modeling Capabilities of Large Language Models for Sequential Decision MakingabstractLarge pretrained models are showing increasingly better performance in reasoning and planning tasks across different modalities, opening the possibility to leverage them for complex sequential decision making problems. In this paper, we investigate the capabilities of Large Language Models (LLMs) for reinforcement learning (RL) across a diversity of interactive domains. We evaluate their ability to produce decision-making policies, either directly, by generating actions, or indirectly, by first generating reward models to train an agent with RL. Our results show that, even without task-specific fine-tuning, LLMs excel at reward modeling. In particular, crafting rewards through artificial intelligence (AI) feedback yields the most generally applicable approach and can enhance performance by improving credit assignment and exploration. Finally, in environments with unfamiliar dynamics, we explore how fine-tuning LLMs with synthetic data can significantly improve their reward modeling capabilities while mitigating catastrophic forgetting, further broadening their utility in sequential decision-making tasks. Martin Klissarov, R. Devon Hjelm, Alexander Toshev, Bogdan Mazoure |
ICLR | 4 |
| 2024 | Large Language Models as Generalizable Policies for Embodied TasksabstractWe show that large language models (LLMs) can be adapted to be generalizable policies for embodied visual tasks. Our approach, called Large LAnguage model Reinforcement Learning Policy (LLaRP), adapts a pre-trained frozen LLM to take as input text instructions and visual egocentric observations and output actions directly in the environment. Using reinforcement learning, we train LLaRP to see and act solely through environmental interactions. We show that LLaRP is robust to complex paraphrasings of task instructions and can generalize to new tasks that require novel optimal behavior. In particular, on 1,000 unseen tasks it achieves 42% success rate, 1.7x the success rate of other common learned baselines or zero-shot applications of LLMs. Finally, to aid the community in studying language conditioned, massively multi-task, embodied AI problems we release a novel benchmark, Language Rearrangement, consisting of 150,000 training and 1,000 testing tasks for language-conditioned rearrangement. Andrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure, Rin Metcalf Susa, Walter Talbott, Natalie Mackraz, R. Devon Hjelm, Alexander Toshev |
ICLR | 4 |
| 2024 | Grounding Multimodal Large Language Models in ActionsabstractMultimodal Large Language Models (MLLMs) have demonstrated a wide range of capabilities across many domains including Embodied AI. In this work, we study how to best ground a MLLM into different embodiments and their associated action spaces, including both continuous and discrete actions. For continuous actions, a set of learned tokenizations that capture an action at various resolutions allows for sufficient modeling precision, yielding the best performance on downstream tasks. For discrete actions, semantically aligning these actions with the native output token space of the MLLM leads to the strongest performance. We arrive at these lessons via a thorough study of seven action grounding approaches on five different environments, encompassing over 114 embodied tasks. Andrew Szot, Bogdan Mazoure, Harsh Agrawal, R. Devon Hjelm, Zsolt Kira, Alexander Toshev |
NeurIPS | 2 |
| 2023 | Learning About Progress From Experts
Jake Bruce, Ankit Anand, Bogdan Mazoure, Rob Fergus |
ICLR | 3 |
| 2022 | Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL
Bogdan Mazoure, Ahmed M. Ahmed 0004, R. Devon Hjelm, Andrey Kolobov, Patrick MacAlpine |
ICLR | 1 |
| 2022 | Improving Zero-Shot Generalization in Offline Reinforcement Learning using Generalized Similarity FunctionsabstractReinforcement learning (RL) agents are widely used for solving complex sequential decision-making tasks, but still exhibit difficulty generalizing to scenarios not seen during training. While prior online approaches demonstrated that using additional signals beyond the reward function can lead to better generalization capabilities in RL agents, i.e. using self-supervised learning (SSL), they struggle in the offline RL setting, i.e. learning from a static dataset. We show that the performance of online algorithms for generalization in RL can be hindered in the offline setting due to poor estimation of similarity between observations. We propose a new theoretically-motivated framework called Generalized Similarity Functions (GSF), which uses contrastive learning to train an offline RL agent to aggregate observations based on the similarity of their expected future behavior, where we quantify this similarity using generalized value functions. We show that GSF is general enough to recover existing SSL objectives while improving zero-shot generalization performance on two complex pixel-based offline RL benchmarks. Bogdan Mazoure, Ilya Kostrikov, Ofir Nachum, Jonathan Tompson |
NeurIPS | 1 |
| 2022 | Low-Rank Representation of Reinforcement Learning PoliciesabstractWe propose a general framework for policy representation for reinforcement learning tasks. This framework involves finding a low-dimensional embedding of the policy on a reproducing kernel Hilbert space (RKHS). The usage of RKHS based methods allows us to derive strong theoretical guarantees on the expected return of the reconstructed policy. Such guarantees are typically lacking in black-box models, but are very desirable in tasks requiring stability and convergence guarantees. We conduct several experiments on classic RL domains. The results confirm that the policies can be robustly represented in a low-dimensional space while the embedded policy incurs almost no decrease in returns. Bogdan Mazoure, Thang Doan, Tianyu Li 0008, Vladimir Makarenkov, Joelle Pineau, Doina Precup, Guillaume Rabusseau |
J. Artif. Intell. Res. | 1 |
| 2021 | A Theoretical Analysis of Catastrophic Forgetting through the NTK Overlap MatrixabstractContinual learning (CL) is a setting in which an agent has to learn from an incoming stream of data during its entire lifetime. Although major advances have been made in the field, one recurring problem which remains unsolved is that of Catastrophic Forgetting (CF). While the issue has been extensively studied empirically, little attention has been paid from a theoretical angle. In this paper, we show that the impact of CF increases as two tasks increasingly align. We introduce a measure of task similarity called the NTK overlap matrix which is at the core of CF. We analyze common projected gradient algorithms and demonstrate how they mitigate forgetting. Then, we propose a variant of Orthogonal Gradient Descent (OGD) which leverages structure of the data through Principal Component Analysis (PCA). Experiments support our theoretical findings and show how our method can help reduce CF on classical CL datasets. Thang Doan, Mehdi Abbana Bennani, Bogdan Mazoure, Guillaume Rabusseau, Pierre Alquier |
AISTATS | 3 |
| 2020 | Efficient Planning under Partial Observability with Unnormalized Q Functions and Spectral LearningabstractLearning and planning in partially-observable domains is one of the most difficult problems in reinforcement learning. Traditional methods consider these two problems as independent, resulting in a classic two-stage paradigm: first learn the environment dynamics and then compute the optimal policy accordingly. This approach, however, disconnects the reward information from the learning of the environment model and can consequently lead to representations that are sample inefficient and time consuming for planning purpose. In this paper, we propose a novel algorithm that incorporate reward information into the representations of the environment to unify these two stages. Our algorithm is closely related to the spectral learning algorithm for predicitive state representations and offers appealing theoretical guarantees and time complexity. We empirically show on two domains that our approach is more sample and time efficient compared to classical methods. Tianyu Li 0008, Bogdan Mazoure, Doina Precup, Guillaume Rabusseau |
AISTATS | 2 |
| 2020 | Deep Reinforcement and InfoMax LearningabstractWe posit that a reinforcement learning (RL) agent will perform better when it uses representations that are better at predicting the future, particularly in terms of few-shot learning and domain adaptation. To test that hypothesis, we introduce an objective based on Deep InfoMax (DIM) which trains the agent to predict the future by maximizing the mutual information between its internal representation of successive timesteps. We provide an intuitive analysis of the convergence properties of our approach from the perspective of Markov chain mixing times, and argue that convergence of the lower bound on mutual information is related to the inverse absolute spectral gap of the transition model. We test our approach in several synthetic settings, where it successfully learns representations that are predictive of the future. Finally, we augment C51, a strong distributional RL agent, with our temporal DIM objective and demonstrate on a continual learning task (inspired by Ms.~PacMan) and on the recently introduced Procgen environment that our approach improves performance, which supports our core hypothesis. Bogdan Mazoure, Remi Tachet des Combes, Thang Doan, Philip Bachman, R. Devon Hjelm |
NeurIPS | 1 |
| 2019 | On-Line Adaptative Curriculum Learning for GANsabstractGenerative Adversarial Networks (GANs) can successfully approximate a probability distribution and produce realistic samples. However, open questions such as sufficient convergence conditions and mode collapse still persist. In this paper, we build on existing work in the area by proposing a novel framework for training the generator against an ensemble of discriminator networks, which can be seen as a one-student/multiple-teachers setting. We formalize this problem within the full-information adversarial bandit framework, where we evaluate the capability of an algorithm to select mixtures of discriminators for providing the generator with feedback during learning. To this end, we propose a reward function which reflects the progress made by the generator and dynamically update the mixture weights allocated to each discriminator. We also draw connections between our algorithm and stochastic optimization methods and then show that existing approaches using multiple discriminators in literature can be recovered from our framework. We argue that less expressive discriminators are smoother and have a general coarse grained view of the modes map, which enforces the generator to cover a wide portion of the data distribution support. On the other hand, highly expressive discriminators ensure samples quality. Finally, experimental results show that our approach improves samples quality and diversity over existing baselines by effectively learning a curriculum. These results also support the claim that weaker discriminators have higher entropy improving modes coverage. Thang Doan, João Monteiro 0002, Isabela Albuquerque, Bogdan Mazoure, Audrey Durand, Joelle Pineau, R. Devon Hjelm |
AAAI | 4 |
| 2019 | Exploring Attention Mechanism for Acoustic-based Classification of Speech Utterances into System-directed and Non-system-directedabstractVoice controlled virtual assistants (VAs) are now available in smartphones, cars, and standalone devices in homes. In most cases, the user needs to first "wake-up" the VA by saying a particular word/phrase every time he/she wants the VA to do something. Eliminating the need for saying the wake-up word for every interaction could improve the user experience. This would require the VA to have the capability of understanding whether the user is talking to it or not. In other words, the challenge is to distinguish between system-directed and non-system-directed speech utterances. In this paper, we present a number of neural network architectures for tackling this classification problem based on using only the acoustic signal. It is shown that a model comprised of convolutional, recurrent, and feed-forward layers can achieve an equal error rate (EER) of below 20% for this task. In addition, we investigate the use of an attention mechanism for helping the model to focus on the more important parts of the signal and to improve handling of variable length inputs sequences. The results show that the proposed attention mechanism significantly improves the model accuracy achieving an EER of 16.25% and 15.62% on two distinct realistic datasets. Atta Norouzian, Bogdan Mazoure, Dermot Connolly, Daniel Willett |
ICASSP | 2 |
| 2017 | Detecting and removing multiplicative spatial bias in high-throughput screening technologiesabstractMOTIVATION: Considerable attention has been paid recently to improve data quality in high-throughput screening (HTS) and high-content screening (HCS) technologies widely used in drug development and chemical toxicity research. However, several environmentally- and procedurally-induced spatial biases in experimental HTS and HCS screens decrease measurement accuracy, leading to increased numbers of false positives and false negatives in hit selection. Although effective bias correction methods and software have been developed over the past decades, almost all of these tools have been designed to reduce the effect of additive bias only. Here, we address the case of multiplicative spatial bias. RESULTS: We introduce three new statistical methods meant to reduce multiplicative spatial bias in screening technologies. We assess the performance of the methods with synthetic and real data affected by multiplicative spatial bias, including comparisons with current bias correction methods. We also describe a wider data correction protocol that integrates methods for removing both assay and plate-specific spatial biases, which can be either additive or multiplicative. CONCLUSIONS: The methods for removing multiplicative spatial bias and the data correction protocol are effective in detecting and cleaning experimental data generated by screening technologies. As our protocol is of a general nature, it can be used by researchers analyzing current or next-generation high-throughput screens. AVAILABILITY AND IMPLEMENTATION: The AssayCorrector program, implemented in R, is available on CRAN. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Iurie Caraus, Bogdan Mazoure, Robert Nadon, Vladimir Makarenkov |
Bioinform. | 2 |