EDBT 2026 Demo / reviewers in the wild / expert
Giuseppe Paolo
dblp:198/1004
· DBLP profile ↗
11ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-4201-5967ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 22% Time series and sequential data · 15% Language models and text generation · 10% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
adapter tuning |
0.9 | 1 | 2025 | AdaPTS: Adapting Univariate Foundation Models to Probabilistic Multivariate Time Series Forecasting · ICML 2025 |
Robotics › Motion planning and robot control
dynamics prediction |
0.9 | 1 | 2025 | Zero-shot Model-based Reinforcement Learning using Large Language Models · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation
foundation model adaptation |
0.9 | 1 | 2025 | AdaPTS: Adapting Univariate Foundation Models to Probabilistic Multivariate Time Series Forecasting · ICML 2025 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.9 | 1 | 2025 | Zero-shot Model-based Reinforcement Learning using Large Language Models · ICLR 2025 |
Machine learning › Time series and sequential data › time series modeling
probabilistic forecasting |
0.9 | 1 | 2025 | AdaPTS: Adapting Univariate Foundation Models to Probabilistic Multivariate Time Series Forecasting · ICML 2025 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.9 | 1 | 2025 | AdaPTS: Adapting Univariate Foundation Models to Probabilistic Multivariate Time Series Forecasting · ICML 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
sharpness-aware minimization |
0.8 | 1 | 2024 | SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention · ICML 2024 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
0.8 | 1 | 2024 | SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention · ICML 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention · ICML 2024 |
Machine learning › Reinforcement learning
exploration |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Machine learning › Reinforcement learning › exploration
novelty-based exploration |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Machine learning › Reinforcement learning
policy learning |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Robotics › Autonomous driving › trajectory prediction
interaction-aware prediction |
0.3 | 1 | 2018 | A Data-driven Model for Interaction-Aware Pedestrian Motion Prediction in Object Cluttered Environments · ICRA 2018 |
Robotics › Autonomous driving › trajectory prediction
pedestrian motion prediction |
0.3 | 1 | 2018 | A Data-driven Model for Interaction-Aware Pedestrian Motion Prediction in Object Cluttered Environments · ICRA 2018 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.3 | 1 | 2025 | Zero-shot Model-based Reinforcement Learning using Large Language Models · ICLR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › cognitive modeling
cognitive architecture |
0.2 | 1 | 2024 | Position: A Call for Embodied AI · ICML 2024 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2024 | Position: A Call for Embodied AI · ICML 2024 |
Robotics › Autonomous driving
trajectory prediction |
0.1 | 1 | 2018 | A Data-driven Model for Interaction-Aware Pedestrian Motion Prediction in Object Cluttered Environments · ICRA 2018 |
Methods — techniques the papers use, named apart from their topics
stochastic modeling · 0.9large language model · 0.9disentangled in-context learning · 0.9bayesian neural network · 0.9adapter · 0.9transformer · 0.8sharpness-aware minimization · 0.8channel-wise attention · 0.8active inference · 0.8autoencoder · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Zero-shot Model-based Reinforcement Learning using Large Language ModelsabstractThe emerging zero-shot capabilities of Large Language Models (LLMs) have led to their applications in areas extending well beyond natural language processing tasks.
In reinforcement learning, while LLMs have been extensively used in text-based environments, their integration with continuous state spaces remains understudied.
In this paper, we investigate how pre-trained LLMs can be leveraged to predict in context the dynamics of continuous Markov decision processes.
We identify handling multivariate data and incorporating the control signal as key challenges that limit the potential of LLMs' deployment in this setup and propose Disentangled In-Context Learning (DICL) to address them.
We present proof-of-concept applications in two reinforcement learning settings: model-based policy evaluation and data-augmented off-policy reinforcement learning, supported by theoretical analysis of the proposed methods.
Our experiments further demonstrate that our approach produces well-calibrated uncertainty estimates. We release the code at https://github.com/abenechehab/dicl. Abdelhakim Benechehab, Youssef Attia El Hili, Ambroise Odonnat, Oussama Zekri, Albert Thomas 0001, Giuseppe Paolo, Maurizio Filippone, Ievgen Redko, Balázs Kégl |
ICLR | 6 |
| 2025 | AdaPTS: Adapting Univariate Foundation Models to Probabilistic Multivariate Time Series ForecastingabstractPre-trained foundation models (FMs) have shown exceptional performance in univariate time series forecasting tasks. However, several practical challenges persist, including managing intricate dependencies among features and quantifying uncertainty in predictions. This study aims to tackle these critical limitations by introducing adapters—feature-space transformations that facilitate the effective use of pre-trained univariate time series FMs for multivariate tasks. Adapters operate by projecting multivariate inputs into a suitable latent space and applying the FM independently to each dimension. Inspired by the literature on representation learning and partially stochastic Bayesian neural networks, we present a range of adapters and optimization/inference strategies. Experiments conducted on both synthetic and real-world datasets confirm the efficacy of adapters, demonstrating substantial enhancements in forecasting accuracy and uncertainty quantification compared to baseline methods. Our framework, AdaPTS, positions adapters as a modular, scalable, and effective solution for leveraging time series FMs in multivariate contexts, thereby promoting their wider adoption in real-world applications. We release the code at https://github.com/abenechehab/AdaPTS. Abdelhakim Benechehab, Vasilii Feofanov, Giuseppe Paolo, Albert Thomas 0001, Maurizio Filippone, Balázs Kégl |
ICML | 3 |
| 2024 | SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise AttentionabstractTransformer-based architectures achieved breakthrough performance in natural language processing and computer vision, yet they remain inferior to simpler linear baselines in multivariate long-term forecasting. To better understand this phenomenon, we start by studying a toy linear forecasting problem for which we show that transformers are incapable of converging to their true solution despite their high expressive power. We further identify the attention of transformers as being responsible for this low generalization capacity. Building upon this insight, we propose a shallow lightweight transformer model that successfully escapes bad local minima when optimized with sharpness-aware optimization. We empirically demonstrate that this result extends to all commonly used real-world multivariate time series datasets. In particular, SAMformer surpasses current state-of-the-art methods and is on par with the biggest foundation model MOIRAI while having significantly fewer parameters. The code is available at https://github.com/romilbert/samformer. Romain Ilbert, Ambroise Odonnat, Vasilii Feofanov, Aladin Virmaux, Giuseppe Paolo, Themis Palpanas, Ievgen Redko |
ICML | 5 |
| 2024 | Position: A Call for Embodied AIabstractWe propose Embodied AI (E-AI) as the next fundamental step in the pursuit of Artificial General Intelligence (AGI), juxtaposing it against current AI advancements, particularly Large Language Models (LLMs). We traverse the evolution of the embodiment concept across diverse fields (philosophy, psychology, neuroscience, and robotics) to highlight how E-AI distinguishes itself from the classical paradigm of static learning. By broadening the scope of E-AI, we introduce a theoretical framework based on cognitive architectures, emphasizing perception, action, memory, and learning as essential components of an embodied agent. This framework is aligned with Friston’s active inference principle, offering a comprehensive approach to E-AI development. Despite the progress made in the field of AI, substantial challenges, such as the formulation of a novel AI learning theory and the innovation of advanced hardware, persist. Our discussion lays down a foundational guideline for future E-AI research. Highlighting the importance of creating E-AI agents capable of seamless communication, collaboration, and coexistence with humans and other intelligent entities within real-world environments, we aim to steer the AI community towards addressing the multifaceted challenges and seizing the opportunities that lie ahead in the quest for AGI. Giuseppe Paolo, Jonas Gonzalez-Billandon, Balázs Kégl |
ICML | 1 |
| 2024 | Discovering and Exploiting Sparse Rewards in a Learned Behavior SpaceabstractLearning optimal policies in sparse rewards settings is difficult as the learning agent has little to no feedback on the quality of its actions. In these situations, a good strategy is to focus on exploration, hopefully leading to the discovery of a reward signal to improve on. A learning algorithm capable of dealing with this kind of setting has to be able to (1) explore possible agent behaviors and (2) exploit any possible discovered reward. Exploration algorithms have been proposed that require the definition of a low-dimension behavior space, in which the behavior generated by the agent's policy can be represented. The need to design a priori this space such that it is worth exploring is a major limitation of these algorithms. In this work, we introduce STAX, an algorithm designed to learn a behavior space on-the-fly and to explore it while optimizing any reward discovered (see Figure 1). It does so by separating the exploration and learning of the behavior space from the exploitation of the reward through an alternating two-step process. In the first step, STAX builds a repertoire of diverse policies while learning a low-dimensional representation of the high-dimensional observations generated during the policies evaluation. In the exploitation step, emitters optimize the performance of the discovered rewarding solutions. Experiments conducted on three different sparse reward environments show that STAX performs comparably to existing baselines while requiring much less prior information about the task as it autonomously builds the behavior space it explores. Giuseppe Paolo, Miranda Coninx, Alban Laflaquière, Stéphane Doncieux |
Evol. Comput. | 1 |
| 2023 | Editorial to the "Evolutionary Reinforcement Learning" Special IssueabstractNo abstract available. Adam Gaier, Giuseppe Paolo, Antoine Cully |
ACM Trans. Evol. Learn. Optim. | 2 |
| 2021 | Sparse reward exploration via novelty search and emittersabstractReward-based optimization algorithms require both exploration, to find rewards, and exploitation, to maximize performance. The need for efficient exploration is even more significant in sparse reward settings, in which performance feedback is given sparingly, thus rendering it unsuitable for guiding the search process. In this work, we introduce the SparsE Reward Exploration via Novelty and Emitters (SERENE) algorithm, capable of efficiently exploring a search space, as well as optimizing rewards found in potentially disparate areas. Contrary to existing emitters-based approaches, SERENE separates the search space exploration and reward exploitation into two alternating processes. The first process performs exploration through Novelty Search, a divergent search algorithm. The second one exploits discovered reward areas through emitters, i.e. local instances of population-based optimization algorithms. A meta-scheduler allocates a global computational budget by alternating between the two processes, ensuring the discovery and efficient exploitation of disjoint reward areas. SERENE returns both a collection of diverse solutions covering the search space and a collection of high-performing solutions for each distinct reward area. We evaluate SERENE on various sparse reward environments and show it compares favorably to existing baselines. Giuseppe Paolo, Miranda Coninx, Stéphane Doncieux, Alban Laflaquière |
GECCO | 1 |
| 2020 | Novelty search makes evolvability inevitableabstractEvolvability is an important feature that impacts the ability of evolutionary processes to find interesting novel solutions and to deal with changing conditions of the problem to solve. The estimation of evolvability is not straight-forward and is generally too expensive to be directly used as selective pressure in the evolutionary process. Indirectly promoting evolvability as a side effect of other easier and faster to compute selection pressures would thus be advantageous. In an unbounded behavior space, it has already been shown that evolvable individuals naturally appear and tend to be selected as they are more likely to invade empty behavior niches. Evolvability is thus a natural byproduct of the search in this context. However, practical agents and environments often impose limits on the reachable behavior space. How do these boundaries impact evolvability? In this context, can evolvability still be promoted without explicitly rewarding it? We show that Novelty Search implicitly creates a pressure for high evolvability even in bounded behavior spaces, and explore the reasons for such a behavior. More precisely we show that, throughout the search, the dynamic evaluation of novelty rewards individuals which are very mobile in the behavior space, which in turn promotes evolvability. Stéphane Doncieux, Giuseppe Paolo, Alban Laflaquière, Miranda Coninx |
GECCO | 2 |
| 2020 | Unsupervised Learning and Exploration of Reachable Outcome SpaceabstractPerforming Reinforcement Learning in sparse rewards settings, with very little prior knowledge, is a challenging problem since there is no signal to properly guide the learning process. In such situations, a good search strategy is fundamental. At the same time, not having to adapt the algorithm to every single problem is very desirable. Here we introduce TAXONS, a Task Agnostic eXploration of Outcome spaces through Novelty and Surprise algorithm. Based on a population-based divergent-search approach, it learns a set of diverse policies directly from high-dimensional observations, without any task-specific information. TAXONS builds a repertoire of policies while training an autoencoder on the high-dimensional observation of the final state of the system to build a low-dimensional outcome space. The learned outcome space, combined with the reconstruction error, is used to drive the search for new policies. Results show that TAXONS can find a diverse set of controllers, covering a good part of the ground-truth outcome space, while having no information about such space. Giuseppe Paolo, Alban Laflaquière, Miranda Coninx, Stéphane Doncieux |
ICRA | 1 |
| 2018 | A Data-driven Model for Interaction-Aware Pedestrian Motion Prediction in Object Cluttered EnvironmentsabstractThis paper reports on a data-driven, interaction-aware motion prediction approach for pedestrians in environments cluttered with static obstacles. When navigating in such workspaces shared with humans, robots need accurate motion predictions of the surrounding pedestrians. Human navigation behavior is mostly influenced by their surrounding pedestrians and by the static obstacles in their vicinity. In this paper we introduce a new model based on Long-Short Term Memory (LSTM) neural networks, which is able to learn human motion behavior from demonstrated data. To the best of our knowledge, this is the first approach using LSTMs, that incorporates both static obstacles and surrounding pedestrians for trajectory forecasting. As part of the model, we introduce a new way of encoding surrounding pedestrians based on a 1d-grid in polar angle space. We evaluate the benefit of interaction-aware motion prediction and the added value of incorporating static obstacles on both simulation and real-world datasets by comparing with state-of-the-art approaches. The results show, that our new approach outperforms the other approaches while being very computationally efficient and that taking into account static obstacles for motion predictions significantly improves the prediction accuracy, especially in cluttered environments. Mark Pfeiffer, Giuseppe Paolo, Hannes Sommer, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma |
ICRA | 2 |
| 2017 | Virtual-to-real deep reinforcement learning: Continuous control of mobile robots for mapless navigationabstractWe present a learning-based mapless motion planner by taking the sparse 10-dimensional range findings and the target position with respect to the mobile robot coordinate frame as input and the continuous steering commands as output. Traditional motion planners for mobile ground robots with a laser range sensor mostly depend on the obstacle map of the navigation environment where both the highly precise laser sensor and the obstacle map building work of the environment are indispensable. We show that, through an asynchronous deep reinforcement learning method, a mapless motion planner can be trained end-to-end without any manually designed features and prior demonstrations. The trained planner can be directly applied in unseen virtual and real environments. The experiments show that the proposed mapless motion planner can navigate the nonholonomic mobile robot to the desired targets without colliding with any obstacles. Lei Tai, Giuseppe Paolo, Ming Liu 0001 |
IROS | 2 |