EDBT 2026 Demo / reviewers in the wild / expert
Maximilien Germain
dblp:243/6182
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 1 | 2025 | Actor-Critic learning for mean-field control in continuous time · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning
continuous-time reinforcement learning |
0.9 | 1 | 2025 | Actor-Critic learning for mean-field control in continuous time · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
mean field control |
0.9 | 1 | 2025 | Actor-Critic learning for mean-field control in continuous time · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.9 | 1 | 2025 | Actor-Critic learning for mean-field control in continuous time · J. Mach. Learn. Res. 2025 |
Methods — techniques the papers use, named apart from their topics
wasserstein space · 0.9entropy regularization · 0.9actor-critic · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Actor-Critic learning for mean-field control in continuous timeabstractWe study policy gradient for mean-field control in continuous time in a reinforcement learning setting. By considering randomised policies with entropy regularisation, we derive a gradient expectation representation of the value function, which is amenable to actor-critic type algorithms, where the value functions and the policies are learnt alternately based on observation samples of the state and model-free estimation of the population state distribution, either by offline or online learning. In the linear-quadratic mean-field framework, we obtain an exact parametrisation of the actor and critic functions defined on the Wasserstein space. Finally, we illustrate the results of our algorithms with some numerical experiments on concrete examples. Noufel Frikha, Maximilien Germain, Mathieu Laurière, Huyên Pham, Xuanye Song |
J. Mach. Learn. Res. | 2 |