VLDB 2026 Research / reviewers in the wild / expert
Antonin Berthon
dblp:256/4996
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 57% Trustworthy machine learning · 27% Language models and text generation · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
exploration |
0.9 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.9 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Natural language and speech › Language models and text generation › LLM agents
LLM-based simulation |
0.9 | 1 | 2025 | G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration · ICML 2025 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
0.9 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
instance-dependent label noise |
0.5 | 1 | 2021 | Confidence Scores Make Instance-dependent Label-noise Learning Possible · ICML 2021 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.5 | 1 | 2021 | Confidence Scores Make Instance-dependent Label-noise Learning Possible · ICML 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Confidence Scores Make Instance-dependent Label-noise Learning Possible · ICML 2021 |
Machine learning › Reinforcement learning › reward design › reward shaping
language-based reward shaping |
0.3 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.3 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
simulation-based inference · 1.7likelihood-free inference · 1.7gradient-free optimization · 1.7large language model · 0.9PPO · 0.9forward correction · 0.5confidence scores · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | G-Sim: Generative Simulations with Large Language Models and Gradient-Free CalibrationabstractConstructing robust simulators is essential for asking "what if?" questions and guiding policy in critical domains like healthcare and logistics. However, existing methods often struggle, either failing to generalize beyond historical data or, when using Large Language Models (LLMs), suffering from inaccuracies and poor empirical alignment. We introduce G-Sim, a hybrid framework that automates simulator construction by synergizing LLM-driven structural design with rigorous empirical calibration. G-Sim employs an LLM in an iterative loop to propose and refine a simulator’s core components and causal relationships, guided by domain knowledge. This structure is then grounded in reality by estimating its parameters using flexible calibration techniques. Specifically, G-Sim can leverage methods that are both likelihood-free and gradient-free with respect to the simulator, such as gradient-free optimization for direct parameter estimation or simulation-based inference for obtaining a posterior distribution over parameters. This allows it to handle non-differentiable and stochastic simulators. By integrating domain priors with empirical evidence, G-Sim produces reliable, causally-informed simulators, mitigating data-inefficiency and enabling robust system-level interventions for complex decision-making. Samuel Holt, Max Ruiz Luyten, Antonin Berthon, Mihaela van der Schaar |
ICML | 3 |
| 2025 | Strategic Planning: A Top-Down Approach to Option GenerationabstractReal-world human decision-making often relies on strategic planning, where high-level goals guide the formulation of sub-goals and subsequent actions, as evidenced by domains such as healthcare, business, and urban policy. Despite notable successes in controlled settings, conventional reinforcement learning (RL) follows a bottom-up paradigm, which can struggle to adapt to real-world complexities such as sparse rewards and limited exploration budgets. While methods like hierarchical RL and environment shaping provide partial solutions, they frequently rely on either ad-hoc designs (e.g. choose the set of high-level actions) or purely data-driven discovery of high-level actions that still requires significant exploration. In this paper, we introduce a top-down framework for RL that explicitly leverages human-like strategy to reduce sample complexity, guide exploration, and enable high-level decision-making. We first formalize the Strategy Problem, which frames policy generation as finding distributions over policies that balance specificity and value. Building on this definition, we propose the Strategist agent—an iterative framework that leverages large language models (LLMs) to synthesize domain knowledge into a structured representation of actionable strategies and sub-goals. We further develop a reward shaping methodology that translates these strategies expressed in natural language into quantitative feedback for RL methods. Empirically, we demonstrate a significantly faster convergence than conventional PPO. Taken together, our findings highlight that top-down strategic exploration opens new avenues for enhancing RL on real-world decision problems. Max Ruiz Luyten, Antonin Berthon, Mihaela van der Schaar |
ICML | 2 |
| 2021 | Confidence Scores Make Instance-dependent Label-noise Learning PossibleabstractIn learning with noisy labels, for every instance, its label can randomly walk to other classes following a transition distribution which is named a noise model. Well-studied noise models are all instance-independent, namely, the transition depends only on the original label but not the instance itself, and thus they are less practical in the wild. Fortunately, methods based on instance-dependent noise have been studied, but most of them have to rely on strong assumptions on the noise models. To alleviate this issue, we introduce confidence-scored instance-dependent noise (CSIDN), where each instance-label pair is equipped with a confidence score. We find that with the help of confidence scores, the transition distribution of each instance can be approximately estimated. Similarly to the powerful forward correction for instance-independent noise, we propose a novel instance-level forward correction for CSIDN. We demonstrate the utility and effectiveness of our method through multiple experiments on datasets with synthetic label noise and real-world unknown noise. Antonin Berthon, Bo Han 0003, Gang Niu 0001, Tongliang Liu, Masashi Sugiyama |
ICML | 1 |