VLDB 2026 Research / reviewers in the wild / expert
Max Ruiz Luyten
dblp:332/2342
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 30% Reinforcement learning · 30% Language models and text generation · 16% | |
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 50% Software testing · 50% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples · ICLR 2025 |
Machine learning › Reinforcement learning
exploration |
0.9 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.9 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Natural language and speech › Language models and text generation › LLM agents
LLM-based simulation |
0.9 | 1 | 2025 | G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration · ICML 2025 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
0.9 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples · ICLR 2025 |
Natural language and speech › Language models and text generation
code generation |
0.8 | 1 | 2024 | L2MAC: Large Language Model Automatic Computer for Extensive Code Generation · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › semantic representation learning
concept-based learning |
0.8 | 1 | 2024 | A theoretical design of concept sets: improving the predictability of concept bottleneck models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability
concept bottleneck model |
0.8 | 1 | 2024 | A theoretical design of concept sets: improving the predictability of concept bottleneck models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | A theoretical design of concept sets: improving the predictability of concept bottleneck models · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems |
0.8 | 1 | 2024 | L2MAC: Large Language Model Automatic Computer for Extensive Code Generation · ICLR 2024 |
Program synthesis and code generation
code generation with language models |
0.8 | 1 | 2024 | L2MAC: Large Language Model Automatic Computer for Extensive Code Generation · ICLR 2024 |
Software testing
model testing |
0.8 | 1 | 2024 | Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models · NeurIPS 2024 |
Machine learning › Reinforcement learning › reward design › reward shaping
language-based reward shaping |
0.3 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.3 | 1 | 2025 | Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.4simulation-based inference · 1.7likelihood-free inference · 1.7gradient-free optimization · 1.7memory-augmented LLM · 1.5instruction registry · 1.5control unit · 1.5stochastic differential equation · 0.9diffusion model · 0.9PPO · 0.9self-falsification · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy SamplesabstractDiffusion models are mainly studied on image data. However, non-image data (e.g., tabular data) are also prevalent in real applications and tend to be noisy due to some inevitable factors in the stage of data collection, degrading the generation quality of diffusion models. In this paper, we consider a novel problem setting where every collected sample is paired with a vector indicating the data quality: risk vector. This setting applies to many scenarios involving noisy data and we propose risk-sensitive SDE, a type of stochastic differential equation (SDE) parameterized by the risk vector, to address it. With some proper coefficients, risk-sensitive SDE can minimize the negative effect of noisy samples on the optimization of diffusion models. We conduct systematic studies for both Gaussian and non-Gaussian noise distributions, providing analytical forms of risk-sensitive SDE. To verify the effectiveness of our method, we have conducted extensive experiments on multiple tabular and time-series datasets, showing that risk-sensitive SDE permits a robust optimization of diffusion models with noisy samples and significantly outperforms previous baselines. Yangming Li, Max Ruiz Luyten, Mihaela van der Schaar |
ICLR | 2 |
| 2025 | G-Sim: Generative Simulations with Large Language Models and Gradient-Free CalibrationabstractConstructing robust simulators is essential for asking "what if?" questions and guiding policy in critical domains like healthcare and logistics. However, existing methods often struggle, either failing to generalize beyond historical data or, when using Large Language Models (LLMs), suffering from inaccuracies and poor empirical alignment. We introduce G-Sim, a hybrid framework that automates simulator construction by synergizing LLM-driven structural design with rigorous empirical calibration. G-Sim employs an LLM in an iterative loop to propose and refine a simulator’s core components and causal relationships, guided by domain knowledge. This structure is then grounded in reality by estimating its parameters using flexible calibration techniques. Specifically, G-Sim can leverage methods that are both likelihood-free and gradient-free with respect to the simulator, such as gradient-free optimization for direct parameter estimation or simulation-based inference for obtaining a posterior distribution over parameters. This allows it to handle non-differentiable and stochastic simulators. By integrating domain priors with empirical evidence, G-Sim produces reliable, causally-informed simulators, mitigating data-inefficiency and enabling robust system-level interventions for complex decision-making. Samuel Holt, Max Ruiz Luyten, Antonin Berthon, Mihaela van der Schaar |
ICML | 2 |
| 2025 | Strategic Planning: A Top-Down Approach to Option GenerationabstractReal-world human decision-making often relies on strategic planning, where high-level goals guide the formulation of sub-goals and subsequent actions, as evidenced by domains such as healthcare, business, and urban policy. Despite notable successes in controlled settings, conventional reinforcement learning (RL) follows a bottom-up paradigm, which can struggle to adapt to real-world complexities such as sparse rewards and limited exploration budgets. While methods like hierarchical RL and environment shaping provide partial solutions, they frequently rely on either ad-hoc designs (e.g. choose the set of high-level actions) or purely data-driven discovery of high-level actions that still requires significant exploration. In this paper, we introduce a top-down framework for RL that explicitly leverages human-like strategy to reduce sample complexity, guide exploration, and enable high-level decision-making. We first formalize the Strategy Problem, which frames policy generation as finding distributions over policies that balance specificity and value. Building on this definition, we propose the Strategist agent—an iterative framework that leverages large language models (LLMs) to synthesize domain knowledge into a structured representation of actionable strategies and sub-goals. We further develop a reward shaping methodology that translates these strategies expressed in natural language into quantitative feedback for RL methods. Empirically, we demonstrate a significantly faster convergence than conventional PPO. Taken together, our findings highlight that top-down strategic exploration opens new avenues for enhancing RL on real-world decision problems. Max Ruiz Luyten, Antonin Berthon, Mihaela van der Schaar |
ICML | 1 |
| 2024 | L2MAC: Large Language Model Automatic Computer for Extensive Code GenerationabstractTransformer-based large language models (LLMs) are constrained by the fixed context window of the underlying transformer architecture, hindering their ability to produce long and coherent outputs. Memory-augmented LLMs are a promising solution, but current approaches cannot handle long output generation tasks since they (1) only focus on reading memory and reduce its evolution to the concatenation of new memories or (2) use very specialized memories that cannot adapt to other domains. This paper presents L2MAC, the first practical LLM-based general-purpose stored-program automatic computer (von Neumann architecture) framework, an LLM-based multi-agent system, for long and consistent output generation. Its memory has two components: the instruction registry, which is populated with a prompt program to solve the user-given task, and a file store, which will contain the final and intermediate outputs. Each instruction in turn is executed by a separate LLM agent, whose context is managed by a control unit capable of precise memory reading and writing to ensure effective interaction with the entire file store. These components enable L2MAC to generate extensive outputs, bypassing the constraints of the finite context window while producing outputs that fulfill a complex user-specified task. We empirically demonstrate that L2MAC achieves state-of-the-art performance in generating large codebases for system design tasks, significantly outperforming other coding methods in implementing the detailed user-specified task; we show that L2MAC works for general-purpose extensive text-based tasks, such as writing an entire book; and we provide valuable insights into L2MAC's performance improvement over existing methods. Samuel Holt, Max Ruiz Luyten, Mihaela van der Schaar |
ICLR | 2 |
| 2024 | A theoretical design of concept sets: improving the predictability of concept bottleneck modelsabstractConcept-based learning, a promising approach in machine learning, emphasizes the value of high-level representations called concepts. However, despite growing interest in concept-bottleneck models (CBMs), there is a lack of clear understanding regarding the properties of concept sets and their impact on model performance. In this work, we define concepts within the machine learning context, highlighting their core properties: 'expressiveness' and 'model-aware inductive bias', and we make explicit the underlying assumption of CBMs. We establish theoretical results for concept-bottleneck models (CBMs), revealing how these properties guide the design of concept sets that optimize model performance. Specifically, we demonstrate that well-chosen concept sets can improve sample efficiency and out-of-distribution robustness in the appropriate regimes. Based on these insights, we propose a method to effectively identify informative and non-redundant concepts. We validate our approach with experiments on CIFAR-10 and MetaShift, showing that concept-bottleneck models outperform the foundational embedding counterpart, particularly in low-data regimes and under distribution shifts. We also examine failure modes and discuss how they can be tackled. Max Ruiz Luyten, Mihaela van der Schaar |
NeurIPS | 1 |
| 2024 | Context-Aware Testing: A New Paradigm for Model Testing with Large Language ModelsabstractThe predominant *de facto* paradigm of testing ML models relies on either using only held-out data to compute aggregate evaluation metrics or by assessing the performance on different subgroups. However, such *data-only testing* methods operate under the restrictive assumption that the available empirical data is the sole input for testing ML models, disregarding valuable contextual information that could guide model testing. In this paper, we challenge the go-to approach of *data-only testing* and introduce *Context-Aware Testing* (CAT) which uses context as an inductive bias to guide the search for meaningful model failures. We instantiate the first CAT system, *SMART Testing*, which employs large language models to hypothesize relevant and likely failures, which are evaluated on data using a *self-falsification mechanism*. Through empirical evaluations in diverse settings, we show that SMART automatically identifies more relevant and impactful failures than alternatives, demonstrating the potential of CAT as a testing paradigm. Paulius Rauba, Nabeel Seedat, Max Ruiz Luyten, Mihaela van der Schaar |
NeurIPS | 3 |