Max Ruiz Luyten

dblp:332/2342 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 30% Reinforcement learning · 30% Language models and text generation · 16%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 50% Software testing · 50%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples · ICLR 2025
Machine learning › Reinforcement learning
exploration
0.912025
Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.912025
Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025
Natural language and speech › Language models and text generation › LLM agents
LLM-based simulation
0.912025
G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration · ICML 2025
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
0.912025
Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples · ICLR 2025
Natural language and speech › Language models and text generation
code generation
0.812024
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › semantic representation learning
concept-based learning
0.812024
A theoretical design of concept sets: improving the predictability of concept bottleneck models · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
concept bottleneck model
0.812024
A theoretical design of concept sets: improving the predictability of concept bottleneck models · NeurIPS 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
A theoretical design of concept sets: improving the predictability of concept bottleneck models · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems
0.812024
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation · ICLR 2024
Program synthesis and code generation
code generation with language models
0.812024
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation · ICLR 2024
Software testing
model testing
0.812024
Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models · NeurIPS 2024
Machine learning › Reinforcement learning › reward design › reward shaping
language-based reward shaping
0.312025
Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025
Machine learning › Reinforcement learning › reward design
reward shaping
0.312025
Strategic Planning: A Top-Down Approach to Option Generation · ICML 2025

Methods — techniques the papers use, named apart from their topics

large language model · 2.4simulation-based inference · 1.7likelihood-free inference · 1.7gradient-free optimization · 1.7memory-augmented LLM · 1.5instruction registry · 1.5control unit · 1.5stochastic differential equation · 0.9diffusion model · 0.9PPO · 0.9self-falsification · 0.8
YearPublicationVenuePosition
2025 Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples
abstract
Diffusion models are mainly studied on image data. However, non-image data (e.g., tabular data) are also prevalent in real applications and tend to be noisy due to some inevitable factors in the stage of data collection, degrading the generation quality of diffusion models. In this paper, we consider a novel problem setting where every collected sample is paired with a vector indicating the data quality: risk vector. This setting applies to many scenarios involving noisy data and we propose risk-sensitive SDE, a type of stochastic differential equation (SDE) parameterized by the risk vector, to address it. With some proper coefficients, risk-sensitive SDE can minimize the negative effect of noisy samples on the optimization of diffusion models. We conduct systematic studies for both Gaussian and non-Gaussian noise distributions, providing analytical forms of risk-sensitive SDE. To verify the effectiveness of our method, we have conducted extensive experiments on multiple tabular and time-series datasets, showing that risk-sensitive SDE permits a robust optimization of diffusion models with noisy samples and significantly outperforms previous baselines.
Yangming Li, Max Ruiz Luyten, Mihaela van der Schaar
ICLR2
2025 G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration
abstract
Constructing robust simulators is essential for asking "what if?" questions and guiding policy in critical domains like healthcare and logistics. However, existing methods often struggle, either failing to generalize beyond historical data or, when using Large Language Models (LLMs), suffering from inaccuracies and poor empirical alignment. We introduce G-Sim, a hybrid framework that automates simulator construction by synergizing LLM-driven structural design with rigorous empirical calibration. G-Sim employs an LLM in an iterative loop to propose and refine a simulator’s core components and causal relationships, guided by domain knowledge. This structure is then grounded in reality by estimating its parameters using flexible calibration techniques. Specifically, G-Sim can leverage methods that are both likelihood-free and gradient-free with respect to the simulator, such as gradient-free optimization for direct parameter estimation or simulation-based inference for obtaining a posterior distribution over parameters. This allows it to handle non-differentiable and stochastic simulators. By integrating domain priors with empirical evidence, G-Sim produces reliable, causally-informed simulators, mitigating data-inefficiency and enabling robust system-level interventions for complex decision-making.
Samuel Holt, Max Ruiz Luyten, Antonin Berthon, Mihaela van der Schaar
ICML2
2025 Strategic Planning: A Top-Down Approach to Option Generation
abstract
Real-world human decision-making often relies on strategic planning, where high-level goals guide the formulation of sub-goals and subsequent actions, as evidenced by domains such as healthcare, business, and urban policy. Despite notable successes in controlled settings, conventional reinforcement learning (RL) follows a bottom-up paradigm, which can struggle to adapt to real-world complexities such as sparse rewards and limited exploration budgets. While methods like hierarchical RL and environment shaping provide partial solutions, they frequently rely on either ad-hoc designs (e.g. choose the set of high-level actions) or purely data-driven discovery of high-level actions that still requires significant exploration. In this paper, we introduce a top-down framework for RL that explicitly leverages human-like strategy to reduce sample complexity, guide exploration, and enable high-level decision-making. We first formalize the Strategy Problem, which frames policy generation as finding distributions over policies that balance specificity and value. Building on this definition, we propose the Strategist agent—an iterative framework that leverages large language models (LLMs) to synthesize domain knowledge into a structured representation of actionable strategies and sub-goals. We further develop a reward shaping methodology that translates these strategies expressed in natural language into quantitative feedback for RL methods. Empirically, we demonstrate a significantly faster convergence than conventional PPO. Taken together, our findings highlight that top-down strategic exploration opens new avenues for enhancing RL on real-world decision problems.
Max Ruiz Luyten, Antonin Berthon, Mihaela van der Schaar
ICML1
2024 L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
abstract
Transformer-based large language models (LLMs) are constrained by the fixed context window of the underlying transformer architecture, hindering their ability to produce long and coherent outputs. Memory-augmented LLMs are a promising solution, but current approaches cannot handle long output generation tasks since they (1) only focus on reading memory and reduce its evolution to the concatenation of new memories or (2) use very specialized memories that cannot adapt to other domains. This paper presents L2MAC, the first practical LLM-based general-purpose stored-program automatic computer (von Neumann architecture) framework, an LLM-based multi-agent system, for long and consistent output generation. Its memory has two components: the instruction registry, which is populated with a prompt program to solve the user-given task, and a file store, which will contain the final and intermediate outputs. Each instruction in turn is executed by a separate LLM agent, whose context is managed by a control unit capable of precise memory reading and writing to ensure effective interaction with the entire file store. These components enable L2MAC to generate extensive outputs, bypassing the constraints of the finite context window while producing outputs that fulfill a complex user-specified task. We empirically demonstrate that L2MAC achieves state-of-the-art performance in generating large codebases for system design tasks, significantly outperforming other coding methods in implementing the detailed user-specified task; we show that L2MAC works for general-purpose extensive text-based tasks, such as writing an entire book; and we provide valuable insights into L2MAC's performance improvement over existing methods.
Samuel Holt, Max Ruiz Luyten, Mihaela van der Schaar
ICLR2
2024 A theoretical design of concept sets: improving the predictability of concept bottleneck models
abstract
Concept-based learning, a promising approach in machine learning, emphasizes the value of high-level representations called concepts. However, despite growing interest in concept-bottleneck models (CBMs), there is a lack of clear understanding regarding the properties of concept sets and their impact on model performance. In this work, we define concepts within the machine learning context, highlighting their core properties: 'expressiveness' and 'model-aware inductive bias', and we make explicit the underlying assumption of CBMs. We establish theoretical results for concept-bottleneck models (CBMs), revealing how these properties guide the design of concept sets that optimize model performance. Specifically, we demonstrate that well-chosen concept sets can improve sample efficiency and out-of-distribution robustness in the appropriate regimes. Based on these insights, we propose a method to effectively identify informative and non-redundant concepts. We validate our approach with experiments on CIFAR-10 and MetaShift, showing that concept-bottleneck models outperform the foundational embedding counterpart, particularly in low-data regimes and under distribution shifts. We also examine failure modes and discuss how they can be tackled.
Max Ruiz Luyten, Mihaela van der Schaar
NeurIPS1
2024 Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models
abstract
The predominant *de facto* paradigm of testing ML models relies on either using only held-out data to compute aggregate evaluation metrics or by assessing the performance on different subgroups. However, such *data-only testing* methods operate under the restrictive assumption that the available empirical data is the sole input for testing ML models, disregarding valuable contextual information that could guide model testing. In this paper, we challenge the go-to approach of *data-only testing* and introduce *Context-Aware Testing* (CAT) which uses context as an inductive bias to guide the search for meaningful model failures. We instantiate the first CAT system, *SMART Testing*, which employs large language models to hypothesize relevant and likely failures, which are evaluated on data using a *self-falsification mechanism*. Through empirical evaluations in diverse settings, we show that SMART automatically identifies more relevant and impactful failures than alternatives, demonstrating the potential of CAT as a testing paradigm.
Paulius Rauba, Nabeel Seedat, Max Ruiz Luyten, Mihaela van der Schaar
NeurIPS3