EDBT 2026 Demo / reviewers in the wild / expert
Natalie Maus
dblp:264/7932
· DBLP profile ↗
9ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-6616-8506ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Optimization for machine learning · 84% Probabilistic and Bayesian machine learning · 14% Representation and self-supervised learning · 2% | |
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
3.2 | 5 | 2025 | Covering Multiple Objectives with a Small Set of Solutions Using Bayesian Optimization · NeurIPS 2025 Approximation-Aware Bayesian Optimization · NeurIPS 2024 Joint Composite Latent Space Bayesian Optimization · ICML 2024 |
Machine learning › Optimization for machine learning
black-box optimization |
3.0 | 4 | 2025 | Covering Multiple Objectives with a Small Set of Solutions Using Bayesian Optimization · NeurIPS 2025 Approximation-Aware Bayesian Optimization · NeurIPS 2024 Joint Composite Latent Space Bayesian Optimization · ICML 2024 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
high-dimensional bayesian optimization |
2.1 | 3 | 2024 | Approximation-Aware Bayesian Optimization · NeurIPS 2024 Joint Composite Latent Space Bayesian Optimization · ICML 2024 Local Latent Space Bayesian Optimization over Structured Inputs · NeurIPS 2022 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
acquisition function |
1.6 | 2 | 2025 | Covering Multiple Objectives with a Small Set of Solutions Using Bayesian Optimization · NeurIPS 2025 Approximation-Aware Bayesian Optimization · NeurIPS 2024 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
multi-objective bayesian optimization |
0.9 | 1 | 2025 | Covering Multiple Objectives with a Small Set of Solutions Using Bayesian Optimization · NeurIPS 2025 |
Query processing and optimization › query optimization
learned query optimization |
0.9 | 1 | 2025 | Learned Offline Query Planning via Bayesian Optimization · Proc. ACM Manag. Data 2025 |
Query processing and optimization
query planning |
0.9 | 1 | 2025 | Learned Offline Query Planning via Bayesian Optimization · Proc. ACM Manag. Data 2025 |
Machine learning › Optimization for machine learning › convex optimization
composite optimization |
0.8 | 1 | 2024 | Joint Composite Latent Space Bayesian Optimization · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.7 | 1 | 2023 | Variational Gaussian Processes with Decoupled Conditionals · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
inducing point methods |
0.7 | 1 | 2023 | Variational Gaussian Processes with Decoupled Conditionals · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.7 | 1 | 2023 | Variational Gaussian Processes with Decoupled Conditionals · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
latent space learning |
0.2 | 1 | 2024 | Joint Composite Latent Space Bayesian Optimization · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
variational autoencoder · 1.7bayesian optimization · 1.7expected improvement · 1.6trust region · 1.3large language model · 0.9LLaVA · 0.9CLIP · 0.9utility-calibrated variational inference · 0.8sparse variational gaussian process · 0.8probabilistic model · 0.8neural network encoder · 0.8latent variable model · 0.8knowledge gradient · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Dataset for Distilling Knowledge Priors from Literature for Therapeutic DesignabstractAI-driven discovery can greatly reduce design time and enhance new therapeutics' effectiveness. Models using simulators explore broad design spaces but risk violating implicit constraints due to a lack of experimental priors. For example, in a new analysis across diverse models on the GuacaMol benchmark using supervised classifiers, over 60\% of molecules proposed had a high probability of being mutagenic. In this work, we introduce Medex, a dataset of priors for design problems extracted from literature describing compounds used in lab settings. It is constructed with LLM pipelines for discovering therapeutic entities in relevant paragraphs and summarizing information in concise fair-use facts. Medex consists of 32.3 million pairs of natural language facts, and appropriate entity representations (i.e. SMILES or RefSeq IDs). To demonstrate the potential of the data, we train LLM, CLIP, and LLaVA architectures to reason jointly about text and design targets and evaluate on tasks from the Therapeutic Data Commons (TDC). Medex is highly effective for creating models with strong priors: in supervised prediction problems that use our data for pretraining, our best models with 15M learnable parameters outperform larger 2B TxGemma on both regression and classification TDC tasks, and perform comparably to 9B models on average. Models built with Medex can be used as constraints while optimizing for novel molecules in GuacaMol, resulting in proposals that are safer and nearly as effective. We release our dataset on HuggingFace at https://huggingface.co/datasets/DocAndDesign/Medex, and will provide expanded versions as the available literature grows. Haydn Thomas Jones, Natalie Maus, Josh Magnus Ludan, Maggie Ziyu Huan, Marcelo Der Torossian Torres, Jiatao Liang, Zachary G. Ives, Yoseph Barash, Cesar de la Fuente-Nunez, Jacob R. Gardner, Mark Yatskar |
NeurIPS | 2 |
| 2025 | Covering Multiple Objectives with a Small Set of Solutions Using Bayesian OptimizationabstractIn multi-objective black-box optimization, the goal is typically to find solutions that optimize a set of $T$ black-box objective functions, $f_1, \ldots f_T$, simultaneously. Traditional approaches often seek a single Pareto-optimal set that balances trade-offs among all objectives. In contrast, we consider a problem setting that departs from this paradigm: finding a small set of $K < T$ solutions, that collectively "cover" the $T$ objectives. A set of solutions is defined as "covering" if, for each objective $f_1, \ldots f_T$, there is at least one good solution. A motivating example for this problem setting occurs in drug design. For example, we may have $T$ pathogens and aim to identify a set of $K < T$ antibiotics such that at least one antibiotic can be used to treat each pathogen. This problem, known as coverage optimization, has yet to be tackled with the Bayesian optimization (BO) framework. To fill this void, we develop Multi-Objective Coverage Bayesian Optimization (MOCOBO), a BO algorithm for solving coverage optimization. Our approach is based on a new acquisition function reminiscent of expected improvement in the vanilla BO setup. We demonstrate the performance of our method on high-dimensional black-box optimization tasks, including applications in peptide and molecular design. Results show that the coverage of the $K < T$ solutions found by MOCOBO matches or nearly matches the coverage of $T$ solutions obtained by optimizing each objective individually. Furthermore, in *in vitro* experiments, the peptides found by MOCOBO exhibited high potency against drug-resistant pathogens, further demonstrating the potential of MOCOBO for drug discovery. All of our code is publicly available at the following link: https://github.com/nataliemaus/mocobo. Natalie Maus, Kyurae Kim, Yimeng Zeng, Haydn Thomas Jones, Fangping Wan, Marcelo Der Torossian Torres, Cesar de la Fuente-Nunez, Jacob R. Gardner |
NeurIPS | 1 |
| 2025 | Learned Offline Query Planning via Bayesian OptimizationabstractAnalytics database workloads often contain queries that are executed repeatedly. Existing optimization techniques generally prioritize keeping optimization cost low, normally well below the time it takes to execute a single instance of a query. If a given query is going to be executed thousands of times, could it be worth investing significantly more optimization time? In contrast to traditional online query optimizers, we propose an offline query optimizer that searches a wide variety of plans and incorporates query execution as a primitive. Our offline query optimizer combines variational auto-encoders with Bayesian optimization to find optimized plans for a given query. We compare our technique to the optimal plans possible with PostgreSQL and recent RL-based systems over several datasets, and show that our technique finds faster query plans. Jeffrey Tao, Natalie Maus, Haydn Thomas Jones, Yimeng Zeng, Jacob R. Gardner, Ryan Marcus |
Proc. ACM Manag. Data | 2 |
| 2024 | Joint Composite Latent Space Bayesian OptimizationabstractBayesian Optimization (BO) is a technique for sample-efficient black-box optimization that employs probabilistic models to identify promising input for evaluation. When dealing with composite-structured functions, such as $f=g \circ h$, evaluating a specific location $x$ yields observations of both the final outcome $f(x) = g(h(x))$ as well as the intermediate output(s) $h(x)$. Previous research has shown that integrating information from these intermediate outputs can enhance BO performance substantially. However, existing methods struggle if the outputs $h(x)$ are high-dimensional. Many relevant problems fall into this setting, including in the context of generative AI, molecular design, or robotics. To effectively tackle these challenges, we introduce Joint Composite Latent Space Bayesian Optimization (JoCo), a novel framework that jointly trains neural network encoders and probabilistic models to adaptively compress high-dimensional input and output spaces into manageable latent representations. This enables effective BO on these compressed representations, allowing JoCo to outperform other state-of-the-art methods in high-dimensional BO on a wide variety of simulated and real-world problems. Natalie Maus, Zhiyuan Jerry Lin, Maximilian Balandat, Eytan Bakshy |
ICML | 1 |
| 2024 | Approximation-Aware Bayesian OptimizationabstractHigh-dimensional Bayesian optimization (BO) tasks such as molecular design often require $>10,$$000$ function evaluations before obtaining meaningful results. While methods like sparse variational Gaussian processes (SVGPs) reduce computational requirements in these settings, the underlying approximations result in suboptimal data acquisitions that slow the progress of optimization. In this paper we modify SVGPs to better align with the goals of BO: targeting informed data acquisition over global posterior fidelity. Using the framework of utility-calibrated variational inference (Lacoste–Julien et al., 2011), we unify GP approximation and data acquisition into a joint optimization problem, thereby ensuring optimal decisions under a limited computational budget. Our approach can be used with any decision-theoretic acquisition function and is readily compatible with trust region methods like TuRBO (Eriksson et al., 2019). We derive efficient joint objectives for the expected improvement (EI) and knowledge gradient (KG) acquisition functions in both the standard and batch BO settings. On a variety of recent high dimensional benchmark tasks in control and molecular design, our approach significantly outperforms standard SVGPs and is capable of achieving comparable rewards with up to $10\times$ fewer function evaluations. Natalie Maus, Kyurae Kim, David Eriksson, Geoff Pleiss, John P. Cunningham, Jacob R. Gardner |
NeurIPS | 1 |
| 2023 | Discovering Many Diverse Solutions with Bayesian OptimizationabstractBayesian optimization (BO) is a popular approach for sample-efficient optimization of black-box objective functions. While BO has been successfully applied to a wide range of scientific applications, traditional approaches to single-objective BO only seek to find a single best solution. This can be a significant limitation in situations where solutions may later turn out to be intractable, for example, a designed molecule may turn out to later violate constraints that can only be evaluated after the optimization process has concluded. To address this issue, we propose rank-ordered Bayesian Optimization with trustregions (ROBOT) which aims to find a portfolio of high-performing solutions that are diverse according to a user-specified diversity measure. We evaluate ROBOT on several real-world applications and show that it can discover large sets of high-performing diverse solutions while requiring few additional function evaluations compared to finding a single best solution. Natalie Maus, Kaiwen Wu, David Eriksson, Jacob R. Gardner |
AISTATS | 1 |
| 2023 | Variational Gaussian Processes with Decoupled ConditionalsabstractVariational Gaussian processes (GPs) approximate exact GP inference by using a small set of inducing points to form a sparse approximation of the true posterior, with the fidelity of the model increasing with additional inducing points. Although the approximation error in principle can be reduced through the use of more inducing points, this leads to scaling optimization challenges and computational complexity. To achieve scalability, inducing point methods typically introduce conditional independencies and then approximations to the training and test conditional distributions. In this paper, we consider an alternative approach to modifying the training and test conditionals, in which we make them more flexible. In particular, we investigate decoupling the parametric form of the predictive mean and covariance in the conditionals, and learn independent parameters for predictive mean and covariance. We derive new evidence lower bounds (ELBO) under these more flexible conditionals, and provide two concrete examples of applying the decoupled conditionals. Empirically, we find this additional flexibility leads to improved model performance on a variety of regression tasks and Bayesian optimization (BO) applications. Xinran Zhu, Kaiwen Wu, Natalie Maus, Jacob R. Gardner, David Bindel |
NeurIPS | 3 |
| 2022 | Local Latent Space Bayesian Optimization over Structured InputsabstractBayesian optimization over the latent spaces of deep autoencoder models (DAEs) has recently emerged as a promising new approach for optimizing challenging black-box functions over structured, discrete, hard-to-enumerate search spaces (e.g., molecules). Here the DAE dramatically simplifies the search space by mapping inputs into a continuous latent space where familiar Bayesian optimization tools can be more readily applied. Despite this simplification, the latent space typically remains high-dimensional. Thus, even with a well-suited latent space, these approaches do not necessarily provide a complete solution, but may rather shift the structured optimization problem to a high-dimensional one. In this paper, we propose LOL-BO, which adapts the notion of trust regions explored in recent work on high-dimensional Bayesian optimization to the structured setting. By reformulating the encoder to function as both an encoder for the DAE globally and as a deep kernel for the surrogate model within a trust region, we better align the notion of local optimization in the latent space with local optimization in the input space. LOL-BO achieves as much as 20 times improvement over state-of-the-art latent space Bayesian optimization methods across six real-world benchmarks, demonstrating that improvement in optimization strategies is as important as developing better DAE models. Natalie Maus, Haydn Thomas Jones, Juston Moore, Matt J. Kusner, John Bradshaw, Jacob R. Gardner |
NeurIPS | 1 |
| 2022 | Estimating heading from optic flow: Comparing deep learning network and human performanceabstractConvolutional neural networks (CNNs) have made significant advances over the past decade with visual recognition, matching or exceeding human performance on certain tasks. Visual recognition is subserved by the ventral stream of the visual system, which, remarkably, CNNs also effectively model. Inspired by this connection, we investigated the extent to which CNNs account for human heading perception, an important function of the complementary dorsal stream. Heading refers to the direction of movement during self-motion, which humans judge with high degrees of accuracy from the streaming pattern of motion on the eye known as optic flow. We examined the accuracy with which CNNs estimate heading from optic flow in a range of situations in which human heading perception has been well studied. These scenarios include heading estimation from sparse optic flow, in the presence of moving objects, and in the presence of rotation. We assessed performance under controlled conditions wherein self-motion was simulated through minimal or realistic scenes. We found that the CNN did not capture the accuracy of heading perception. The addition of recurrent processing to the network, however, closed the gap in performance with humans substantially in many situations. Our work highlights important self-motion scenarios in which recurrent processing supports heading estimation that approaches human-like accuracy. Natalie Maus, Oliver W. Layton |
Neural Networks | 1 |