VLDB 2026 Research / reviewers in the wild / expert
Aaron Klein
dblp:178/3281
· DBLP profile ↗
13ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Optimization for machine learning · 53% Trustworthy machine learning · 14% Probabilistic and Bayesian machine learning · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 60% Hardware accelerators and domain-specific architectures · 40% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
2.2 | 5 | 2025 | Hyperband-based Bayesian Optimization for Black-box Prompt Selection · ICML 2025 BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 BOHB: Robust and Efficient Hyperparameter Optimization at Scale · ICML 2018 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
1.6 | 4 | 2023 | Optimizing Hyperparameters with Conformal Quantile Regression · ICML 2023 Meta-Surrogate Benchmarking for Hyperparameter Optimization · NeurIPS 2019 BOHB: Robust and Efficient Hyperparameter Optimization at Scale · ICML 2018 |
Performance modeling and evaluation
benchmarking |
1.1 | 2 | 2024 | HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models · NeurIPS 2024 NAS-Bench-101: Towards Reproducible Neural Architecture Search · ICML 2019 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.0 | 2 | 2023 | Optimizing Hyperparameters with Conformal Quantile Regression · ICML 2023 Uncertainty Estimates and Multi-hypotheses Networks for Optical Flow · ECCV (7) 2018 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
multi-fidelity bayesian optimization |
0.9 | 1 | 2025 | Hyperband-based Bayesian Optimization for Black-box Prompt Selection · ICML 2025 |
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt selection |
0.9 | 1 | 2025 | Hyperband-based Bayesian Optimization for Black-box Prompt Selection · ICML 2025 |
Hardware accelerators and domain-specific architectures › neural architecture search
hardware-aware neural architecture search |
0.8 | 1 | 2024 | HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
0.7 | 1 | 2023 | Optimizing Hyperparameters with Conformal Quantile Regression · ICML 2023 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
acquisition function |
0.5 | 1 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 |
Machine learning › Optimization for machine learning
black-box optimization |
0.5 | 1 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
density ratio estimation |
0.5 | 1 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
expected improvement |
0.5 | 1 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.4 | 1 | 2019 | NAS-Bench-101: Towards Reproducible Neural Architecture Search · ICML 2019 |
Machine learning › Reinforcement learning › bandit
bandit optimization |
0.3 | 1 | 2018 | BOHB: Robust and Efficient Hyperparameter Optimization at Scale · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
multi-hypothesis estimation |
0.3 | 1 | 2018 | Uncertainty Estimates and Multi-hypotheses Networks for Optical Flow · ECCV (7) 2018 |
Computer vision › 3D vision › motion estimation
optical flow |
0.3 | 1 | 2018 | Uncertainty Estimates and Multi-hypotheses Networks for Optical Flow · ECCV (7) 2018 |
Machine learning › Learning theory › learning dynamics
learning curve prediction |
0.3 | 1 | 2017 | Learning Curve Prediction with Bayesian Neural Networks · ICLR (Poster) 2017 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo › hamiltonian monte carlo
stochastic gradient hamiltonian monte carlo |
0.2 | 1 | 2016 | Bayesian Optimization with Robust Bayesian Neural Networks · NIPS 2016 |
Machine learning › Efficient and distributed learning
automated machine learning |
0.2 | 1 | 2015 | Efficient and Robust Automated Machine Learning · NIPS 2015 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.2 | 1 | 2015 | Efficient and Robust Automated Machine Learning · NIPS 2015 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks |
0.1 | 1 | 2017 | Learning Curve Prediction with Bayesian Neural Networks · ICLR (Poster) 2017 |
Methods — techniques the papers use, named apart from their topics
gaussian process · 1.5hyperband · 1.2deep kernel · 0.9weight sharing · 0.8surrogate prediction · 0.8multi-objective optimization · 0.8graph isomorphism · 0.8architecture optimization · 0.8conformalized quantile regression · 0.7bayesian optimization · 0.5bayesian neural network · 0.5density ratio estimation · 0.5binary classification · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hyperband-based Bayesian Optimization for Black-box Prompt SelectionabstractOptimal prompt selection is crucial for maximizing large language model (LLM) performance on downstream tasks, especially in black-box settings where models are only accessible via APIs. Black-box prompt selection is challenging due to potentially large, combinatorial search spaces, absence of gradient information, and high evaluation cost of prompts on a validation set. We propose HbBoPs, a novel method that combines a structural-aware deep kernel Gaussian Process with Hyperband as a multi-fidelity scheduler to efficiently select prompts. HbBoPs uses embeddings of instructions and few-shot exemplars, treating them as modular components within prompts. This enhances the surrogate model’s ability to predict which prompt to evaluate next in a sample-efficient manner. Hyperband improves query-efficiency by adaptively allocating resources across different fidelity levels, reducing the number of validation instances required for evaluating prompts. Extensive experiments across ten diverse benchmarks and three LLMs demonstrate that HbBoPs outperforms state-of-the-art methods in both performance and efficiency. Lennart Schneider, Martin Wistuba, Aaron Klein, Jacek Golebiowski, Giovanni Zappella, Felice Antonio Merra |
ICML | 3 |
| 2024 | HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language ModelsabstractThe increasing size of language models necessitates a thorough analysis across multiple dimensions to assess trade-offs among crucial hardware metrics such as latency, energy consumption, GPU memory usage, and performance. Identifying optimal model configurations under specific hardware constraints is becoming essential but remains challenging due to the computational load of exhaustive training and evaluation on multiple devices. To address this, we introduce HW-GPT-Bench, a hardware-aware benchmark that utilizes surrogate predictions to approximate various hardware metrics across 13 devices of architectures in the GPT-2 family, with architectures containing up to 1.55B parameters. Our surrogates, via calibrated predictions and reliable uncertainty estimates, faithfully model the heteroscedastic noise inherent in the energy and latency measurements. To estimate perplexity, we employ weight-sharing techniques from Neural Architecture Search (NAS), inheriting pretrained weights from the largest GPT-2 model. Finally, we demonstrate the utility of HW-GPT-Bench by simulating optimization trajectories of various multi-objective optimization algorithms in just a few seconds. Rhea Sanjay Sukthanker, Arber Zela, Benedikt Staffler, Aaron Klein, Lennart Purucker, Jörg K. H. Franke, Frank Hutter |
NeurIPS | 4 |
| 2023 | Optimizing Hyperparameters with Conformal Quantile RegressionabstractMany state-of-the-art hyperparameter optimization (HPO) algorithms rely on model-based optimizers that learn surrogate models of the target function to guide the search. Gaussian processes are the de facto surrogate model due to their ability to capture uncertainty. However, they make strong assumptions about the observation noise, which might not be warranted in practice. In this work, we propose to leverage conformalized quantile regression which makes minimal assumptions about the observation noise and, as a result, models the target function in a more realistic and robust fashion which translates to quicker HPO convergence on empirical benchmarks. To apply our method in a multi-fidelity setting, we propose a simple, yet effective, technique that aggregates observed results across different resource levels and outperforms conventional methods across many empirical tasks. David Salinas, Jacek Golebiowski, Aaron Klein, Matthias W. Seeger, Cédric Archambeau |
ICML | 3 |
| 2021 | Hyperparameter Transfer Learning with Adaptive ComplexityabstractBayesian optimization (BO) is a data-efficient approach to automatically tune the hyperparameters of machine learning models. In practice, one frequently has to solve similar hyperparameter tuning problems sequentially. For example, one might have to tune a type of neural network learned across a series of different classification problems. Recent work on multi-task BO exploits knowledge gained from previous hyperparameter tuning tasks to speed up a new tuning task. However, previous approaches do not account for the fact that BO is a sequential decision making procedure. Hence, there is in general a mismatch between the number of evaluations collected in the current tuning task compared to the number of evaluations accumulated in all previously completed tasks. In this work, we enable multi-task BO to compensate for this mismatch, such that the transfer learning procedure is able to handle different data regimes in a principled way. We propose a new multi-task BO method that learns a set of ordered, non-linear basis functions of increasing complexity via nested drop-out and automatic relevance determination. Experiments on a variety of hyperparameter tuning problems show that our method improves the sample efficiency of recently published multi-task BO methods. Samuel Horváth, Aaron Klein, Peter Richtárik, Cédric Archambeau |
AISTATS | 2 |
| 2021 | BORE: Bayesian Optimization by Density-Ratio EstimationabstractBayesian optimization (BO) is among the most effective and widely-used blackbox optimization methods. BO proposes solutions according to an explore-exploit trade-off criterion encoded in an acquisition function, many of which are computed from the posterior predictive of a probabilistic surrogate model. Prevalent among these is the expected improvement (EI). The need to ensure analytical tractability of the predictive often poses limitations that can hinder the efficiency and applicability of BO. In this paper, we cast the computation of EI as a binary classification problem, building on the link between class-probability estimation and density-ratio estimation, and the lesser-known link between density-ratios and EI. By circumventing the tractability constraints, this reformulation provides numerous advantages, not least in terms of expressiveness, versatility, and scalability. Louis C. Tiao, Aaron Klein, Matthias W. Seeger, Edwin V. Bonilla, Cédric Archambeau, Fabio Ramos 0001 |
ICML | 2 |
| 2019 | NAS-Bench-101: Towards Reproducible Neural Architecture SearchabstractRecent advances in neural architecture search (NAS) demand tremendous computational resources, which makes it difficult to reproduce experiments and imposes a barrier-to-entry to researchers without access to large-scale computation. We aim to ameliorate these problems by introducing NAS-Bench-101, the first public architecture dataset for NAS research. To build NAS-Bench-101, we carefully constructed a compact, yet expressive, search space, exploiting graph isomorphisms to identify 423k unique convolutional architectures. We trained and evaluated all of these architectures multiple times on CIFAR-10 and compiled the results into a large dataset of over 5 million trained models. This allows researchers to evaluate the quality of a diverse range of models in milliseconds by querying the pre-computed dataset. We demonstrate its utility by analyzing the dataset as a whole and by benchmarking a range of architecture optimization algorithms. Chris Ying, Aaron Klein, Eric Christiansen, Esteban Real, Kevin Murphy 0002, Frank Hutter |
ICML | 2 |
| 2019 | Meta-Surrogate Benchmarking for Hyperparameter OptimizationabstractDespite the recent progress in hyperparameter optimization (HPO), available benchmarks that resemble real-world scenarios consist of a few and very large problem instances that are expensive to solve. This blocks researchers and practitioners no only from systematically running large-scale comparisons that are needed to draw statistically significant results but also from reproducing experiments that were conducted before. This work proposes a method to alleviate these issues by means of a meta-surrogate model for HPO tasks trained on off-line generated data. The model combines a probabilistic encoder with a multi-task model such that it can generate inexpensive and realistic tasks of the class of problems of interest. We demonstrate that benchmarking HPO methods on samples of the generative model allows us to draw more coherent and statistically significant conclusions that can be reached orders of magnitude faster than using the original tasks. We provide evidence of our findings for various HPO methods on a wide class of problems. Aaron Klein, Zhenwen Dai, Frank Hutter, Neil D. Lawrence, Javier González 0002 |
NeurIPS | 1 |
| 2018 | Uncertainty Estimates and Multi-hypotheses Networks for Optical Flow
Eddy Ilg, Özgün Çiçek, Silvio Galesso, Aaron Klein, Osama Makansi, Frank Hutter, Thomas Brox |
ECCV (7) | 4 |
| 2018 | BOHB: Robust and Efficient Hyperparameter Optimization at ScaleabstractModern deep learning methods are very sensitive to many hyperparameters, and, due to the long training times of state-of-the-art models, vanilla Bayesian hyperparameter optimization is typically computationally infeasible. On the other hand, bandit-based configuration evaluation approaches based on random search lack guidance and do not converge to the best configurations as quickly. Here, we propose to combine the benefits of both Bayesian optimization and bandit-based methods, in order to achieve the best of both worlds: strong anytime performance and fast convergence to optimal configurations. We propose a new practical state-of-the-art hyperparameter optimization method, which consistently outperforms both Bayesian optimization and Hyperband on a wide range of problem types, including high-dimensional toy functions, support vector machines, feed-forward neural networks, Bayesian neural networks, deep reinforcement learning, and convolutional neural networks. Our method is robust and versatile, while at the same time being conceptually simple and easy to implement. Stefan Falkner, Aaron Klein, Frank Hutter |
ICML | 2 |
| 2017 | Fast Bayesian Optimization of Machine Learning Hyperparameters on Large DatasetsabstractBayesian optimization has become a successful tool for hyperparameter optimization of machine learning algorithms, such as support vector machines or deep neural networks. Despite its success, for large datasets, training and validating a single configuration often takes hours, days, or even weeks, which limits the achievable performance. To accelerate hyperparameter optimization, we propose a generative model for the validation error as a function of training set size, which is learned during the optimization process and allows exploration of preliminary configurations on small subsets, by extrapolating to the full dataset. We construct a Bayesian optimization procedure, dubbed FABOLAS, which models loss and training time as a function of dataset size and automatically trades off high information gain about the global optimum against computational cost. Experiments optimizing support vector machines and deep neural networks show that FABOLAS often finds high-quality solutions 10 to 100 times faster than other state-of-the-art Bayesian optimization methods or the recently proposed bandit strategy Hyperband. Aaron Klein, Stefan Falkner, Simon Bartels, Philipp Hennig, Frank Hutter |
AISTATS | 1 |
| 2017 | Learning Curve Prediction with Bayesian Neural Networks
Aaron Klein, Stefan Falkner, Jost Tobias Springenberg, Frank Hutter |
ICLR (Poster) | 1 |
| 2016 | Bayesian Optimization with Robust Bayesian Neural NetworksabstractBayesian optimization is a prominent method for optimizing expensive to evaluate black-box functions that is prominently applied to tuning the hyperparameters of machine learning algorithms. Despite its successes, the prototypical Bayesian optimization approach - using Gaussian process models - does not scale well to either many hyperparameters or many function evaluations. Attacking this lack of scalability and flexibility is thus one of the key challenges of the field. We present a general approach for using flexible parametric models (neural networks) for Bayesian optimization, staying as close to a truly Bayesian treatment as possible. We obtain scalability through stochastic gradient Hamiltonian Monte Carlo, whose robustness we improve via a scale adaptation. Experiments including multi-task Bayesian optimization with 21 tasks, parallel optimization of deep neural networks and deep reinforcement learning show the power and flexibility of this approach. Jost Tobias Springenberg, Aaron Klein, Stefan Falkner, Frank Hutter |
NIPS | 2 |
| 2015 | Efficient and Robust Automated Machine LearningabstractThe success of machine learning in a broad range of applications has led to an ever-growing demand for machine learning systems that can be used off the shelf by non-experts. To be effective in practice, such systems need to automatically choose a good algorithm and feature preprocessing steps for a new dataset at hand, and also set their respective hyperparameters. Recent work has started to tackle this automated machine learning (AutoML) problem with the help of efficient Bayesian optimization methods. In this work we introduce a robust new AutoML system based on scikit-learn (using 15 classifiers, 14 feature preprocessing methods, and 4 data preprocessing methods, giving rise to a structured hypothesis space with 110 hyperparameters). This system, which we dub auto-sklearn, improves on existing AutoML methods by automatically taking into account past performance on similar datasets, and by constructing ensembles from the models evaluated during the optimization. Our system won the first phase of the ongoing ChaLearn AutoML challenge, and our comprehensive analysis on over 100 diverse datasets shows that it substantially outperforms the previous state of the art in AutoML. We also demonstrate the performance gains due to each of our contributions and derive insights into the effectiveness of the individual components of auto-sklearn. Matthias Feurer 0001, Aaron Klein, Katharina Eggensperger, Jost Tobias Springenberg, Manuel Blum 0002, Frank Hutter |
NIPS | 2 |