Steven Adriaensen

dblp:148/1033 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Optimization for machine learning · 46% Deep learning architectures and training · 25% Probabilistic and Bayesian machine learning · 21%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
hyperparameter optimization
2.842025
Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning · NeurIPS 2025
In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter Optimization · ICML 2024
Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks · NeurIPS 2023
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
1.022025
In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter Optimization · ICML 2024
Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › neural processes
prior-data fitted networks
0.922025
Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks · NeurIPS 2023
Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks · ICML 2025
Machine learning › Deep learning architectures and training
activation function
0.912025
Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning Dynamics · NeurIPS 2025
Machine learning › Learning theory › generalization
extrapolation
0.912025
Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks · ICML 2025
Machine learning › Optimization for machine learning
gradient flow
0.912025
Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning Dynamics · NeurIPS 2025
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
multi-fidelity bayesian optimization
0.912025
Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning · NeurIPS 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks · ICML 2025
Machine learning › Deep learning architectures and training › activation function
self-gated activation
0.912025
Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning Dynamics · NeurIPS 2025
Machine learning › Deep learning architectures and training
training dynamics
0.912025
Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning Dynamics · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.712023
Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.712023
Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks · NeurIPS 2023
Machine learning › Optimization for machine learning
learning curve extrapolation
0.712023
Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks · NeurIPS 2023
Machine learning › Optimization for machine learning › hyperparameter optimization
dynamic algorithm configuration
0.512021
DACBench: A Benchmark Library for Dynamic Algorithm Configuration · IJCAI 2021
Performance modeling and evaluation
benchmarking
0.512021
DACBench: A Benchmark Library for Dynamic Algorithm Configuration · IJCAI 2021
Mathematical optimization › black-box optimization
algorithm configuration
0.212016
Towards a White Box Approach to Automated Algorithm Design · IJCAI 2016
Mathematical optimization
automated algorithm design
0.212016
Towards a White Box Approach to Automated Algorithm Design · IJCAI 2016
Machine learning › Deep learning architectures and training › regularization
early stopping
0.212023
Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks · NeurIPS 2023
Machine learning › Learning theory
model selection
0.212023
Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

prior-data fitted networks · 2.3transformer · 1.4transfer learning · 0.9surrogate modeling · 0.9meta-learning · 0.9gompertz function · 0.9bayesian active learning · 0.9acquisition function design · 0.9in-context learning · 0.8acquisition function · 0.8dynamic algorithm configuration · 0.5benchmarking · 0.5optimization · 0.2metaheuristic · 0.2
YearPublicationVenuePosition
2025 Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
abstract
Scaling has been a major driver of recent advancements in deep learning. Numerous empirical studies have found that scaling laws often follow the power-law and proposed several variants of power-law functions to predict the scaling behavior at larger scales. However, existing methods mostly rely on point estimation and do not quantify uncertainty, which is crucial for real-world applications involving decision-making problems such as determining the expected performance improvements achievable by investing additional computational resources. In this work, we explore a Bayesian framework based on Prior-data Fitted Networks (PFNs) for neural scaling law extrapolation. Specifically, we design a prior distribution that enables the sampling of infinitely many synthetic functions resembling real-world neural scaling laws, allowing our PFN to meta-learn the extrapolation. We validate the effectiveness of our approach on real-world neural scaling laws, comparing it against both the existing point estimation methods and Bayesian approaches. Our method demonstrates superior performance, particularly in data-limited scenarios such as Bayesian active learning, underscoring its potential for reliable, uncertainty-aware extrapolation in practical applications.
Dong Bok Lee, Steven Adriaensen, Sung Ju Hwang, Frank Hutter, Seon Joo Kim, Hae Beom Lee
ICML3
2025 Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning Dynamics
abstract
Activation functions are fundamental elements of deep learning architectures as they significantly influence training dynamics. ReLU, while widely used, is prone to the dying neuron problem, which has been mitigated by variants such as LeakyReLU, PReLU, and ELU that better handle negative neuron outputs. Recently, self-gated activations like GELU and Swish have emerged as state-of-the-art alternatives, leveraging their smoothness to ensure stable gradient flow and prevent neuron inactivity. In this work, we introduce the Gompertz Linear Unit (GoLU), a novel self-gated activation function defined as $\mathrm{GoLU}(x) = x \\, \mathrm{Gompertz}(x)$, where $\mathrm{Gompertz}(x) = e^{-e^{-x}}$. The GoLU activation leverages the right-skewed asymmetry in the Gompertz function to reduce variance in the latent space more effectively compared to GELU and Swish, while preserving robust gradient flow. Extensive experiments across diverse tasks, including Image Classification, Language Modeling, Semantic Segmentation, Object Detection, Instance Segmentation, and Diffusion, highlight GoLU's superior performance relative to state-of-the-art activation functions, establishing GoLU as a robust alternative to existing activation functions.
Indrashis Das, Mahmoud Safari, Steven Adriaensen, Frank Hutter
NeurIPS3
2025 Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
abstract
In this paper, we address the problem of cost-sensitive hyperparameter optimization (HPO) built upon freeze-thaw Bayesian optimization (BO). Specifically, we assume a scenario where users want to early-stop the HPO process when the expected performance improvement is not satisfactory with respect to the additional computational cost. Motivated by this scenario, we introduce \emph{utility} in the freeze-thaw framework, a function describing the trade-off between the cost and performance that can be estimated from the user's preference data. This utility function, combined with our novel acquisition function and stopping criterion, allows us to dynamically continue training the configuration that we expect to maximally improve the utility in the future, and also automatically stop the HPO process around the maximum utility. Further, we improve the sample efficiency of existing freeze-thaw methods with transfer learning to develop a specialized surrogate model for the cost-sensitive HPO problem. We validate our algorithm on established multi-fidelity HPO benchmarks and show that it outperforms all the previous freeze-thaw BO and transfer-BO baselines we consider, while achieving a significantly better trade-off between the cost and performance.
Dong Bok Lee, Aoxuan Silvia Zhang, Byungjoo Kim, Junhyeon Park, Steven Adriaensen, Juho Lee 0001, Sung Ju Hwang, Hae Beom Lee
NeurIPS5
2024 In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter Optimization
abstract
With the increasing computational costs associated with deep learning, automated hyperparameter optimization methods, strongly relying on black-box Bayesian optimization (BO), face limitations. Freeze-thaw BO offers a promising grey-box alternative, strategically allocating scarce resources incrementally to different configurations. However, the frequent surrogate model updates inherent to this approach pose challenges for existing methods, requiring retraining or fine-tuning their neural network surrogates online, introducing overhead, instability, and hyper-hyperparameters. In this work, we propose FT-PFN, a novel surrogate for Freeze-thaw style BO. FT-PFN is a prior-data fitted network (PFN) that leverages the transformers' in-context learning ability to efficiently and reliably do Bayesian learning curve extrapolation in a single forward pass. Our empirical analysis across three benchmark suites shows that the predictions made by FT-PFN are more accurate and 10-100 times faster than those of the deep Gaussian process and deep ensemble surrogates used in previous work. Furthermore, we show that, when combined with our novel acquisition mechanism (MFPI-random), the resulting in-context freeze-thaw BO method (ifBO), yields new state-of-the-art performance in the same three families of deep learning HPO benchmarks considered in prior work.
Herilalaina Rakotoarison, Steven Adriaensen, Neeratyoy Mallik, Samir Garibov, Eddie Bergman, Frank Hutter
ICML2
2023 Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks
abstract
Learning curve extrapolation aims to predict model performance in later epochs of training, based on the performance in earlier epochs. In this work, we argue that, while the inherent uncertainty in the extrapolation of learning curves warrants a Bayesian approach, existing methods are (i) overly restrictive, and/or (ii) computationally expensive. We describe the first application of prior-data fitted neural networks (PFNs) in this context. A PFN is a transformer, pre-trained on data generated from a prior, to perform approximate Bayesian inference in a single forward pass. We propose LC-PFN, a PFN trained to extrapolate 10 million artificial right-censored learning curves generated from a parametric prior proposed in prior art using MCMC. We demonstrate that LC-PFN can approximate the posterior predictive distribution more accurately than MCMC, while being over 10 000 times faster. We also show that the same LC-PFN achieves competitive performance extrapolating a total of 20 000 real learning curves from four learning curve benchmarks (LCBench, NAS-Bench-201, Taskset, and PD1) that stem from training a wide range of model architectures (MLPs, CNNs, RNNs, and Transformers) on 53 different datasets with varying input modalities (tabular, image, text, and protein data). Finally, we investigate its potential in the context of model selection and find that a simple LC-PFN based predictive early stopping criterion obtains 2 - 6x speed-ups on 45 of these datasets, at virtually no overhead.
Steven Adriaensen, Herilalaina Rakotoarison, Samuel Müller 0005, Frank Hutter
NeurIPS1
2022 Automated Dynamic Algorithm Configuration
abstract
The performance of an algorithm often critically depends on its parameter configuration. While a variety of automated algorithm configuration methods have been proposed to relieve users from the tedious and error-prone task of manually tuning parameters, there is still a lot of untapped potential as the learned configuration is static, i.e., parameter settings remain fixed throughout the run. However, it has been shown that some algorithm parameters are best adjusted dynamically during execution. Thus far, this is most commonly achieved through hand-crafted heuristics. A promising recent alternative is to automatically learn such dynamic parameter adaptation policies from data. In this article, we give the first comprehensive account of this new field of automated dynamic algorithm configuration (DAC), present a series of recent advances, and provide a solid foundation for future research in this field. Specifically, we (i) situate DAC in the broader historical context of AI research; (ii) formalize DAC as a computational problem; (iii) identify the methods used in prior art to tackle this problem; and (iv) conduct empirical case studies for using DAC in evolutionary optimization, AI planning, and machine learning.
Steven Adriaensen, André Biedenkapp, Gresa Shala, Noor H. Awad, Theresa Eimer, Marius Lindauer, Frank Hutter
J. Artif. Intell. Res.1
2021 DACBench: A Benchmark Library for Dynamic Algorithm Configuration
abstract
Dynamic Algorithm Configuration (DAC) aims to dynamically control a target algorithm's hyperparameters in order to improve its performance. Several theoretical and empirical results have demonstrated the benefits of dynamically controlling hyperparameters in domains like evolutionary computation, AI Planning or deep learning. Replicating these results, as well as studying new methods for DAC, however, is difficult since existing benchmarks are often specialized and incompatible with the same interfaces. To facilitate benchmarking and thus research on DAC, we propose DACBench, a benchmark library that seeks to collect and standardize existing DAC benchmarks from different AI domains, as well as provide a template for new ones. For the design of DACBench, we focused on important desiderata, such as (i) flexibility, (ii) reproducibility, (iii) extensibility and (iv) automatic documentation and visualization. To show the potential, broad applicability and challenges of DAC, we explore how a set of six initial benchmarks compare in several dimensions of difficulty.
Theresa Eimer, André Biedenkapp, Maximilian Reimer, Steven Adriaensen, Frank Hutter, Marius Lindauer
IJCAI4
2020 Learning Step-Size Adaptation in CMA-ES
Gresa Shala, André Biedenkapp, Noor H. Awad, Steven Adriaensen, Marius Lindauer, Frank Hutter
PPSN (1)4
2016 Case study: An analysis of accidental complexity in a state-of-the-art hyper-heuristic for HyFlex
abstract
While simplicity is an important factor affecting algorithm re-usability, it is often overlooked in algorithm design, which has a tendency to produce overly complex methods. In this paper we demonstrate Accidental Complexity Analysis (ACA), a research practice targeted at detecting and eliminating accidental complexity, without loss of performance (c.f. refactoring in software engineering), using it to analyze the presence of accidental complexity in GIHH, a state-of-the-art selection hyper-heuristic for HyFlex. We identify various algorithmic sub-mechanisms contributing little to GIHH's overall performance, and validate many other. As an outcome we present Lean-GIHH, a simplified, re-implementation of GIHH.
Steven Adriaensen, Ann Nowé
CEC1
2016 Towards a White Box Approach to Automated Algorithm Design
Steven Adriaensen, Ann Nowé
IJCAI1
2015 A benchmark set extension and comparative study for the HyFlex framework
abstract
In this work we conduct a comparative study of several publicly available, state-of-the-art hyper-heuristics for HyFlex in order to assess their generality across domains. To this purpose we extend the HyFlex benchmark set with 3 new problem domains: The 0-1 Knap Sack, Quadratic Assignment and Max-Cut Problem. To our knowledge, this is the first public extension of the benchmark since the CHeSC 2011 competition. In addition, this is the first study testing the Fair-Share Iterated Local Search (FS-ILS) method, designed in prior research, using a semi-automated design approach, on new unseen problem domains. We show that, of the methods compared, Adap-HH (CHeSC 2011 winner) clearly perfoms the most consistently, overall. In addition, we identify a weakness of, as well as a way to further simplify the FS-ILS method. Finally, we found that, overall, the state-of-the-art methods compared, generalized much better than a naive baseline.
Steven Adriaensen, Gabriela Ochoa, Ann Nowé
CEC1
2014 Designing reusable metaheuristic methods: A semi-automated approach
abstract
Many interesting optimization problems cannot be solved efficiently. Recently, a lot of work has been done on meta-heuristic optimization methods that quickly find approximate solutions to otherwise intractable problems. While successful, the field suffers from a notable lack of reuse of methods, both in practical applications as in research. In this paper, we describe a semi-automated approach to design more re-usable methods, based on key principles of re-usability such as simplicity, modularity and generality. We illustrate this methodology by designing general metaheuristics (using hyperheuristics) and show that the methods obtained are competitive with the contestants of the Cross-Domain Heuristic Search Competition (2011). In particular, we find a method performing better than the competition's winner, which can be considered the state-of-the-art in domain-independent metaheuristic search.
Steven Adriaensen, Tim Brys, Ann Nowé
IEEE Congress on Evolutionary Computation1
2014 Fair-share ILS: a simple state-of-the-art iterated local search hyperheuristic
abstract
In this work we present a simple state-of-the-art selection hyperheuristic called Fair-Share Iterated Local Search (FS-ILS). FS-ILS is an iterated local search method using a conservative restart condition. Each iteration, a perturbation heuristic is selected proportionally to the acceptance rate of its previously proposed candidate solutions (after iterative improvement) by a domain-independent variant of the Metropolis condition. FS-ILS was developed in prior work using a semi-automated design approach. That work focused on how the method was found, rather than the method itself. As a result, it lacked a detailed explanation and analysis of the method, which will be the main contribution of this work. In our experiments we analyze FS-ILS's parameter sensitivity, accidental complexity and compare it to the contestants of the CHeSC (2011) competition.
Steven Adriaensen, Tim Brys, Ann Nowé
GECCO1