Arlind Kadra

dblp:252/5295 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Optimization for machine learning · 42% Deep learning architectures and training · 17% Trustworthy machine learning · 14%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
hyperparameter optimization
2.032024
Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and How · ICLR 2024
Scaling Laws for Hyperparameter Optimization · NeurIPS 2023
Supervising the Multi-Fidelity Race of Hyperparameter Configurations · NeurIPS 2022
Machine learning › Representation and self-supervised learning
tabular data
1.322024
Interpretable Mesomorphic Networks for Tabular Data · NeurIPS 2024
Well-tuned Simple Nets Excel on Tabular Datasets · NeurIPS 2021
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
1.222023
Scaling Laws for Hyperparameter Optimization · NeurIPS 2023
Supervising the Multi-Fidelity Race of Hyperparameter Configurations · NeurIPS 2022
Machine learning › Deep learning architectures and training
hypernetwork
0.812024
Interpretable Mesomorphic Networks for Tabular Data · NeurIPS 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Interpretable Mesomorphic Networks for Tabular Data · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › explainable AI
interpretable neural network
0.812024
Interpretable Mesomorphic Networks for Tabular Data · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › pre-trained models
pre-trained model selection
0.812024
Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and How · ICLR 2024
Machine learning › Optimization for machine learning
learning curve extrapolation
0.712023
Scaling Laws for Hyperparameter Optimization · NeurIPS 2023
Machine learning › Learning theory › neural network theory
power-law scaling
0.712023
Scaling Laws for Hyperparameter Optimization · NeurIPS 2023
Machine learning › Optimization for machine learning › hyperparameter optimization
multi-fidelity hyperparameter tuning
0.612022
Supervising the Multi-Fidelity Race of Hyperparameter Configurations · NeurIPS 2022
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron
0.512021
Well-tuned Simple Nets Excel on Tabular Datasets · NeurIPS 2021
Machine learning › Deep learning architectures and training
regularization
0.512021
Well-tuned Simple Nets Excel on Tabular Datasets · NeurIPS 2021
Machine learning and data management › machine learning lifecycle management
experiment management
0.512021
OpenML-Python: an extensible Python API for OpenML · J. Mach. Learn. Res. 2021
Machine learning and data management › machine learning systems
machine learning platform
0.512021
OpenML-Python: an extensible Python API for OpenML · J. Mach. Learn. Res. 2021
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.212022
Supervising the Multi-Fidelity Race of Hyperparameter Configurations · NeurIPS 2022
Software maintenance and evolution
software ecosystems
0.112021
OpenML-Python: an extensible Python API for OpenML · J. Mach. Learn. Res. 2021

Methods — techniques the papers use, named apart from their topics

learning curve modeling · 1.3scikit-learn extension · 1.0API design · 1.0performance prediction · 0.8per-instance linear models · 0.8deep hypernetworks · 0.8gray-box evaluation · 0.7ensemble of neural networks · 0.7deep power laws · 0.7gaussian process · 0.6acquisition function · 0.6gradient-boosted decision trees · 0.5
YearPublicationVenuePosition
2024 Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and How
abstract
With the ever-increasing number of pretrained models, machine learning practitioners are continuously faced with which pretrained model to use, and how to finetune it for a new dataset. In this paper, we propose a methodology that jointly searches for the optimal pretrained model and the hyperparameters for finetuning it. Our method transfers knowledge about the performance of many pretrained models with multiple hyperparameter configurations on a series of datasets. To this aim, we evaluated over 20k hyperparameter configurations for finetuning 24 pretrained image classification models on 87 datasets to generate a large-scale meta-dataset. We meta-learn a gray-box performance predictor on the learning curves of this meta-dataset and use it for fast hyperparameter optimization on new datasets. We empirically demonstrate that our resulting approach can quickly select an accurate pretrained model for a new dataset together with its optimal hyperparameters.
Sebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter, Josif Grabocka
ICLR3
2024 Interpretable Mesomorphic Networks for Tabular Data
abstract
Even though neural networks have been long deployed in applications involving tabular data, still existing neural architectures are not explainable by design. In this paper, we propose a new class of interpretable neural networks for tabular data that are both deep and linear at the same time (i.e. mesomorphic). We optimize deep hypernetworks to generate explainable linear models on a per-instance basis. As a result, our models retain the accuracy of black-box deep networks while offering free-lunch explainability for tabular data by design. Through extensive experiments, we demonstrate that our explainable deep networks have comparable performance to state-of-the-art classifiers on tabular data and outperform current existing methods that are explainable by design.
Arlind Kadra, Sebastian Pineda-Arango, Josif Grabocka
NeurIPS1
2023 Scaling Laws for Hyperparameter Optimization
abstract
Hyperparameter optimization is an important subfield of machine learning that focuses on tuning the hyperparameters of a chosen algorithm to achieve peak performance. Recently, there has been a stream of methods that tackle the issue of hyperparameter optimization, however, most of the methods do not exploit the dominant power law nature of learning curves for Bayesian optimization. In this work, we propose Deep Power Laws (DPL), an ensemble of neural network models conditioned to yield predictions that follow a power-law scaling pattern. Our method dynamically decides which configurations to pause and train incrementally by making use of gray-box evaluations. We compare our method against 7 state-of-the-art competitors on 3 benchmarks related to tabular, image, and NLP datasets covering 59 diverse tasks. Our method achieves the best results across all benchmarks by obtaining the best any-time results compared to all competitors.
Arlind Kadra, Maciej Janowski, Martin Wistuba, Josif Grabocka
NeurIPS1
2022 Supervising the Multi-Fidelity Race of Hyperparameter Configurations
abstract
Multi-fidelity (gray-box) hyperparameter optimization techniques (HPO) have recently emerged as a promising direction for tuning Deep Learning methods. However, existing methods suffer from a sub-optimal allocation of the HPO budget to the hyperparameter configurations. In this work, we introduce DyHPO, a Bayesian Optimization method that learns to decide which hyperparameter configuration to train further in a dynamic race among all feasible configurations. We propose a new deep kernel for Gaussian Processes that embeds the learning curve dynamics, and an acquisition function that incorporates multi-budget information. We demonstrate the significant superiority of DyHPO against state-of-the-art hyperparameter optimization methods through large-scale experiments comprising 50 datasets (Tabular, Image, NLP) and diverse architectures (MLP, CNN/NAS, RNN).
Martin Wistuba, Arlind Kadra, Josif Grabocka
NeurIPS2
2021 Well-tuned Simple Nets Excel on Tabular Datasets
abstract
Tabular datasets are the last "unconquered castle" for deep learning, with traditional ML methods like Gradient-Boosted Decision Trees still performing strongly even against recent specialized neural architectures. In this paper, we hypothesize that the key to boosting the performance of neural networks lies in rethinking the joint and simultaneous application of a large set of modern regularization techniques. As a result, we propose regularizing plain Multilayer Perceptron (MLP) networks by searching for the optimal combination/cocktail of 13 regularization techniques for each dataset using a joint optimization over the decision on which regularizers to apply and their subsidiary hyperparameters. We empirically assess the impact of these regularization cocktails for MLPs in a large-scale empirical study comprising 40 tabular datasets and demonstrate that (i) well-regularized plain MLPs significantly outperform recent state-of-the-art specialized neural network architectures, and (ii) they even outperform strong traditional ML methods, such as XGBoost.
Arlind Kadra, Marius Lindauer, Frank Hutter, Josif Grabocka
NeurIPS1
2021 OpenML-Python: an extensible Python API for OpenML
abstract
OpenML is an online platform for open science collaboration in machine learning, used to share datasets and results of machine learning experiments. In this paper, we introduce OpenML-Python, a client API for Python, which opens up the OpenML platform for a wide range of Python-based machine learning tools. It provides easy access to all datasets, tasks and experiments on OpenML from within Python. It also provides functionality to conduct machine learning experiments, upload the results to OpenML, and reproduce results which are stored on OpenML. Furthermore, it comes with a scikit-learn extension and an extension mechanism to easily integrate other machine learning libraries written in Python into the OpenML ecosystem. Source code and documentation are available at https://github.com/openml/openml-python/.
Matthias Feurer 0001, Jan N. van Rijn, Arlind Kadra, Pieter Gijsbers, Neeratyoy Mallik, Sahithya Ravi, Andreas C. Müller 0001, Joaquin Vanschoren, Frank Hutter
J. Mach. Learn. Res.3