EDBT 2026 Demo / reviewers in the wild / expert
Felix Mohr
dblp:140/2172
· DBLP profile ↗
26ranked-venue papers
13as first author
14since 2021 · last 2025
0000-0002-9293-2424ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Credal Prediction based on Relative LikelihoodabstractPredictions in the form of sets of probability distributions, so-called credal sets, provide a suitable means to represent a learner's epistemic uncertainty. In this paper, we propose a theoretically grounded approach to credal prediction based on the statistical notion of relative likelihood: The target of prediction is the set of all (conditional) probability distributions produced by the collection of plausible models, namely those models whose relative likelihood exceeds a specified threshold. This threshold has an intuitive interpretation and allows for controlling the trade-off between correctness and precision of credal predictions. We tackle the problem of approximating credal sets defined in this way by means of suitably modified ensemble learning techniques.
To validate our approach, we illustrate its effectiveness by experiments on benchmark datasets demonstrating superior uncertainty representation without compromising predictive performance. We also compare our method against several state-of-the-art baselines in credal prediction. Timo Löhr, Paul Hofman, Felix Mohr, Eyke Hüllermeier |
NeurIPS | 3 |
| 2025 | LCDB 1.1: A Database Illustrating Learning Curves Are More Ill-Behaved Than Previously ThoughtabstractSample-wise learning curves plot performance versus training set size. They are useful for studying scaling laws and speeding up hyperparameter tuning and model selection. Learning curves are often assumed to be well-behaved: monotone (i.e. improving with more data) and convex. By constructing the Learning Curves Database 1.1 (LCDB 1.1), a large-scale database with high-resolution learning curves including more modern learners (CatBoost, TabNet, RealMLP, and TabPFN), we show that learning curves are less often well-behaved than previously thought. Using statistically rigorous methods, we observe significant ill-behavior in approximately 15% of the learning curves, almost twice as much as in previous estimates. We also identify which learners are to blame and show that specific learners are more ill-behaved than others. Additionally, we demonstrate that different feature scalings rarely resolve ill-behavior. We evaluate the impact of ill-behavior on downstream tasks, such as learning curve fitting and model selection, and find it poses significant challenges, underscoring the relevance and potential of LCDB 1.1 as a challenging benchmark for future research. Felix Mohr, Tom J. Viering |
NeurIPS | 2 |
| 2024 | Learning Curve Extrapolation Methods Across Extrapolation Settings
Lionel Kielhöfer, Felix Mohr, Jan N. van Rijn |
IDA (2) | 2 |
| 2024 | The unreasonable effectiveness of early discarding after one epoch in neural network hyperparameter optimizationabstractTo reach high performance with deep learning, hyperparameter optimization (HPO) is essential. This process is usually time-consuming due to costly evaluations of neural networks. Early discarding techniques limit the resources granted to unpromising candidates by observing the empirical learning curves and canceling neural network training as soon as the lack of competitiveness of a candidate becomes evident. Despite two decades of research, little is understood about the trade-off between the aggressiveness of discarding and the loss of predictive performance. Our paper studies this trade-off for several commonly used discarding techniques such as successive halving and learning curve extrapolation. Our surprising finding is that these commonly used techniques offer minimal to no added value compared to the simple strategy of discarding after a constant number of epochs of training. The chosen number of epochs mostly depends on the available compute budget. We call this approach i-Epoch (i being the constant number of epochs with which neural networks are trained) and suggest to assess the quality of early discarding techniques by comparing how their Pareto-Front (in consumed training epochs and predictive performance) complement the Pareto-Front of i-Epoch. Romain Egele, Felix Mohr, Tom J. Viering, Prasanna Balaprakash |
Neurocomputing | 2 |
| 2024 | Learning curves for decision making in supervised machine learning: a survey
Felix Mohr, Jan N. van Rijn |
Mach. Learn. | 1 |
| 2023 | RRR-Net: Reusing, Reducing, and Recycling a Deep Backbone NetworkabstractIt has become mainstream in computer vision and other machine learning domains to reuse backbone networks pretrained on large datasets as preprocessors. Typically, the last layer is replaced by a shallow learning machine of sorts; the newly-added classification head and (optionally) deeper layers are fine-tuned on a new task. Due to its strong performance and simplicity, a common pre-trained backbone network is ResNet152. However, ResNet152 is relatively large and induces inference latency. In many cases, a compact and efficient backbone with similar performance would be preferable over a larger, slower one. This paper investigates techniques to reuse a pre-trained backbone with the objective of creating a smaller and faster model. Starting from a large ResNet152 backbone pre-trained on ImageNet, we first reduce it from 51 blocks to 5 blocks, reducing its number of parameters and FLOPs by more than 6 times, without significant performance degradation. Then, we split the model after 3 blocks into several branches, while preserving the same number of parameters and FLOPs, to create an ensemble of sub-networks to improve performance. Our experiments on a large benchmark of 40 image classification datasets from various domains suggest that our techniques match the performance (if not better) of “classical backbone fine-tuning” while achieving a smaller model size and faster inference speed. Haozhe Sun, Isabelle Guyon, Felix Mohr, Hedi Tabia |
IJCNN | 3 |
| 2023 | Towards Green Automated Machine Learning: Status Quo and Future DirectionsabstractAutomated machine learning (AutoML) strives for the automatic configuration of machine learning algorithms and their composition into an overall (software) solution — a machine learning pipeline — tailored to the learning task (dataset) at hand. Over the last decade, AutoML has developed into an independent research field with hundreds of contributions. At the same time, AutoML is being criticized for its high resource consumption as many approaches rely on the (costly) evaluation of many machine learning pipelines, as well as the expensive large-scale experiments across many datasets and approaches. In the spirit of recent work on Green AI, this paper proposes Green AutoML, a paradigm to make the whole AutoML process more environmentally friendly. Therefore, we first elaborate on how to quantify the environmental footprint of an AutoML tool. Afterward, different strategies on how to design and benchmark an AutoML tool w.r.t. their “greenness”, i.e., sustainability, are summarized. Finally, we elaborate on how to be transparent about the environmental footprint and what kind of research incentives could direct the community in a more sustainable AutoML research direction. As part of this, we propose a sustainability checklist to be attached to every AutoML paper featuring all core aspects of Green AutoML. Tanja Tornede, Alexander Tornede, Jonas Hanselle, Felix Mohr, Marcel Wever, Eyke Hüllermeier |
J. Artif. Intell. Res. | 4 |
| 2023 | Naive automated machine learningabstractAbstract An essential task of automated machine learning ( $$\text {AutoML}$$ AutoML ) is the problem of automatically finding the pipeline with the best generalization performance on a given dataset. This problem has been addressed with sophisticated $$\text {black-box}$$ black-box optimization techniques such as Bayesian optimization, grammar-based genetic algorithms, and tree search algorithms. Most of the current approaches are motivated by the assumption that optimizing the components of a pipeline in isolation may yield sub-optimal results. We present $$\text {Naive AutoML}$$ Naive AutoML , an approach that precisely realizes such an in-isolation optimization of the different components of a pre-defined pipeline scheme. The returned pipeline is obtained by just taking the best algorithm of each slot. The isolated optimization leads to substantially reduced search spaces, and, surprisingly, this approach yields comparable and sometimes even better performance than current state-of-the-art optimizers. Felix Mohr, Marcel Wever |
Mach. Learn. | 1 |
| 2023 | Fast and Informative Model Selection Using Learning Curve Cross-ValidationabstractCommon cross-validation (CV) methods like k-fold cross-validation or Monte Carlo cross-validation estimate the predictive performance of a learner by repeatedly training it on a large portion of the given data and testing it on the remaining data. These techniques have two major drawbacks. First, they can be unnecessarily slow on large datasets. Second, beyond an estimation of the final performance, they give almost no insights into the learning process of the validated algorithm. In this article, we present a new approach for validation based on learning curves (LCCV). Instead of creating train-test splits with a large portion of training data, LCCV iteratively increases the number of instances used for training. In the context of model selection, it discards models that are unlikely to become competitive. In a series of experiments on 75 datasets, we could show that in over 90% of the cases using LCCV leads to the same performance as using 5/10-fold CV while substantially reducing the runtime (median runtime reductions of over 50%); the performance using LCCV never deviated from CV by more than 2.5%. We also compare it to a racing-based method and successive halving, a multi-armed bandit method. Additionally, it provides important insights, which for example allows assessing the benefits of acquiring more data. Felix Mohr, Jan N. van Rijn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Meta-Album: Multi-domain Meta-Dataset for Few-Shot Image ClassificationabstractWe introduce Meta-Album, an image classification meta-dataset designed to facilitate few-shot learning, transfer learning, meta-learning, among other tasks. It includes 40 open datasets, each having at least 20 classes with 40 examples per class, with verified licences. They stem from diverse domains, such as ecology (fauna and flora), manufacturing (textures, vehicles), human actions, and optical character recognition, featuring various image scales (microscopic, human scales, remote sensing). All datasets are preprocessed, annotated, and formatted uniformly, and come in 3 versions (Micro $\subset$ Mini $\subset$ Extended) to match users’ computational resources. We showcase the utility of the first 30 datasets on few-shot learning problems. The other 10 will be released shortly after. Meta-Album is already more diverse and larger (in number of datasets) than similar efforts, and we are committed to keep enlarging it via a series of competitions. As competitions terminate, their test data are released, thus creating a rolling benchmark, available through OpenML.org. Our website https://meta-album.github.io/ contains the source code of challenge winning methods, baseline methods, data loaders, and instructions for contributing either new datasets or algorithms to our expandable meta-dataset. Dustin Carrión-Ojeda, Sergio Escalera, Isabelle Guyon, Mike Huisman, Felix Mohr, Jan N. van Rijn, Haozhe Sun, Joaquin Vanschoren, Phan Anh Vu |
NeurIPS | 6 |
| 2022 | LCDB 1.0: An Extensive Learning Curves Database for Classification Tasks
Felix Mohr, Tom J. Viering, Marco Loog, Jan N. van Rijn |
ECML/PKDD (5) | 1 |
| 2021 | Single Player Monte-Carlo Tree Search Based on the Plackett-Luce Model
Felix Mohr, Viktor Bengs, Eyke Hüllermeier |
AAAI | 1 |
| 2021 | Predicting Machine Learning Pipeline Runtimes in the Context of Automated Machine LearningabstractAutomated machine learning (AutoML) seeks to automatically find so-called machine learning pipelines that maximize the prediction performance when being used to train a model on a given dataset. One of the main and yet open challenges in AutoMLis an effective use of computational resources: An AutoML process involves the evaluation of many candidate pipelines, which are costly but often ineffective because they are canceled due to a timeout. In this paper, we present an approach to predict the runtime of two-step machine learning pipelines with up to one pre-processor, which can be used to anticipate whether or not a pipeline will time out. Separate runtime models are trained offline for each algorithm that may be used in a pipeline, and an overall prediction is derived from these models. We empirically show that the approach increases successful evaluations made by an AutoML tool while preserving or even improving on the previously best solutions. Felix Mohr, Marcel Wever, Alexander Tornede, Eyke Hüllermeier |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | AutoML for Multi-Label Classification: Overview and Empirical EvaluationabstractAutomated machine learning (AutoML) supports the algorithmic construction and data-specific customization of machine learning pipelines, including the selection, combination, and parametrization of machine learning algorithms as main constituents. Generally speaking, AutoML approaches comprise two major components: a search space model and an optimizer for traversing the space. Recent approaches have shown impressive results in the realm of supervised learning, most notably (single-label) classification (SLC). Moreover, first attempts at extending these approaches towards multi-label classification (MLC) have been made. While the space of candidate pipelines is already huge in SLC, the complexity of the search space is raised to an even higher power in MLC. One may wonder, therefore, whether and to what extent optimizers established for SLC can scale to this increased complexity, and how they compare to each other. This paper makes the following contributions: First, we survey existing approaches to AutoML for MLC. Second, we augment these approaches with optimizers not previously tried for MLC. Third, we propose a benchmarking framework that supports a fair and systematic comparison. Fourth, we conduct an extensive experimental study, evaluating the methods on a suite of MLC problems. We find a grammar-based best-first search to compare favorably to other optimizers. Marcel Wever, Alexander Tornede, Felix Mohr, Eyke Hüllermeier |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Run2Survive: A Decision-theoretic Approach to Algorithm Selection based on Survival AnalysisabstractAlgorithm selection (AS) deals with the automatic selection of an algorithm from a fixed set of candidate algorithms most suitable for a specific instance of an algorithmic problem class, where “suitability” often refers to an algorithm’s runtime. Due to possibly extremely long runtimes of candidate algorithms, training data for algorithm selection models is usually generated under time constraints in the sense that not all algorithms are run to completion on all instances. Thus, training data usually comprises censored information, as the true runtime of algorithms timed out remains unknown. However, many standard AS approaches are not able to handle such information in a proper way. On the other side, survival analysis (SA) naturally supports censored data and offers appropriate ways to use such data for learning distributional models of algorithm runtime, as we demonstrate in this work. We leverage such models as a basis of a sophisticated decision-theoretic approach to algorithm selection, which we dub Run2Survive. Moreover, taking advantage of a framework of this kind, we advocate a risk-averse approach to algorithm selection, in which the avoidance of a timeout is given high priority. In an extensive experimental study with the standard benchmark ASlib, our approach is shown to be highly competitive and in many cases even superior to state-of-the-art AS approaches. Alexander Tornede, Marcel Wever, Felix Mohr, Eyke Hüllermeier |
ACML | 4 |
| 2020 | LiBRe: Label-Wise Selection of Base Learners in Binary Relevance for Multi-label ClassificationabstractIn multi-label classification (MLC), each instance is associated with a set of class labels, in contrast to standard classification, where an instance is assigned a single label. Binary relevance (BR) learning, which reduces a multi-label to a set of binary classification problems, one per label, is arguably the most straight-forward approach to MLC. In spite of its simplicity, BR proved to be competitive to more sophisticated MLC methods, and still achieves state-of-the-art performance for many loss functions. Somewhat surprisingly, the optimal choice of the base learner for tackling the binary classification problems has received very little attention so far. Taking advantage of the label independence assumption inherent to BR, we propose a label-wise base learner selection method optimizing label-wise macro averaged performance measures. In an extensive experimental evaluation, we find that or approach, called LiBRe, can significantly improve generalization performance. Marcel Wever, Alexander Tornede, Felix Mohr, Eyke Hüllermeier |
IDA | 3 |
| 2018 | Ensembles of evolved nested dichotomies for classificationabstractIn multinomial classification, reduction techniques are commonly used to decompose the original learning problem into several simpler problems. For example, by recursively bisecting the original set of classes, so-called nested dichotomies define a set of binary classification problems that are organized in the structure of a binary tree. In contrast to the existing one-shot heuristics for constructing nested dichotomies and motivated by recent work on algorithm configuration, we propose a genetic algorithm for optimizing the structure of such dichotomies. A key component of this approach is the proposed genetic representation that facilitates the application of standard genetic operators, while still supporting the exchange of partial solutions under recombination. We evaluate the approach in an extensive experimental study, showing that it yields classifiers with superior generalization performance. Marcel Wever, Felix Mohr, Eyke Hüllermeier |
GECCO | 2 |
| 2018 | Reduction Stumps for Multi-class Classification
Felix Mohr, Marcel Wever, Eyke Hüllermeier |
IDA | 1 |
| 2018 | ML-Plan: Automated machine learning via hierarchical planning
Felix Mohr, Marcel Wever, Eyke Hüllermeier |
Mach. Learn. | 1 |
| 2015 | A Metric for Functional Reusability of Services
Felix Mohr |
ICSR | 1 |
| 2015 | Template-Based Generation of Semantic Services
Felix Mohr, Sven Walther |
ICSR | 1 |
| 2015 | Market-Specific Service Compositions: Specification and MatchingabstractThe Collaborative Research Centre "On-The-Fly Computing" works on foundations and principles for the vision of the Future Internet. It proposes the paradigm of On-The-Fly Computing, which tackles emerging worldwide service markets. In these markets, service providers trade software, platform, and infrastructure as a service. Service requesters state requirements on services. To satisfy these requirements, the new role of brokers, who are (human) actors building service compositions on the fly, is introduced. Brokers have to specify service compositions formally and comprehensively using a domain-specific language (DSL), and to use service matching for the discovery of the constituent services available in the market. The broker's choice of the DSL and matching approaches influences her success of building compositions as distinctive properties of different service markets play a significant role. In this paper, we propose a new approach of engineering a situation-specific DSL by customizing a comprehensive, modular DSL and its matching for given service market properties. This enables the broker to create market-specific composition specifications and to perform market-specific service matching. As a result, the broker builds service compositions satisfying the requester's requirements more accurately. We evaluated the presented concepts using case studies in service markets for tourism and university management. Svetlana Arifulina, Felix Mohr, Gregor Engels, Marie Platenius-Mohr, Wilhelm Schäfer |
SERVICES | 2 |
| 2014 | Estimating Functional Reusability of Services
Felix Mohr |
ICSOC | 1 |
| 2014 | Issues of automated software composition in AI planningabstractAutomated programming aims at automatically assembling a new software artifact from existing software modules. Although automated programming was revitalized through automated software composition in the last decade, the problem cannot be considered solved. Automated software composition is widely accepted as being a planning task, but the problem is that it has very special properties that other planning problems do not have and that are commonly overseen. These properties usually imply that the composition problem cannot be solved with standard planning tools. This paper gives a brief and intuitive description of the planning problem that most approaches are based on. It points out special properties of this problem and explains why it is not adequate to solve the problem with classical planning tools as done by most existing approaches. Felix Mohr |
ASE | 1 |
| 2014 | Combining Automatic Service Composition with Adaptive Service Recommendation for Dynamic Markets of ServicesabstractAutomatic service composition is still a challenging task. It is even more challenging when dealing with a dynamic market of services for end users. New services may enter the market while other services are completely removed. Furthermore, end users are typically no experts in the domain in which they formulate a request. As a consequence, ambiguous user requests will inevitably emerge and have to be taken into account. To meet these challenges, we propose a new approach that combines automatic service composition with adaptive service recommendation. A best first backward search algorithm produces solutions that are functional correct with respect to user requests. An adaptive recommendation system supports the search algorithm in decision-making. Reinforcement Learning techniques enable the system to adjust its recommendation strategy over time based on user ratings. The integrated approach is described on a conceptional level and demonstrated by means of an illustrative example from the image processing domain. Alexander Jungmann, Felix Mohr, Bernd Kleinjohann |
SERVICES | 2 |
| 2013 | Semi-Automated Software Composition Through Generated ComponentsabstractSoftware composition has been studied as a subject of state based planning for decades. Existing composition approaches that are efficient enough to be used in practice are limited to sequential arrangements of software components. This restriction dramatically reduces the number of composition problems that can be solved. However, there are many composition problems that could be solved by existing approaches if they had a possibility to combine components in very simple non-sequential ways. Felix Mohr, Hans Kleine Büning |
iiWAS | 1 |