Raphael Fischer 0001

dblp:249/4056 · DBLP profile ↗
← Back
6ranked-venue papers in the field
5as first author
5since 2021 · last 2026
0000-0002-1808-5773ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 6 (5 first)
YearPublicationVenuePosition
2026 Lift what you can: green online learning with heterogeneous ensembles
abstract
Abstract Ensemble methods for stream mining necessitate managing multiple models and updating them as data distributions evolve. Considering the calls for more sustainability, established methods are however not sufficiently considerate of ensemble members’ computational expenses and instead overly focus on predictive capabilities. To address these challenges and enable green online learning, we propose heterogeneous online ensembles (HEROS). For every training step, HEROS chooses a subset of models from a pool of models initialized with diverse hyperparameter choices under resource constraints to train. We introduce a Markov decision process to theoretically capture the trade-offs between predictive performance and sustainability constraints. Based on this framework, we present different policies for choosing which models to train on incoming data. Most notably, we propose the novel $$\zeta $$ -policy, which focuses on training near-optimal models at reduced costs. Using a stochastic model, we theoretically prove that our $$\zeta $$ -policy achieves near optimal performance while using fewer resources compared to the best performing policy. In our experiments across 11 benchmark datasets, we find empiric evidence that our $$\zeta $$ -policy is a strong contribution to the state-of-the-art, demonstrating highly accurate performance, in some cases even outperforming competitors, and simultaneously being much more resource-friendly.
Kirsten Köbschall, Sebastian Buschjäger, Raphael Fischer 0001, Lisa Hartung, Stefan Kramer 0001
Data Min. Knowl. Discov.3
2024 AutoXPCR: Automated Multi-Objective Model Selection for Time Series Forecasting
abstract
Automated machine learning (AutoML) streamlines the creation of ML models, but few specialized methods have approached the challenging domain of time series forecasting. Deep neural networks (DNNs) often deliver state-of-the-art predictive performance for forecasting data, however these models are also criticized for being computationally intensive black boxes. As a result, when searching for the "best" model, it is crucial to also acknowledge other aspects, such as interpretability and resource consumption. In this paper, we propose AutoXPCR - a novel method that produces DNNs for forecasting under consideration of multiple objectives in an automated and explainable fashion. Our approach leverages meta-learning to estimate any model's performance along PCR criteria, which encompass (P)redictive error, (C)omplexity, and (R)esource demand. Explainability is addressed on multiple levels, as AutoXPCR pro-vides by-product explanations of recommendations and allows to interactively control the desired PCR criteria importance and trade-offs. We demonstrate the practical feasibility AutoXPCR across 108 forecasting data sets from various domains. Notably, our method outperforms competing AutoML approaches - on average, it only requires 20% of computation costs for recommending highly efficient models with 85% of the empirical best quality.
Raphael Fischer 0001, Amal Saadallah
KDD1
2024 MetaQuRe: Meta-learning from Model Quality and Resource Consumption
Raphael Fischer 0001, Marcel Wever, Sebastian Buschjäger, Thomas Liebig
ECML/PKDD (7)1
2024 Towards more sustainable and trustworthy reporting in machine learning
abstract
Abstract With machine learning (ML) becoming a popular tool across all domains, practitioners are in dire need of comprehensive reporting on the state-of-the-art. Benchmarks and open databases provide helpful insights for many tasks, however suffer from several phenomena: Firstly, they overly focus on prediction quality, which is problematic considering the demand for more sustainability in ML. Depending on the use case at hand, interested users might also face tight resource constraints and thus should be allowed to interact with reporting frameworks, in order to prioritize certain reported characteristics. Furthermore, as some practitioners might not yet be well-skilled in ML, it is important to convey information on a more abstract, comprehensible level. Usability and extendability are key for moving with the state-of-the-art and in order to be trustworthy, frameworks should explicitly address reproducibility. In this work, we analyze established reporting systems under consideration of the aforementioned issues. Afterwards, we propose STREP, our novel framework that aims at overcoming these shortcomings and paves the way towards more sustainable and trustworthy reporting. We use STREP’s (publicly available) implementation to investigate various existing report databases. Our experimental results unveil the need for making reporting more resource-aware and demonstrate our framework’s capabilities of overcoming current reporting limitations. With our work, we want to initiate a paradigm shift in reporting and help with making ML advances more considerate of sustainability and trustworthiness.
Raphael Fischer 0001, Thomas Liebig, Katharina Morik
Data Min. Knowl. Discov.1
2023 Prioritization of Identified Data Science Use Cases in Industrial Manufacturing via C-EDIF Scoring
abstract
While data science and artificial intelligence (AI) can be highly beneficial for industrial manufacturers, it is not yet readily usable. Therefore, putting it to good use requires to understand the domain challenges and identify opportunities for deploying AI. Our work aims at solving this task by proposing a generalized framework for ❨1❩ exploring companies for use cases and (2) prioritizing them via C-EDIF scoring. This novel approach allows to determine the business importance of any use case by considering the underlying evaluability, data situation, impact and infeasibility. Besides the theoretical framework, our work also provides real-world insights from applying C-EDIF scoring in an extensive use case exploration phase. These results stem from a strategic partnership between data scientists and Wilo SE, a renowned pump manufacturing company, where we successfully identified and rated opportunities for AI.
Raphael Fischer 0001, Andreas Pauly, Rahel Wilking, Anoop Kini, David Graurock
DSAA1
2020 No Cloud on the Horizon: Probabilistic Gap Filling in Satellite Image Series
abstract
Spatio-temporal data sets such as satellite image series are of utmost importance for understanding global developments like climate change or urbanization. However, incompleteness of data can greatly impact usability and knowledge discovery. In fact, there are many cases where not a single data point in the set is fully observed. For filling gaps, we introduce a novel approach that utilizes Markov random fields (MRFs). We extend the probabilistic framework to also consider empirical prior information, which allows to train even on highly incomplete data. Moreover, we devise a way to make discrete MRFs predict continuous values via state superposition. Experiments on real-world remote sensing imagery suffering from cloud cover show that the proposed approach outperforms state-of-the-art gap filling techniques.
Raphael Fischer 0001, Nico Piatkowski, Charlotte Pelletier, Geoffrey I. Webb, François Petitjean, Katharina Morik
DSAA1