Sebastian Buschjäger

dblp:210/4306 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-2780-3618ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 6 first-author · 5 since 2021Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Lift what you can: green online learning with heterogeneous ensembles
abstract
Abstract Ensemble methods for stream mining necessitate managing multiple models and updating them as data distributions evolve. Considering the calls for more sustainability, established methods are however not sufficiently considerate of ensemble members’ computational expenses and instead overly focus on predictive capabilities. To address these challenges and enable green online learning, we propose heterogeneous online ensembles (HEROS). For every training step, HEROS chooses a subset of models from a pool of models initialized with diverse hyperparameter choices under resource constraints to train. We introduce a Markov decision process to theoretically capture the trade-offs between predictive performance and sustainability constraints. Based on this framework, we present different policies for choosing which models to train on incoming data. Most notably, we propose the novel $$\zeta $$ -policy, which focuses on training near-optimal models at reduced costs. Using a stochastic model, we theoretically prove that our $$\zeta $$ -policy achieves near optimal performance while using fewer resources compared to the best performing policy. In our experiments across 11 benchmark datasets, we find empiric evidence that our $$\zeta $$ -policy is a strong contribution to the state-of-the-art, demonstrating highly accurate performance, in some cases even outperforming competitors, and simultaneously being much more resource-friendly.
Kirsten Köbschall, Sebastian Buschjäger, Raphael Fischer 0001, Lisa Hartung, Stefan Kramer 0001
Data Min. Knowl. Discov.2
2026 STRATA2.0: A Serverless Middleware for Machine Learning Training
abstract
Serverless computing has gained increasing interest in recent years for enabling large-scale machine learning tasks. However, training a machine learning model in a serverless setting is a complex task and several challenges need to be addressed particularly in data distribution, result aggregation, resource heterogeneity, failures, container ephemerality, network and execution cost. These difficulties stem from the inherent complexity of distributed computation and the coordination demands of the machine learning algorithms. We propose STRATA2.0, a serverless middleware for Machine Learning training in serverless environments. STRATA2.0 provides a comprehensive suite of mechanisms designed to support efficient training of machine learning models on serverless infrastructures and address key challenges related to efficient data communication, coordination and synchronization, scalable and time efficient training of ML models using heterogeneous containers. Our extensive experimental results demonstrate that STRATA2.0 achieves the same level of accuracy with fewer data points, is on average three times faster in training time compared to centralized approaches, reduces energy consumption by up to 50%, and remains resilient when up to 60% of training instances fail.
Dimitrios Tomaras, Sebastian Buschjäger, Vana Kalogeraki, Katharina Morik, Dimitrios Gunopulos
IEEE Trans. Parallel Distributed Syst.2
2025 TrackScorer: Skyrmion Logic-in-Memory Accelerator for Tree-Based Ranking Models
abstract
Racetrack memories (RTMs) have been shown to have lower leakage power and higher density compared to traditional DRAM/SRAM technologies. However, their efficiency is often hindered by the need to shift the targeted data to access ports for read and write operations. Suitable mapping approaches are therefore essential to unleash their potential. In this work, we explore the mapping of the popular tree-based document ranking algorithm, Quickscorer, onto Skyrmion-based racetrack memories (SK-RTMs). Our approach leverages a Logic-in-Memory (LiM) accelerator, specifically designed to execute simple logic operations directly within SK-RTMs, enabling an efficient mapping of Quickscorer by exploiting its bitvector representation and inter-leaved traversal scheme of tree structures through bitwise logical operations. We present several mapping strategies, including one based on a quadratic assignment problem (QAP) optimization algorithm for optimal data placement of Quickscorer onto the racetracks. Our results demonstrate a significant reduction in read and write operations and, in certain cases, a decrease in the time spent shifting data during Quickscorer inference.
Elijah Cishugi, Sebastian Buschjäger, Martijn Noorlander, Marco Ottavi, Kuan-Hsun Chen
DATE2
2025 Splitting stump forests: tree ensemble compression for edge devices (extended version)
abstract
Abstract We introduce Splitting Stump Forests—small ensembles of weak learners extracted from a trained random forest. The high memory consumption of random forests renders them unfit for resource-constrained devices. We show empirically that we can significantly reduce the model size and inference time by selecting nodes that evenly split the arriving training data and applying a linear model on the resulting representation. Our extensive empirical evaluation indicates that Splitting Stump Forests outperform random forests and state-of-the-art compression methods on memory-limited embedded devices.
Fouad Alkhoury, Sebastian Buschjäger, Pascal Welke
Mach. Learn.2
2024 Federated Time Series Classification with ROCKET features
abstract
This paper proposes FROCKS, a federated time series classification method using ROCKET features.Our approach dynamically adapts the models' features by selecting and exchanging the bestperforming ROCKET kernels from a federation of clients.Specifically, the server gathers the best-performing kernels of the clients together with the associated model parameters, and it performs a weighted average if a kernel is best-performing for more than one client.We compare the proposed method with state-of-the-art approaches on the UCR archive binary classification datasets and show superior performance on most datasets.
Bruno Casella, Matthias Jakobs, Marco Aldinucci, Sebastian Buschjäger
ESANN4
2024 Stress-Testing USB Accelerators for Efficient Edge Inference
abstract
Several manufacturers sell specialized USB devices for accelerating machine learning (ML) on the edge. While being generally promoted as a versatile solution for more efficient edge inference with deep learning models, extensive practical insights on their usability and performance are hard to find. In order to make ML deployment on the edge more sustainable, our work investigates how operational and resource-efficient these USB accelerators really are. For that, we first introduce a novel and theoretically motivated methodology. It allows for comparing intricate model performance in terms of quality and resource consumption across different execution environments. We then put it into practice by studying the usability and efficiency of Google's Coral edge tensor processing unit (TPU) and Intel's neural compute stick 2 (NCS). In total, we benchmark over 30 models across nine hardware configurations, which reveals significant trade-offs. Our work demonstrates that USB accelerators are indeed capable of reducing the energy consumption by a factor up to ten, however this improvement cannot be observed for all configurations - more than 50% of the investigated models cannot be run on accelerator hardware and in several other cases, the power draw is only marginally improved. Our experiments show that the Intel NCS improves efficiency in a more stable way, while the TPU shows further benefits in specific cases but performs less predictably. We conclude that the tested USB accelerators are only beneficial for deploying standardized deep learning models more efficiently and cannot live up to the claims of being standardized out-of-the-box solutions.
Alexander van der Staay, Raphael Fischer 0001, Sebastian Buschjäger
SEC3
2024 Language-Based Deployment Optimization for Random Forests (Invited Paper)
abstract
Arising popularity for resource-efficient machine learning models makes random forests and decision trees famous models in recent years. Naturally, these models are tuned, optimized, and transformed to feature maximally low-resource consumption. A subset of these strategies targets the model structure and model logic and therefore induces a trade-off between resource-efficiency and prediction performance. An orthogonal set of approaches targets hardware-specific optimizations, which can improve performance without changing the behavior of the model. Since such hardware-specific optimizations are usually hardware-dependent and inflexible in their realizations, this paper envisions a more general application of such optimization strategies at the level of programming languages. We therefore discuss a set of suitable optimization strategies first in general and envision their application in LLVM IR, i.e. a flexible and hardware-independent ecosystem.
Jannik Malcher, Daniel Biebert, Kuan-Hsun Chen, Sebastian Buschjäger, Christian Hakert, Jian-Jia Chen
LCTES4
2024 STRATA: Random Forests going Serverless
abstract
Serverless computing has received growing interest in recent years for supporting large-scale machine learning tasks. However, training a machine learning model in a serverless environment is a nontrivial procedure and several challenges still need to be addressed in the data distribution and result aggregation steps as well as the cost of execution due to the inherent complexity of the distributed computation and the coordination required in the learning algorithm. In this work, we focus on Random Forests, a state-of-the-art technique in many Machine Learning applications. We propose STRATA, a cost-effective middleware to train Random Forests atop a serverless environment that successfully addresses these training challenges. As we show in our extensive experimental evaluation STRATA achieves 3X better training times on average compared to a centralized approach and can withstand up to 70% of failures during training.
Dimitrios Tomaras, Sebastian Buschjäger, Vana Kalogeraki, Katharina Morik, Dimitrios Gunopulos
Middleware2
2024 Rejection Ensembles with Online Calibration
Sebastian Buschjäger
ECML/PKDD (6)1
2024 MetaQuRe: Meta-learning from Model Quality and Resource Consumption
Raphael Fischer 0001, Marcel Wever, Sebastian Buschjäger, Thomas Liebig
ECML/PKDD (7)3
2023 Joint leaf-refinement and ensemble pruning through L1 regularization
abstract
Abstract Ensembles are among the state-of-the-art in many machine learning applications. With the ongoing integration of ML models into everyday life, e.g., in the form of the Internet of Things, the deployment and continuous application of models become more and more an important issue. Therefore, small models that offer good predictive performanceanduse small amounts of memory are required. Ensemble pruning is a standard technique for removing unnecessary classifiers from a large ensemble that reduces the overall resource consumption and sometimes improves the performance of the original ensemble. Similarly, leaf-refinement is a technique that improves the performance of a tree ensemble by jointly re-learning the probability estimates in the leaf nodes of the trees, thereby allowing for smaller ensembles while preserving their predictive performance. In this paper, we develop a new method that combines both approaches into a single algorithm. To do so, we introduce $$L_1$$ L1 regularization into the leaf-refinement objective, which allows us to jointly prune and refine trees at the same time. In an extensive experimental evaluation, we show that our approach not only offers statistically significantly better performance than the state-of-the-art but also offers a better accuracy-memory trade-off. We conclude our experimental evaluation with a case study showing the effectiveness of our method in a real-world setting.
Sebastian Buschjäger, Katharina Morik
Data Min. Knowl. Discov.1
2022 Shrub Ensembles for Online Classification
abstract
Online learning algorithms have become a ubiquitous tool in the machine learning toolbox and are frequently used in small, resource-constraint environments. Among the most successful online learning methods are Decision Tree (DT) ensembles. DT ensembles provide excellent performance while adapting to changes in the data, but they are not resource efficient. Incremental tree learners keep adding new nodes to the tree but never remove old ones increasing the memory consumption over time. Gradient-based tree learning, on the other hand, requires the computation of gradients over the entire tree which is costly for even moderately sized trees. In this paper, we propose a novel memory-efficient online classification ensemble called shrub ensembles for resource-constraint systems. Our algorithm trains small to medium-sized decision trees on small windows and uses stochastic proximal gradient descent to learn the ensemble weights of these `shrubs'. We provide a theoretical analysis of our algorithm and include an extensive discussion on the behavior of our approach in the online setting. In a series of 2~959 experiments on 12 different datasets, we compare our method against 8 state-of-the-art methods. Our Shrub Ensembles retain an excellent performance even when only little memory is available. We show that SE offers a better accuracy-memory trade-off in 7 of 12 cases, while having a statistically significant better performance than most other methods. Our implementation is available under https://github.com/sbuschjaeger/se-online .
Sebastian Buschjäger, Sibylle Hess, Katharina Morik
AAAI1
2022 FeFET-Based Binarized Neural Networks Under Temperature-Dependent Bit Errors
abstract
Ferroelectric FET (FeFET) is a highly promising emerging non-volatile memory (NVM) technology, especially for binarized neural network (BNN) inference on the low-power edge. The reliability of such devices, however, inherently depends on temperature. Hence, changes in temperature during run time manifest themselves as changes in bit error rates. In this work, we reveal the temperature-dependent bit error model of FeFET memories, evaluate its effect on BNN accuracy, and propose countermeasures. We begin on the transistor level and accurately model the impact of temperature on bit error rates of FeFET. This analysis reveals temperature-dependent asymmetric bit error rates. Afterwards, on the application level, we evaluate the impact of the temperature-dependent bit errors on the accuracy of BNNs. Under such bit errors, the BNN accuracy drops to unacceptable levels when no countermeasures are employed. We propose two countermeasures: (1) Training BNNs for bit error tolerance by injecting bit flips into the BNN data, and (2) applying a bit error rate assignment algorithm (BERA) which operates in a layer-wise manner and does not inject bit flips during training. In experiments, the BNNs, to which the countermeasures are applied to, effectively tolerate temperature-dependent bit errors for the entire range of operating temperature.
Mikail Yayla, Sebastian Buschjäger, Aniket Gupta, Jian-Jia Chen, Jörg Henkel, Katharina Morik, Kuan-Hsun Chen, Hussam Amrouch
IEEE Trans. Computers2
2022 Reliable Binarized Neural Networks on Unreliable Beyond Von-Neumann Architecture
abstract
Specialized hardware accelerators beyond von-Neumann, that offer processing capability in where the data resides without moving it, become inevitable in data-centric computing. Emerging non-volatile memories, like Ferroelectric Field-Effect Transistor (FeFET), are able to build compact Logic-in-Memory (LiM). In this work, we investigate the probability of error (Perror) in FeFET-based XNOR LiM, demonstrating the new trade-off between the speed and reliability. Using our reliability model, we present how Binarized Neural Networks (BNNs) can be proactively trained in the presence of XNOR-induced errors towards obtaining robust BNNs at the design time. Furthermore, leveraging the trade-off between Perror and speed, we present a run-time adaptation technique, that selectively trades-off Perror and XNOR speed for every BNN layer. Our results demonstrate that when a small loss (e.g., 1%) in inference accuracy could be accepted, our design-time and run-time techniques provide error-resilient BNNs that exhibit 75% and 50% (FashionMNIST) and 38% and 24% (CIFAR10) XNOR speedups, respectively.
Mikail Yayla, Simon Thomann, Sebastian Buschjäger, Katharina Morik, Jian-Jia Chen, Hussam Amrouch
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Efficient Realization of Decision Trees for Real-Time Inference
abstract
For timing-sensitive edge applications, the demand for efficient lightweight machine learning solutions has increased recently. Tree ensembles are among the state-of-the-art in many machine learning applications. While single decision trees are comparably small, an ensemble of trees can have a significant memory footprint leading to cache locality issues, which are crucial to performance in terms of execution time. In this work, we analyze memory-locality issues of the two most common realizations of decision trees, i.e., native and if-else trees. We highlight that both realizations demand a more careful memory layout to improve caching behavior and maximize performance. We adopt a probabilistic model of decision tree inference to find the best memory layout for each tree at the application layer. Further, we present an efficient heuristic to take architecture-dependent information into account thereby optimizing the given ensemble for a target computer architecture. Our code-generation framework, which is freely available on an open-source repository, produces optimized code sessions while preserving the structure and accuracy of the trees. With several real-world data sets, we evaluate the elapsed time of various tree realizations on server hardware as well as embedded systems for Intel and ARM processors. Our optimized memory layout achieves a reduction in execution time up to 75 % execution for server-class systems, and up to 70 % for embedded systems, respectively.
Kuan-Hsun Chen, Chiahui Su, Christian Hakert, Sebastian Buschjäger, Chao-Lin Lee, Jenq Kuen Lee, Katharina Morik, Jian-Jia Chen
ACM Trans. Embed. Comput. Syst.4
2021 Margin-Maximization in Binarized Neural Networks for Optimizing Bit Error Tolerance
abstract
To overcome the memory wall in neural network (NN) inference systems, recent studies have proposed to use approximate memory, in which the supply voltage and access latency parameters are tuned, for lower energy consumption and faster access at the cost of reliability. To tolerate the occuring bit errors, the state-of-the-art approaches apply bit flip injections to the NNs during training, which require high overheads and do not scale well for large NNs and high bit error rates. In this work, we focus on binarized NNs (BNNs), whose simpler structure allows better exploration of bit error tolerance metrics based on margins. We provide formal proofs to quantify the maximum number of bit flips that can be tolerated. With the proposed margin-based metrics and the well-known hinge loss for maximum margin classification in support vector machines (SVMs), we are able to construct a modified hinge loss (MHL) to train BNNs for bit error tolerance without any bit flip injections. Our experimental results indicate that the MHL enables the possibility for BNNs to tolerate higher bit error rates than with bit flip training and, therefore, allows to further lower the requirements on approximate memories used for BNNs.
Sebastian Buschjäger, Jian-Jia Chen, Kuan-Hsun Chen, Mario Günzel, Christian Hakert, Katharina Morik, Rodion Novkin, Lukas Pfahler, Mikail Yayla
DATE1
2021 Very Fast Streaming Submodular Function Maximization
Sebastian Buschjäger, Philipp-Jan Honysz, Lukas Pfahler, Katharina Morik
ECML/PKDD (3)1
2020 Generalized Isolation Forest: Some Theory and More Applications Extended Abstract
abstract
Isolation Forest is a popular outlier detection algorithm that isolates outlier observations from regular observations by building multiple random decision trees. Multiple extensions enhance the original Isolation Forest algorithm including the Extended Isolation Forest which allows for non-rectangular splits and the SCiForest which improves the fitting of individual trees. All these approaches rate the outlierness of an observation by its average path-length. However, we find a lack of theoretical explanation on why these isolation-based algorithms offer such good practical performance. In this paper, we present a theoretical framework that describes the effectiveness of isolation-based approaches from a distributional viewpoint. We show that these algorithms fit a mixture of distributions, where the average path length of an observation can be viewed as a (somewhat crude) approximation of the mixture coefficient. Using this framework, we derive the Generalized Isolation Forest (GIF) which also trains random trees, but combining them moves beyond using the average path-length. In an extensive evaluation of over 350, 000 experiments, we show that GIF outperforms the other methods on a variety of datasets while having comparable runtime.
Sebastian Buschjäger, Philipp-Jan Honysz, Katharina Morik
DSAA1
2020 On-Site Gamma-Hadron Separation with Deep Learning on FPGAs
Sebastian Buschjäger, Lukas Pfahler, Jens Buß, Katharina Morik, Wolfgang Rhode
ECML/PKDD (4)1
2019 Gaussian Model Trees for Traffic Imputation
abstract
Traffic congestion is one of the most pressing issues for smart cities. Information on traffic flow can be used to reduce congestion by predicting vehicle counts at unmonitored locations so that counter-measures can be applied before congestion appears. To do so pricy sensors must be distributed sparsely in the city and at important roads in the city center to collect road and vehicle information throughout the city in real-time. Then, Machine Learning models can be applied to predict vehicle counts at unmonitored locations. To be fault-tolerant and increase coverage of the traffic predictions to the suburbs, rural regions, or even neighboring villages, these Machine Learning models should not operate at a central traffic control room but rather be distributed across the city. Gaussian Processes (GP) work well in the context of traffic count prediction, but cannot capitalize on the vast amount of data available in an entire city. Furthermore, Gaussian Processes are a global and centralized model, which requires all measurements to be available at a central computation node. Product of Expert (PoE) models have been proposed as a scalable alternative to Gaussian Processes. A PoE model trains multiple, independent GPs on different subsets of the data and weight individual predictions based on each experts uncertainty. These methods work well, but they assume that experts are independent even though they may share data points. Furthermore, PoE models require exhaustive communication bandwidth between the individual experts to form the final prediction. In this paper we propose a hierarchical Product of Expert model, which consist of multiple layers of small, independent and local GP experts. We view Gaussian Process induction as regularized optimization procedure and utilize this view to derive an efficient algorithm which selects independent regions of the data. Then, we train local expert models on these regions, so that each expert is responsible for a given region. The resulting algorithm scales well for large amounts of data and outperforms flat PoE models in terms of communication cost, model size and predictive performance. Last, we discuss how to deploy these local expert models onto small devices.
Sebastian Buschjäger, Thomas Liebig, Katharina Morik
ICPRAM1
2018 Realization of Random Forest for Real-Time Evaluation through Tree Framing
abstract
The optimization of learning has always been of particular concern for big data analytics. However, the ongoing integration of machine learning models into everyday life also demand the evaluation to be extremely fast and in real-time. Moreover, in the Internet of Things, the computing facilities that run the learned model are restricted. Hence, the implementation of the model application must take the characteristics of the executing platform into account Although there exist some heuristics that optimize the code, principled approaches for fast execution of learned models are rare. In this paper, we introduce a method that optimizes the execution of Decision Trees (DT). Decision Trees form the basis of many ensemble methods, such as Random Forests (RF) or Extremely Randomized Trees (ET). For these methods to work best, trees should be as large as possible. This challenges the data and the instruction cache of modern CPUs and thus demand a more careful memory layout. Based on a probabilistic view of decision tree execution, we optimize the two most common implementation schemes of decision trees. We discuss the advantages and disadvantages of both implementations and present a theoretically well-founded memory layout which maximizes locality during execution in both cases. The method is applied to three computer architectures, namely ARM (RISC), PPC (Extended RISC) and Intel (CISC) and is automatically adopted to the specific architecture by a code generator. We perform over 1800 experiments on several real-world data sets and report an average speed-up of 2 to 4 across all three architectures by using the proposed memory layout. Moreover, we find that our implementation outperforms sklearn, which was used to train the models by a factor of 1500.
Sebastian Buschjäger, Kuan-Hsun Chen, Jian-Jia Chen, Katharina Morik
ICDM1