Martin Wistuba

dblp:122/8131 · DBLP profile ↗
← Back
37ranked-venue papers
15as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 13 first-author · 11 since 2021Databases, data management, data science and information retrieval · 19 · 11 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Theory of computation · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Hyperband-based Bayesian Optimization for Black-box Prompt Selection
abstract
Optimal prompt selection is crucial for maximizing large language model (LLM) performance on downstream tasks, especially in black-box settings where models are only accessible via APIs. Black-box prompt selection is challenging due to potentially large, combinatorial search spaces, absence of gradient information, and high evaluation cost of prompts on a validation set. We propose HbBoPs, a novel method that combines a structural-aware deep kernel Gaussian Process with Hyperband as a multi-fidelity scheduler to efficiently select prompts. HbBoPs uses embeddings of instructions and few-shot exemplars, treating them as modular components within prompts. This enhances the surrogate model’s ability to predict which prompt to evaluate next in a sample-efficient manner. Hyperband improves query-efficiency by adaptively allocating resources across different fidelity levels, reducing the number of validation instances required for evaluating prompts. Extensive experiments across ten diverse benchmarks and three LLMs demonstrate that HbBoPs outperforms state-of-the-art methods in both performance and efficiency.
Lennart Schneider, Martin Wistuba, Aaron Klein, Jacek Golebiowski, Giovanni Zappella, Felice Antonio Merra
ICML2
2023 Variational Boosted Soft Trees
abstract
Gradient boosting machines (GBMs) based on decision trees consistently demonstrate state-of-the-art results on regression and classification tasks with tabular data, often outperforming deep neural networks. However, these models do not provide well-calibrated predictive uncertainties, which prevents their use for decision making in high-risk applications. The Bayesian treatment is known to improve predictive uncertainty calibration, but previously proposed Bayesian GBM methods are either computationally expensive, or resort to crude approximations. Variational inference is often used to implement Bayesian neural networks, but is difficult to apply to GBMs, because the decision trees used as weak learners are non-differentiable. In this paper, we propose to implement Bayesian GBMs using variational inference with soft decision trees, a fully differentiable alternative to standard decision trees introduced by Irsoy et al. Our experiments demonstrate that variational soft trees and variational soft GBMs provide useful uncertainty estimates, while retaining good predictive performance. The proposed models show higher test likelihoods when compared to the state-of-the-art Bayesian GBMs in 7/10 tabular regression datasets and improved out-of-distribution detection in 5/10 datasets.
Tristan Cinquin, Tammo Rukat, Philipp Schmidt 0002, Martin Wistuba, Artur Bekasov
AISTATS4
2023 PASHA: Efficient HPO and NAS with Progressive Resource Allocation
Ondrej Bohdal, Lukas Balles, Martin Wistuba, Beyza Ermis, Cédric Archambeau, Giovanni Zappella
ICLR3
2023 Scaling Laws for Hyperparameter Optimization
abstract
Hyperparameter optimization is an important subfield of machine learning that focuses on tuning the hyperparameters of a chosen algorithm to achieve peak performance. Recently, there has been a stream of methods that tackle the issue of hyperparameter optimization, however, most of the methods do not exploit the dominant power law nature of learning curves for Bayesian optimization. In this work, we propose Deep Power Laws (DPL), an ensemble of neural network models conditioned to yield predictions that follow a power-law scaling pattern. Our method dynamically decides which configurations to pause and train incrementally by making use of gray-box evaluations. We compare our method against 7 state-of-the-art competitors on 3 benchmarks related to tabular, image, and NLP datasets covering 59 diverse tasks. Our method achieves the best results across all benchmarks by obtaining the best any-time results compared to all competitors.
Arlind Kadra, Maciej Janowski, Martin Wistuba, Josif Grabocka
NeurIPS3
2022 Bandit Limited Discrepancy Search and Application to Machine Learning Pipeline Optimization
abstract
Optimizing a machine learning (ML) pipeline has been an important topic of AI and ML. Despite recent progress, pipeline optimization remains a challenging problem, due to potentially many combinations to consider as well as slow training and validation. We present the BLDS algorithm for optimized algorithm selection (ML operations) in a fixed ML pipeline structure. BLDS performs multi-fidelity optimization for selecting ML algorithms trained with smaller computational overhead, while controlling its pipeline search based on multi-armed bandit and limited discrepancy search. Our experiments on well-known classification benchmarks show that BLDS is superior to competing algorithms. We also combine BLDS with hyperparameter optimization, empirically showing the advantage of BLDS.
Akihiro Kishimoto, Djallel Bouneffouf 0001, Radu Marinescu 0002, Parikshit Ram, Ambrish Rawat, Martin Wistuba, Paulito P. Palmes, Adi Botea
AAAI6
2022 Memory Efficient Continual Learning with Transformers
abstract
In many real-world scenarios, data to train machine learning models becomes available over time. Unfortunately, these models struggle to continually learn new concepts without forgetting what has been learnt in the past. This phenomenon is known as catastrophic forgetting and it is difficult to prevent due to practical constraints. For instance, the amount of data that can be stored or the computational resources that can be used might be limited. Moreover, applications increasingly rely on large pre-trained neural networks, such as pre-trained Transformers, since compute or data might not be available in sufficiently large quantities to practitioners to train from scratch. In this paper, we devise a method to incrementally train a model on a sequence of tasks using pre-trained Transformers and extending them with Adapters. Different than the existing approaches, our method is able to scale to a large number of tasks without significant overhead and allows sharing information across tasks. On both image and text classification tasks, we empirically demonstrate that our method maintains a good predictive performance without retraining the model or increasing the number of model parameters over time. The resulting model is also significantly faster at inference time compared to Adapter-based state-of-the-art methods.
Beyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal, Cédric Archambeau
NeurIPS3
2022 Supervising the Multi-Fidelity Race of Hyperparameter Configurations
abstract
Multi-fidelity (gray-box) hyperparameter optimization techniques (HPO) have recently emerged as a promising direction for tuning Deep Learning methods. However, existing methods suffer from a sub-optimal allocation of the HPO budget to the hyperparameter configurations. In this work, we introduce DyHPO, a Bayesian Optimization method that learns to decide which hyperparameter configuration to train further in a dynamic race among all feasible configurations. We propose a new deep kernel for Gaussian Processes that embeds the learning curve dynamics, and an acquisition function that incorporates multi-budget information. We demonstrate the significant superiority of DyHPO against state-of-the-art hyperparameter optimization methods through large-scale experiments comprising 50 datasets (Tabular, Image, NLP) and diverse architectures (MLP, CNN/NAS, RNN).
Martin Wistuba, Arlind Kadra, Josif Grabocka
NeurIPS1
2021 Searching for Machine Learning Pipelines Using a Context-Free Grammar
abstract
AutoML automatically selects, composes and parameterizes machine learning algorithms into a workflow or pipeline of operations that aims at maximizing performance on a given dataset. Although current methods for AutoML achieved impressive results they mostly concentrate on optimizing fixed linear workflows. In this paper, we take a different approach and focus on generating and optimizing pipelines of complex directed acyclic graph shapes. These complex pipeline structure may lead to discovering hidden features and thus boost performance considerably. We explore the power of heuristic search and context-free grammars to search and optimize these kinds of pipelines. Experiments on various benchmark datasets show that our approach is highly competitive and often outperforms existing AutoML systems.
Radu Marinescu 0002, Akihiro Kishimoto, Parikshit Ram, Ambrish Rawat, Martin Wistuba, Paulito P. Palmes, Adi Botea
AAAI5
2021 AutoText: An End-to-End AutoAI Framework for Text
abstract
Building models for natural language processing (NLP) tasks remains a daunting task for many, requiring significant technical expertise, efforts, and resources. In this demonstration, we present AutoText, an end-to-end AutoAI framework for text, to lower the barrier of entry in building NLP models. AutoText combines state-of-the-art AutoAI optimization techniques and learning algorithms for NLP tasks into a single extensible framework. Through its simple, yet powerful UI, non-AI experts (e.g., domain experts) can quickly generate performant NLP models with support to both control (e.g., via specifying constraints) and understand learned models.
Arunima Chaudhary, Alayt Issak, Kiran Kate, Yannis Katsis, Abel N. Valente, Dakuo Wang, Alexandre V. Evfimievski, Sairam Gurajada, Ban Kawas, Cristiano Malossi, Lucian Popa 0001, Tejaswini Pedapati, Horst Samulowitz, Martin Wistuba, Yunyao Li 0001
AAAI14
2021 Automated Data Science for Relational Data
abstract
Feature engineering is a crucial but tedious task that requires up to 80% of the total time in data science projects. A significant challenge is when data consists of tables from different data sources, thus data scientists need to wisely aggregate and join tables while performing feature engineering task. In this work, we demonstrate a novel system called OneBM (One Button Machine), that enables data scientists to increase their efficiency with automated feature engineering for relational data. OneBM takes as input a relational dataset with multiple tables and its entity relation diagram (ERD) which can be declared with a novel, easy-to-use drag-and-drop graphical user interface. The system then automatically identifies and executes relevant joins and aggregates in the data, and generates new features with a rich set of transformations for various types of data including but not limited to time-series, sequences, number sets and itemsets, etc. The generated features then can be used by automated model selection and hyper-parameter optimization algorithms to complete a fully end-to-end automated data science (or AutoDS) workflow. A follow-up user evaluation illustrated how data scientists can perform multi-table feature engineering tasks in minutes using our system, compared to repeatedly coding SQL-like queries to transform and aggregate relational data requiring weeks of manual labor for comparable performance. In the live demos we plan to show two use cases with real-world datasets (video demos are available at the links in the footnote): sale prediction1and call center user experience2. Pre-registered partcipants can play with these use-cases and the given datasets via Watson Studio on the cloud.
Hoang Thanh Lam, Beat Buesser, Hong Min, Tran Ngoc Minh, Martin Wistuba, Udayan Khurana, Gregory Bramble, Theodoros Salonidis, Dakuo Wang, Horst Samulowitz
ICDE5
2021 Few-Shot Bayesian Optimization with Deep Kernel Surrogates
Martin Wistuba, Josif Grabocka
ICLR1
2021 Hardware-Aware Neural Architecture Search: Survey and Taxonomy
abstract
There is no doubt that making AI mainstream by bringing powerful, yet power hungry deep neural networks (DNNs) to resource-constrained devices would required an efficient co-design of algorithms, hardware and software. The increased popularity of DNN applications deployed on a wide variety of platforms, from tiny microcontrollers to data-centers, have resulted in multiple questions and challenges related to constraints introduced by the hardware. In this survey on hardware-aware neural architecture search (HW-NAS), we present some of the existing answers proposed in the literature for the following questions: "Is it possible to build an efficient DL model that meets the latency and energy constraints of tiny edge devices?", "How can we reduce the trade-off between the accuracy of a DL model and its ability to be deployed in a variety of platforms?". The survey provides a new taxonomy of HW-NAS and assesses the hardware cost estimation strategies. We also highlight the challenges and limitations of existing approaches and potential future directions. We hope that this survey will help to fuel the research towards efficient deep learning.
Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar, Martin Wistuba, Naigang Wang
IJCAI5
2020 Learning to Rank Learning Curves
abstract
Many automated machine learning methods, such as those for hyperparameter and neural architecture optimization, are computationally expensive because they involve training many different model configurations. In this work, we present a new method that saves computational budget by terminating poor configurations early on in the training. In contrast to existing methods, we consider this task as a ranking and transfer learning problem. We qualitatively show that by optimizing a pairwise ranking loss and leveraging learning curves from other data sets, our model is able to effectively rank learning curves without having to observe many or very long learning curves. We further demonstrate that our method can be used to accelerate a neural architecture search by a factor of up to 100 without a significant performance degradation of the discovered architecture. In further experiments we analyze the quality of ranking, the influence of different model components as well as the predictive behavior of the model.
Martin Wistuba, Tejaswini Pedapati
ICML1
2020 Survey on Automated End-to-End Data Science?
abstract
Data science is labor-intensive and human experts are scarce but heavily involved in every aspect of it. This makes data science time consuming and restricted to experts with the resulting quality heavily dependent on their experience and skills. To make data science more accessible and scalable, we need its democratization. Automated Data Science (AutoDS) is aimed towards that goal and is emerging as an important research and business topic. We introduce and define the AutoDS challenge, followed by a proposal of a general AutoDS framework that covers existing approaches but also provides guidance for the development of new methods. We categorize and review the existing literature from multiple aspects of the problem setup and employed techniques. Then we provide several views on how AI could succeed in automating end-to-end AutoDS. We hope this survey can serve as insightful guideline for the AutoDS field and provide inspiration for future research.
Djallel Bouneffouf 0001, Charu C. Aggarwal, Thanh Hoang, Udayan Khurana, Horst Samulowitz, Beat Buesser, Sijia Liu 0001, Tejaswini Pedapati, Parikshit Ram, Ambrish Rawat, Martin Wistuba, Alexander G. Gray
IJCNN11
2020 Automation of Deep Learning - Theory and Practice
abstract
The growing interest in both the automation of machine learning and deep learning has inevitably led to the development of a wide variety of methods to automate deep learning. The choice of network architecture has proven critical, and many improvements in deep learning are due to new structuring of it. However, deep learning techniques are computationally intensive and their use requires a high level of domain knowledge. Even a partial automation of this process therefore helps to make deep learning more accessible for everyone. In this tutorial we present a uniform formalism that enables different methods to be categorized and compare the different approaches in terms of their performance. We achieve this through a comprehensive discussion of the commonly used architecture search spaces and architecture optimization algorithms based on reinforcement learning and evolutionary algorithms as well as approaches that include surrogate and one-shot models. In addition, we discuss approaches to accelerate the search for neural architectures based on early termination and transfer learning and address the new research directions, which include constrained and multi-objective architecture search as well as the automated search for data augmentation, optimizers, and activation functions.
Martin Wistuba, Ambrish Rawat, Tejaswini Pedapati
ICMR1
2020 XferNAS: Transfer Neural Architecture Search
Martin Wistuba
ECML/PKDD (3)1
2019 Optimal Exploitation of Clustering and History Information in Multi-armed Bandit
abstract
We consider the stochastic multi-armed bandit problem and the contextual bandit problem with historical observations and pre-clustered arms. The historical observations can contain any number of instances for each arm, and the pre-clustering information is a fixed clustering of arms provided as part of the input. We develop a variety of algorithms which incorporate this offline information effectively during the online exploration phase and derive their regret bounds. In particular, we develop the META algorithm which effectively hedges between two other algorithms: one which uses both historical observations and clustering, and another which uses only the historical observations. The former outperforms the latter when the clustering quality is good, and vice-versa. Extensive experiments on synthetic and real world datasets on Warafin drug dosage and web server selectionfor latency minimization validate our theoretical insights and demonstrate that META is a robust strategy for optimally exploiting the pre-clustering information.
Djallel Bouneffouf 0001, Srinivasan Parthasarathy 0002, Horst Samulowitz, Martin Wistuba
IJCAI4
2019 Scalable Large Margin Gaussian Process Classification
Martin Wistuba, Ambrish Rawat
ECML/PKDD (2)1
2018 Practical Deep Learning Architecture Optimization
abstract
The design of neural network architectures for a new data set is a laborious task which requires human deep learning expertise. In order to make deep learning available for a broader audience, automated methods for finding a neural network architecture are vital. Recently proposed methods can already achieve human expert level performances. However, these methods have run times of months or even years of GPU computing time, ignoring hardware constraints as faced by many researchers and companies. We propose the use of Monte Carlo planning in combination with two different UCT (upper confidence bound applied to trees) derivations to search for network architectures. We adapt the UCT algorithm to the needs of network architecture search by proposing two ways of sharing information between different branches of the search tree. In an empirical study we are able to demonstrate that this method is able to find competitive networks for MNIST, SVHN and CIFAR-10 in just a single GPU day. Extending the search time to five GPU days, we are able to outperform man-made architectures and our competitors which consider the same types of layers.
Martin Wistuba
DSAA1
2018 Deep Learning Architecture Search by Neuro-Cell-Based Evolution with Function-Preserving Mutations
Martin Wistuba
ECML/PKDD (2)1
2018 Scalable Gaussian process-based transfer surrogates for hyperparameter optimization
Martin Wistuba, Nicolas Schilling, Lars Schmidt-Thieme
Mach. Learn.1
2017 Personalized Deep Learning for Tag Recommendation
Hanh T. H. Nguyen, Martin Wistuba, Josif Grabocka, Lucas Drumond, Lars Schmidt-Thieme
PAKDD (1)2
2017 Personalized Tag Recommendation for Images Using Deep Transfer Learning
Hanh T. H. Nguyen, Martin Wistuba, Lars Schmidt-Thieme
ECML/PKDD (2)2
2017 Automatic Frankensteining: Creating Complex Ensembles Autonomously
abstract
Automating machine learning by providing techniques that autonomously find the best algorithm, hyperparameter configuration and preprocessing is helpful for both researchers and practitioners. Therefore, it is not surprising that automated machine learning has become a very interesting field of research. While current research is mainly focusing on finding good pairs of algorithms and hyperparameter configurations, we will present an approach that automates the process of creating a top performing ensemble of several layers, different algorithms and hyperparameter configurations. These kinds of ensembles are called jokingly Frankenstein ensembles and proved their benefit on versatile data sets in many machine learning challenges. We compare our approach Automatic Frankensteining with the current state of the art for automated machine learning on 80 different data sets and can show that it outperforms them on the majority using the same training time. Furthermore, we compare Automatic Frankensteining on a large scale data set to more than 3,500 machine learning expert teams and are able to outperform more than 3,000 of them within 12 CPU hours.
Martin Wistuba, Nicolas Schilling, Lars Schmidt-Thieme
SDM1
2016 Hyperparameter Optimization Machines
abstract
Algorithm selection and hyperparameter tuning are omnipresent problems for researchers and practitioners. Hence, it is not surprising that the efforts in automatizing this process using various meta-learning approaches have been increased. Sequential model-based optimization (SMBO) is ne of the most popular frameworks for finding optimal hyperparameter configurations. Originally designed for black-box optimization, researchers have contributed different meta-learning approaches to speed up the optimization process. We create a generalized framework of SMBO and its recent additions which gives access to adaptive hyperparameter transfer learning with simple surrogates (AHT), a new class of hyperparameter optimization strategies. AHT provides less time-overhead for the optimization process by replacing time-and space-consuming transfer surrogate models with simple surrogates that employ adaptive transfer learning. In an empirical comparison on two different meta-data sets, we can show that AHT outperforms various instances of the SMBO framework in the scenarios of hyperparameter tuning and algorithm selection.
Martin Wistuba, Nicolas Schilling, Lars Schmidt-Thieme
DSAA1
2016 Scalable Hyperparameter Optimization with Products of Gaussian Process Experts
Nicolas Schilling, Martin Wistuba, Lars Schmidt-Thieme
ECML/PKDD (1)2
2016 Two-Stage Transfer Surrogate Model for Automatic Hyperparameter Optimization
Martin Wistuba, Nicolas Schilling, Lars Schmidt-Thieme
ECML/PKDD (1)1
2016 Fast classification of univariate and multivariate time series through shapelet discovery
Josif Grabocka, Martin Wistuba, Lars Schmidt-Thieme
Knowl. Inf. Syst.2
2015 Learning hyperparameter optimization initializations
abstract
Hyperparameter optimization is often done manually or by using a grid search. However, recent research has shown that automatic optimization techniques are able to accelerate this optimization process and find hyperparameter configurations that lead to better models. Currently, transferring knowledge from previous experiments to a new experiment is of particular interest because it has been shown that it allows to further improve the hyperparameter optimization. We propose to transfer knowledge by means of an initialization strategy for hyperparameter optimization. In contrast to the current state of the art initialization strategies, our strategy is neither limited to hyperparameter configurations that have been evaluated on previous experiments nor does it need meta-features. The initial hyperparameter configurations are derived by optimizing for a meta-loss formally defined in this paper. This loss depends on the hyperparameter response function of the data sets that were investigated in past experiments. Since this function is unknown and only few observations are given, the meta-loss is not differentiable. We propose to approximate the response function by a differentiable plug-in estimator. Then, we are able to learn the initial hyperparameter configuration sequence by applying gradient-based optimization techniques. Extensive experiments are conducted on two meta-data sets. Our initialization strategy is compared to the state of the art for initialization strategies and further methods that are able to transfer knowledge between data sets. We give empirical evidence that our work provides an improvement over the state of the art.
Martin Wistuba, Nicolas Schilling, Lars Schmidt-Thieme
DSAA1
2015 Sequential Model-Free Hyperparameter Tuning
abstract
Hyperparameter tuning is often done manually but current research has proven that automatic tuning yields effective hyperparameter configurations even faster and does not require any expertise. To further improve the search, recent publications propose transferring knowledge from previous experiments to new experiments. We adapt the sequential model-based optimization by replacing its surrogate model and acquisition function with one policy that is optimized for the task of hyperparameter tuning. This policy generalizes over previous experiments but neither uses a model nor uses meta-features, nevertheless, outperforms the state of the art. We show that a static ranking of hyperparameter combinations yields competitive results and substantially outperforms a random hyperparameter search. Thus, it is a fast and easy alternative to complex hyperparameter tuning strategies and allows practitioners to tune their hyperparameters by simply using a look-up table. We made look-up tables for two classifiers publicly available: SVM and AdaBoost. Furthermore, we propose a similarity measure for data sets that yields more comprehensible results than those using meta-features. We show how this similarity measure can be applied to surrogate models in the SMBO framework and empirically show that this change leads to better hyperparameter configurations in less trials.
Martin Wistuba, Nicolas Schilling, Lars Schmidt-Thieme
ICDM1
2015 Joint Model Choice and Hyperparameter Optimization with Factorized Multilayer Perceptrons
abstract
Recent work has demonstrated that hyperparameter optimization within the sequential model-based optimization (SMBO) framework is generally possible. This approach replaces the expensive-to-evaluate function that maps hyperparameters to the performance of a learned model on validation data by a surrogate model which is much cheaper to evaluate. The current state of the art in hyperparameter optimization learns these surrogate models across a variety of solved data sets where a grid search has already been employed. In this way, surrogate models are learned across data sets, and thus able to generalize better. However, meta features that describe characteristics of a data set are usually needed in order for the surrogate model to differentiate between same hyperparameter configurations on different data sets. Another research area that is closely related focuses on model choice, i.e. picking the right model for a given task, which is also a problem that many practitioners face in machine learning. In this paper, we aim to solve both of these problems with a unified surrogate model that learns across different data sets, different classifiers and their respective hyperparameters. We employ factorized multilayer perceptrons, a surrogate model that consists of a multilayer perceptron architecture, but offers the prediction of a factorization machine in the first layer. In this way, data sets, models and hyperparameters are being represented in a joint lower dimensional latent feature space. Experiments on a publicly available meta data set containing 59 individual data sets and 19 prediction models demonstrate the efficiency of our approach.
Nicolas Schilling, Martin Wistuba, Lucas Drumond, Lars Schmidt-Thieme
ICTAI2
2015 Hyperparameter Optimization with Factorized Multilayer Perceptrons
Nicolas Schilling, Martin Wistuba, Lucas Drumond, Lars Schmidt-Thieme
ECML/PKDD (2)2
2015 Hyperparameter Search Space Pruning - A New Component for Sequential Model-Based Hyperparameter Optimization
Martin Wistuba, Nicolas Schilling, Lars Schmidt-Thieme
ECML/PKDD (2)1
2015 Scalable Classification of Repetitive Time Series Through Frequencies of Local Polynomials
abstract
Time-series classification has attracted considerable research attention due to the various domains where time-series data are observed, ranging from medicine to econometrics. Traditionally, the focus of time-series classification has been on short time-series data composed of a few patterns exhibiting variabilities, while recently there have been attempts to focus on longer series composed of multiple local patrepeating with an arbitrary irregularity. The primary contribution of this paper relies on presenting a method which can detect local patterns in repetitive time-series via fitting local polynomial functions of a specified degree. We capture the repetitiveness degrees of time-series datasets via a new measure. Furthermore, our method approximates local polynomials in linear time and ensures an overall linear running time complexity. The coefficients of the polynomial functions are converted to symbolic words via equi-area discretizations of the coefficients' distributions. The symbolic polynomial words enable the detection of similar local patterns by assigning the same word to similar polynomials. Moreover, a histogram of the frequencies of the words is constructed from each time-series' bag of words. Each row of the histogram enables a new representation for the series and symbolizes the occurrence of local patterns and their frequencies. In an experimental comparison against state-of-the-art baselines on repetitive datasets, our method demonstrates significant improvements in terms of prediction accuracy.
Josif Grabocka, Martin Wistuba, Lars Schmidt-Thieme
IEEE Trans. Knowl. Data Eng.2
2014 Minimal Invasive Integration of Learning Analytics Services in Intelligent Tutoring Systems
abstract
A common problem when trying to apply data mining techniques to improve educational systems is the disconnection between those who have the expertise (e.g. Universities) and those who have access to the data (e.g. Small companies). Bringing expertise into educational in-production systems is complicated because companies are reluctant to invest a lot of effort into integrating new technology that they do not fully trust, while the technology cannot prove its worth without access to real, valid data. In this paper we explore the requirements that machine learning systems have to be applied to specific learning problems (sequencing and performance prediction), and then propose a minimally invasive protocol for sequencing (based on web services) to easily integrate Learning Analytics Services into e-learning systems.
Carlotta Schatten, Martin Wistuba, Lars Schmidt-Thieme, Sergio Gutiérrez Santos
ICALT2
2014 Learning time-series shapelets
abstract
Shapelets are discriminative sub-sequences of time series that best predict the target variable. For this reason, shapelet discovery has recently attracted considerable interest within the time-series research community. Currently shapelets are found by evaluating the prediction qualities of numerous candidates extracted from the series segments. In contrast to the state-of-the-art, this paper proposes a novel perspective in terms of learning shapelets. A new mathematical formalization of the task via a classification objective function is proposed and a tailored stochastic gradient learning algorithm is applied. The proposed method enables learning near-to-optimal shapelets directly without the need to try out lots of candidates. Furthermore, our method can learn true top-K shapelets by capturing their interaction. Extensive experimentation demonstrates statistically significant improvement in terms of wins and ranks against 13 baselines over 28 time-series datasets.
Josif Grabocka, Nicolas Schilling, Martin Wistuba, Lars Schmidt-Thieme
KDD3
2013 Factorized Decision Trees for Active Learning in Recommender Systems
abstract
A key challenge in recommender systems is how to profile new users. A well-known solution for this problem is to use active learning techniques and ask the new user to rate a few items to reveal her preferences. The sequence of queries should not be static, i.e in each step the best query depends on the responses of the new user to the previous queries. Decision trees have been proposed to capture the dynamic aspect of this process. In this paper we improve decision trees in two ways. First, we propose the Most Popular Sampling (MPS) method to increase the speed of the tree construction. In each node, instead of checking all candidate items, only those which are popular among users associated with the node are examined. Second, we develop a new algorithm to build decision trees. It is called Factorized Decision Trees (FDT) and exploits matrix factorization to predict the ratings at nodes of the tree. The experimental results on the Netflix dataset show that both contributions are successful. The MPS increases the speed of the tree construction without harming the accuracy. And FDT improves the accuracy of rating predictions especially in the last queries.
Rasoul Karimi, Martin Wistuba, Alexandros Nanopoulos, Lars Schmidt-Thieme
ICTAI2