Thomas Bäck

dblp:b/ThomasBack · also Thomas H. W. Bäck · DBLP profile ↗
← Back
19ranked-venue papers in the field
1as first author
5since 2021 · last 2025
0000-0001-6768-1478ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 8Big Data, Cloud & Distributed Data Systems · 4Knowledge Engineering, Semantic Web & Information Systems · 4 (1 first)Other / Interdisciplinary · 3
YearPublicationVenuePosition
2025 Gradient Free Multi-Objective Counterfactual Explainability for Multivariate Time Series Classification
abstract
NWO
Sofoklis Kitharidis, Furong Ye, Marius Ottolini, Thomas Bäck, Niki van Stein
IEEE Big Data5
2024 Optimizing Causal Interventions in Hybrid Bayesian Networks - A Discretization, Knowledge Compilation, and Heuristic Optimization Approach
Maarten C. Vonk, Diederick Vermetten, Jacob de Nobel, Sebastiaan Brand, Ninoslav Malekovic, Thomas Bäck, Alfons Laarman, Anna V. Kononova
IPMU (1)6
2022 Robust subgroup discovery
abstract
Abstract We introduce the problem ofrobust subgroup discovery, i.e., finding a set of interpretable descriptions of subsets that 1) stand out with respect to one or more target attributes, 2) are statistically robust, and 3) non-redundant. Many attempts have been made to mine eitherlocallyrobust subgroups or to tackle the pattern explosion, but we are the first to address both challenges at the same time from aglobalmodelling perspective. First, we formulate the broad model class of subgroup lists, i.e., ordered sets of subgroups, for univariate and multivariate targets that can consist of nominal or numeric variables, including traditional top-1 subgroup discovery in its definition. This novel model class allows us to formalise the problem of optimal robust subgroup discovery using the Minimum Description Length (MDL) principle, where we resort to optimal Normalised Maximum Likelihood and Bayesian encodings for nominal and numeric targets, respectively. Second, finding optimal subgroup lists is NP-hard. Therefore, we propose SSD++, a greedy heuristic that finds good subgroup lists and guarantees that the most significant subgroup found according to the MDL criterion is added in each iteration. In fact, the greedy gain is shown to be equivalent to a Bayesian one-sample proportion, multinomial, or t-test between the subgroup and dataset marginal target distributions plus a multiple hypothesis testing penalty. Furthermore, we empirically show on 54 datasets that SSD++ outperforms previous subgroup discovery methods in terms of quality, generalisation on unseen data, and subgroup list size.
Hugo Manuel Proença, Peter Grünwald, Thomas Bäck, Matthijs van Leeuwen
Data Min. Knowl. Discov.3
2021 Improved Automated CASH Optimization with Tree Parzen Estimators for Class Imbalance Problems
abstract
The imbalanced classification problem is very relevant in both academic and industrial applications. The task of finding the best machine learning model to use for a specific imbalanced dataset is complicated due to a large number of existing algorithms, each with its own hyperparameters. The Combined Algorithm Selection and Hyperparameter optimization (CASH) has been introduced to tackle both aspects at the same time. However, CASH has not been studied in detail in the class imbalance domain, where the best combination of resampling technique and classification algorithm is searched for, together with their optimized hyperparameters. Thus, we target the CASH problem for imbalanced classification. We experiment with a search space of 5 classification algorithms, 21 resampling approaches and 64 relevant hyperparameters in total. Moreover, we investigate performance of 2 well-known optimization approaches: Random search and Tree Parzen Estimators approach which is a kind of Bayesian optimization. For comparison, we also perform grid search on all combinations of resampling techniques and classification algorithms with their default hyperparameters. Our experimental results show that a Bayesian optimization approach outperforms the other approaches for CASH in this application domain.
Jiawen Kong, Hao Wang 0025, Stefan Menzel, Bernhard Sendhoff, Anna V. Kononova, Thomas Bäck
DSAA7
2021 Differential evolution outside the box
Anna V. Kononova, Fabio Caraffini, Thomas Bäck
Inf. Sci.3
2020 Automated Machine Learning for the Classification of Normal and Abnormal Electromyography Data
abstract
Needle electromyography (EMG) is a common technique used in clinical neurophysiology to record the electrical activity of muscles at different levels of activation. It can be used to diagnose various neurological/muscular disorders, as the EMG signals of patients with both nerve diseases (neuropathies) and muscle diseases (myopathies) differ from the signal in healthy controls. A major drawback of this examination is that it relies on visual inspection and as such, it is highly subjective and prone to errors. Based on EMG time series of 65 individuals (40 with ALS/IBM and 25 healthy), we aim to develop an automated machine-learning pipeline for the classification of EMG recordings of muscles in either disease or healthy (muscle-level). The automated pipeline consists of feature extraction, feature selection, modelling algorithm, and optimization, in which the most significant features are automatically selected from the feature space and the hyperparameters of the model are optimized by a Bayesian technique as part of the automated approach. Aside from the muscle-level approach, we also explore a patient-level approach, which uses the output of the muscle-level automated pipeline in a post-processing manner to classify patients in being either disease or healthy, based on their muscle recordings. The resulting two approaches yield an AUC score of 81.7% (muscle-level) and 81.5% (patient-level), indicating that such approaches can assist clinicians in diagnosing if a patient has a neuropathy/myopathy or is healthy.
Marios Kefalas, Milan Koch, Victor Geraedts, Hao Wang 0025, Martijn Tannemaat, Thomas Bäck
IEEE BigData6
2020 On the Performance of Oversampling Techniques for Class Imbalance Problems
Jiawen Kong, Thiago Rios, Wojtek Kowalczyk, Stefan Menzel, Thomas Bäck
PAKDD (2)5
2020 Discovering Outstanding Subgroup Lists for Numeric Targets Using MDL
Hugo Manuel Proença, Peter Grünwald, Thomas Bäck, Matthijs van Leeuwen
ECML/PKDD (1)3
2019 Automated Machine Learning for EEG-Based Classification of Parkinson's Disease Patients
abstract
The treatment of Parkinson’s Disease (PD) with Deep Brain Stimulation (DBS) can provide a constant level of motor functioning. Several patients, however, may suffer from postoperative cognitive deterioration. The DBS screening therefore includes an assessment of cognitive functioning prior to DBS surgery. However, these assessments may be influenced by factors such as fatigue or motivation and there is a need for novel biomarkers of cognitive dysfunction to complement the DBS screening. Electroencephalography (EEG) has been previously suggested to identify potential cognitive impairment in PD patients and may have utility during the DBS screening. A limited set of biomarkers (features) from the EEG has been identified for this purpose. Finding new biomarkers is time-consuming and there is no driving hypothesis on which new biomarkers may be important. Based on EEG time series of 40 DBS candidates, this research focuses on automated machine learning techniques to develop EEG-based algorithms for the evaluation of the cognitive function of PD patients. The automated pipeline consists of feature extraction, feature selection, modelling algorithm and optimization. With this approach we extract 794 features from each of the 21 EEG channels which results in a massive feature space. From this feature space the most significant features are selected and used for modelling. The hyperparameters of the model are optimized by a Bayesian technique as part of the automated approach. Aside from the automatically computed features, we also explore the use of features commonly used during clinical evaluation of the EEG, with the result that the model based on automatically computed features achieves a significant higher accuracy (84.0%). The newly identified features are potentially new biomarkers. We used the knowledge gathered from our automated approach to build a hand-crafted model resulting in an accuracy of 91.0%.
Milan Koch, Victor Geraedts, Hao Wang 0025, Martijn Tannemaat, Thomas Bäck
IEEE BigData5
2018 A Novel Uncertainty Quantification Method for Efficient Global Optimization
Niki van Stein, Hao Wang 0025, Wojtek Kowalczyk, Thomas Bäck
IPMU (3)4
2017 Corrigendum to 'Multiobjective optimization of classifiers by means of 3D convex-hull-based evolutionary algorithms' [Information Sciences volumes 367-368 (2016) 80-104]
Jiaqi Zhao 0001, Vitor Basto-Fernandes, Licheng Jiao, Iryna Yevseyeva, Asep Maulana, Rui Li 0001, Thomas Bäck, Ke Tang 0001, Michael T. M. Emmerich
Inf. Sci.7
2016 Local subspace-based outlier detection using global neighbourhoods
abstract
Outlier detection in high-dimensional data is a challenging yet important task, as it has applications in, e.g., fraud detection and quality control. State-of-the-art density-based algorithms perform well because they 1) take the local neighbourhoods of data points into account and 2) consider feature subspaces. In highly complex and high-dimensional data, however, existing methods are likely to overlook important outliers because they do not explicitly take into account that the data is often a mixture distribution of multiple components. We therefore introduce GLOSS, an algorithm that performs local subspace outlier detection using global neighbourhoods. Experiments on synthetic data demonstrate that GLOSS more accurately detects local outliers in mixed data than its competitors. Moreover, experiments on real-world data show that our approach identifies relevant outliers overlooked by existing methods, confirming that one should keep an eye on the global perspective even when doing local outlier detection.
Niki van Stein, Matthijs van Leeuwen, Thomas Bäck
IEEE BigData3
2016 Analysis and Visualization of Missing Value Patterns
Niki van Stein, Wojtek Kowalczyk, Thomas Bäck
IPMU (2)3
2016 Multiobjective optimization of classifiers by means of 3D convex-hull-based evolutionary algorithms
Jiaqi Zhao 0001, Vitor Basto-Fernandes, Licheng Jiao, Iryna Yevseyeva, Asep Maulana, Rui Li 0001, Thomas Bäck, Ke Tang 0001, Michael T. M. Emmerich
Inf. Sci.7
2015 Optimally Weighted Cluster Kriging for Big Data Regression
Niki van Stein, Hao Wang 0025, Wojtek Kowalczyk, Thomas Bäck, Michael T. M. Emmerich
IDA4
2005 Reliable Hierarchical Clustering with the Self-organizing Map
Elena V. Samsonova, Thomas Bäck, Joost N. Kok, Adriaan P. IJzerman
IDA2
2003 Combining and Comparing Cluster Methods in a Receptor Database
Elena V. Samsonova, Thomas Bäck, Margot W. Beukers, Adriaan P. IJzerman, Joost N. Kok
IDA2
2002 Adaptive business intelligence based on evolution strategies: some application examples of self-adaptive software
Thomas Bäck
Inf. Sci.1
1993 An Overview of Evolutionary Computation
William M. Spears, Kenneth A. De Jong, Thomas Bäck, David B. Fogel, Hugo de Garis
ECML3