EDBT 2026 Demo / reviewers in the wild / expert
Felix Bießmann
dblp:05/3961 · also Felix Biessmann
· DBLP profile ↗
19ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-3422-1026ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | From Data Imputation to Data Cleaning - Automated Cleaning of Tabular Data Improves Downstream Predictive PerformanceabstractThe translation of Machine Learning (ML) research innovations to real-world applications and the maintenance of ML components are hindered by reoccurring challenges, such as reaching high predictive performance, robustness, complying with regulatory constraints, or meeting ethical standards. Many of these challenges are related to data quality and, in particular, to the lack of automation in data pipelines upstream of ML components. Automated data cleaning remains challenging since many approaches neglect the dependency structure of the data errors and require task-specific heuristics or human input for calibration. In this study, we develop and evaluate an application-agnostic ML-based data cleaning approach using well-established imputation techniques for automated detection and cleaning of erroneous values. To improve the degree of automation, we combine imputation techniques with conformal prediction (CP), a model-agnostic and distribution-free method to quantify and calibrate the uncertainty of ML models. Extensive empirical evaluations demonstrate that Conformal Data Cleaning (CDC) improves predictive performance in downstream ML tasks in the majority of cases. Our code is available on GitHub: \url{https://github.com/se-jaeger/conformal-data-cleaning}. Sebastian Jäger, Felix Bießmann |
AISTATS | 2 |
| 2023 | Automated Extraction of Fine-Grained Standardized Product Information from Unstructured Multilingual Web Data
Alexander Flick, Sebastian Jäger, Ivana Trajanovska, Felix Bießmann |
ECIR (3) | 4 |
| 2023 | CycleSense: Detecting near miss incidents in bicycle traffic from mobile motion sensors
Ahmet-Serdar Karakaya, Thomas Ritter, Felix Bießmann, David Bermbach |
Pervasive Mob. Comput. | 3 |
| 2021 | JENGA - A Framework to Study the Impact of Data Errors on the Predictions of Machine Learning Models
Sebastian Schelter, Tammo Rukat, Felix Bießmann |
EDBT | 3 |
| 2021 | Enforcing Constraints for Machine Learning Systems via Declarative Feature Selection: An Experimental StudyabstractResponsible usage of Machine Learning (ML) systems in practice does not only require enforcing high prediction quality, but also accounting for other constraints, such as fairness, privacy, or execution time. One way to address multiple user-specified constraints on ML systems is feature selection. Yet, optimizing feature selection strategies for multiple metrics is difficult to implement and has been underrepresented in previous experimental studies. Here, we propose Declarative Feature Selection (DFS) to simplify the design and validation of ML systems satisfying diverse user-specified constraints. We benchmark and evaluate a representative series of feature selection algorithms. From our extensive experimental results, we derive concrete suggestions on when to use which strategy and show that a meta-learning-driven optimizer can accurately predict the right strategy for an ML task at hand. These results demonstrate that feature selection can help to build ML systems that meet combinations of user-specified constraints, independent of the ML methods used. Felix Neutatz, Felix Bießmann, Ziawasch Abedjan |
SIGMOD Conference | 2 |
| 2021 | An Interactive Garment for Orchestra Conducting: IoT-enabled Textile & Machine Learning to Direct Musical PerformanceabstractWe present an overview and initial results from a project bringing together orchestra conducting, e-textile material studies, costume tailoring, low power computing and machine learning (ML). We describe a wearable interactive system comprising of textile sensors embedded into a suit, low-power transmission and gesture recognition using creative computing tools. We introduce first observations made during the semi-participatory approach, which placed the conductor’s movements and personal performative expressiveness at the centre for technical and conceptual development. The project is a two-month collaboration between the Verworner-Krause Kammerorchester (VKKO), technical and design researchers, currently still running. Preliminary analyses of the data recorded while the conductor is wearing the prototype demonstrate that the developed system can be used to robustly decode a large number of conducting and performative movements. In particular the user interface of the ML system is designed such that the training of the algorithms can be intuitively controlled by the conductor, in sync with the MIDI clock. Berit Greinke, Giorgia Petri, Pauline Vierne, Paul Bießmann, Alexandra Börner, Kaspar Schleiser, Emmanuel Baccelli, Claas Krause, Christopher Verworner, Felix Bießmann |
TEI | 10 |
| 2020 | Calibrating Human-AI Collaboration: Impact of Risk, Ambiguity and Transparency on Algorithmic Bias
Philipp Schmidt 0002, Felix Bießmann |
CD-MAKE | 2 |
| 2020 | Learning to Validate the Predictions of Black Box Classifiers on Unseen DataabstractMachine Learning (ML) models are difficult to maintain in production settings. In particular, deviations of the unseen serving data (for which we want to compute predictions) from the source data (on which the model was trained) pose a central challenge, especially when model training and prediction are outsourced via cloud services. Errors or shifts in the serving data can affect the predictive quality of a model, but are hard to detect for engineers operating ML deployments. Sebastian Schelter, Tammo Rukat, Felix Bießmann |
SIGMOD Conference | 3 |
| 2019 | Differential Data Quality Verification on Partitioned DataabstractModern companies and institutions rely on data to guide every single decision. Missing or incorrect information seriously compromises any decision process. In previous work, we presented Deequ, a Spark-based library for automating the verification of data quality at scale. Deequ provides a declarative API, which combines common quality constraints with user-defined validation code, and thereby enables "unit tests for data". However, we found that the previous computational model of Deequ is not flexible enough for many scenarios in modern data pipelines, which handle large, partitioned datasets. Such scenarios require the evaluation of dataset-level quality constraints after individual partition updates, without having to re-read already processed partitions. Additionally, such scenarios often require the verification of data quality on select combinations of partitions. We therefore present a differential generalization of the computational model of Deequ, based on algebraic states with monoid properties. We detail how to efficiently implement the corresponding operators and aggregation functions in Apache Spark. Furthermore, we show how to optimize the resulting workloads to minimize the required number of passes over the data, and empirically validate that our approach decreases the runtimes for updating data metrics under data changes and for different combinations of partitions. Sebastian Schelter, Stefan Grafberger, Philipp Schmidt 0002, Tammo Rukat, Mario Kießling, Andrey Taptunov, Felix Bießmann, Dustin Lange |
ICDE | 7 |
| 2019 | Unit Testing Data with DeequabstractModern companies and institutions rely on data to guide every single decision. Missing or incorrect information seriously compromises any decision process. We demonstrate "Deequ", an Apache Spark-based library for automating the verification of data quality at scale. This library provides a declarative API, which combines common quality constraints with user-defined validation code, and thereby enables "unit tests for data". Deequ is available as open source, meets the requirements of production use cases at Amazon, and scales to datasets with billions of records if the constraints to evaluate are chosen carefully. Our demonstration walks attendees through a fictitious business use case of validating daily product reviews from a public dataset, and is executed in a proprietary interactive notebook environment. We show attendees how to define data unit tests from automatically suggested constraints and how to create customized tests. Additionally, we demonstrate how to apply Deequ to validate incrementally growing datasets, and give examples of how to configure anomaly detection algorithms on time series of data quality metrics to further automate the data validation. Sebastian Schelter, Felix Bießmann, Dustin Lange, Tammo Rukat, Philipp Schmidt 0002, Stephan Seufert, Pierre Brunelle, Andrey Taptunov |
SIGMOD Conference | 2 |
| 2019 | DataWig: Missing Value Imputation for TablesabstractWith the growing importance of machine learning (ML) algorithms for practical applications, reducing data quality problems in ML pipelines has become a major focus of research. In many cases missing values can break data pipelines which makes completeness one of the most impactful data quality challenges. Current missing value imputation methods are focusing on numerical or categorical data and can be difficult to scale to datasets with millions of rows. We release DataWig, a robust and scalable approach for missing value imputation that can be applied to tables with heterogeneous data types, including unstructured text. DataWig combines deep learning feature extractors with automatic hyperparameter tuning. This enables users without a machine learning background, such as data engineers, to impute missing values with minimal effort in tables with more heterogeneous data types than supported in existing libraries, while requiring less glue code for feature engineering and offering more flexible modelling options. We demonstrate that DataWig compares favourably to existing imputation packages. Source code, documentation, and unit tests for this package are available at: https://github.com/awslabs/datawig Felix Bießmann, Tammo Rukat, Philipp Schmidt 0002, Prathik Naidu, Sebastian Schelter, Andrey Taptunov, Dustin Lange, David Salinas |
J. Mach. Learn. Res. | 1 |
| 2018 | "Deep" Learning for Missing Value Imputationin Tables with Non-Numerical DataabstractThe success of applications that process data critically depends on the quality of the ingested data. Completeness of a data source is essential in many cases. Yet, most missing value imputation approaches suffer from severe limitations. They are almost exclusively restricted to numerical data, and they either offer only simple imputation methods or are difficult to scale and maintain in production. Here we present a robust and scalable approach to imputation that extends to tables with non-numerical values, including unstructured text data in diverse languages. Experiments on public data sets as well as data sets sampled from a large product catalog in different languages (English and Japanese) demonstrate that the proposed approach is both scalable and yields more accurate imputations than previous approaches. Training on data sets with several million rows is a matter of minutes on a single machine. With a median imputation F1 score of 0.93 across a broad selection of data sets our approach achieves on average a 23-fold improvement compared to mode imputation. While our system allows users to apply state-of-the-art deep learning models if needed, we find that often simple linear n-gram models perform on par with deep learning methods at a much lower operational cost. The proposed method learns all parameters of the entire imputation pipeline automatically in an end-to-end fashion, rendering it attractive as a generic plugin both for engineers in charge of data pipelines where data completeness is relevant, as well as for practitioners without expertise in machine learning who need to impute missing values in tables with non-numerical data. Felix Bießmann, David Salinas, Sebastian Schelter, Philipp Schmidt 0002, Dustin Lange |
CIKM | 1 |
| 2018 | Automating Large-Scale Data Quality VerificationabstractModern companies and institutions rely on data to guide every single business process and decision. Missing or incorrect information seriously compromises any decision process downstream. Therefore, a crucial, but tedious task for everyone involved in data processing is to verify the quality of their data. We present a system for automating the verification of data quality at scale, which meets the requirements of production use cases. Our system provides a declarative API, which combines common quality constraints with user-defined validation code, and thereby enables 'unit tests' for data. We efficiently execute the resulting constraint validation workload by translating it to aggregation queries on Apache Spark. Our platform supports the incremental validation of data quality on growing datasets, and leverages machine learning, e.g., for enhancing constraint suggestions, for estimating the 'predictability' of a column, and for detecting anomalies in historic data quality time series. We discuss our design decisions, describe the resulting system architecture, and present an experimental evaluation on various datasets. Sebastian Schelter, Dustin Lange, Philipp Schmidt 0002, Meltem Celikel, Felix Bießmann, Andreas Grafberger |
Proc. VLDB Endow. | 5 |
| 2015 | Multivariate Machine Learning Methods for Fusing Multimodal Functional Neuroimaging DataabstractMultimodal data are ubiquitous in engineering, communications, robotics, computer vision, or more generally speaking in industry and the sciences. All disciplines have developed their respective sets of analytic tools to fuse the information that is available in all measured modalities. In this paper, we provide a review of classical as well as recent machine learning methods (specifically factor models) for fusing information from functional neuroimaging techniques such as: LFP, EEG, MEG, fNIRS, and fMRI. Early and late fusion scenarios are distinguished, and appropriate factor models for the respective scenarios are presented along with example applications from selected multimodal neuroimaging studies. Further emphasis is given to the interpretability of the resulting model parameters, in particular by highlighting how factor models relate to physical models needed for source localization. The methods we discuss allow for the extraction of information from neural data, which ultimately contributes to 1) better neuroscientific understanding; 2) enhance diagnostic performance; and 3) discover neural signals of interest that correlate maximally with a given cognitive paradigm. While we clearly study the multimodal functional neuroimaging challenge, the discussed machine learning techniques have a wide applicability, i.e., in general data fusion, and may thus be informative to the general interested reader. Sven Dähne, Felix Bießmann, Wojciech Samek, Stefan Haufe, Dominique Goltz, Christopher Gundlach, Arno Villringer, Siamac Fazli, Klaus-Robert Müller |
Proc. IEEE | 2 |
| 2015 | Learning From More Than One Data Source: Data Fusion Techniques for Sensorimotor Rhythm-Based Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) are successfully used in scientific, therapeutic and other applications. Remaining challenges are among others a low signal-to-noise ratio of neural signals, lack of robustness for decoders in the presence of inter-trial and inter-subject variability, time constraints on the calibration phase and the use of BCIs outside a controlled lab environment. Recent advances in BCI research addressed these issues by novel combinations of complementary analysis as well as recording techniques, so called hybrid BCIs. In this paper, we review a number of data fusion techniques for BCI along with hybrid methods for BCI that have recently emerged. Our focus will be on sensorimotor rhythm-based BCIs. We will give an overview of the three main lines of research in this area, integration of complementary features of neural activation, integration of multiple previous sessions and of multiple subjects, and show how these techniques can be used to enhance modern BCI systems. Siamac Fazli, Sven Dähne, Wojciech Samek, Felix Bießmann, Klaus-Robert Müller |
Proc. IEEE | 4 |
| 2013 | Integration of Multivariate Data Streams With Bandpower SignalsabstractThe urge to further our understanding of multimodal neural data has recently become an important topic due to the ever increasing availability of simultaneously recorded data from different neural imaging modalities. In case where EEG is one of the modalities, it is of interest to relate a nonlinear function of the raw EEG time-domain signal, say, EEG band power, to another modality such as the hemodynamic response, as measured with NIRS or fMRI. In this work we tackle exactly this problem defining a novel algorithm that we denote multimodal source power correlation analysis (mSPoC). The validity and high performance of the mSPoC framework is demonstrated for simulated and real-world multimodal data. Sven Dähne, Felix Bießmann, Frank C. Meinecke, Jan Mehnert, Siamac Fazli, Klaus-Robert Müller |
IEEE Trans. Multim. | 2 |
| 2012 | Canonical Trends: Detecting Trend Setters in Web Data
Felix Bießmann, Jens-Michalis Papaioannou, Mikio L. Braun, Andreas Harth |
ICML | 1 |
| 2010 | Temporal kernel CCA and its application in multimodal neuronal data analysisabstractData recorded from multiple sources sometimes exhibit non-instantaneous couplings. For simple data sets, cross-correlograms may reveal the coupling dynamics. But when dealing with high-dimensional multivariate data there is no such measure as the cross-correlogram. We propose a simple algorithm based on Kernel Canonical Correlation Analysis (kCCA) that computes a multivariate temporal filter which links one data modality to another one. The filters can be used to compute a multivariate extension of the cross-correlogram, the canonical correlogram, between data sources that have different dimensionalities and temporal resolutions. The canonical correlogram reflects the coupling dynamics between the two sources. The temporal filter reveals which features in the data give rise to these couplings and when they do so. We present results from simulations and neuroscientific experiments showing that tkCCA yields easily interpretable temporal filters and correlograms. In the experiments, we simultaneously performed electrode recordings and functional magnetic resonance imaging (fMRI) in primary visual cortex of the non-human primate. While electrode recordings reflect brain activity directly, fMRI provides only an indirect view of neural activity via the Blood Oxygen Level Dependent (BOLD) response. Thus it is crucial for our understanding and the interpretation of fMRI signals in general to relate them to direct measures of neural activity acquired with electrodes. The results computed by tkCCA confirm recent models of the hemodynamic response to neural activity and allow for a more detailed analysis of neurovascular coupling dynamics. Felix Bießmann, Frank C. Meinecke, Arthur Gretton, Alexander Rauch, Gregor Rainer, Nikos K. Logothetis, Klaus-Robert Müller |
Mach. Learn. | 1 |
| 2008 | Effects of Stimulus Type and of Error-Correcting Code Design on BCI Speller PerformanceabstractFrom an information-theoretic perspective, a noisy transmission system such as a visual Brain-Computer Interface (BCI) speller could benefit from the use of error-correcting codes. However, optimizing the code solely according to the maximal minimum-Hamming-distance criterion tends to lead to an overall increase in target frequency of target stimuli, and hence a significantly reduced average target-to-target interval (TTI), leading to difficulties in classifying the individual event-related potentials (ERPs) due to overlap and refractory effects. Clearly any change to the stimulus setup must also respect the possible psychophysiological consequences. Here we report new EEG data from experiments in which we explore stimulus types and codebooks in a within-subject design, finding an interaction between the two factors. Our data demonstrate that the traditional, row-column code has particular spatial properties that lead to better performance than one would expect from its TTIs and Hamming-distances alone, but nonetheless error-correcting codes can improve performance provided the right stimulus type is used. N. Jeremy Hill, Jason Farquhar, Suzanna Martens, Felix Bießmann, Bernhard Schölkopf |
NIPS | 4 |