Jörg Wicker

dblp:56/3110 · also Jörg S. Wicker, Jörg Simon Wicker · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-0533-3368ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 13 · 5 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Self-purification: Enhancing adversarial defense by leveraging local relative robustness
Rui Zhang 0070, Jörg Wicker, Katharina Dost, Qinli Yang, Junming Shao
Expert Syst. Appl.2
2025 Understanding Rumen Methanogen Interactions in Sheep Using Machine Learning
Katharina Dost, Steffen Albrecht, Paul H. Maclean, Jörg Wicker
ECML/PKDD (8)4
2025 Assessing the risk of discriminatory bias in classification datasets
abstract
Abstract Bias in machine learning models remains a critical challenge, particularly in datasets with numeric features where discrimination may be subtle and hard to detect. Existing fairness frameworks rely on expert knowledge of marginalized groups, such as specific racial groups, and categorical features defining them. Furthermore, most frameworks evaluate bias in models rather than datasets, despite the fact that model bias can often be traced back to dataset shortcomings. Our research aims to remedy this gap by capturing dataset flaws in a set of meta-features at the dataset level, and to warn practitioners of bias risk when using such datasets for model training. We neither restrict the feature type nor expect domain knowledge. To this end, we develop methods to synthesize biased datasets and extend current fairness metrics to continuous features in order to quantify dataset-level discrimination risks. Our approach constructs a meta-database of diverse datasets, from which we derive transferable meta-features that capture dataset properties indicative of bias risk. Our findings demonstrate that dataset-level characteristics can serve as cost-effective indicators of bias risk, providing a novel method for data auditing that does not rely on expert knowledge. This work lays the foundation for early-warning systems, moving beyond model-focused assessments toward a data-centric approach.
Kejun Dai, Jonathan Kim, Saso Dzeroski, Jörg Wicker, Gillian Dobbie, Katharina Dost
Mach. Learn.4
2024 Resource-Constrained Binary Image Classification
Sean Park, Jörg Wicker, Katharina Dost
DS (2)2
2024 Remote Sensing for Water Quality: A Multi-Task, Metadata-Driven Hypernetwork Approach
Olivier Graffeuille, Yun Sing Koh, Jörg Wicker, Moritz K. Lehmann
IJCAI3
2024 Regional bias in monolingual English language models
abstract
Abstract In Natural Language Processing (NLP), pre-trained language models (LLMs) are widely employed and refined for various tasks. These models have shown considerable social and geographic biases creating skewed or even unfair representations of certain groups. Research focuses on biases toward L2 (English as a second language) regions but neglects bias within L1 (first language) regions. In this work, we ask if there is regional bias within L1 regions already inherent in pre-trained LLMs and, if so, what the consequences are in terms of downstream model performance. We contribute an investigation framework specifically tailored for low-resource regions, offering a method to identify bias without imposing strict requirements for labeled datasets. Our research reveals subtle geographic variations in the word embeddings of BERT, even in cultures traditionally perceived as similar. These nuanced features, once captured, have the potential to significantly impact downstream tasks. Generally, models exhibit comparable performance on datasets that share similarities, and conversely, performance may diverge when datasets differ in their nuanced features embedded within the language. It is crucial to note that estimating model performance solely based on standard benchmark datasets may not necessarily apply to the datasets with distinct features from the benchmark datasets. Our proposed framework plays a pivotal role in identifying and addressing biases detected in word embeddings, particularly evident in low-resource regions such as New Zealand.
Jiachen Lyu, Katharina Dost, Yun Sing Koh, Jörg Wicker
Mach. Learn.4
2024 Hitting the target: stopping active learning at the cost-based optimum
abstract
Abstract Active learning allows machine learning models to be trained using fewer labels while retaining similar performance to traditional supervised learning. An active learner selects the most informative data points, requests their labels, and retrains itself. While this approach is promising, it raises the question of how to determine when the model is ‘good enough’ without the additional labels required for traditional evaluation. Previously, different stopping criteria have been proposed aiming to identify the optimal stopping point. Yet, optimality can only be expressed as a domain-dependent trade-off between accuracy and the number of labels, and no criterion is superior in all applications. As a further complication, a comparison of criteria for a particular real-world application would require practitioners to collect additional labelled data they are aiming to avoid by using active learning in the first place. This work enables practitioners to employ active learning by providing actionable recommendations for which stopping criteria are best for a given real-world scenario. We contribute the first large-scale comparison of stopping criteria for pool-based active learning, using a cost measure to quantify the accuracy/label trade-off, public implementations of all stopping criteria we evaluate, and an open-source framework for evaluating stopping criteria. Our research enables practitioners to substantially reduce labelling costs by utilizing the stopping criterion which best suits their domain.
Zac Pullar-Strecker, Katharina Dost, Eibe Frank, Jörg Wicker
Mach. Learn.4
2024 From What You See to What We Smell: Linking Human Emotions to Bio-Markers in Breath
abstract
Research has shown that the composition of breath can differ based on the human's behavioral patterns and mental and physical states immediately before being collected. These breath-collection techniques have also been extended to observe the general processes occurring in groups of humans and can link them to what those groups are collectively experiencing. In this research, we applied machine learning techniques to the breath data collected from cinema audiences. These techniques included XGBOOST Regression, Hierarchical Clustering, and Item Basket analyses created using the Apriori algorithm. They were conducted to find associations between the biomarkers in the crowd's breath and the movie's audio-visual stimuli and thematic events. This analysis enabled us to directly link what the group was experiencing and their biological response to that experience. We first extracted visual and auditory features from a movie to achieve this. We compared it to the biomarkers in the crowd's breath using regression and pattern mining techniques. Our results supported the theory that a crowd's collective experience directly correlates to the biomarkers in the crowd's breath. Consequently, these findings suggest that visual and auditory experiences have predictable effects on the human body that can be monitored without requiring expensive or invasive neuroimaging techniques.
Joshua Bensemann, Hasnain Cheena, David Tse Jung Huang, Elizabeth Broadbent, Jörg Wicker
IEEE Trans. Affect. Comput.6
2023 BAARD: Blocking Adversarial Examples by Testing for Applicability, Reliability and Decidability
Xinglong Chang, Katharina Dost, Kaiqi Zhao 0001, Ambra Demontis, Fabio Roli, Gillian Dobbie, Jörg Wicker
PAKDD (1)7
2023 Targeted Attacks on Time Series Forecasting
Katharina Dost, Xinglong Chang, Gillian Dobbie, Jörg Wicker
PAKDD (4)6
2022 Semi-supervised Conditional Density Estimation with Wasserstein Laplacian Regularisation
abstract
Conditional Density Estimation (CDE) has wide-reaching applicability to various real-world problems, such as spatial density estimation and environmental modelling. CDE estimates the probability density of a random variable rather than a single value and can thus model uncertainty and inverse problems. This task is inherently more complex than regression, and many algorithms suffer from overfitting, particularly when modelled with few labelled data points. For applications where unlabelled data is abundant but labelled data is scarce, we propose Wasserstein Laplacian Regularisation, a semi-supervised learning framework that allows CDE algorithms to leverage these unlabelled data. The framework minimises an objective function which ensures that the learned model is smooth along the manifold of the underlying data, as measured by Wasserstein distance. When applying our framework to Mixture Density Networks, the resulting semi-supervised algorithm can achieve similar performance to a supervised model with up to three times as many labelled data points on baseline datasets. We additionally apply our technique to the problem of remote sensing for chlorophyll-a estimation in inland waters.
Olivier Graffeuille, Yun Sing Koh, Jörg Wicker, Moritz K. Lehmann
AAAI3
2022 Closing the Loop: Graph Networks to Unify Semantic Objects and Visual Features for Multi-object Scenes
abstract
In Simultaneous Localization and Mapping (SLAM), Loop Closure Detection (LCD) is essential to minimize drift when recognizing previously visited places. Visual Bag- of-Words (vBoW) has been an LCD algorithm of choice for many state-of-the-art SLAM systems. It uses a set of visual features to provide robust place recognition but fails to perceive the semantics or spatial relationship between feature points. Previous work has mainly focused on addressing these issues by combining vBoW with semantic and spatial information from objects in the scene. However, they are unable to exploit spatial information of local visual features and lack a structure that unifies semantic objects and visual features, therefore limiting the symbiosis between the two components. This paper proposes SymbioLCD2, which creates a unified graph structure to integrate semantic objects and visual features symbiotically. Our novel graph-based LCD system utilizes the unified graph structure by applying a Weisfeiler-Lehman graph kernel with temporal constraints to robustly predict loop closure candidates. Evaluation of the proposed system shows that having a unified graph structure incorporating semantic objects and visual features improves LCD prediction accuracy, illustrating that the proposed graph structure provides a strong symbiosis between these two complementary components. It also outperforms other Machine Learning algorithms - such as SVM, Decision Tree, Random Forest, Neural Network and GNN based Graph Matching Networks. Furthermore, it has shown good performance in detecting loop closure candidates earlier than state-of-the-art SLAM systems, demonstrating that extended semantic and spatial awareness from the unified graph structure significantly impacts LCD performance.
Jonathan J. Y. Kim, Martin Urschler, Patricia J. Riddle, Jörg Wicker
IROS4
2022 Divide and Imitate: Multi-cluster Identification and Mitigation of Selection Bias
Katharina Dost, Hamish Duncanson, Ioannis Ziogas, Patricia J. Riddle, Jörg Wicker
PAKDD (2)5
2021 SymbioLCD: Ensemble-Based Loop Closure Detection using CNN-Extracted Objects and Visual Bag-of-Words
abstract
Loop closure detection is an essential tool of Simultaneous Localization and Mapping (SLAM) to minimize drift in its localization. Many state-of-the-art loop closure detection (LCD) algorithms use visual Bag-of-Words (vBoW), which is robust against partial occlusions in a scene but cannot perceive the semantics or spatial relationships between feature points. CNN object extraction can address those issues, by providing semantic labels and spatial relationships between objects in a scene. Previous work has mainly focused on replacing vBoW with CNN derived features. In this paper we propose SymbioLCD, a novel ensemble-based LCD that utilizes both CNN-extracted objects and vBoW features for LCD candidate prediction. When used in tandem, the added elements of object semantics and spatial-awareness creates a more robust and symbiotic loop closure detection system. The proposed SymbioLCD uses scale-invariant spatial and semantic matching, Hausdorff distance with temporal constraints, and a Random Forest that utilizes combined information from both CNN-extracted objects and vBoW features for predicting accurate loop closure candidates. Evaluation of the proposed method shows it outperforms other Machine Learning (ML) algorithms - such as SVM, Decision Tree and Neural Network, and demonstrates that there is a strong symbiosis between CNN-extracted object information and vBoW features which assists accurate LCD candidate prediction. Furthermore, it is able to perceive loop closure candidates earlier than state-of-the-art SLAM algorithms, utilizing added spatial and semantic information from CNN-extracted objects.
Jonathan J. Y. Kim, Martin Urschler, Patricia J. Riddle, Jörg Wicker
IROS4
2020 Your Best Guess When You Know Nothing: Identification and Mitigation of Selection Bias
abstract
Machine Learning typically assumes that training and test sets are independently drawn from the same distribution, but this assumption is often violated in practice which creates a bias. Many attempts to identify and mitigate this bias have been proposed, but they usually rely on ground-truth information. But what if the researcher is not even aware of the bias? In contrast to prior work, this paper introduces a new method, Imitate, to identify and mitigate Selection Bias in the case that we may not know if (and where) a bias is present, and hence no ground-truth information is available. Imitate investigates the dataset's probability density, then adds generated points in order to smooth out the density and have it resemble a Gaussian, the most common density occurring in real-world applications. If the artificial points focus on certain areas and are not widespread, this could indicate a Selection Bias where these areas are underrepresented in the sample. We demonstrate the effectiveness of the proposed method in both, synthetic and real-world datasets. We also point out limitations and future research directions.
Katharina Dost, Katerina Tashkova, Patricia J. Riddle, Jörg Wicker
ICDM4
2019 XOR-Based Boolean Matrix Decomposition
abstract
Boolean matrix factorization (BMF) is a data summarizing and dimension-reduction technique. Existing BMF methods build on matrix properties defined by Boolean algebra, where the addition operator is the logical inclusive OR and the multiplication operator the logical AND. As a consequence, this leads to the lack of an additive inverse in all Boolean matrix operations, which produces an indelible type of approximation error. Previous research adopted various methods to address such an issue and produced reasonably accurate approximation. However, an exact factorization is rarely found in the literature. In this paper, we introduce a new algorithm named XBMaD (XOR-based Boolean Matrix Decomposition) where the addition operator is defined as the exclusive OR (XOR). This change completely removes the error-mitigation issue of OR-based BMF methods, and allows for an exact error-free factorization. An evaluation comparing XBMaD and classic OR-based methods suggested that XBMAD performed equal or in most cases more accurately and faster.
Jörg Wicker, Yan Cathy Hua, Rayner Rebello, Bernhard Pfahringer
ICDM1
2017 The best privacy defense is a good privacy offense: obfuscating a search engine user's profile
Jörg Wicker, Stefan Kramer 0001
Data Min. Knowl. Discov.1
2016 A Nonlinear Label Compression and Transformation Method for Multi-label Classification Using Autoencoders
Jörg Wicker, Andrey Tyukin, Stefan Kramer 0001
PAKDD (1)1
2015 Cinema Data Mining: The Smell of Fear
abstract
While the physiological response of humans to emotional events or stimuli is well-investigated for many modalities (like EEG, skin resistance, ...), surprisingly little is known about the exhalation of so-called Volatile Organic Compounds (VOCs) at quite low concentrations in response to such stimuli. VOCs are molecules of relatively small mass that quickly evaporate or sublimate and can be detected in the air that surrounds us. The paper introduces a new field of application for data mining, where trace gas responses of people reacting on-line to films shown in cinemas (or movie theaters) are related to the semantic content of the films themselves. To do so, we measured the VOCs from a movie theater over a whole month in intervals of thirty seconds, and annotated the screened films by a controlled vocabulary compiled from multiple sources. To gain a better understanding of the data and to reveal unknown relationships, we have built prediction models for so-called forward prediction (the prediction of future VOCs from the past), backward prediction (the prediction of past scene labels from future VOCs), which is some form of abductive reasoning, and Granger causality. Experimental results show that some VOCs and some labels can be predicted with relatively low error, and that hint for causality with low p-values can be detected in the data. The data set is publicly available at: https://github.com/jorro/smelloffear.
Jörg Wicker, Nicolas Krauter, Bettina Derstorff, Christof Stönner, Efstratios Bourtsoukidis, Thomas Klüpfel, Stefan Kramer 0001
KDD1
2015 Scavenger - A Framework for Efficient Evaluation of Dynamic and Modular Algorithms
Andrey Tyukin, Stefan Kramer 0001, Jörg Wicker
ECML/PKDD (3)3
2014 BMaD - A Boolean Matrix Decomposition Framework
Andrey Tyukin, Stefan Kramer 0001, Jörg Wicker
ECML/PKDD (3)3
2010 Predicting biodegradation products and pathways: a hybrid knowledge- and machine learning-based approach
abstract
MOTIVATION: Current methods for the prediction of biodegradation products and pathways of organic environmental pollutants either do not take into account domain knowledge or do not provide probability estimates. In this article, we propose a hybrid knowledge- and machine learning-based approach to overcome these limitations in the context of the University of Minnesota Pathway Prediction System (UM-PPS). The proposed solution performs relative reasoning in a machine learning framework, and obtains one probability estimate for each biotransformation rule of the system. As the application of a rule then depends on a threshold for the probability estimate, the trade-off between recall (sensitivity) and precision (selectivity) can be addressed and leveraged in practice. RESULTS: Results from leave-one-out cross-validation show that a recall and precision of approximately 0.8 can be achieved for a subset of 13 transformation rules. Therefore, it is possible to optimize precision without compromising recall. We are currently integrating the results into an experimental version of the UM-PPS server. AVAILABILITY: The program is freely available on the web at http://wwwkramer.in.tum.de/research/applications/biodegradation/data. CONTACT: [email protected].
Jörg Wicker, Kathrin Fenner, Lynda B. M. Ellis, Lawrence P. Wackett, Stefan Kramer 0001
Bioinform.1
2008 An inductive database and query language in the relational model
abstract
In the demonstration, we will present the concepts and an implementation of an inductive database -- as proposed by Imielinski and Mannila -- in the relational model. The goal is to support all steps of the knowledge discovery process, from pre-processing via data mining to post-processing, on the basis of queries to a database system. The query language SIQL (structured inductive query language), an SQL extension, offers query primitives for feature selection, discretization, pattern mining, clustering, instance-based learning and rule induction. A prototype system processing such queries was implemented as part of the SINDBAD (structured inductive database development) project. Key concepts of this system, among others, are the closure of operators and distances between objects. To support the analysis of multi-relational data, we incorporated multi-relational distance measures based on set distances and recursive descent. The inclusion of rule-based classification models made it necessary to extend the data model and the software architecture significantly. The prototype is applied to three different applications: gene expression analysis, gene regulation prediction and structure-activity relationships (SARs) of small molecules.
Lothar Richter, Jörg Wicker, Kristina Kessler, Stefan Kramer 0001
EDBT2
2008 SINDBAD and SiQL: An Inductive Database and Query Language in the Relational Model
Jörg Wicker, Lothar Richter, Kristina Kessler, Stefan Kramer 0001
ECML/PKDD (2)1