Katharina Dost

dblp:285/3277 · DBLP profile ↗
← Back
6ranked-venue papers in the field
3as first author
5since 2021 · last 2025
0000-0002-1514-0685ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 6 (3 first)
YearPublicationVenuePosition
2025 Understanding Rumen Methanogen Interactions in Sheep Using Machine Learning
Katharina Dost, Steffen Albrecht, Paul H. Maclean, Jörg Wicker
ECML/PKDD (8)1
2023 BAARD: Blocking Adversarial Examples by Testing for Applicability, Reliability and Decidability
Xinglong Chang, Katharina Dost, Kaiqi Zhao 0001, Ambra Demontis, Fabio Roli, Gillian Dobbie, Jörg Wicker
PAKDD (1)2
2023 Targeted Attacks on Time Series Forecasting
Katharina Dost, Xinglong Chang, Gillian Dobbie, Jörg Wicker
PAKDD (4)2
2023 Interpretability Meets Generalizability: A Hybrid Machine Learning System to Identify Nonlinear Granger Causality in Global Stock Indices
Yixiao Lu, Yokiu Lee, Johnathan Chi-Ho Leung, Alvin Cheung, Katharina Dost, Katerina Tashkova, Thomas Lacombe
PAKDD (2)6
2022 Divide and Imitate: Multi-cluster Identification and Mitigation of Selection Bias
Katharina Dost, Hamish Duncanson, Ioannis Ziogas, Patricia J. Riddle, Jörg Wicker
PAKDD (2)1
2020 Your Best Guess When You Know Nothing: Identification and Mitigation of Selection Bias
abstract
Machine Learning typically assumes that training and test sets are independently drawn from the same distribution, but this assumption is often violated in practice which creates a bias. Many attempts to identify and mitigate this bias have been proposed, but they usually rely on ground-truth information. But what if the researcher is not even aware of the bias? In contrast to prior work, this paper introduces a new method, Imitate, to identify and mitigate Selection Bias in the case that we may not know if (and where) a bias is present, and hence no ground-truth information is available. Imitate investigates the dataset's probability density, then adds generated points in order to smooth out the density and have it resemble a Gaussian, the most common density occurring in real-world applications. If the artificial points focus on certain areas and are not widespread, this could indicate a Selection Bias where these areas are underrepresented in the sample. We demonstrate the effectiveness of the proposed method in both, synthetic and real-world datasets. We also point out limitations and future research directions.
Katharina Dost, Katerina Tashkova, Patricia J. Riddle, Jörg Wicker
ICDM1