Roman Kern

dblp:85/6799 · DBLP profile ↗
← Back
30ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0003-0202-6100ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 19 · 8 since 2021Artificial intelligence and machine learning · 10 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Saga++: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications
abstract
In the exploratory data science lifecycle, data scientists often spent the majority of their time finding, integrating, validating, and cleaning relevant datasets. Despite recent work on data validation, and numerous error detection and correction algorithms, in practice, data cleaning for ML remains largely a manual, unpleasant, and labor-intensive trial and error process, especially in large-scale, distributed computation settings. The target ML application—such as classification or regression models—can be used as a signal of valuable feedback though, for selecting effective data cleaning strategies. In this article, we introduce Saga++ , a framework for automatically generating the top-K most effective data cleaning pipelines. Saga++ adopts ideas from Auto-ML, feature selection, and hyper-parameter tuning. Our framework is extensible for user-provided constraints, new data cleaning primitives, and ML applications; automatically generates hybrid runtime plans of local and distributed operations; and performs pruning by interesting properties (e.g., monotonicity). Furthermore, we exploit guided sampling on the input dataset to enable enumeration on a smaller subset, reducing the time required to discover the top-K pipelines. As a post-processing step, we also perform pipeline pruning on the selected top-K pipelines, removing redundant and less effective cleaning primitives. Instead of full automation—which is rather unrealistic— Saga++ simplifies the mechanical aspects of data cleaning. Our experiments show that Saga++ yields robust accuracy improvements over state-of-the-art, and good scalability regarding increasing data sizes and number of evaluated pipelines.
Shafaq Siddiqi, Arnab Phani, Roman Kern, Matthias Boehm 0001
ACM Trans. Database Syst.3
2025 Ensemble Watermarks for Large Language Models
abstract
As large language models (LLMs) reach human-like fluency, reliably distinguishing AIgenerated text from human authorship becomes increasingly difficult.While watermarks already exist for LLMs, they often lack flexibility and struggle with attacks such as paraphrasing.To address these issues, we propose a multi-feature method for generating watermarks that combines multiple distinct watermark features into an ensemble watermark.Concretely, we combine acrostica and sensorimotor norms with the established red-green watermark to achieve a 98% detection rate.After a paraphrasing attack, the performance remains high with 95% detection rate.In comparison, the red-green feature alone as a baseline achieves a detection rate of 49% after paraphrasing.The evaluation of all feature combinations reveals that the ensemble of all three consistently has the highest detection rate across several LLMs and watermark strength settings.Due to the flexibility of combining features in the ensemble, various requirements and tradeoffs can be addressed.Additionally, the same detection function can be used without adaptations for all ensemble configurations.This method is particularly of interest to facilitate accountability and prevent societal harm.
Georg Niess, Roman Kern
ACL (1)2
2025 Detecting abrupt changes in missing time series data
abstract
When time series data contain missing values, it is common practice to substitute them using missing value imputation. However, if there is an unobserved abrupt change in the missing values, then standard imputation techniques are insufficient since they are biased towards normal data. Likewise, standard detectors cannot find abrupt changes that “hide” in missing data. To address these shortcomings, we propose Interval Forecast Imputation (IFI), which is a simple and intuitive combination of uncertainty intervals, forecasting, and anomaly detection that detects abrupt changes in missing time series data. A further advantage of IFI is that it is compatible with every state of the art forecasting technique—ranging from simple exponential smoothing over neural network-assisted forecasts to the popular Prophet library—while requiring only O ( 1 ) additional time and space. In our experiments, we observe that IFI can detect abrupt changes in missing data and improves the imputation accuracy of all forecasting methods it is combined with.
Maximilian Toller, Bernhard C. Geiger, Roman Kern
Inf. Sci.3
2025 Assessing the impact of differential privacy in transfer learning with deep neural networks and transformer language models
Samuel Sousa 0001, Andreas Trügler, Roman Kern
Neural Comput. Appl.3
2024 Exploring the Capabilities of GPT4-Vision as OCR Engine
Alex Ghiriti, Wolfgang Göderle, Roman Kern
TPDL (2)3
2024 Assessing trustworthy AI: Technical and legal perspectives of fairness in AI
abstract
Artificial Intelligence systems are used more and more nowadays, from the application of decision support systems to autonomous vehicles. Hence, the widespread use of AI systems in various fields raises concerns about their potential impact on human safety and autonomy, especially regarding fair decision-making. In our research, we primarily concentrate on aspects of non-discrimination, encompassing both group and individual fairness. Therefore, it must be ensured that decisions made by such systems are fair and unbiased. Although there are many different methods for bias mitigation, few of them meet existing legal requirements. Unclear legal frameworks further worsen this problem. To address this issue, this paper investigates current state-of-the-art methods for bias mitigation and contrasts them with the legal requirements, with the scope limited to the European Union and with a particular focus on the AI Act. Moreover, the paper initially examines state-of-the-art approaches to ensure AI fairness, and subsequently, outlines various fairness measures. Challenges of defining fairness and the need for a comprehensive legal methodology to address fairness in AI systems are discussed. The paper contributes to the ongoing discussion on fairness in AI and highlights the importance of meeting legal requirements to ensure fairness and non-discrimination for all data subjects.
Markus Kattnig, Alessa Angerschmid, Thomas Reichel, Roman Kern
Comput. Law Secur. Rev.4
2023 SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications
abstract
In the exploratory data science lifecycle, data scientists often spent the majority of their time finding, integrating, validating and cleaning relevant datasets. Despite recent work on data validation, and numerous error detection and correction algorithms, in practice, data cleaning for ML remains largely a manual, unpleasant, and labor-intensive trial and error process, especially in large-scale, distributed computation. The target ML application---such as classification or regression models---can be used as a signal of valuable feedback though, for selecting effective data cleaning strategies. In this paper, we introduce SAGA, a framework for automatically generating the top-K most effective data cleaning pipelines. SAGA adopts ideas from Auto-ML, feature selection, and hyper-parameter tuning. Our framework is extensible for user-provided constraints, new data cleaning primitives, and ML applications; automatically generates hybrid runtime plans of local and distributed operations; and performs pruning by interesting properties (e.g., monotonicity). Instead of full automation---which is rather unrealistic---SAGA simplifies the mechanical aspects of data cleaning. Our experiments show that SAGA yields robust accuracy improvements over state-of-the-art, and good scalability regarding increasing data sizes and number of evaluated pipelines.
Shafaq Siddiqi, Roman Kern, Matthias Boehm 0001
Proc. ACM Manag. Data2
2023 Cluster Purging: Efficient Outlier Detection Based on Rate-Distortion Theory
abstract
Rate-distortion theory-based outlier detection builds upon the rationale that a good data compression will encode outliers with unique symbols. Based on this rationale, we propose Cluster Purging, which is an extension of clustering-based outlier detection. This extension allows one to assess the representivity of clusterings, and to find data that are best represented by individual unique clusters. We propose two efficient algorithms for performing Cluster Purging, one being parameter-free, while the other algorithm has a parameter that controls representivity estimations, allowing it to be tuned in supervised setups. In an experimental evaluation, we show that Cluster Purging improves upon outliers detected from raw clusterings, and that Cluster Purging competes strongly against state-of-the-art alternatives.
Maximilian Toller, Bernhard C. Geiger, Roman Kern
IEEE Trans. Knowl. Data Eng.3
2022 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines
Patrick Damme, Marius Birkenbach, Constantinos Bitsakos, Matthias Boehm 0001, Philippe Bonnet, Florina M. Ciorba, Mark Dokter, Pawel Dowgiallo, Ahmed Eleliemy, Christian Färber, Georgios I. Goumas, Dirk Habich, Niclas Hedam, Marlies Hofer, Kevin Innerebner, Vasileios Karakostas, Roman Kern, Tomaz Kosar, Alexander Krause 0001, Daniel Krems, Andreas Laber, Wolfgang Lehner, Eric Mier, Marcus Paradies, Bernhard Peischl, Gabrielle Poerwawinata, Stratos Psomadakis, Tilmann Rabl, Piotr Ratuszniak, Pedro Silva 0011, Nikolai Skuppin, Andreas Starzacher, Benjamin Steinwender, Ilin Tolovski, Pinar Tözün, Wojciech Ulatowski, Yuanyuan Wang 0002, Izajasz P. Wrosz, Ales Zamuda, Ce Zhang 0001, Xiao Xiang Zhu 0001
CIDR18
2022 Adversarial Inter-Group Link Injection Degrades the Fairness of Graph Neural Networks
abstract
We present evidence for the existence and effectiveness of adversarial attacks on graph neural networks (GNNs) that aim to degrade fairness. These attacks can disadvantage a particular subgroup of nodes in GNN-based node classification, where nodes of the underlying network have sensitive attributes, such as race or gender. We conduct qualitative and experimental analyses explaining how adversarial link injection impairs the fairness of GNN predictions. For example, an attacker can compromise the fairness of GNN-based node classification by injecting adversarial links between nodes belonging to opposite subgroups and opposite class labels. Our experiments on empirical datasets demonstrate that adversarial fairness attacks can significantly degrade the fairness of GNN predictions (attacks are effective) with a low perturbation rate (attacks are efficient) and without a significant drop in accuracy (attacks are deceptive). This work demonstrates the vulnerability of GNN models to adversarial fairness attacks. We hope our findings raise awareness about this issue in our community and lay a foundation for the future development of GNN models that are more robust to such attacks.
Hussain Hussain, Sandipan Sikdar, Denis Helic, Elisabeth Lex, Markus Strohmaier, Roman Kern
ICDM7
2022 Causal Investigation of Public Opinion during the COVID-19 Pandemic via Social Media Text
abstract
Understanding the needs and fears of citizens, especially during a pandemic such as COVID-19, is essential for any government or legislative entity. An effective COVID-19 strategy further requires that the public understand and accept the restriction plans imposed by these entities. In this paper, we explore a causal mediation scenario in which we want to emphasize the use of NLP methods in combination with methods from economics and social sciences. Based on sentiment analysis of Tweets towards the current COVID-19 situation in the UK and Sweden, we conduct several causal inference experiments and attempt to decouple the effect of government restrictions on mobility behavior from the effect that occurs due to public perception of the COVID-19 strategy in a country. To avoid biased results we control for valid country specific epidemiological and time-varying confounders. Comprehensive experiments show that not all changes in mobility are caused by countries implemented policies but also by the support of individuals in the fight against this pandemic. We find that social media texts are an important source to capture citizens’ concerns and trust in policy makers and are suitable to evaluate the success of government policies.
Michael Jantscher, Roman Kern
LREC2
2022 Constructing robust health indicators from complex engineered systems via anticausal learning
Georgios Koutroulis, Belgin Mutlu, Roman Kern
Eng. Appl. Artif. Intell.3
2021 KOMPOS: Connecting Causal Knots in Large Nonlinear Time Series with Non-Parametric Regression Splines
abstract
Recovering causality from copious time series data beyond mere correlations has been an important contributing factor in numerous scientific fields. Most existing works assume linearity in the data that may not comply with many real-world scenarios. Moreover, it is usually not sufficient to solely infer the causal relationships. Identifying the correct time delay of cause-effect is extremely vital for further insight and effective policies in inter-disciplinary domains. To bridge this gap, we propose KOMPOS, a novel algorithmic framework that combines a powerful concept from causal discovery of additive noise models with graphical ones. We primarily build our structural causal model from multivariate adaptive regression splines with inherent additive local nonlinearities, which render the underlying causal structure more easily identifiable. In contrast to other methods, our approach is not restricted to Gaussian or non-Gaussian noise due to the non-parametric attribute of the regression method. We conduct extensive experiments on both synthetic and real-world datasets, demonstrating the superiority of the proposed algorithm over existing causal discovery methods, especially for the challenging cases of autocorrelated and non-stationary time series.
Georgios Koutroulis, Leo Botler, Belgin Mutlu, Konrad Diwold, Kay Römer, Roman Kern
ACM Trans. Intell. Syst. Technol.6
2019 Chatbots Assisting German Business Management Applications
Florian Steinbauer, Roman Kern, Mark Kröll
IEA/AIE2
2019 A Health Factor for Process Patterns Enhancing Semiconductor Manufacturing by Pattern Recognition in Analog Wafermaps
abstract
Electrical measurement data at the end of semi-conductor frontend production, so-called wafer test data, provide deep insight into the preceding manufacturing process. Patterns in these datasets, such as spatial regularities on the wafer, frequently indicate that deviations occurred during production, potentially leading to failures in the produced devices. As such patterns of interest differ w.r.t. their shapes and equally important their intensities, pattern recognition is challenging, but crucial as a prerequisite for production environments in Industry 4.0. In this work, we propose an indicator for the presence and development of process patterns, a so-called “Health Factor for Process Patterns”, embedded in a framework of statistical decision theory. We provide adequate machine learning components, focusing on the recognition and assessment of known patterns in analog wafer test data. Finally, we conduct experiments using simulated as well as real-world datasets to demonstrate that our method yields competitive results and can be extended to a decision support system for industrial usage.
Stefan Schrunner, Anna Jenul, Michael Scheiber, Anja Zernig, Andre Kästner, Roman Kern
SMC6
2019 Self- and Cross-Excitation in Stack Exchange Question & Answer Communities
abstract
In this paper, we quantify the impact of self- and cross-excitation on the temporal development of user activity in Stack Exchange Question & Answer (Q&A) communities. We study differences in user excitation between growing and declining Stack Exchange communities, and between those dedicated to STEM and humanities topics by leveraging Hawkes processes. We find that growing communities exhibit early stage, high cross-excitation by a small core of power users reacting to the community as a whole, and strong long-term self-excitation in general and cross-excitation by casual users in particular, suggesting community openness towards less active users. Further, we observe that communities in the humanities exhibit long-term power user cross-excitation, whereas in STEM communities activity is more evenly distributed towards casual user self-excitation. We validate our findings via permutation tests and quantify the impact of these excitation effects with a range of prediction experiments. Our work enables researchers to quantitatively assess the evolution and activity potential of Q&A communities.
Simon Walk, Roman Kern, Markus Strohmaier, Denis Helic
WWW3
2019 SAZED: parameter-free domain-agnostic season length estimation in time series data
abstract
Season length estimation is the task of identifying the number of observations in the dominant repeating pattern of seasonal time series data. As such, it is a common pre-processing task crucial for various downstream applications. Inferring season length from a real-world time series is often challenging due to phenomena such as slightly varying period lengths and noise. These issues may, in turn, lead practitioners to dedicate considerable effort to preprocessing of time series data since existing approaches either require dedicated parameter-tuning or their performance is heavily domain-dependent. Hence, to address these challenges, we propose SAZED: spectral and average autocorrelation zero distance density. SAZED is a versatile ensemble of multiple, specialized time series season length estimation approaches. The combination of various base methods selected with respect to domain-agnostic criteria and a novel seasonality isolation technique, allow a broad applicability to real-world time series of varied properties. Further, SAZED is theoretically grounded and parameter-free, with a computational complexity of $$\mathcal {O}(n\log n)$$ , which makes it applicable in practice. In our experiments, SAZED was statistically significantly better than every other method on at least one dataset. The datasets we used for the evaluation consist of time series data from various real-world domains, sterile synthetic test cases and synthetic data that were designed to be seasonal and yet have no finite statistical moments of any order.
Maximilian Toller, Roman Kern
Data Min. Knowl. Discov.3
2018 Understanding wafer patterns in semiconductor production with variational auto-encoders
Roman Kern
ESANN2
2018 A Comparison of Supervised Approaches for Process Pattern Recognition in Analog Semiconductor Wafer Test Data
abstract
The semiconductor industry is currently leveraging to exploit machine learning techniques to improve and automate the manufacturing process. An essential step is the wafer test, where each single device is measured electrically, resulting in an image of the wafer. Our work is based on the hypothesis that deviations of production processes can be detected via spatial patterns on these wafermaps. Supervised learning methods are one possibility to recognize such patterns in an automated way - however, the training sample size is very low. In our work, we present and compare several methods for multiclass classification, which can deal with this limitation: multiclass decision trees, as well as decomposition methods like round robin and error-correcting output coding (ECOC). As elementary classifiers, we compare binary decision trees and logistic regression using an elastic net regularization. The evaluation shows that the decomposition methods outperform the multiclass decision tree regarding both, accuracy and practical demands.
Stefan Schrunner, Olivia Bluder, Anja Zernig, Andre Kästner, Roman Kern
ICMLA5
2017 Big data as a promoter of industry 4.0: Lessons of the semiconductor industry
abstract
The catchphrase “Industry 4.0” is widely regarded as a methodology for succeeding in modern manufacturing. This paper provides an overview of the history, technologies and concepts of Industry 4.0. One of the biggest challenges to implementing the Industry 4.0 paradigms in manufacturing are the heterogeneity of system landscapes and integrating data from various sources, such as different suppliers and different data formats. These issues have been addressed in the semiconductor industry since the early 1980s and some solutions have become well-established standards. Hence, the semiconductor industry can provide guidelines for a transition towards Industry 4.0 in other manufacturing domains. In this work, the methodologies of Industry 4.0, cyber-physical systems and Big data processes are discussed. Based on a thorough literature review and experiences from the semiconductor industry, we offer implementation recommendations for Industry 4.0 using the manufacturing process of an electronics manufacturer as an example.
David Cemernek, Heimo Gursch, Roman Kern
INDIN3
2016 Do Ambiguous Words Improve Probing for Federated Search?
Günter Urak, Hermann Ziak, Roman Kern
TPDL3
2016 Generating Tailored Classification Schemas for German Patents
Oliver Pimas, Stefan Klampfl, Thomas Kohl, Roman Kern, Mark Kröll
NLDB4
2016 An Information Retrieval Based Approach for Multilingual Ontology Matching
Andi Rexha, Mauro Dragoni, Roman Kern, Mark Kröll
NLDB3
2013 An Unsupervised Machine Learning Approach to Body Text and Table of Contents Extraction from Digital Scientific Articles
Stefan Klampfl, Roman Kern
TPDL2
2012 Evaluation of Folksonomy Induction Algorithms
abstract
Algorithms for constructing hierarchical structures from user-generated metadata have caught the interest of the academic community in recent years. In social tagging systems, the output of these algorithms is usually referred to as folksonomies (from folk-generated taxonomies). Evaluation of folksonomies and folksonomy induction algorithms is a challenging issue complicated by the lack of golden standards, lack of comprehensive methods and tools as well as a lack of research and empirical/simulation studies applying these methods. In this article, we report results from a broad comparative study of state-of-the-art folksonomy induction algorithms that we have applied and evaluated in the context of five social tagging systems. In addition to adopting semantic evaluation techniques, we present and adopt a new technique that can be used to evaluate the usefulness of folksonomies for navigation . Our work sheds new light on the properties and characteristics of state-of-the-art folksonomy induction algorithms and introduces a new pragmatic approach to folksonomy evaluation, while at the same time identifying some important limitations and challenges of folksonomy evaluation. Our results show that folksonomy induction algorithms specifically developed to capture intuitions of social tagging systems outperform traditional hierarchical clustering techniques. To the best of our knowledge, this work represents the largest and most comprehensive evaluation study of state-of-the-art folksonomy induction algorithms to date.
Markus Strohmaier, Denis Helic, Dominik Benz, Christian Körner, Roman Kern
ACM Trans. Intell. Syst. Technol.5
2012 Understanding why users tag: A survey of tagging motivation literature and results from an empirical study
abstract
While recent progress has been achieved in understanding the structure and dynamics of social tagging systems, we know little about the underlying user motivations for tagging, and how they influence resulting folksonomies and tags. This paper addresses three issues related to this question. (1) What distinctions of user motivations are identified by previous research, and in what ways are the motivations of users amenable to quantitative analysis? (2) To what extent does tagging motivation vary across different social tagging systems? (3) How does variability in user motivation influence resulting tags and folksonomies? In this paper, we present measures to detect whether a tagger is primarily motivated by categorizing or describing resources, and apply these measures to datasets from seven different tagging systems. Our results show that (a) users’ motivation for tagging varies not only across, but also within tagging systems, and that (b) tag agreement among users who are motivated by categorizing resources is significantly lower than among users who are motivated by describing resources. Our findings are relevant for (1) the development of tag-based user interfaces, (2) the analysis of tag semantics and (3) the design of search algorithms for social tagging systems.
Markus Strohmaier, Christian Körner, Roman Kern
J. Web Semant.3
2010 Why do Users Tag? Detecting Users' Motivation for Tagging in Social Tagging Systems
Markus Strohmaier, Christian Körner, Roman Kern
ICWSM3
2010 Analysis of structural relationships for hierarchical cluster labeling
abstract
Cluster label quality is crucial for browsing topic hierarchies obtained via document clustering. Intuitively, the hierarchical structure should influence the labeling accuracy. However, most labeling algorithms ignore such structural properties and therefore, the impact of hierarchical structures on the labeling accuracy is yet unclear. In our work we integrate hierarchical information, i.e. sibling and parent-child relations, in the cluster labeling process. We adapt standard labeling approaches, namely Maximum Term Frequency, Jensen-Shannon Divergence, Chi Square Test, and Information Gain, to take use of those relationships and evaluate their impact on 4 different datasets, namely the Open Directory Project, Wikipedia, TREC Ohsumed and the CLEF IP European Patent dataset. We show, that hierarchical relationships can be exploited to increase labeling accuracy especially on high-level nodes.
Markus Muhr, Roman Kern, Michael Granitzer
SIGIR2
2009 Efficient linear text segmentation based on information retrieval techniques
abstract
The task of linear text segmentation is to split a large text document into shorter fragments, usually blocks of consecutive sentences. The algorithms that demonstrated the best performance for this task come at the price of high computational complexity. In our work we present an algorithm that has a computational complexity of O(n) with n being the number of sentences in a document. The performance of our approach is evaluated against algorithms of higher complexity using standard benchmark data sets and we demonstrate that our approach provides comparable accuracy.
Roman Kern, Michael Granitzer
MEDES1
2008 Recommending Tags for Pictures Based on Text, Visual Content and User Context
abstract
Abstract—Imagine you are member of an online social system and want to upload a picture into the community pool. In current social software systems, you can probably tag your photo, share it or send it to a photo printing service and multiple other stuff. The system creates around you a space full of pictures, other interesting content (descriptions, comments) and full of users as well. The one thing current systems do not do, is understand what your pictures are about. We present here a collection of functionalities that make a step in that direction when put together to be consumed by a tag recommendation system for pictures. We use the data richness inherent in social online environments for recommending tags by analysing different aspects of the same data (text, visual contentand user context). We also give an assessment of the quality of thus recommended tags.
Stefanie N. Lindstaedt, Viktoria Pammer-Schindler, Roland Mörzinger, Roman Kern, Helmut Mülner, Claudia Wagner 0001
ICIW4