Christoph Jansen

dblp:148/6871 · DBLP profile ↗
← Back
25ranked-venue papers
15as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 first-author · 11 since 2021Systems, architecture and hardware · 9 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Empirical decision theory
Christoph Jansen, Georg Schollmeyer, Thomas Augustin 0001, Julian Rodemann
Inf. Sci.1
2025 Consensus in Motion: A Case of Dynamic Rationality of Sequential Learning in Probability Aggregation
Polina Gordienko, Christoph Jansen, Thomas Augustin 0001, Martin Rechenauer
ECSQARU2
2025 Statistical Multicriteria Evaluation of LLM-Generated Text
abstract
Assessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplistic aggregations that fail to capture the nuanced trade-offs between coherence, diversity, fluency, and other relevant indicators of text quality. In this work, we adapt a recently proposed framework for statistical inference based on Generalized Stochastic Dominance (GSD) that addresses three critical limitations in existing benchmarking methodologies: the inadequacy of single-metric evaluation, the incompatibility between cardinal automatic metrics and ordinal human judgments, and the lack of inferential statistical guarantees. The GSD-front approach enables simultaneous evaluation across multiple quality dimensions while respecting their different measurement scales, building upon partial orders of decoding strategies, thus avoiding arbitrary weighting of the involved metrics. By applying this framework to evaluate common decoding strategies against human-generated text, we demonstrate its ability to identify statistically significant performance differences while accounting for potential deviations from the i.i.d. assumption of the sampling design.
Esteban Garces Arias, Hannah Blocher, Julian Rodemann, Matthias Aßenmacher, Christoph Jansen
INLG5
2024 Statistical Multicriteria Benchmarking via the GSD-Front
abstract
Given the vast number of classifiers that have been (and continue to be) proposed, reliable methods for comparing them are becoming increasingly important. The desire for reliability is broken down into three main aspects: (1) Comparisons should allow for different quality metrics simultaneously. (2) Comparisons should take into account the statistical uncertainty induced by the choice of benchmark suite. (3) The robustness of the comparisons under small deviations in the underlying assumptions should be verifiable. To address (1), we propose to compare classifiers using a generalized stochastic dominance ordering (GSD) and present the GSD-front as an information-efficient alternative to the classical Pareto-front. For (2), we propose a consistent statistical estimator for the GSD-front and construct a statistical test for whether a (potentially new) classifier lies in the GSD-front of a set of state-of-the-art classifiers. For (3), we relax our proposed test using techniques from robust statistics and imprecise probabilities. We illustrate our concepts on the benchmark suite PMLB and on the platform OpenML.
Christoph Jansen, Georg Schollmeyer, Julian Rodemann, Hannah Blocher, Thomas Augustin 0001
NeurIPS1
2024 Reciprocal Learning
abstract
We demonstrate that numerous machine learning algorithms are specific instances of one single paradigm: reciprocal learning. These instances range from active learning over multi-armed bandits to self-training. We show that all these algorithms not only learn parameters from data but also vice versa: They iteratively alter training data in a way that depends on the current model fit. We introduce reciprocal learning as a generalization of these algorithms using the language of decision theory. This allows us to study under what conditions they converge. The key is to guarantee that reciprocal learning contracts such that the Banach fixed-point theorem applies. In this way, we find that reciprocal learning converges at linear rates to an approximately optimal model under some assumptions on the loss function, if their predictions are probabilistic and the sample adaption is both non-greedy and either randomized or regularized. We interpret these findings and provide corollaries that relate them to active learning, self-training, and bandits.
Julian Rodemann, Christoph Jansen, Georg Schollmeyer
NeurIPS2
2024 Comparing machine learning algorithms by union-free generic depth
abstract
We propose a framework for descriptively analyzing sets of partial orders based on the concept of depth functions. Despite intensive studies in linear and metric spaces, there is very little discussion on depth functions for non-standard data types such as partial orders. We introduce an adaptation of the well-known simplicial depth to the set of all partial orders, the union-free generic (ufg) depth. Moreover, we utilize our ufg depth for a comparison of machine learning algorithms based on multidimensional performance measures. Concretely, we provide two examples of classifier comparisons on samples of standard benchmark data sets. Our results demonstrate promisingly the wide variety of different analysis approaches based on ufg methods. Furthermore, the examples outline that our approach differs substantially from existing benchmarking approaches, and thus adds a new perspective to the vivid debate on classifier comparison.1
Hannah Blocher, Georg Schollmeyer, Malte Nalenz, Christoph Jansen
Int. J. Approx. Reason.4
2023 Multi-target Decision Making Under Conditions of Severe Uncertainty
Christoph Jansen, Georg Schollmeyer, Thomas Augustin 0001
MDAI1
2023 Robust statistical comparison of random variables with locally varying scale of measurement
abstract
Spaces with locally varying scale of measurement, like multidimensional structures with differently scaled dimensions, are pretty common in statistics and machine learning. Nevertheless, it is still understood as an open question how to exploit the entire information encoded in them properly. We address this problem by considering an order based on (sets of) expectations of random variables mapping into such non-standard spaces. This order contains stochastic dominance and expectation order as extreme cases when no, or respectively perfect, cardinal structure is given. We derive a (regularized) statistical test for our proposed generalized stochastic dominance (GSD) order, operationalize it by linear optimization, and robustify it by imprecise probability models. Our findings are illustrated with data from multidimensional poverty measurement, finance, and medicine.
Christoph Jansen, Georg Schollmeyer, Hannah Blocher, Julian Rodemann, Thomas Augustin 0001
UAI1
2023 The vendor-agnostic EMPAIA platform for integrating AI applications into digital pathology infrastructures
abstract
209
Christoph Jansen, Björn Lindequist, Klaus Strohmenger, Daniel Romberg, Tobias Küster, Nick Weiss, Michael Franz, Lars Ole Schwen, Theodore Evans, André Homeyer, Norman Zerbe
Future Gener. Comput. Syst.1
2023 Statistical Comparisons of Classifiers by Generalized Stochastic Dominance
abstract
Although being a crucial question for the development of machine learning algorithms, there is still no consensus on how to compare classifiers over multiple data sets with respect to several criteria. Every comparison framework is confronted with (at least) three fundamental challenges: the multiplicity of quality criteria, the multiplicity of data sets and the randomness of the selection of data sets. In this paper, we add a fresh view to the vivid debate by adopting recent developments in decision theory. Based on so-called preference systems, our framework ranks classifiers by a generalized concept of stochastic dominance, which powerfully circumvents the cumbersome, and often even self-contradictory, reliance on aggregates. Moreover, we show that generalized stochastic dominance can be operationalized by solving easy-to-handle linear programs and moreover statistically tested employing an adapted two-sample observation-randomization test. This yields indeed a powerful framework for the statistical comparison of classifiers over multiple data sets with respect to multiple quality criteria simultaneously. We illustrate and investigate our framework in a simulation study and with a set of standard benchmark data sets.
Christoph Jansen, Malte Nalenz, Georg Schollmeyer, Thomas Augustin 0001
J. Mach. Learn. Res.1
2022 The EMPAIA Platform: Vendor-neutral integration of AI applications into digital pathology infrastructures
abstract
Automated image analysis and artificial intelligence (AI) are a growing market in digital pathology. While various proprietary pathology systems exist, there are no fully vendor-agnostic integration approaches for AI apps. This makes it difficult for vendors of AI solutions to integrate their products into the multitude of non-standard software systems in pathology. The EMPAIA Consortium (EcosysteM for Pathology Diagnostics with AI Assistance) develops an open and decentralized platform allowing AI-based apps of different vendors to be integrated with existing lab IT infrastructures. This is intended to lower the barriers to entry for AI vendors and provide pathologists with access to advanced AI tools. The EMPAIA platform is based on web technologies that can be deployed both on-premises and in the cloud. There are open-source reference implementations for core platform services that can be integrated with or replaced by proprietary alternatives as long as they conform to open API specifications. Apps can be obtained through a central marketplace so pathologists can use them in their daily workflow. In this paper, we provide an overview of the EMPAIA platform architecture. We identify critical use cases and requirements for AI-based software platforms in pathology and explain how these are fulfilled by the EMPAIA platform. Finally, we evaluate the efficiency of routing image data through the platform.
Christoph Jansen, Klaus Strohmenger, Daniel Romberg, Tobias Küster, Nick Weiss, Björn Lindequist, Michael Franz, André Homeyer, Norman Zerbe
CCGRID1
2022 Statistical Models for Partial Orders Based on Data Depth and Formal Concept Analysis
Hannah Blocher, Georg Schollmeyer, Christoph Jansen
IPMU (2)3
2022 Decision Making with State-Dependent Preference Systems
Christoph Jansen, Thomas Augustin 0001
IPMU (1)1
2022 Information efficient learning of complexly structured preferences: Elicitation procedures and their application to decision making under uncertainty
Christoph Jansen, Hannah Blocher, Thomas Augustin 0001, Georg Schollmeyer
Int. J. Approx. Reason.1
2020 Curious Containers: A framework for computational reproducibility in life sciences with support for Deep Learning applications
Christoph Jansen, Jonas Annuscheit, Bruno Schilling, Klaus Strohmenger, Michael Witt 0001, Felix Bartusch, Christian Herta, Peter Hufnagl, Dagmar Krefting
Future Gener. Comput. Syst.1
2019 Reproducibility and Performance of Deep Learning Applications for Cancer Detection in Pathological Images
abstract
Convolutional Neural Networks (CNN) are used for automatic cancer detection in pathological images. These data-driven experiments are difficult to reproduce, because the CNNs may require CUDA-enabled Nvidia GPUs for acceleration and training is often performed on a large dataset stored on a researcher's computer, inaccessible to others. We introduce the RED file format for reproducible experiment description, where executable programs are packaged and referenced as Docker container images. Data inputs and outputs are described as network resources using standard transmission and authentication protocols instead of local file paths. Following the FAIR guiding principles, the RED format is based on and compatible with the established Common Workflow Language specification. RED files are interpreted by the accompanying Curious Containers (CC) software. Arbitrarily large datasets are mounted inside containers via FUSE network filesystems like SSHFS. SSHFS is compared to NFS and a local SSD in artificial benchmarks and in the context of a CNN training scenario, where SSHFS introduces a performance decrease by a factor of 1.8. We are convinced that RED can greatly improve the reproducibility of deep learning workloads and data-driven experiments. This is in particular important in clinical scenarios where the result of an analysis may contribute to a patient's treatment.
Christoph Jansen, Bruno Schilling, Klaus Strohmenger, Michael Witt 0001, Jonas Annuscheit, Dagmar Krefting
CCGRID1
2018 Sandboxing of biomedical applications in Linux containers based on system call evaluation
abstract
Summary Applications for biomedical data processing often integrate external libraries and frameworks for common algorithmic tasks. It typically reduces development time and increases overall code quality. With the introduction of lightweight container‐based virtualization, the bundling of applications and their required dependencies has become feasible, and containers can be transferred and executed in distributed environments. However, the incorporation of unreviewed code poses a security threat as it might contain malicious components. In this paper, measures to minimize risks of untrusted application execution are presented. Based on the system calls issued during sample execution of the application, both the container itself and the container runtime configuration are restricted to the set of actions the application requires. It is shown that the employed security measures are suited to counteract different attacks while application runtime is not affected.
Michael Witt 0001, Christoph Jansen, Dagmar Krefting, Achim Streit
Concurr. Comput. Pract. Exp.2
2018 Concepts for decision making under severe uncertainty with partial ordinal and partial cardinal preferences
Christoph Jansen, Georg Schollmeyer, Thomas Augustin 0001
Int. J. Approx. Reason.1
2017 Fine-grained Supervision and Restriction of Biomedical Applications in Linux Containers
abstract
Applications for data analysis of biomedical data are complex programs and often consist of multiple components. Re-usage of existing solutions from external code repositories or program libraries is common in algorithm development. To ease reproducibility as well as transfer of algorithms and required components into distributed infrastructures Linux containers are increasingly used in those environments, that are at least partly connected to the internet. However concerns about the untrusted application remain and are of high interest when medical data is processed. Additionally, the portability of the containers needs to be ensured by using only security technologies, that do not require additional kernel modules. In this paper we describe measures and a solution to secure the execution of an example biomedical application for normalization of multidimensional biosignal recordings. This application, the required runtime environment and the security mechanisms are installed in a Docker-based container. A fine-grained restricted environment (sandbox) for the execution of the application and the prevention of unwanted behaviour is created inside the container. The sandbox is based on the filtering of system calls, as they are required to interact with the operating system to access potentially restricted resources e.g. the filesystem or network. Due to the low-level character of system calls, the creation of an adequate rule set for the sandbox is challenging. Therefore the presented solution includes a monitoring component to collect required data for defining the rules for the application sandbox. Performance evaluation of the application execution shows no significant impact of the resulting sandbox, while detailed monitoring may increase runtime up to over 420%.
Michael Witt 0001, Christoph Jansen, Dagmar Krefting, Achim Streit
CCGrid2
2017 Decision Theory Meets Linear Optimization Beyond Computation
Christoph Jansen, Thomas Augustin 0001, Georg Schollmeyer
ECSQARU1
2017 Multicenter data sharing for collaboration in sleep medicine
Maximilian Beier, Christoph Jansen, Geert Mayer, Thomas Penzel, Andrea Rodenbeck, René Siewert, Michael Witt 0001, Jie Wu 0014, Dagmar Krefting
Future Gener. Comput. Syst.2
2016 Employing Docker Swarm on OpenStack for Biomedical Analysis
Christoph Jansen, Michael Witt 0001, Dagmar Krefting
ICCSA (2)1
2015 Multicenter Data Sharing for Collaboration in Sleep Medicine
abstract
Clinical Sleep Research is an inherent multidisciplinary field, as many health issues may affect a person's sleep conditions and sleep disorders may cause several health problems. Many patients with chronic sleep disorders suffer from different further medical conditions - called multimorbidity. Due to the high variety of the reasons and the courses of sleep disorders, individual cases are difficult to compare. Therefore there is a high demand for sleep researchers to collaborate with each other to reach necessary participant numbers and multidisciplinary expertise. To date, inter-institutional sleep research is poorly supported by IT systems. In particular the heterogeneity and the quality variations within the acquired bio signal data - caused by different bio signal recorders or different measurement procedures - are impeding common bio signal data processing. In this manuscript we introduce a virtual research platform supporting inter-institutional data sharing and processing. The infrastructure is based on XNAT - a free and open-source neuroimaging research platform - a loosely coupled service oriented architecture and scalable virtualization in the backend. The system is capable of local pseudonymization of bio signal data, mapping to a standardized set of parameters and automatic quality assessment. Terms and quality measures are derived from the "Manual for the Scoring of Sleep and Associated Events" of the American Academy of Sleep Medicine, the de-facto standard for diagnostic bio signal analysis in sleep medicine.
Maximilian Beier, Christoph Jansen, Geert Mayer, Thomas Penzel, Andrea Rodenbeck, René Siewert, Jie Wu 0014, Dagmar Krefting
CCGRID2
2015 Reconstructing Missing Areas in Facial Images
abstract
In this paper, we present a novel approach to reconstruct missing areas in facial images by using a series of Restricted Boltzman Machines (RBMs). RBMs created with a low number of hidden neurons generalize well and are able to reconstruct basic structures in the missing areas. On the other hand networks with many hidden neurons tend to emphasize details, when using the reconstruction of the previous, more generalized RBMs, as their input. Since trained RBMs are fast in encoding and decoding data by design, our method is also suitable for processing video streams.
Christoph Jansen, Radek Mackowiak, Nico Hezel, Moritz Ufer, Gregor Altstadt, Kai Uwe Barthel
ISM1
2014 Extending XNAT towards a Cloud-Based Quality Assessment Platform for Retinal Optical Coherence Tomographies
abstract
Neurosciencific research is increasingly based on image analysis methods. Large sets of imaging data are processed using complex image analysis tools. While today magnetic resonance imaging (MRI) is widely used for both functional and anatomical analysis of the human brain, new imaging modalities are beginning to prove their capabilities for neurological research. Among them, optical coherence tomography (OCT) allows for noninvasive visualization of anatomical structures on a micrometer scale. Becoming a standard diagnostic tool in ophthalmology, it is of rising interest for neurological research. Crucial to all data analysis methods is the quality of the input data. The platform presented in this paper is designed for automatic quality assessment of retinal OCTs. It extends the image management platform XNAT by services to calculate and store quality measures. It is also extensible regarding new quality measure algorithms, allowing the developer to upload Matlab code, compile it for the infrastructure's hardware architecture and test it in the system. The image processing tools to calculate the quality measures are provided as a cloud-based service employing Open Stack as underlying IT infrastructure. The prototype implementation encompassing security and performance aspects are presented.
Jie Wu 0014, Christoph Jansen, Maximilian Beier, Michael Witt 0001, Dagmar Krefting
CCGRID2