Johannes Textor

dblp:63/4217 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0002-0459-9458ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Classifier Systems as Linear Probability Models
abstract
Classifier systems solve regression and classification problems in high-dimensional spaces by generating and evolving large populations of simple rules; examples include learning classifier systems and artificial immune systems. The properties of these flexible adaptable systems are less well understood than those of more classical machine learning algorithms. Here, we reveal a deep connection between classifier systems and probabilistic models such as Naïve Bayes and Markov chains by showing that all of these can be expressed as generalized linear probability models. This connection shows that any probability distribution can in principle be expressed by a classifier system. We then harness this new perspective to investigate the tradeoff between model complexity and calibration — i.e., the ability to accurately fit the sequence probabilities observed in the training set — for classifier systems applied to sequence probability modeling. Contrasting our results to Markov chains of varying order, we find that a simple model classifier system has a broadly similar complexity-calibration tradeoff. We hope that our approach paves the way for further systematic investigation of the fundamental properties of classifier systems, which could make them more accessible for the machine learning community.
Gijs Schröder, Johannes Textor
GECCO2
2025 Expert-In-The-Loop Causal Discovery: Iterative Model Refinement Using Expert Knowledge
abstract
Many researchers construct directed acyclic graph (DAG) models manually based on domain knowledge. Although numerous causal discovery algorithms were developed to automatically learn DAGs and other causal models from data, these remain challenging to use due to their tendency to produce results that contradict domain knowledge, among other issues. Here we propose a hybrid, iterative structure learning approach that combines domain knowledge with data-driven insights to assist researchers in constructing DAGs. Our method leverages conditional independence testing to iteratively identify variable pairs where an edge is either missing or superfluous. Based on this information, we can choose to add missing edges with appropriate orientation based on domain knowledge or remove unnecessary ones. We also give a method to rank these missing edges based on their impact on the overall model fit. In a simulation study, we find that this iterative approach to leverage domain knowledge already starts outperforming purely data-driven structure learning if the orientation of new edge is correctly determined in at least two out of three cases. We present a proof-of-concept implementation using a large language model as a domain expert and a graphical user interface designed to assist human experts with DAG construction.
Ankur Ankan, Johannes Textor
UAI2
2024 Fitting Stochastic Lattice Models Using Approximate Gradients
abstract
Stochastic lattice models (sLMs) are computational tools for simulating spatiotemporal dynamics in physics, computational biology, chemistry, ecology, and other fields. Despite their widespread use, it is challenging to fit sLMs to data, as their likelihood function is commonly intractable and the models non-differentiable. The adjacent field of agent-based modelling (ABM), faced with similar challenges, has recently introduced an approach to approximate gradients in network-controlled ABMs via reparameterization tricks. This approach enables efficient gradient-based optimization with automatic differentiation (AD), which allows for a directed local search of suitable parameters rather than estimation via black-box sampling. In this study, we investigate the feasibility of using similar reparameterization tricks to fit sLMs through backpropagation of approximate gradients. We consider three common scenarios: fitting to single-state transitions, fitting to trajectories, and identification of stable lattice configurations. We demonstrate that all tasks can be solved by AD using three example sLMs from sociology, biophysics, and physical chemistry. Our results show that AD via approximate gradients is a promising method to fit sLMs to data for a wide variety of models and tasks.
an Schering, Sander Keemink, Johannes Textor
ECMS3
2024 A Uniformly Bounded Correlation Function for Spatial Point Patterns
abstract
A point pattern is a dataset of coordinates, typically in 2D or 3D space.Point patterns are ubiquitous in diverse applications including Geographic Information Systems, Astronomy, Ecology, Biology and Medicine.Among the statistics used to quantify point patterns, most are based on Ripley's 𝐾-function, which measures the deviation of the observed pattern from a completely random arrangement of points.This approach is useful for constructing null hypothesis tests, but Ripley's 𝐾 and its variants are less suitable as quantitative effect sizes because their ranges and expected values generally depend on the scale or the size of the region in which the pattern is observed.To address this, we propose a new function that behaves like a correlation coefficient for point patterns: it is tightly bounded by -1 and 1, with a value of -1 corresponding to a maximally dispersed arrangement of points, 0 indicating complete spatial randomness, and 1 representing maximal clustering.These properties are independent of scale and observation window size assuming appropriate edge correction.Evaluating our function on simulated data, we show that it has comparable statistical calibration and power to 𝐾-based baselines.We hope that the ease of interpretation of our bounded function will facilitate the analysis of spatial data across multiple fields.
Evgenia Martynova, Johannes Textor
KDD2
2024 Population-Based Algorithms Built on Weighted Automata
Gijs Schröder, Inge M. N. Wortel, Johannes Textor
PPSN (3)3
2024 pgmpy: A Python Toolkit for Bayesian Networks
abstract
Bayesian Networks (BNs) are used in various fields for modeling, prediction, and decision making. pgmpy is a python package that provides a collection of algorithms and tools to work with BNs and related models. It implements algorithms for structure learning, parameter estimation, approximate and exact inference, causal inference, and simulations. These implementations focus on modularity and easy extensibility to allow users to quickly modify/add to existing algorithms, or to implement new algorithms for different use cases. pgmpy is released under the MIT License; the source code is available at: https://github.com/pgmpy/pgmpy, and the documentation at: https://pgmpy.org.
Ankur Ankan, Johannes Textor
J. Mach. Learn. Res.2
2023 A Simple Unified Approach to Testing High-Dimensional Conditional Independences for Categorical and Ordinal Data
abstract
Conditional independence (CI) tests underlie many approaches to model testing and structure learning in causal inference. Most existing CI tests for categorical and ordinal data stratify the sample by the conditioning variables, perform simple independence tests in each stratum, and combine the results. Unfortunately, the statistical power of this approach degrades rapidly as the number of conditioning variables increases. Here we propose a simple unified CI test for ordinal and categorical data that maintains reasonable calibration and power in high dimensions. We show that our test outperforms existing baselines in model testing and structure learning for dense directed graphical models while being comparable for sparse models. Our approach could be attractive for causal model testing because it is easy to implement, can be used with non-parametric or parametric probability models, has the symmetry property, and has reasonable computational requirements.
Ankur Ankan, Johannes Textor
AAAI2
2023 Combining Graphical and Algebraic Approaches for Parameter Identification in Latent Variable Structural Equation Models
abstract
Measurement error is ubiquitous in many variables “latent-to-observed” (L2O) transformation from the MIIV approach and develop an equivalent graphical L2O transformation that allows applying existing graphical criteria to latent parameters in SEMs. We combine L2O transformation with graphical instrumental variable criteria to obtain an efficient algorithm for non-iterative parameter identification in SEMs with latent variables. We prove that this graphical L2O transformation with the instrumental set criterion is equivalent to the state-of-the-art MIIV approach for SEMs, and show that it can lead to novel identification strategies when combined with other graphical criteria.
Ankur Ankan, Inge M. N. Wortel, Kenneth Bollen, Johannes Textor
AISTATS4
2023 Corrigendum to "Separators and adjustment sets in causal graphs: Complete criteria and an algorithmic framework" [Artif. Intell. 270 (2019) 1-40]
Benito van der Zander, Maciej Liskiewicz, Johannes Textor
Artif. Intell.3
2023 Interpreting T-cell search "strategies" in the light of evolution under constraints
abstract
Two decades of in vivo imaging have revealed how diverse T-cell motion patterns can be. Such recordings have sparked the notion of search "strategies": T cells may have evolved ways to search for antigen efficiently depending on the task at hand. Mathematical models have indeed confirmed that several observed T-cell migration patterns resemble a theoretical optimum; for example, frequent turning, stop-and-go motion, or alternating short and long motile runs have all been interpreted as deliberately tuned behaviours, optimising the cell's chance of finding antigen. But the same behaviours could also arise simply because T cells cannot follow a straight, regular path through the tight spaces they navigate. Even if T cells do follow a theoretically optimal pattern, the question remains: which parts of that pattern have truly been evolved for search, and which merely reflect constraints from the cell's migration machinery and surroundings? We here employ an approach from the field of evolutionary biology to examine how cells might evolve search strategies under realistic constraints. Using a cellular Potts model (CPM), where motion arises from intracellular dynamics interacting with cell shape and a constraining environment, we simulate evolutionary optimization of a simple task: explore as much area as possible. We find that our simulated cells indeed evolve their motility patterns. But the evolved behaviors are not shaped solely by what is functionally optimal; importantly, they also reflect mechanistic constraints. Cells in our model evolve several motility characteristics previously attributed to search optimisation-even though these features are not beneficial for the task given here. Our results stress that search patterns may evolve for other reasons than being "optimal". In part, they may be the inevitable side effects of interactions between cell shape, intracellular dynamics, and the diverse environments T cells face in vivo.
Inge M. N. Wortel, Johannes Textor
PLoS Comput. Biol.2
2019 Separators and adjustment sets in causal graphs: Complete criteria and an algorithmic framework
Benito van der Zander, Maciej Liskiewicz, Johannes Textor
Artif. Intell.3
2017 Complete Graphical Characterization and Construction of Adjustment Sets in Markov Equivalence Classes of Ancestral Graphs
Emilija Perkovic, Johannes Textor, Markus Kalisch, Marloes H. Maathuis
J. Mach. Learn. Res.2
2015 Efficiently Finding Conditional Instruments for Causal Inference
Benito van der Zander, Johannes Textor, Maciej Liskiewicz
IJCAI2
2015 A Complete Generalized Adjustment Criterion
Emilija Perkovic, Johannes Textor, Markus Kalisch, Marloes H. Maathuis
UAI2
2015 Learning from Pairwise Marginal Independencies
Johannes Textor, Alexander Idelberger, Maciej Liskiewicz
UAI1
2015 Crawling and Gliding: A Computational Model for Shape-Driven Cell Migration
abstract
Cell migration is a complex process involving many intracellular and extracellular factors, with different cell types adopting sometimes strikingly different morphologies. Modeling realistically behaving cells in tissues is computationally challenging because it implies dealing with multiple levels of complexity. We extend the Cellular Potts Model with an actin-inspired feedback mechanism that allows small stochastic cell rufflings to expand to cell protrusions. This simple phenomenological model produces realistically crawling and deforming amoeboid cells, and gliding half-moon shaped keratocyte-like cells. Both cell types can migrate randomly or follow directional cues. They can squeeze in between other cells in densely populated environments or migrate collectively. The model is computationally light, which allows the study of large, dense and heterogeneous tissues containing cells with realistic shapes and migratory properties.
Ioana Niculescu, Johannes Textor, Rob J. De Boer
PLoS Comput. Biol.2
2014 A generic finite automata based approach to implementing lymphocyte repertoire models
abstract
Artificial immune systems (AIS) inspired by lymphocyte repertoires include negative and positive selection, clonal selection, and B~cell algorithms. Such AISs are used in computer science for machine learning and optimization, and in biology for modeling of fundamental immunological processes. In both cases, the necessary size of repertoire models can be huge. Here, we show that when lymphocyte repertoire models based on string patterns can be compactly represented as finite automata (FA), this allows to efficiently perform negative selection, positive selection, insertion into, deletion from, uniform sampling from, and counting the repertoire. Specifically, for r-contiguous pattern matching, all these tasks can be performed in polynomial time. But even in NP-hard cases like Hamming distance matching, the FA representation can still lead to practically important efficiency gains. We demonstrate the feasibility and flexibility of this approach by implementing T~cell positive selection simulations based on human genomic data using four different pattern rules. Hence, FA-based repertoire models generalize previous efficient negative selection algorithms to perform several related algorithmic tasks, are easy to implement and customize, and are applicable to real-world bioinformatic problems.
Johannes Textor, Katharina Dannenberg, Maciej Liskiewicz
GECCO1
2014 Constructing Separators and Adjustment Sets in Ancestral Graphs
Benito van der Zander, Maciej Liskiewicz, Johannes Textor
UAI3
2014 Random Migration and Signal Integration Promote Rapid and Robust T Cell Recruitment
abstract
To fight infections, rare T cells must quickly home to appropriate lymph nodes (LNs), and reliably localize the antigen (Ag) within them. The first challenge calls for rapid trafficking between LNs, whereas the second may require extensive search within each LN. Here we combine simulations and experimental data to investigate which features of random T cell migration within and between LNs allow meeting these two conflicting demands. Our model indicates that integrating signals from multiple random encounters with Ag-presenting cells permits reliable detection of even low-dose Ag, and predicts a kinetic feature of cognate T cell arrest in LNs that we confirm using intravital two-photon data. Furthermore, we obtain the most reliable retention if T cells transit through LNs stochastically, which may explain the long and widely distributed LN dwell times observed in vivo. Finally, we demonstrate that random migration, both between and within LNs, allows recruiting the majority of cognate precursors within a few days for various realistic infection scenarios. Thus, the combination of two-scale stochastic migration and signal integration is an efficient and robust strategy for T cell immune surveillance.
Johannes Textor, Sarah E. Henrickson, Judith N. Mandl, Ulrich H. von Andrian, Jürgen Westermann, Rob J. De Boer, Joost B. Beltman
PLoS Comput. Biol.1
2013 Analytical results on the Beauchemin model of lymphocyte migration
abstract
The Beauchemin model is a simple particle-based description of stochastic lymphocyte migration in tissue, which has been successfully applied to studying immunological questions. In addition to being easy to implement, the model is also to a large extent mathematically tractable. This article provides a comprehensive overview of both existing and new analytical results on the Beauchemin model within a common mathematical framework. Specifically, we derive the motility coefficient, the mean square displacement, and the confinement ratio, and discuss four different methods for simulating biased migration of pre-defined speed. The results provide new insight into published studies and a reference point for future research based on this simple and popular lymphocyte migration model.
Johannes Textor, Mathieu Sinn, Rob J. De Boer
BMC Bioinform.1
2012 Efficient Negative Selection Algorithms by Sampling and Approximate Counting
Johannes Textor
PPSN (1)1
2011 Adjustment Criteria in Causal Diagrams: An Algorithmic Perspective
Johannes Textor, Maciej Liskiewicz
UAI1
2011 Negative selection algorithms on strings with efficient training and linear-time classification
Michael Elberfeld, Johannes Textor
Theor. Comput. Sci.2
2010 Negative selection algorithms without generating detectors
abstract
Negative selection algorithms are immune-inspired classifiers that are trained on negative examples only. Classification is performed by generating detectors that match none of the negative examples, and these detectors are then matched against the elements to be classified. This can be a performance bottleneck: A large number of detectors may be required for acceptable sensitivity, or finding detectors that match none of the negative examples may be difficult. In this paper, we show how negative selection can be implemented without generating detectors explicitly, which for many detector types leads to polynomial time algorithms whereas the common approach to sample detectors randomly takes exponential time in the worst case.
Maciej Liskiewicz, Johannes Textor
GECCO2
2009 Hybrid Simulation Algorithms for an Agent-Based Model of the Immune Response
abstract
The immune system is of central interest for the life sciences, but its high complexity makes it a challenging system to study. Computational models of the immune system can help to improve our understanding of its fundamental principles. In this article, we analyze and extend the Celada-Seiden model, a simple and elegant agent-based model of the entire immune response, which, however, lacks biophysically sound simulation methodology. We extend the stochastic model to a stochastic-deterministic hybrid, and link the deterministic version to continuous physical and chemical laws. This gives precise meaning to all simulation processes, and helps to increase performance. To demonstrate an application for the model, we implement and study two different hypotheses about T cell-mediated immune memory.
Johannes Textor, Björn Hansen
Cybern. Syst.1