VLDB 2026 Research / reviewers in the wild / expert
Fabian J. Theis
dblp:t/FabianJTheis
· DBLP profile ↗
109ranked-venue papers
22as first author
22since 2021 · last 2025
0000-0002-2419-1943ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 16 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 46 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MAGNet: Motif-Agnostic Generation of Molecules from ScaffoldsabstractRecent advances in machine learning for molecules exhibit great potential for facilitating drug discovery from in silico predictions.
Most models for molecule generation rely on the decomposition of molecules into frequently occurring substructures (motifs), from which they generate novel compounds.
While motif representations greatly aid in learning molecular distributions, such methods fail to represent substructures beyond their known motif set, posing a fundamental limitation for discovering novel compounds.
To address this limitation and enhance structural expressivity, we propose to separate structure from features by abstracting motifs to scaffolds and, subsequently, allocating atom and bond types.
To this end, we introduce a novel factorisation of the molecules' data distribution that considers the entire molecular context and facilitates learning adequate assignments of atoms and bonds to scaffolds. Complementary to this, we propose MAGNet, the first model to freely learn motifs. Importantly, we demonstrate that MAGNet's improved expressivity leads to molecules with more structural diversity and, at the same time, diverse atom and bond assignments. Leon Hetzel, Johanna Sommer, Bastian Rieck, Fabian J. Theis, Stephan Günnemann |
ICLR | 4 |
| 2025 | Multi-Modal and Multi-Attribute Generation of Single Cells with CFGenabstractGenerative modeling of single-cell RNA-seq data is crucial for tasks like trajectory inference, batch effect removal, and simulation of realistic cellular data. However, recent deep generative models simulating synthetic single cells from noise operate on pre-processed continuous gene expression approximations, overlooking the discrete nature of single-cell data, which limits their effectiveness and hinders the incorporation of robust noise models. Additionally, aspects like controllable multi-modal and multi-label generation of cellular data remain underexplored. This work introduces CellFlow for Generation (CFGen), a flow-based conditional generative model that preserves the inherent discreteness of single-cell data. CFGen reliably generates whole-genome, multi-modal, single-cell data, improving the recovery of crucial biological data characteristics while tackling relevant generative tasks such as rare cell type augmentation and batch correction. We also introduce a novel framework for compositional data generation using Flow Matching. By showcasing CFGen on a diverse set of biological datasets and settings, we provide evidence of its value to the fields of computational biology and deep generative models. Alessandro Palma, Till Richter, Hanyi Zhang, Manuel Lubetzki, Alexander Tong 0001, Andrea Dittadi, Fabian J. Theis |
ICLR | 7 |
| 2025 | Disentangled Representation Learning with the Gromov-Monge GapabstractLearning disentangled representations from unlabelled data is a fundamental challenge in machine learning. Solving it may unlock other problems, such as generalization, interpretability, or fairness. Although remarkably challenging to solve in theory, disentanglement is often achieved in practice through prior matching. Furthermore, recent works have shown that prior matching approaches can be enhanced by leveraging geometrical considerations, e.g., by learning representations that preserve geometric features of the data, such as distances or angles between points. However, matching the prior while preserving geometric features is challenging, as a mapping that *fully* preserves these features while aligning the data distribution with the prior does not exist in general. To address these challenges, we introduce a novel approach to disentangled representation learning based on quadratic optimal transport. We formulate the problem using Gromov-Monge maps that transport one distribution onto another with minimal distortion of predefined geometric features, preserving them *as much as can be achieved*. To compute such maps, we propose the Gromov-Monge-Gap (GMG), a regularizer quantifying whether a map moves a reference distribution with minimal geometry distortion. We demonstrate the effectiveness of our approach for disentanglement across four standard benchmarks, outperforming other methods leveraging geometric considerations. Théo Uscidda, Luca Eyring, Karsten Roth, Fabian J. Theis, Zeynep Akata, Marco Cuturi |
ICLR | 4 |
| 2025 | Enforcing Latent Euclidean Geometry in Single-Cell VAEs for Manifold InterpolationabstractLatent space interpolations are a powerful tool for navigating deep generative models in applied settings. An example is single-cell RNA sequencing, where existing methods model cellular state transitions as latent space interpolations with variational autoencoders, often assuming linear shifts and Euclidean geometry. However, unless explicitly enforced, linear interpolations in the latent space may not correspond to geodesic paths on the data manifold, limiting methods that assume Euclidean geometry in the data representations. We introduce FlatVI, a novel training framework that regularises the latent manifold of discrete-likelihood variational autoencoders towards Euclidean geometry, specifically tailored for modelling single-cell count data. By encouraging straight lines in the latent space to approximate geodesic interpolations on the decoded single-cell manifold, FlatVI enhances compatibility with downstream approaches that assume Euclidean latent geometry. Experiments on synthetic data support the theoretical soundness of our approach, while applications to time-resolved single-cell RNA sequencing data demonstrate improved trajectory reconstruction and manifold interpolation. Alessandro Palma, Sergei Rybakov, Leon Hetzel, Stephan Günnemann, Fabian J. Theis |
ICML | 5 |
| 2025 | Modeling Microenvironment Trajectories on Spatial Transcriptomics with NicheFlowabstractUnderstanding the evolution of cellular microenvironments in spatiotemporal data is essential for deciphering tissue development and disease progression. While experimental techniques like spatial transcriptomics now enable high-resolution mapping of tissue organization across space and time, current methods that model cellular evolution operate at the single-cell level, overlooking the coordinated development of cellular states in a tissue. We introduce NicheFlow, a flow-based generative model that infers the temporal trajectory of cellular microenvironments across sequential spatial slides. By representing local cell neighborhoods as point clouds, NicheFlow jointly models the evolution of cell states and spatial coordinates using optimal transport and Variational Flow Matching. Our approach successfully recovers both global spatial architecture and local microenvironment composition across diverse spatiotemporal datasets, from embryonic to brain development. Kristiyan Sakalyan, Alessandro Palma, Filippo Guerranti, Fabian J. Theis, Stephan Günnemann |
NeurIPS | 4 |
| 2025 | TarDis: Achieving Robust and Structured Disentanglement of Multiple Covariates
Kemal Inecik, Aleyna Kara, Antony Rose, Muzlifah Haniffa, Fabian J. Theis |
RECOMB | 5 |
| 2025 | Integration and Querying of Multimodal Single-Cell Data with PoE-VAE
Anastasia Litinetskaya, Maiia Schulman, Fabiola Curion, Artur Szalata, Alireza Omidi, Mohammad Lotfollahi, Fabian J. Theis |
RECOMB | 7 |
| 2025 | QuiCAT: a scalable and flexible framework for mapping synthetic sequencesabstractMOTIVATION: Synthetic cellular tagging technologies play a crucial role in cell fate and lineage-tracing studies. Their integration with single-cell and spatial transcriptomics assays has heightened the need for scalable software solutions to analyze such data. However, previous methods are either designed for a subset of tagging technologies, or lack the performance needed for large-scale applications. RESULTS: To address these challenges, we developed Quick Clonal Analysis Toolkit (QuiCAT), an end-to-end Python-based package that streamlines the extraction, clustering, and analysis of synthetic tags from sequencing data. QuiCAT outperforms existing pipelines in both speed and accuracy. Its outputs are widely compatible with the Python ecosystem for single-cell and spatial transcriptomics data analysis packages allowing seamless integrations and downstream analyses. QuiCAT provides users with two workflows: a reference-free approach for extracting and mapping synthetic tags, and a reference-based approach for aligning tags against known sequences. We validate QuiCAT across diverse datasets, including population-level data, single-cell and spatially resolved transcriptomics, and benchmarked it against the two most recently published tools. Our computational optimizations enhance performance while improving accuracy. AVAILABILITY: QuiCAT is available as a Python package to be installed. The source code is available at https://github.com/theislab/quicat. Daniele Lucarelli, Tina Kos, Caylie Shull, Sara Jiménez, Rupert Öllinger, Roland Rad, Dieter Saur, Fabian J. Theis |
Bioinform. | 8 |
| 2025 | Spatial transcriptomics deconvolution methods generalize well to spatial chromatin accessibility dataabstractMOTIVATION: Spatially resolved chromatin accessibility profiling offers the potential to investigate gene regulatory processes within the spatial context of tissues. However, current methods typically work at spot resolution, aggregating measurements from multiple cells, thereby obscuring cell-type-specific spatial patterns of accessibility. Spot deconvolution methods have been developed and extensively benchmarked for spatial transcriptomics, yet no dedicated methods exist for spatial chromatin accessibility, and it is unclear if RNA-based approaches are applicable to that modality. RESULTS: Here, we demonstrate that these RNA-based approaches can be applied to spot-based chromatin accessibility data by a systematic evaluation of five top-performing spatial transcriptomics deconvolution methods. To assess performance, we developed a simulation framework that generates both transcriptomic and accessibility spot data from dissociated single-cell and targeted multiomic datasets, enabling direct comparisons across both data modalities. Our results show that Cell2location and RCTD, in contrast to other methods, exhibit robust performance on spatial chromatin accessibility data, achieving accuracy comparable to RNA-based deconvolution. Generally, we observed that RNA-based deconvolution exhibited slightly better performance compared to chromatin accessibility-based deconvolution, especially for resolving rare cell types, indicating room for future development of specialized methods. In conclusion, our findings demonstrate that existing deconvolution methods can be readily applied to chromatin accessibility-based spatial data. Our work provides a simulation framework and establishes a performance baseline to guide the development and evaluation of methods optimized for spatial epigenomics. AVAILABILITY AND IMPLEMENTATION: All methods, simulation frameworks, peak selection strategies, analysis notebooks and scripts are available at https://github.com/theislab/deconvATAC. Sarah Ouologuem, Laura D. Martens, Anna C. Schaar, Maiia Shulman, Julien Gagneur, Fabian J. Theis |
Bioinform. | 6 |
| 2024 | Mixed Models with Multiple Instance Learning
Jan P. Engelmann, Alessandro Palma, Jakub M. Tomczak, Fabian J. Theis, Francesco Paolo Casale |
AISTATS | 4 |
| 2024 | Unbalancedness in Neural Monge Maps Improves Unpaired Domain TranslationabstractIn optimal transport (OT), a Monge map is known as a mapping that transports a source distribution to a target distribution in the most cost-efficient way. Recently, multiple neural estimators for Monge maps have been developed and applied in diverse unpaired domain translation tasks, e.g. in single-cell biology and computer vision. However, the classic OT framework enforces mass conservation, which
makes it prone to outliers and limits its applicability in real-world scenarios. The latter can be particularly harmful in OT domain translation tasks, where the relative position of a sample within a distribution is explicitly taken into account. While unbalanced OT tackles this challenge in the discrete setting, its integration into neural Monge map estimators has received limited attention. We propose a theoretically
grounded method to incorporate unbalancedness into any Monge map estimator. We improve existing estimators to model cell trajectories over time and to predict cellular responses to perturbations. Moreover, our approach seamlessly integrates with the OT flow matching (OT-FM) framework. While we show that OT-FM performs competitively in image translation, we further improve performance by
incorporating unbalancedness (UOT-FM), which better preserves relevant features. We hence establish UOT-FM as a principled method for unpaired image translation. Luca Eyring, Dominik Klein 0005, Théo Uscidda, Giovanni Palla, Niki Kilbertus, Zeynep Akata, Fabian J. Theis |
ICLR | 7 |
| 2024 | Unified Guidance for Geometry-Conditioned Molecular GenerationabstractEffectively designing molecular geometries is essential to advancing pharmaceutical innovations, a domain, which has experienced great attention through the success of generative models and, in particular, diffusion models. However, current molecular diffusion models are tailored towards a specific downstream task and lack adaptability. We introduce UniGuide, a framework for controlled geometric guidance of unconditional diffusion models that allows flexible conditioning during inference without the requirement of extra training or networks. We show how applications such as structure-based, fragment-based, and ligand-based drug design are formulated in the UniGuide framework and demonstrate on-par or superior performance compared to specialised models. Offering a more versatile approach, UniGuide has the potential to streamline the development of molecular generative models, allowing them to be readily used in diverse application scenarios. Sirine Ayadi, Leon Hetzel, Johanna Sommer, Fabian J. Theis, Stephan Günnemann |
NeurIPS | 4 |
| 2024 | GENOT: Entropic (Gromov) Wasserstein Flow Matching with Applications to Single-Cell GenomicsabstractSingle-cell genomics has significantly advanced our understanding of cellular behavior, catalyzing innovations in treatments and precision medicine. However,
single-cell sequencing technologies are inherently destructive and can only measure a limited array of data modalities simultaneously. This limitation underscores
the need for new methods capable of realigning cells. Optimal transport (OT)
has emerged as a potent solution, but traditional discrete solvers are hampered by
scalability, privacy, and out-of-sample estimation issues. These challenges have
spurred the development of neural network-based solvers, known as neural OT
solvers, that parameterize OT maps. Yet, these models often lack the flexibility
needed for broader life science applications. To address these deficiencies, our
approach learns stochastic maps (i.e. transport plans), allows for any cost function,
relaxes mass conservation constraints and integrates quadratic solvers to tackle the
complex challenges posed by the (Fused) Gromov-Wasserstein problem. Utilizing
flow matching as a backbone, our method offers a flexible and effective framework.
We demonstrate its versatility and robustness through applications in cell development studies, cellular drug response modeling, and cross-modality cell translation,
illustrating significant potential for enhancing therapeutic strategies. Dominik Klein 0005, Théo Uscidda, Fabian J. Theis, Marco Cuturi |
NeurIPS | 3 |
| 2024 | A benchmark for prediction of transcriptomic responses to chemical perturbations across cell typesabstractSingle-cell transcriptomics has revolutionized our understanding of cellular heterogeneity and drug perturbation effects. However, its high cost and the vast chemical space of potential drugs present barriers to experimentally characterizing the effect of chemical perturbations in all the myriad cell types of the human body. To overcome these limitations, several groups have proposed using machine learning methods to directly predict the effect of chemical perturbations either across cell contexts or chemical space. However, advances in this field have been hindered by a lack of well-designed evaluation datasets and benchmarks. To drive innovation in perturbation modeling, the Open Problems Perturbation Prediction (OP3) benchmark introduces a framework for predicting the effects of small molecule perturbations on cell type-specific gene expression. OP3 leverages the Open Problems in Single-cell Analysis benchmarking infrastructure and is enabled by a new single-cell perturbation dataset, encompassing 146 compounds tested on human blood cells. The benchmark includes diverse data representations, evaluation metrics, and winning methods from our "Single-cell perturbation prediction: generalizing experimental interventions to unseen contexts" competition at NeurIPS 2023. We envision that the OP3 benchmark and competition will drive innovation in single-cell perturbation prediction by improving the accessibility, visibility, and feasibility of this challenge, thereby promoting the impact of machine learning in drug discovery. Artur Szalata, Andrew Benz, Robrecht Cannoodt, Mauricio Cortes, Jason Fong, Sunil Kuppasani, Richard Lieberman, Javier Mas-Rosario, Rico Meinl, Jalil Nourisa, Jared Tumiel, Tin M. Tunjic, Mengbo Wang 0001, Noah Weber, Benedict Anchang, Fabian J. Theis, Malte Lücken, Daniel Burkhardt |
NeurIPS | 18 |
| 2024 | GraphCompass: spatial metrics for differential analyses of cell organization across conditionsabstractSUMMARY: Spatial omics technologies are increasingly leveraged to characterize how disease disrupts tissue organization and cellular niches. While multiple methods to analyze spatial variation within a sample have been published, statistical and computational approaches to compare cell spatial organization across samples or conditions are mostly lacking. We present GraphCompass, a comprehensive set of omics-adapted graph analysis methods to quantitatively evaluate and compare the spatial arrangement of cells in samples representing diverse biological conditions. GraphCompass builds upon the Squidpy spatial omics toolbox and encompasses various statistical approaches to perform cross-condition analyses at the level of individual cell types, niches, and samples. Additionally, GraphCompass provides custom visualization functions that enable effective communication of results. We demonstrate how GraphCompass can be used to address key biological questions, such as how cellular organization and tissue architecture differ across various disease states and which spatial patterns correlate with a given pathological condition. GraphCompass can be applied to various popular omics techniques, including, but not limited to, spatial proteomics (e.g. MIBI-TOF), spot-based transcriptomics (e.g. 10× Genomics Visium), and single-cell resolved transcriptomics (e.g. Stereo-seq). In this work, we showcase the capabilities of GraphCompass through its application to three different studies that may also serve as benchmark datasets for further method development. With its easy-to-use implementation, extensive documentation, and comprehensive tutorials, GraphCompass is accessible to biologists with varying levels of computational expertise. By facilitating comparative analyses of cell spatial organization, GraphCompass promises to be a valuable asset in advancing our understanding of tissue function in health and disease. . Mayar Ali, Merel Kuijs, Soroor Hediyeh-Zadeh, Tim Treis, Karin Hrovatin, Giovanni Palla, Anna C. Schaar, Fabian J. Theis |
Bioinform. | 8 |
| 2024 | longmixr: a tool for robust clustering of high-dimensional cross-sectional and longitudinal variables of mixed data typesabstractSUMMARY: Accurate clustering of mixed data, encompassing binary, categorical, and continuous variables, is vital for effective patient stratification in clinical questionnaire analysis. To address this need, we present longmixr, a comprehensive R package providing a robust framework for clustering mixed longitudinal data using finite mixture modeling techniques. By incorporating consensus clustering, longmixr ensures reliable and stable clustering results. Moreover, the package includes a detailed vignette that facilitates cluster exploration and visualization. AVAILABILITY AND IMPLEMENTATION: The R package is freely available at https://cran.r-project.org/package=longmixr with detailed documentation, including a case vignette, at https://cellmapslab.github.io/longmixr/. Jonas Hagenberg, Monika Budde, Teodora Pandeva, Ivan Kondofersky, Sabrina K. Schaupp, Fabian J. Theis, Thomas G. Schulze, Nikola S. Müller, Urs Heilbronner, Richa Batra, Janine Knauer-Arloth |
Bioinform. | 6 |
| 2023 | B-Cos Aligned Transformers Learn Human-Interpretable Features
Manuel Tran, Amal Lahiani, Yashin Dicente Cid, Melanie Boxberg, Peter Lienemann, Christian Matek, Sophia J. Wagner, Fabian J. Theis, Eldad Klaiman, Tingying Peng |
MICCAI (8) | 8 |
| 2023 | Training Transitive and Commutative Multimodal Transformers with LoReTTaabstractTraining multimodal foundation models is challenging due to the limited availability of multimodal datasets. While many public datasets pair images with text, few combine images with audio or text with audio. Even rarer are datasets that align all three modalities at once. Critical domains such as healthcare, infrastructure, or transportation are particularly affected by missing modalities. This makes it difficult to integrate all modalities into a large pre-trained neural network that can be used out-of-the-box or fine-tuned for different downstream tasks. We introduce LoReTTa ($\textbf{L}$inking m$\textbf{O}$dalities with a t$\textbf{R}$ansitive and commutativ$\textbf{E}$ pre-$\textbf{T}$raining s$\textbf{T}$r$\textbf{A}$tegy) to address this understudied problem. Our self-supervised framework unifies causal modeling and masked modeling with the rules of commutativity and transitivity. This allows us to transition within and between modalities. As a result, our pre-trained models are better at exploring the true underlying joint probability distribution. Given a dataset containing only the disjoint combinations $(A, B)$ and $(B, C)$, LoReTTa can model the relation $A \leftrightarrow C$ with $A \leftrightarrow B \leftrightarrow C$. In particular, we show that a transformer pre-trained with LoReTTa can handle any mixture of modalities at inference time, including the never-seen pair $(A, C)$ and the triplet $(A, B, C)$. We extensively evaluate our approach on a synthetic, medical, and reinforcement learning dataset. Across different domains, our universal multimodal transformer consistently outperforms strong baselines such as GPT, BERT, and CLIP on tasks involving the missing modality tuple. Manuel Tran, Yashin Dicente Cid, Amal Lahiani, Fabian J. Theis, Tingying Peng, Eldad Klaiman |
NeurIPS | 4 |
| 2023 | Multi-omics regulatory network inference in the presence of missing dataabstractA key problem in systems biology is the discovery of regulatory mechanisms that drive phenotypic behaviour of complex biological systems in the form of multi-level networks. Modern multi-omics profiling techniques probe these fundamental regulatory networks but are often hampered by experimental restrictions leading to missing data or partially measured omics types for subsets of individuals due to cost restrictions. In such scenarios, in which missing data is present, classical computational approaches to infer regulatory networks are limited. In recent years, approaches have been proposed to infer sparse regression models in the presence of missing information. Nevertheless, these methods have not been adopted for regulatory network inference yet. In this study, we integrated regression-based methods that can handle missingness into KiMONo, a Knowledge guided Multi-Omics Network inference approach, and benchmarked their performance on commonly encountered missing data scenarios in single- and multi-omics studies. Overall, two-step approaches that explicitly handle missingness performed best for a wide range of random- and block-missingness scenarios on imbalanced omics-layers dimensions, while methods implicitly handling missingness performed best on balanced omics-layers dimensions. Our results show that robust multi-omics network inference in the presence of missing data with KiMONo is feasible and thus allows users to leverage available multi-omics data to its full extent. Juan D. Henao, Michael Lauber, Manuel Azevedo, Anastasiia Grekova, Fabian J. Theis, Markus List, Christoph Ogris, Benjamin Schubert |
Briefings Bioinform. | 5 |
| 2022 | Noise Transfer for Unsupervised Domain Adaptation of Retinal OCT Images
Valentin Koch, Olle G. Holmberg, Hannah Spitzer, Johannes Schiefelbein, Ben Asani, Michael Hafner, Fabian J. Theis |
MICCAI (2) | 7 |
| 2022 | Sparsity in Continuous-Depth Neural NetworksabstractNeural Ordinary Differential Equations (NODEs) have proven successful in learning dynamical systems in terms of accurately recovering the observed trajectories. While different types of sparsity have been proposed to improve robustness, the generalization properties of NODEs for dynamical systems beyond the observed data are underexplored. We systematically study the influence of weight and feature sparsity on forecasting as well as on identifying the underlying dynamical laws. Besides assessing existing methods, we propose a regularization technique to sparsify ``input-output connections'' and extract relevant features during training. Moreover, we curate real-world datasets including human motion capture and human hematopoiesis single-cell RNA-seq data to realistically analyze different levels of out-of-distribution (OOD) generalization in forecasting and dynamics identification respectively. Our extensive empirical evaluation on these challenging benchmarks suggests that weight sparsity improves generalization in the presence of noise or irregular sampling. However, it does not prevent learning spurious feature dependencies in the inferred dynamics, rendering them impractical for predictions under interventions, or for inferring the true underlying dynamics. Instead, feature sparsity can indeed help with recovering sparse ground-truth dynamics compared to unregularized NODEs. Hananeh Aliee, Till Richter, Mikhail Solonin, Ignacio Ibarra, Fabian J. Theis, Niki Kilbertus |
NeurIPS | 5 |
| 2022 | Predicting Cellular Responses to Novel Drug Perturbations at a Single-Cell ResolutionabstractSingle-cell transcriptomics enabled the study of cellular heterogeneity in response to perturbations at the resolution of individual cells. However, scaling high-throughput screens (HTSs) to measure cellular responses for many drugs remains a challenge due to technical limitations and, more importantly, the cost of such multiplexed experiments. Thus, transferring information from routinely performed bulk RNA HTS is required to enrich single-cell data meaningfully.We introduce chemCPA, a new encoder-decoder architecture to study the perturbational effects of unseen drugs. We combine the model with an architecture surgery for transfer learning and demonstrate how training on existing bulk RNA HTS datasets can improve generalisation performance. Better generalisation reduces the need for extensive and costly screens at single-cell resolution. We envision that our proposed method will facilitate more efficient experiment designs through its ability to generate in-silico hypotheses, ultimately accelerating drug discovery. Leon Hetzel, Simon Böhm, Niki Kilbertus, Stephan Günnemann, Mohammad Lotfollahi, Fabian J. Theis |
NeurIPS | 6 |
| 2020 | Learning Tn5 Sequence Bias from ATAC-seq on Naked Chromatin
Meshal Ansari, David S. Fischer, Fabian J. Theis |
ICANN (1) | 3 |
| 2020 | Copy number aberrations from Affymetrix SNP 6.0 genotyping data - how accurate are commonly used prediction approaches?abstractCopy number aberrations (CNAs) are known to strongly affect oncogenes and tumour suppressor genes. Given the critical role CNAs play in cancer research, it is essential to accurately identify CNAs from tumour genomes. One particular challenge in finding CNAs is the effect of confounding variables. To address this issue, we assessed how commonly used CNA identification algorithms perform on SNP 6.0 genotyping data in the presence of confounding variables. We simulated realistic synthetic data with varying levels of three confounding variables-the tumour purity, the length of a copy number region and the CNA burden (the percentage of CNAs present in a profiled genome)-and evaluated the performance of OncoSNP, ASCAT, GenoCNA, GISTIC and CGHcall. Furthermore, we implemented and assessed CGHcall*, an adjusted version of CGHcall accounting for high CNA burden. Our analysis on synthetic data indicates that tumour purity and the CNA burden strongly influence the performance of all the algorithms. No algorithm can correctly find lost and gained genomic regions across all tumour purities. The length of CNA regions influenced the performance of ASCAT, CGHcall and GISTIC. OncoSNP, GenoCNA and CGHcall* showed little sensitivity. Overall, CGHcall* and OncoSNP showed reasonable performance, particularly in samples with high tumour purity. Our analysis on the HapMap data revealed a good overlap between CGHcall, CGHcall* and GenoCNA results and experimentally validated data. Our exploratory analysis on the TCGA HNSCC data revealed plausible results of CGHcall, CGHcall* and GISTIC in consensus HNSCC CNA regions. Code is available at https://github.com/adspit/PASCAL. Adriana Pitea, Ivan Kondofersky, Steffen Sass, Fabian J. Theis, Nikola S. Müller, Kristian Unger |
Briefings Bioinform. | 4 |
| 2020 | Automatic identification of relevant genes from low-dimensional embeddings of single-cell RNA-seq dataabstractMOTIVATION: Dimensionality reduction is a key step in the analysis of single-cell RNA-sequencing data. It produces a low-dimensional embedding for visualization and as a calculation base for downstream analysis. Nonlinear techniques are most suitable to handle the intrinsic complexity of large, heterogeneous single-cell data. However, with no linear relation between gene and embedding coordinate, there is no way to extract the identity of genes driving any cell's position in the low-dimensional embedding, making it difficult to characterize the underlying biological processes. RESULTS: In this article, we introduce the concepts of local and global gene relevance to compute an equivalent of principal component analysis loadings for non-linear low-dimensional embeddings. Global gene relevance identifies drivers of the overall embedding, while local gene relevance identifies those of a defined sub-region. We apply our method to single-cell RNA-seq datasets from different experimental protocols and to different low-dimensional embedding techniques. This shows our method's versatility to identify key genes for a variety of biological processes. AVAILABILITY AND IMPLEMENTATION: To ensure reproducibility and ease of use, our method is released as part of destiny 3.0, a popular R package for building diffusion maps from single-cell transcriptomic data. It is readily available through Bioconductor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Philipp Angerer, David S. Fischer, Fabian J. Theis, Antonio Scialdone, Carsten Marr |
Bioinform. | 3 |
| 2020 | Conditional out-of-distribution generation for unpaired data using transfer VAEabstractMOTIVATION: While generative models have shown great success in sampling high-dimensional samples conditional on low-dimensional descriptors (stroke thickness in MNIST, hair color in CelebA, speaker identity in WaveNet), their generation out-of-distribution poses fundamental problems due to the difficulty of learning compact joint distribution across conditions. The canonical example of the conditional variational autoencoder (CVAE), for instance, does not explicitly relate conditions during training and, hence, has no explicit incentive of learning such a compact representation. RESULTS: We overcome the limitation of the CVAE by matching distributions across conditions using maximum mean discrepancy in the decoder layer that follows the bottleneck. This introduces a strong regularization both for reconstructing samples within the same condition and for transforming samples across conditions, resulting in much improved generalization. As this amount to solving a style-transfer problem, we refer to the model as transfer VAE (trVAE). Benchmarking trVAE on high-dimensional image and single-cell RNA-seq, we demonstrate higher robustness and higher accuracy than existing approaches. We also show qualitatively improved predictions by tackling previously problematic minority classes and multiple conditions in the context of cellular perturbation response to treatment and disease based on high-dimensional single-cell gene expression data. For generic tasks, we improve Pearson correlations of high-dimensional estimated means and variances with their ground truths from 0.89 to 0.97 and 0.75 to 0.87, respectively. We further demonstrate that trVAE learns cell-type-specific responses after perturbation and improves the prediction of most cell-type-specific genes by 65%. AVAILABILITY AND IMPLEMENTATION: The trVAE implementation is available via github.com/theislab/trvae. The results of this article can be reproduced via github.com/theislab/trvae_reproducibility. Mohammad Lotfollahi, Mohsen Naghipourfar, Fabian J. Theis, F. Alexander Wolf |
Bioinform. | 3 |
| 2020 | DeepWAS: Multivariate genotype-phenotype associations by directly integrating regulatory information using deep learningabstractGenome-wide association studies (GWAS) identify genetic variants associated with traits or diseases. GWAS never directly link variants to regulatory mechanisms. Instead, the functional annotation of variants is typically inferred by post hoc analyses. A specific class of deep learning-based methods allows for the prediction of regulatory effects per variant on several cell type-specific chromatin features. We here describe "DeepWAS", a new approach that integrates these regulatory effect predictions of single variants into a multivariate GWAS setting. Thereby, single variants associated with a trait or disease are directly coupled to their impact on a chromatin feature in a cell type. Up to 61 regulatory SNPs, called dSNPs, were associated with multiple sclerosis (MS, 4,888 cases and 10,395 controls), major depressive disorder (MDD, 1,475 cases and 2,144 controls), and height (5,974 individuals). These variants were mainly non-coding and reached at least nominal significance in classical GWAS. The prediction accuracy was higher for DeepWAS than for classical GWAS models for 91% of the genome-wide significant, MS-specific dSNPs. DSNPs were enriched in public or cohort-matched expression and methylation quantitative trait loci and we demonstrated the potential of DeepWAS to generate testable functional hypotheses based on genotype data alone. DeepWAS is available at https://github.com/cellmapslab/DeepWAS. Janine Knauer-Arloth, Gökcen Eraslan, Till F. M. Andlauer, Jade Martins, Stella Iurato, Brigitte Kühnel, Melanie Waldenberger, Josef Frank, Ralf Gold, Bernhard Hemmer, Felix Luessi, Sandra Nischwitz, Friedemann Paul, Heinz Wiendl, Christian Gieger, Stefanie Heilmann-Heimbach, Tim Kacprowski, Matthias Laudes, Thomas Meitinger, Annette Peters, Rajesh Rawal, Konstantin Strauch, Susanne Lucae, Bertram Müller-Myhsok, Marcella Rietschel, Fabian J. Theis, Elisabeth B. Binder, Nikola S. Müller |
PLoS Comput. Biol. | 26 |
| 2020 | Model-based analysis of response and resistance factors of cetuximab treatment in gastric cancer cell linesabstractTargeted cancer therapies are powerful alternatives to chemotherapies or can be used complementary to these. Yet, the response to targeted treatments depends on a variety of factors, including mutations and expression levels, and therefore their outcome is difficult to predict. Here, we develop a mechanistic model of gastric cancer to study response and resistance factors for cetuximab treatment. The model captures the EGFR, ERK and AKT signaling pathways in two gastric cancer cell lines with different mutation patterns. We train the model using a comprehensive selection of time and dose response measurements, and provide an assessment of parameter and prediction uncertainties. We demonstrate that the proposed model facilitates the identification of causal differences between the cell lines. Furthermore, our study shows that the model provides predictions for the responses to different perturbations, such as knockdown and knockout experiments. Among other results, the model predicted the effect of MET mutations on cetuximab sensitivity. These predictive capabilities render the model a basis for the assessment of gastric cancer signaling and possibly for the development and discovery of predictive biomarkers. Elba Raimúndez-Álvarez, Simone Keller, Gwen Zwingenberger, Karolin Ebert, Sabine Hug, Fabian J. Theis, Dieter Maier, Birgit Luber, Jan Hasenauer |
PLoS Comput. Biol. | 6 |
| 2018 | Bayesian parameter estimation for biochemical reaction networks using region-based adaptive parallel temperingabstractMotivation: Mathematical models have become standard tools for the investigation of cellular processes and the unraveling of signal processing mechanisms. The parameters of these models are usually derived from the available data using optimization and sampling methods. However, the efficiency of these methods is limited by the properties of the mathematical model, e.g. non-identifiabilities, and the resulting posterior distribution. In particular, multi-modal distributions with long valleys or pronounced tails are difficult to optimize and sample. Thus, the developement or improvement of optimization and sampling methods is subject to ongoing research. Results: We suggest a region-based adaptive parallel tempering algorithm which adapts to the problem-specific posterior distributions, i.e. modes and valleys. The algorithm combines several established algorithms to overcome their individual shortcomings and to improve sampling efficiency. We assessed its properties for established benchmark problems and two ordinary differential equation models of biochemical reaction networks. The proposed algorithm outperformed state-of-the-art methods in terms of calculation efficiency and mixing. Since the algorithm does not rely on a specific problem structure, but adapts to the posterior distribution, it is suitable for a variety of model classes. Availability and implementation: The code is available both as Supplementary Material and in a Git repository written in MATLAB. Supplementary information: Supplementary data are available at Bioinformatics online. Benjamin Ballnus, Steffen Schaper, Fabian J. Theis, Jan Hasenauer |
Bioinform. | 3 |
| 2018 | netReg: network-regularized linear models for biological association studiesabstractSummary: Modelling biological associations or dependencies using linear regression is often complicated when the analyzed data-sets are high-dimensional and less observations than variables are available (n ≪ p). For genomic data-sets penalized regression methods have been applied settling this issue. Recently proposed regression models utilize prior knowledge on dependencies, e.g. in the form of graphs, arguing that this information will lead to more reliable estimates for regression coefficients. However, none of the proposed models for multivariate genomic response variables have been implemented as a computationally efficient, freely available library. In this paper we propose netReg, a package for graph-penalized regression models that use large networks and thousands of variables. netReg incorporates a priori generated biological graph information into linear models yielding sparse or smooth solutions for regression coefficients. Availability and implementation: netReg is implemented as both R-package and C ++ commandline tool. The main computations are done in C ++, where we use Armadillo for fast matrix calculations and Dlib for optimization. The R package is freely available on Bioconductorhttps://bioconductor.org/packages/netReg. The command line tool can be installed using the conda channel Bioconda. Installation details, issue reports, development versions, documentation and tutorials for the R and C ++ versions and the R package vignette can be found on GitHub https://dirmeier.github.io/netReg/. The GitHub page also contains code for benchmarking and example datasets used in this paper. Contact: [email protected]. Simon Dirmeier, Christiane Fuchs, Nikola S. Müller, Fabian J. Theis |
Bioinform. | 4 |
| 2017 | Model-based branching point detection in single-cell data by K-branches clusteringabstractMOTIVATION: The identification of heterogeneities in cell populations by utilizing single-cell technologies such as single-cell RNA-Seq, enables inference of cellular development and lineage trees. Several methods have been proposed for such inference from high-dimensional single-cell data. They typically assign each cell to a branch in a differentiation trajectory. However, they commonly assume specific geometries such as tree-like developmental hierarchies and lack statistically sound methods to decide on the number of branching events. RESULTS: We present K-Branches, a solution to the above problem by locally fitting half-lines to single-cell data, introducing a clustering algorithm similar to K-Means. These halflines are proxies for branches in the differentiation trajectory of cells. We propose a modified version of the GAP statistic for model selection, in order to decide on the number of lines that best describe the data locally. In this manner, we identify the location and number of subgroups of cells that are associated with branching events and full differentiation, respectively. We evaluate the performance of our method on single-cell RNA-Seq data describing the differentiation of myeloid progenitors during hematopoiesis, single-cell qPCR data of mouse blastocyst development, single-cell qPCR data of human myeloid monocytic leukemia and artificial data. AVAILABILITY AND IMPLEMENTATION: An R implementation of K-Branches is freely available at https://github.com/theislab/kbranches. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nikolaos-Kosmas Chlis, F. Alexander Wolf, Fabian J. Theis |
Bioinform. | 3 |
| 2017 | Parameter estimation for dynamical systems with discrete events and logical operationsabstractMotivation: Ordinary differential equation (ODE) models are frequently used to describe the dynamic behaviour of biochemical processes. Such ODE models are often extended by events to describe the effect of fast latent processes on the process dynamics. To exploit the predictive power of ODE models, their parameters have to be inferred from experimental data. For models without events, gradient based optimization schemes perform well for parameter estimation, when sensitivity equations are used for gradient computation. Yet, sensitivity equations for models with parameter- and state-dependent events and event-triggered observations are not supported by existing toolboxes. Results: In this manuscript, we describe the sensitivity equations for differential equation models with events and demonstrate how to estimate parameters from event-resolved data using event-triggered observations in parameter estimation. We consider a model for GFP expression after transfection and a model for spiking neurons and demonstrate that we can improve computational efficiency and robustness of parameter estimation by using sensitivity equations for systems with events. Moreover, we demonstrate that, by using event-outputs, it is possible to consider event-resolved data, such as time-to-event data, for parameter estimation with ODE models. By providing a user-friendly, modular implementation in the toolbox AMICI, the developed methods are made publicly available and can be integrated in other systems biology toolboxes. Availability and Implementation: We implement the methods in the open-source toolbox Advanced MATLAB Interface for CVODES and IDAS (AMICI, https://github.com/ICB-DCM/AMICI ). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Fabian Fröhlich, Fabian J. Theis, Joachim O. Rädler, Jan Hasenauer |
Bioinform. | 2 |
| 2017 | fastER: a user-friendly tool for ultrafast and robust cell segmentation in large-scale microscopyabstractMOTIVATION: Quantitative large-scale cell microscopy is widely used in biological and medical research. Such experiments produce huge amounts of image data and thus require automated analysis. However, automated detection of cell outlines (cell segmentation) is typically challenging due to, e.g. high cell densities, cell-to-cell variability and low signal-to-noise ratios. RESULTS: Here, we evaluate accuracy and speed of various state-of-the-art approaches for cell segmentation in light microscopy images using challenging real and synthetic image data. The results vary between datasets and show that the tested tools are either not robust enough or computationally expensive, thus limiting their application to large-scale experiments. We therefore developed fastER, a trainable tool that is orders of magnitude faster while producing state-of-the-art segmentation quality. It supports various cell types and image acquisition modalities, but is easy-to-use even for non-experts: it has no parameters and can be adapted to specific image sets by interactively labelling cells for training. As a proof of concept, we segment and count cells in over 200 000 brightfield images (1388 × 1040 pixels each) from a six day time-lapse microscopy experiment; identification of over 46 000 000 single cells requires only about two and a half hours on a desktop computer. AVAILABILITY AND IMPLEMENTATION: C ++ code, binaries and data at https://www.bsse.ethz.ch/csd/software/faster.html . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Oliver Hilsenbeck, Michael Schwarzfischer, Dirk Loeffler, Sotiris Dimopoulos, Simon Hastreiter, Carsten Marr, Fabian J. Theis, Timm Schroeder |
Bioinform. | 7 |
| 2017 | A scalable moment-closure approximation for large-scale biochemical reaction networksabstractMOTIVATION: Stochastic molecular processes are a leading cause of cell-to-cell variability. Their dynamics are often described by continuous-time discrete-state Markov chains and simulated using stochastic simulation algorithms. As these stochastic simulations are computationally demanding, ordinary differential equation models for the dynamics of the statistical moments have been developed. The number of state variables of these approximating models, however, grows at least quadratically with the number of biochemical species. This limits their application to small- and medium-sized processes. RESULTS: In this article, we present a scalable moment-closure approximation (sMA) for the simulation of statistical moments of large-scale stochastic processes. The sMA exploits the structure of the biochemical reaction network to reduce the covariance matrix. We prove that sMA yields approximating models whose number of state variables depends predominantly on local properties, i.e. the average node degree of the reaction network, instead of the overall network size. The resulting complexity reduction is assessed by studying a range of medium- and large-scale biochemical reaction networks. To evaluate the approximation accuracy and the improvement in computational efficiency, we study models for JAK2/STAT5 signalling and NF κ B signalling. Our method is applicable to generic biochemical reaction networks and we provide an implementation, including an SBML interface, which renders the sMA easily accessible. AVAILABILITY AND IMPLEMENTATION: The sMA is implemented in the open-source MATLAB toolbox CERENA and is available from https://github.com/CERENADevelopers/CERENA . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Atefeh Kazeroonian, Fabian J. Theis, Jan Hasenauer |
Bioinform. | 2 |
| 2017 | pulver: an R package for parallel ultra-rapid p-value computation for linear regression interaction termsabstractBACKGROUND: Genome-wide association studies allow us to understand the genetics of complex diseases. Human metabolism provides information about the disease-causing mechanisms, so it is usual to investigate the associations between genetic variants and metabolite levels. However, only considering genetic variants and their effects on one trait ignores the possible interplay between different "omics" layers. Existing tools only consider single-nucleotide polymorphism (SNP)-SNP interactions, and no practical tool is available for large-scale investigations of the interactions between pairs of arbitrary quantitative variables. RESULTS: We developed an R package called pulver to compute p-values for the interaction term in a very large number of linear regression models. Comparisons based on simulated data showed that pulver is much faster than the existing tools. This is achieved by using the correlation coefficient to test the null-hypothesis, which avoids the costly computation of inversions. Additional tricks are a rearrangement of the order, when iterating through the different "omics" layers, and implementing this algorithm in the fast programming language C++. Furthermore, we applied our algorithm to data from the German KORA study to investigate a real-world problem involving the interplay among DNA methylation, genetic variants, and metabolite levels. CONCLUSIONS: The pulver package is a convenient and rapid tool for screening huge numbers of linear regression models for significant interaction terms in arbitrary pairs of quantitative variables. pulver is written in R and C++, and can be downloaded freely from CRAN at https://cran.r-project.org/web/packages/pulver/ . Sophie Molnos, Clemens Baumbach, Simone Wahl, Martina Müller-Nurasyid, Konstantin Strauch, Rui Wang-Sattler, Melanie Waldenberger, Thomas Meitinger, Jerzy Adamski, Gabi Kastenmüller, Karsten Suhre, Annette Peters, Harald Grallert, Fabian J. Theis, Christian Gieger |
BMC Bioinform. | 14 |
| 2017 | Scalable Parameter Estimation for Genome-Scale Biochemical Reaction NetworksabstractMechanistic mathematical modeling of biochemical reaction networks using ordinary differential equation (ODE) models has improved our understanding of small- and medium-scale biological processes. While the same should in principle hold for large- and genome-scale processes, the computational methods for the analysis of ODE models which describe hundreds or thousands of biochemical species and reactions are missing so far. While individual simulations are feasible, the inference of the model parameters from experimental data is computationally too intensive. In this manuscript, we evaluate adjoint sensitivity analysis for parameter estimation in large scale biochemical reaction networks. We present the approach for time-discrete measurement and compare it to state-of-the-art methods used in systems and computational biology. Our comparison reveals a significantly improved computational efficiency and a superior scalability of adjoint sensitivity analysis. The computational complexity is effectively independent of the number of parameters, enabling the analysis of large- and genome-scale models. Our study of a comprehensive kinetic model of ErbB signaling shows that parameter estimation using adjoint sensitivity analysis requires a fraction of the computation time of established methods. The proposed method will facilitate mechanistic modeling of genome-scale cellular processes, as required in the age of omics. Fabian Fröhlich, Barbara Kaltenbacher, Fabian J. Theis, Jan Hasenauer |
PLoS Comput. Biol. | 3 |
| 2016 | destiny: diffusion maps for large-scale single-cell data in RabstractUNLABELLED: : Diffusion maps are a spectral method for non-linear dimension reduction and have recently been adapted for the visualization of single-cell expression data. Here we present destiny, an efficient R implementation of the diffusion map algorithm. Our package includes a single-cell specific noise model allowing for missing and censored values. In contrast to previous implementations, we further present an efficient nearest-neighbour approximation that allows for the processing of hundreds of thousands of cells and a functionality for projecting new data on existing diffusion maps. We exemplarily apply destiny to a recent time-resolved mass cytometry dataset of cellular reprogramming. AVAILABILITY AND IMPLEMENTATION: destiny is an open-source R/Bioconductor package "bioconductor.org/packages/destiny" also available at www.helmholtz-muenchen.de/icb/destiny A detailed vignette describing functions and workflows is provided with the package. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Philipp Angerer, Laleh Haghverdi, Maren Büttner, Fabian J. Theis, Carsten Marr, Florian A. Büttner |
Bioinform. | 4 |
| 2016 | MEMO: multi-experiment mixture model analysis of censored dataabstractMOTIVATION: The statistical analysis of single-cell data is a challenge in cell biological studies. Tailored statistical models and computational methods are required to resolve the subpopulation structure, i.e. to correctly identify and characterize subpopulations. These approaches also support the unraveling of sources of cell-to-cell variability. Finite mixture models have shown promise, but the available approaches are ill suited to the simultaneous consideration of data from multiple experimental conditions and to censored data. The prevalence and relevance of single-cell data and the lack of suitable computational analytics make automated methods, that are able to deal with the requirements posed by these data, necessary. RESULTS: We present MEMO, a flexible mixture modeling framework that enables the simultaneous, automated analysis of censored and uncensored data acquired under multiple experimental conditions. MEMO is based on maximum-likelihood inference and allows for testing competing hypotheses. MEMO can be applied to a variety of different single-cell data types. We demonstrate the advantages of MEMO by analyzing right and interval censored single-cell microscopy data. Our results show that an examination of censoring and the simultaneous consideration of different experimental conditions are necessary to reveal biologically meaningful subpopulation structures. MEMO allows for a stringent analysis of single-cell data and enables researchers to avoid misinterpretation of censored data. Therefore, MEMO is a valuable asset for all fields that infer the characteristics of populations by looking at single individuals such as cell biology and medicine. AVAILABILITY AND IMPLEMENTATION: MEMO is implemented in MATLAB and freely available via github (https://github.com/MEMO-toolbox/MEMO). CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Eva-Maria Geissen, Jan Hasenauer, Stephanie Heinrich, Silke Hauf, Fabian J. Theis, Nicole Radde |
Bioinform. | 5 |
| 2016 | Inference for Stochastic Chemical Kinetics Using Moment Equations and System Size ExpansionabstractQuantitative mechanistic models are valuable tools for disentangling biochemical pathways and for achieving a comprehensive understanding of biological systems. However, to be quantitative the parameters of these models have to be estimated from experimental data. In the presence of significant stochastic fluctuations this is a challenging task as stochastic simulations are usually too time-consuming and a macroscopic description using reaction rate equations (RREs) is no longer accurate. In this manuscript, we therefore consider moment-closure approximation (MA) and the system size expansion (SSE), which approximate the statistical moments of stochastic processes and tend to be more precise than macroscopic descriptions. We introduce gradient-based parameter optimization methods and uncertainty analysis methods for MA and SSE. Efficiency and reliability of the methods are assessed using simulation examples as well as by an application to data for Epo-induced JAK/STAT signaling. The application revealed that even if merely population-average data are available, MA and SSE improve parameter identifiability in comparison to RRE. Furthermore, the simulation examples revealed that the resulting estimates are more reliable for an intermediate volume regime. In this regime the estimation error is reduced and we propose methods to determine the regime boundaries. These results illustrate that inference using MA and SSE is feasible and possesses a high sensitivity. Fabian Fröhlich, Philipp Thomas, Atefeh Kazeroonian, Fabian J. Theis, Ramon Grima, Jan Hasenauer |
PLoS Comput. Biol. | 4 |
| 2015 | Diffusion maps for high-dimensional single-cell analysis of differentiation dataabstractMOTIVATION: Single-cell technologies have recently gained popularity in cellular differentiation studies regarding their ability to resolve potential heterogeneities in cell populations. Analyzing such high-dimensional single-cell data has its own statistical and computational challenges. Popular multivariate approaches are based on data normalization, followed by dimension reduction and clustering to identify subgroups. However, in the case of cellular differentiation, we would not expect clear clusters to be present but instead expect the cells to follow continuous branching lineages. RESULTS: Here, we propose the use of diffusion maps to deal with the problem of defining differentiation trajectories. We adapt this method to single-cell data by adequate choice of kernel width and inclusion of uncertainties or missing measurement values, which enables the establishment of a pseudotemporal ordering of single cells in a high-dimensional gene expression space. We expect this output to reflect cell differentiation trajectories, where the data originates from intrinsic diffusion-like dynamics. Starting from a pluripotent stage, cells move smoothly within the transcriptional landscape towards more differentiated states with some stochasticity along their path. We demonstrate the robustness of our method with respect to extrinsic noise (e.g. measurement noise) and sampling density heterogeneities on simulated toy data as well as two single-cell quantitative polymerase chain reaction datasets (i.e. mouse haematopoietic stem cells and mouse embryonic stem cells) and an RNA-Seq data of human pre-implantation embryos. We show that diffusion maps perform considerably better than Principal Component Analysis and are advantageous over other techniques for non-linear dimension reduction such as t-distributed Stochastic Neighbour Embedding for preserving the global structures and pseudotemporal ordering of cells. AVAILABILITY AND IMPLEMENTATION: The Matlab implementation of diffusion maps for single-cell data is available at https://www.helmholtz-muenchen.de/icb/single-cell-diffusion-map. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Laleh Haghverdi, Florian A. Büttner, Fabian J. Theis |
Bioinform. | 3 |
| 2015 | Reconstructing gene regulatory dynamics from high-dimensional single-cell snapshot dataabstractMOTIVATION: High-dimensional single-cell snapshot data are becoming widespread in the systems biology community, as a mean to understand biological processes at the cellular level. However, as temporal information is lost with such data, mathematical models have been limited to capture only static features of the underlying cellular mechanisms. RESULTS: Here, we present a modular framework which allows to recover the temporal behaviour from single-cell snapshot data and reverse engineer the dynamics of gene expression. The framework combines a dimensionality reduction method with a cell time-ordering algorithm to generate pseudo time-series observations. These are in turn used to learn transcriptional ODE models and do model selection on structural network features. We apply it on synthetic data and then on real hematopoietic stem cells data, to reconstruct gene expression dynamics during differentiation pathways and infer the structure of a key gene regulatory network. AVAILABILITY AND IMPLEMENTATION: C++ and Matlab code available at https://www.helmholtz-muenchen.de/fileadmin/ICB/software/inferenceSnapshot.zip. Andrea Ocone, Laleh Haghverdi, Nikola S. Müller, Fabian J. Theis |
Bioinform. | 4 |
| 2015 | Data2Dynamics: a modeling environment tailored to parameter estimation in dynamical systemsabstractUNLABELLED: Modeling of dynamical systems using ordinary differential equations is a popular approach in the field of systems biology. Two of the most critical steps in this approach are to construct dynamical models of biochemical reaction networks for large datasets and complex experimental conditions and to perform efficient and reliable parameter estimation for model fitting. We present a modeling environment for MATLAB that pioneers these challenges. The numerically expensive parts of the calculations such as the solving of the differential equations and of the associated sensitivity system are parallelized and automatically compiled into efficient C code. A variety of parameter estimation algorithms as well as frequentist and Bayesian methods for uncertainty analysis have been implemented and used on a range of applications that lead to publications. AVAILABILITY AND IMPLEMENTATION: The Data2Dynamics modeling environment is MATLAB based, open source and freely available at http://www.data2dynamics.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andreas Raue, Bernhard Steiert, Max Schelker, Clemens Kreutz, Tim Maiwald, Helge Hass, Joep Vanlier, Christian Tönsing, Lorenz Adlung, Raphael Engesser, Wolfgang Mader 0002, Tim Heinemann, Jan Hasenauer, Marcel Schilling, Thomas Höfer, Edda Klipp, Fabian J. Theis, Ursula Klingmüller, Birgit Schoeberl, Jens Timmer |
Bioinform. | 17 |
| 2015 | RAMONA: a Web application for gene set analysis on multilevel omics dataabstractSUMMARY: Decreasing costs of modern high-throughput experiments allow for the simultaneous analysis of altered gene activity on various molecular levels. However, these multi-omics approaches lead to a large amount of data, which is hard to interpret for a non-bioinformatician. Here, we present the remotely accessible multilevel ontology analysis (RAMONA). It offers an easy-to-use interface for the simultaneous gene set analysis of combined omics datasets and is an extension of the previously introduced MONA approach. RAMONA is based on a Bayesian enrichment method for the inference of overrepresented biological processes among given gene sets. Overrepresentation is quantified by interpretable term probabilities. It is able to handle data from various molecular levels, while in parallel coping with redundancies arising from gene set overlaps and related multiple testing problems. The comprehensive output of RAMONA is easy to interpret and thus allows for functional insight into the affected biological processes. With RAMONA, we provide an efficient implementation of the Bayesian inference problem such that ontologies consisting of thousands of terms can be processed in the order of seconds. AVAILABILITY AND IMPLEMENTATION: RAMONA is implemented as ASP.NET Web application and publicly available at http://icb.helmholtz-muenchen.de/ramona. Steffen Sass, Florian A. Büttner, Nikola S. Müller, Fabian J. Theis |
Bioinform. | 4 |
| 2015 | Model selection using limiting distributions of second-order blind source separation algorithms
Katrin Illner, Jari Miettinen, Christiane Fuchs, Sara Taskinen, Klaus Nordhausen, Hannu Oja, Fabian J. Theis |
Signal Process. | 7 |
| 2014 | Probabilistic PCA of censored data: accounting for uncertainties in the visualization of high-throughput single-cell qPCR dataabstractMOTIVATION: High-throughput single-cell quantitative real-time polymerase chain reaction (qPCR) is a promising technique allowing for new insights in complex cellular processes. However, the PCR reaction can be detected only up to a certain detection limit, whereas failed reactions could be due to low or absent expression, and the true expression level is unknown. Because this censoring can occur for high proportions of the data, it is one of the main challenges when dealing with single-cell qPCR data. Principal component analysis (PCA) is an important tool for visualizing the structure of high-dimensional data as well as for identifying subpopulations of cells. However, to date it is not clear how to perform a PCA of censored data. We present a probabilistic approach that accounts for the censoring and evaluate it for two typical datasets containing single-cell qPCR data. RESULTS: We use the Gaussian process latent variable model framework to account for censoring by introducing an appropriate noise model and allowing a different kernel for each dimension. We evaluate this new approach for two typical qPCR datasets (of mouse embryonic stem cells and blood stem/progenitor cells, respectively) by performing linear and non-linear probabilistic PCA. Taking the censoring into account results in a 2D representation of the data, which better reflects its known structure: in both datasets, our new approach results in a better separation of known cell types and is able to reveal subpopulations in one dataset that could not be resolved using standard PCA. AVAILABILITY AND IMPLEMENTATION: The implementation was based on the existing Gaussian process latent variable model toolbox (https://github.com/SheffieldML/GPmat); extensions for noise models and kernels accounting for censoring are available at http://icb.helmholtz-muenchen.de/censgplvm. Florian A. Büttner, Victoria Moignard, Berthold Göttgens, Fabian J. Theis |
Bioinform. | 4 |
| 2014 | MCA: Multiresolution Correlation Analysis, a graphical tool for subpopulation identification in single-cell gene expression dataabstractBACKGROUND: Biological data often originate from samples containing mixtures of subpopulations, corresponding e.g. to distinct cellular phenotypes. However, identification of distinct subpopulations may be difficult if biological measurements yield distributions that are not easily separable. RESULTS: We present Multiresolution Correlation Analysis (MCA), a method for visually identifying subpopulations based on the local pairwise correlation between covariates, without needing to define an a priori interaction scale. We demonstrate that MCA facilitates the identification of differentially regulated subpopulations in simulated data from a small gene regulatory network, followed by application to previously published single-cell qPCR data from mouse embryonic stem cells. We show that MCA recovers previously identified subpopulations, provides additional insight into the underlying correlation structure, reveals potentially spurious compartmentalizations, and provides insight into novel subpopulations. CONCLUSIONS: MCA is a useful method for the identification of subpopulations in low-dimensional expression data, as emerging from qPCR or FACS measurements. With MCA it is possible to investigate the robustness of covariate correlations with respect subpopulations, graphically identify outliers, and identify factors contributing to differential regulation between pairs of covariates. MCA thus provides a framework for investigation of expression correlations for genes of interests and biological hypothesis generation. Justin Feigelman, Fabian J. Theis, Carsten Marr |
BMC Bioinform. | 2 |
| 2014 | ODE Constrained Mixture Modelling: A Method for Unraveling Subpopulation Structures and DynamicsabstractFunctional cell-to-cell variability is ubiquitous in multicellular organisms as well as bacterial populations. Even genetically identical cells of the same cell type can respond differently to identical stimuli. Methods have been developed to analyse heterogeneous populations, e.g., mixture models and stochastic population models. The available methods are, however, either incapable of simultaneously analysing different experimental conditions or are computationally demanding and difficult to apply. Furthermore, they do not account for biological information available in the literature. To overcome disadvantages of existing methods, we combine mixture models and ordinary differential equation (ODE) models. The ODE models provide a mechanistic description of the underlying processes while mixture models provide an easy way to capture variability. In a simulation study, we show that the class of ODE constrained mixture models can unravel the subpopulation structure and determine the sources of cell-to-cell variability. In addition, the method provides reliable estimates for kinetic rates and subpopulation characteristics. We use ODE constrained mixture modelling to study NGF-induced Erk1/2 phosphorylation in primary sensory neurones, a process relevant in inflammatory and neuropathic pain. We propose a mechanistic pathway model for this process and reconstructed static and dynamical subpopulation characteristics across experimental conditions. We validate the model predictions experimentally, which verifies the capabilities of ODE constrained mixture models. These results illustrate that ODE constrained mixture models can reveal novel mechanistic insights and possess a high sensitivity. Jan Hasenauer, Christine Hasenauer, Tim Hucho, Fabian J. Theis |
PLoS Comput. Biol. | 4 |
| 2013 | Visualizing edge-edge relations in graphsabstractGraphs are used to model relations between sets of objects. Objects are represented by vertices and relations by edges of the graph. Besides vertex-vertex relations, in some application domains also relations between edges exist. Our new visualization approach supports the investigation of both relation types in one diagram. Edge-edge relations are visualized as curves that are directly integrated into the node-link diagram that represents the object-relation structure. In contrast, vertex-vertex relations are illustrated distinguishably from edge-edge relations using straight links as representations. While the shape of links is used to differentiate between the relation types, the weights of the edge-edge relations are mapped to the width and color of the curves. To facilitate an extensive analysis of interrelations, our approach incorporates several interaction techniques that can be used for filtering and highlighting. The usability of our visualization is demonstrated with two case studies in the application domains of bioinformatics and financial services. Corinna Vehlow, Jan Hasenauer, Fabian J. Theis, Daniel Weiskopf |
PacificVis | 3 |
| 2013 | An automatic method for robust and fast cell detection in bright field images from high-throughput microscopyabstractBACKGROUND: In recent years, high-throughput microscopy has emerged as a powerful tool to analyze cellular dynamics in an unprecedentedly high resolved manner. The amount of data that is generated, for example in long-term time-lapse microscopy experiments, requires automated methods for processing and analysis. Available software frameworks are well suited for high-throughput processing of fluorescence images, but they often do not perform well on bright field image data that varies considerably between laboratories, setups, and even single experiments. RESULTS: In this contribution, we present a fully automated image processing pipeline that is able to robustly segment and analyze cells with ellipsoid morphology from bright field microscopy in a high-throughput, yet time efficient manner. The pipeline comprises two steps: (i) Image acquisition is adjusted to obtain optimal bright field image quality for automatic processing. (ii) A concatenation of fast performing image processing algorithms robustly identifies single cells in each image. We applied the method to a time-lapse movie consisting of ∼315,000 images of differentiating hematopoietic stem cells over 6 days. We evaluated the accuracy of our method by comparing the number of identified cells with manual counts. Our method is able to segment images with varying cell density and different cell types without parameter adjustment and clearly outperforms a standard approach. By computing population doubling times, we were able to identify three growth phases in the stem cell population throughout the whole movie, and validated our result with cell cycle times from single cell tracking. CONCLUSIONS: Our method allows fully automated processing and analysis of high-throughput bright field microscopy data. The robustness of cell detection and fast computation time will support the analysis of high-content screening experiments, on-line analysis of time-lapse experiments as well as development of methods to automatically track single-cell genealogies. Felix Buggenthin, Carsten Marr, Michael Schwarzfischer, Philipp S. Hoppe, Oliver Hilsenbeck, Timm Schroeder, Fabian J. Theis |
BMC Bioinform. | 7 |
| 2013 | Modeling of 2D diffusion processes based on microscopy data: parameter estimation and practical identifiability analysisabstractBACKGROUND: Diffusion is a key component of many biological processes such as chemotaxis, developmental differentiation and tissue morphogenesis. Since recently, the spatial gradients caused by diffusion can be assessed in-vitro and in-vivo using microscopy based imaging techniques. The resulting time-series of two dimensional, high-resolutions images in combination with mechanistic models enable the quantitative analysis of the underlying mechanisms. However, such a model-based analysis is still challenging due to measurement noise and sparse observations, which result in uncertainties of the model parameters. METHODS: We introduce a likelihood function for image-based measurements with log-normal distributed noise. Based upon this likelihood function we formulate the maximum likelihood estimation problem, which is solved using PDE-constrained optimization methods. To assess the uncertainty and practical identifiability of the parameters we introduce profile likelihoods for diffusion processes. RESULTS AND CONCLUSION: As proof of concept, we model certain aspects of the guidance of dendritic cells towards lymphatic vessels, an example for haptotaxis. Using a realistic set of artificial measurement data, we estimate the five kinetic parameters of this model and compute profile likelihoods. Our novel approach for the estimation of model parameters from image data as well as the proposed identifiability analysis approach is widely applicable to diffusion processes. The profile likelihood based method provides more rigorous uncertainty bounds in contrast to local approximation methods. Sabrina Hock, Jan Hasenauer, Fabian J. Theis |
BMC Bioinform. | 3 |
| 2013 | iVUN: interactive Visualization of Uncertain biochemical reaction NetworksabstractBACKGROUND: Mathematical models are nowadays widely used to describe biochemical reaction networks. One of the main reasons for this is that models facilitate the integration of a multitude of different data and data types using parameter estimation. Thereby, models allow for a holistic understanding of biological processes. However, due to measurement noise and the limited amount of data, uncertainties in the model parameters should be considered when conclusions are drawn from estimated model attributes, such as reaction fluxes or transient dynamics of biological species. METHODS AND RESULTS: We developed the visual analytics system iVUN that supports uncertainty-aware analysis of static and dynamic attributes of biochemical reaction networks modeled by ordinary differential equations. The multivariate graph of the network is visualized as a node-link diagram, and statistics of the attributes are mapped to the color of nodes and links of the graph. In addition, the graph view is linked with several views, such as line plots, scatter plots, and correlation matrices, to support locating uncertainties and the analysis of their time dependencies. As demonstration, we use iVUN to quantitatively analyze the dynamics of a model for Epo-induced JAK2/STAT5 signaling. CONCLUSION: Our case study showed that iVUN can be used to perform an in-depth study of biochemical reaction networks, including attribute uncertainties, correlations between these attributes and their uncertainties as well as the attribute dynamics. In particular, the linking of different visualization options turned out to be highly beneficial for the complex analysis tasks that come with the biological systems as presented here. Corinna Vehlow, Jan Hasenauer, Andrei Kramer, Andreas Raue, Sabine Hug, Jens Timmer, Nicole Radde, Fabian J. Theis, Daniel Weiskopf |
BMC Bioinform. | 8 |
| 2012 | A novel approach for resolving differences in single-cell gene expression patterns from zygote to blastocystabstractMOTIVATION: Single-cell experiments of cells from the early mouse embryo yield gene expression data for different developmental stages from zygote to blastocyst. To better understand cell fate decisions during differentiation, it is desirable to analyse the high-dimensional gene expression data and assess differences in gene expression patterns between different developmental stages as well as within developmental stages. Conventional methods include univariate analyses of distributions of genes at different stages or multivariate linear methods such as principal component analysis (PCA). However, these approaches often fail to resolve important differences as each lineage has a unique gene expression pattern which changes gradually over time yielding different gene expressions both between different developmental stages as well as heterogeneous distributions at a specific stage. Furthermore, to date, no approach taking the temporal structure of the data into account has been presented. RESULTS: We present a novel framework based on Gaussian process latent variable models (GPLVMs) to analyse single-cell qPCR expression data of 48 genes from mouse zygote to blastocyst as presented by (Guo et al., 2010). We extend GPLVMs by introducing gene relevance maps and gradient plots to provide interpretability as in the linear case. Furthermore, we take the temporal group structure of the data into account and introduce a new factor in the GPLVM likelihood which ensures that small distances are preserved for cells from the same developmental stage. Using our novel framework, it is possible to resolve differences in gene expressions for all developmental stages. Furthermore, a new subpopulation of cells within the 16-cell stage is identified which is significantly more trophectoderm-like than the rest of the population. The trophectoderm-like subpopulation was characterized by considerable differences in the expression of Id2, Gata4 and, to a smaller extent, Klf4 and Hand1. The relevance of Id2 as early markers for TE cells is consistent with previously published results. AVAILABILITY: The mappings were implemented based on Prof. Neil Lawrence's FGPLVM toolbox(1); extensions for relevance analysis and including the structure of the data can be obtained from one of the authors' homepage.(2) CONTACT: [email protected]. Florian A. Büttner, Fabian J. Theis |
Bioinform. | 2 |
| 2012 | On the hypothesis-free testing of metabolite ratios in genome-wide and metabolome-wide association studiesabstractBACKGROUND: Genome-wide association studies (GWAS) with metabolic traits and metabolome-wide association studies (MWAS) with traits of biomedical relevance are powerful tools to identify the contribution of genetic, environmental and lifestyle factors to the etiology of complex diseases. Hypothesis-free testing of ratios between all possible metabolite pairs in GWAS and MWAS has proven to be an innovative approach in the discovery of new biologically meaningful associations. The p-gain statistic was introduced as an ad-hoc measure to determine whether a ratio between two metabolite concentrations carries more information than the two corresponding metabolite concentrations alone. So far, only a rule of thumb was applied to determine the significance of the p-gain. RESULTS: Here we explore the statistical properties of the p-gain through simulation of its density and by sampling of experimental data. We derive critical values of the p-gain for different levels of correlation between metabolite pairs and show that B/(2*α) is a conservative critical value for the p-gain, where α is the level of significance and B the number of tested metabolite pairs. CONCLUSIONS: We show that the p-gain is a well defined measure that can be used to identify statistically significant metabolite ratios in association studies and provide a conservative significance cut-off for the p-gain for use in future association studies with metabolic traits. Ann-Kristin Petersen, Jan Krumsiek, Brigitte Wägele, Fabian J. Theis, Heinz-Erich Wichmann, Christian Gieger, Karsten Suhre |
BMC Bioinform. | 4 |
| 2012 | ICA over finite fields - Separability and algorithms
Harold W. Gutch, Peter Gruber 0002, Arie Yeredor, Fabian J. Theis |
Signal Process. | 4 |
| 2012 | The signal separation evaluation campaign (2007-2010): Achievements and remaining challenges
Emmanuel Vincent 0001, Shoko Araki, Fabian J. Theis, Guido Nolte, Pau Bofill, Hiroshi Sawada, Alexey Ozerov, Vikrham Gowreesunker, Dominik Lutter, Ngoc Q. K. Duong |
Signal Process. | 3 |
| 2010 | Structuring heterogeneous biological information using fuzzy clustering of k-partite graphsabstractBACKGROUND: Extensive and automated data integration in bioinformatics facilitates the construction of large, complex biological networks. However, the challenge lies in the interpretation of these networks. While most research focuses on the unipartite or bipartite case, we address the more general but common situation of k-partite graphs. These graphs contain k different node types and links are only allowed between nodes of different types. In order to reveal their structural organization and describe the contained information in a more coarse-grained fashion, we ask how to detect clusters within each node type. RESULTS: Since entities in biological networks regularly have more than one function and hence participate in more than one cluster, we developed a k-partite graph partitioning algorithm that allows for overlapping (fuzzy) clusters. It determines for each node a degree of membership to each cluster. Moreover, the algorithm estimates a weighted k-partite graph that connects the extracted clusters. Our method is fast and efficient, mimicking the multiplicative update rules commonly employed in algorithms for non-negative matrix factorization. It facilitates the decomposition of networks on a chosen scale and therefore allows for analysis and interpretation of structures on various resolution levels. Applying our algorithm to a tripartite disease-gene-protein complex network, we were able to structure this graph on a large scale into clusters that are functionally correlated and biologically meaningful. Locally, smaller clusters enabled reclassification or annotation of the clusters' elements. We exemplified this for the transcription factor MECP2. CONCLUSIONS: In order to cope with the overwhelming amount of information available from biomedical literature, we need to tackle the challenge of finding structures in large networks with nodes of multiple types. To this end, we presented a novel fuzzy k-partite graph partitioning algorithm that allows the decomposition of these objects in a comprehensive fashion. We validated our approach both on artificial and real-world data. It is readily applicable to any further problem. Mara L. Hartsperger, Florian Blöchl, Volker Stümpflen, Fabian J. Theis |
BMC Bioinform. | 4 |
| 2010 | Knowledge-based matrix factorization temporally resolves the cellular responses to IL-6 stimulationabstractBACKGROUND: External stimulations of cells by hormones, cytokines or growth factors activate signal transduction pathways that subsequently induce a re-arrangement of cellular gene expression. The analysis of such changes is complicated, as they consist of multi-layered temporal responses. While classical analyses based on clustering or gene set enrichment only partly reveal this information, matrix factorization techniques are well suited for a detailed temporal analysis. In signal processing, factorization techniques incorporating data properties like spatial and temporal correlation structure have shown to be robust and computationally efficient. However, such correlation-based methods have so far not be applied in bioinformatics, because large scale biological data rarely imply a natural order that allows the definition of a delayed correlation function. RESULTS: We therefore develop the concept of graph-decorrelation. We encode prior knowledge like transcriptional regulation, protein interactions or metabolic pathways in a weighted directed graph. By linking features along this underlying graph, we introduce a partial ordering of the features (e.g. genes) and are thus able to define a graph-delayed correlation function. Using this framework as constraint to the matrix factorization task allows us to set up the fast and robust graph-decorrelation algorithm (GraDe). To analyze alterations in the gene response in IL-6 stimulated primary mouse hepatocytes, we performed a time-course microarray experiment and applied GraDe. In contrast to standard techniques, the extracted time-resolved gene expression profiles showed that IL-6 activates genes involved in cell cycle progression and cell division. Genes linked to metabolic and apoptotic processes are down-regulated indicating that IL-6 mediated priming renders hepatocytes more responsive towards cell proliferation and reduces expenditures for the energy metabolism. CONCLUSIONS: GraDe provides a novel framework for the decomposition of large-scale 'omics' data. We were able to show that including prior knowledge into the separation task leads to a much more structured and detailed separation of the time-dependent responses upon IL-6 stimulation compared to standard methods. A Matlab implementation of the GraDe algorithm is freely available at http://cmb.helmholtz-muenchen.de/grade. Andreas Kowarsch, Florian Blöchl, Sebastian Bohl, Maria Saile, Norbert Gretz, Ursula Klingmüller, Fabian J. Theis |
BMC Bioinform. | 7 |
| 2010 | Odefy -- From discrete to continuous modelsabstractBACKGROUND: Phenomenological information about regulatory interactions is frequently available and can be readily converted to Boolean models. Fully quantitative models, on the other hand, provide detailed insights into the precise dynamics of the underlying system. In order to connect discrete and continuous modeling approaches, methods for the conversion of Boolean systems into systems of ordinary differential equations have been developed recently. As biological interaction networks have steadily grown in size and complexity, a fully automated framework for the conversion process is desirable. RESULTS: We present Odefy, a MATLAB- and Octave-compatible toolbox for the automated transformation of Boolean models into systems of ordinary differential equations. Models can be created from sets of Boolean equations or graph representations of Boolean networks. Alternatively, the user can import Boolean models from the CellNetAnalyzer toolbox, GINSim and the PBN toolbox. The Boolean models are transformed to systems of ordinary differential equations by multivariate polynomial interpolation and optional application of sigmoidal Hill functions. Our toolbox contains basic simulation and visualization functionalities for both, the Boolean as well as the continuous models. For further analyses, models can be exported to SQUAD, GNA, MATLAB script files, the SB toolbox, SBML and R script files. Odefy contains a user-friendly graphical user interface for convenient access to the simulation and exporting functionalities. We illustrate the validity of our transformation approach as well as the usage and benefit of the Odefy toolbox for two biological systems: a mutual inhibitory switch known from stem cell differentiation and a regulatory network giving rise to a specific spatial expression pattern at the mid-hindbrain boundary. CONCLUSIONS: Odefy provides an easy-to-use toolbox for the automatic conversion of Boolean models to systems of ordinary differential equations. It can be efficiently connected to a variety of input and output formats for further analysis and investigations. The toolbox is open-source and can be downloaded at http://cmb.helmholtz-muenchen.de/odefy. Jan Krumsiek, Sebastian Pölsterl, Dominik M. Wittmann, Fabian J. Theis |
BMC Bioinform. | 4 |
| 2010 | Patterns of Subnet Usage Reveal Distinct Scales of Regulation in the Transcriptional Regulatory Network of Escherichia coliabstractThe set of regulatory interactions between genes, mediated by transcription factors, forms a species' transcriptional regulatory network (TRN). By comparing this network with measured gene expression data, one can identify functional properties of the TRN and gain general insight into transcriptional control. We define the subnet of a node as the subgraph consisting of all nodes topologically downstream of the node, including itself. Using a large set of microarray expression data of the bacterium Escherichia coli, we find that the gene expression in different subnets exhibits a structured pattern in response to environmental changes and genotypic mutation. Subnets with fewer changes in their expression pattern have a higher fraction of feed-forward loop motifs and a lower fraction of small RNA targets within them. Our study implies that the TRN consists of several scales of regulatory organization: (1) subnets with more varying gene expression controlled by both transcription factors and post-transcriptional RNA regulation and (2) subnets with less varying gene expression having more feed-forward loops and less post-transcriptional RNA regulation. Carsten Marr, Fabian J. Theis, Larry S. Liebovitch, Marc-Thorsten Hütt |
PLoS Comput. Biol. | 2 |
| 2009 | ICA, kernel methods and nonnegativity: New paradigms for dynamical component analysis of fMRI data
Peter Gruber 0002, Anke Meyer-Bäse, Simon Y. Foo, Fabian J. Theis |
Eng. Appl. Artif. Intell. | 4 |
| 2009 | Hypergraphs and Cellular Networksabstract3,41Max Planck Institute for Dynamics of Complex Technical Systems, Magdeburg, Germany, 2Institute for Mathematical Optimization, Faculty of Mathematics, Otto-von-Guericke University Magdeburg, Magdeburg, Germany, 3Institute for Bioinformatics and Systems Biology, Helmholtz Zentrum Mu¨nchen—German Research Center forEnvironmental Health, Neuherberg, Germany, 4Max Planck Institute for Dynamics and Self-Organization, Go¨ttingen, Germany Steffen Klamt, Utz-Uwe Haus, Fabian J. Theis |
PLoS Comput. Biol. | 3 |
| 2009 | Spatial Analysis of Expression Patterns Predicts Genetic Interactions at the Mid-Hindbrain BoundaryabstractThe isthmic organizer mediating differentiation of mid- and hindbrain during vertebrate development is characterized by a well-defined pattern of locally restricted gene expression domains around the mid-hindbrain boundary (MHB). This pattern is established and maintained by a regulatory network between several transcription and secreted factors that is not yet understood in full detail. In this contribution we show that a Boolean analysis of the characteristic spatial gene expression patterns at the murine MHB reveals key regulatory interactions in this network. Our analysis employs techniques from computational logic for the minimization of Boolean functions. This approach allows us to predict also the interplay of the various regulatory interactions. In particular, we predict a maintaining, rather than inducing, effect of Fgf8 on Wnt1 expression, an issue that remained unclear from published data. Using mouse anterior neural plate/tube explant cultures, we provide experimental evidence that Fgf8 in fact only maintains but does not induce ectopic Wnt1 expression in these explants. In combination with previously validated interactions, this finding allows for the construction of a regulatory network between key transcription and secreted factors at the MHB. Analyses of Boolean, differential equation and reaction-diffusion models of this network confirm that it is indeed able to explain the stable maintenance of the MHB as well as time-courses of expression patterns both under wild-type and various knock-out conditions. In conclusion, we demonstrate that similar to temporal also spatial expression patterns can be used to gain information about the structure of regulatory networks. We show, in particular, that the spatial gene expression patterns around the MHB help us to understand the maintenance of this boundary on a systems level. Dominik M. Wittmann, Florian Blöchl, Dietrich Trümbach, Wolfgang Wurst, Nilima Prakash, Fabian J. Theis |
PLoS Comput. Biol. | 6 |
| 2009 | Reconstruction of graphs based on random walks
Dominik M. Wittmann, Daniel Schmidl, Florian Blöchl, Fabian J. Theis |
Theor. Comput. Sci. | 4 |
| 2008 | Knowledge-based gene expression classification via matrix factorizationabstractMOTIVATION: Modern machine learning methods based on matrix decomposition techniques, like independent component analysis (ICA) or non-negative matrix factorization (NMF), provide new and efficient analysis tools which are currently explored to analyze gene expression profiles. These exploratory feature extraction techniques yield expression modes (ICA) or metagenes (NMF). These extracted features are considered indicative of underlying regulatory processes. They can as well be applied to the classification of gene expression datasets by grouping samples into different categories for diagnostic purposes or group genes into functional categories for further investigation of related metabolic pathways and regulatory networks. RESULTS: In this study we focus on unsupervised matrix factorization techniques and apply ICA and sparse NMF to microarray datasets. The latter monitor the gene expression levels of human peripheral blood cells during differentiation from monocytes to macrophages. We show that these tools are able to identify relevant signatures in the deduced component matrices and extract informative sets of marker genes from these gene expression profiles. The methods rely on the joint discriminative power of a set of marker genes rather than on single marker genes. With these sets of marker genes, corroborated by leave-one-out or random forest cross-validation, the datasets could easily be classified into related diagnostic categories. The latter correspond to either monocytes versus macrophages or healthy vs Niemann Pick C disease patients. Reinhard Schachtner, Dominik Lutter, P. Knollmüller, Ana Maria Tomé, Fabian J. Theis, Gerd Schmitz 0001, Martin Stetter, Pedro Gómez-Vilda, Elmar Wolfgang Lang |
Bioinform. | 5 |
| 2008 | Analyzing M-CSF dependent monocyte/macrophage differentiation: Expression modes and meta-modes derived from an independent component analysisabstractBACKGROUND: The analysis of high-throughput gene expression data sets derived from microarray experiments still is a field of extensive investigation. Although new approaches and algorithms are published continuously, mostly conventional methods like hierarchical clustering algorithms or variance analysis tools are used. Here we take a closer look at independent component analysis (ICA) which is already discussed widely as a new analysis approach. However, deep exploration of its applicability and relevance to concrete biological problems is still missing. In this study, we investigate the relevance of ICA in gaining new insights into well characterized regulatory mechanisms of M-CSF dependent macrophage differentiation. RESULTS: Statistically independent gene expression modes (GEM) were extracted from observed gene expression signatures (GES) through ICA of different microarray experiments. From each GEM we deduced a group of genes, henceforth called sub-mode. These sub-modes were further analyzed with different database query and literature mining tools and then combined to form so called meta-modes. With them we performed a knowledge-based pathway analysis and reconstructed a well known signal cascade. CONCLUSION: We show that ICA is an appropriate tool to uncover underlying biological mechanisms from microarray data. Most of the well known pathways of M-CSF dependent monocyte to macrophage differentiation can be identified by this unsupervised microarray data analysis. Moreover, recent research results like the involvement of proliferation associated cellular mechanisms during macrophage differentiation can be corroborated. Dominik Lutter, Peter Ugocsai, Margot Grandl, Evelyn Orso, Fabian J. Theis, Elmar Wolfgang Lang, Gerd Schmitz 0001 |
BMC Bioinform. | 5 |
| 2008 | Hybridizing sparse component analysis with genetic algorithms for microarray analysis
Kurt Stadlthanner, Fabian J. Theis, Elmar Wolfgang Lang, Ana Maria Tomé, Carlos García Puntonet, Juan Manuel Górriz |
Neurocomputing | 2 |
| 2008 | A robust model for spatiotemporal dependencies
Fabian J. Theis, Peter Gruber 0002, Ingo R. Keck, Elmar Wolfgang Lang |
Neurocomputing | 1 |
| 2007 | Exploiting Blind Matrix Decomposition Techniques to Identify Diagnostic Marker Genes
Reinhard Schachtner, Dominik Lutter, Fabian J. Theis, Elmar Wolfgang Lang, Ana Maria Tomé, Gerd Schmitz 0001 |
ICANN (2) | 3 |
| 2007 | Blind Matrix Decomposition Via Genetic Optimization of Sparseness and Nonnegativity Constraints
Kurt Stadlthanner, Fabian J. Theis, Elmar Wolfgang Lang, Ana Maria Tomé, Carlos García Puntonet |
ICANN (1) | 2 |
| 2007 | Statistical Analysis of Sample-Size Effects in ICA
J. Michael Herrmann, Fabian J. Theis |
IDEAL | 2 |
| 2007 | Sparse Nonnegative Matrix Factorization with Genetic Algorithms for Microarray AnalysisabstractNonnegative Matrix Factorization (NMF) has proven to be a useful tool for the analysis of nonnegative multivariate data. Gene expression profiles naturally conform to assumptions about data formats raised by NMF. However, it is known not to lead to unique results concerning the component signals extracted. In this paper we consider an extension of the NMF algorithm which provides unique solutions whenever the underlying component signals are sufficiently sparse. A new sparseness measure is proposed most appropriate to suitably transformed gene expression profiles. The resulting fitness function is discontinuous and exhibits many local minima, hence we use a genetic algorithm for its optimization. The algorithm is applied to toy data to investigate its properties as well as to a microarray data set related to Pseudo-Xanthoma Elasticum (PXE). Kurt Stadlthanner, Dominik Lutter, Fabian J. Theis, Elmar Wolfgang Lang, Ana Maria Tomé, Petia Georgieva, Carlos García Puntonet |
IJCNN | 3 |
| 2007 | Joint low-rank approximation for extracting non-Gaussian subspaces
Motoaki Kawanabe, Fabian J. Theis |
Signal Process. | 2 |
| 2006 | Sparseness by Iterative Projections Onto SpheresabstractMany interesting signals share the property of being sparsely active. The search for such sparse components within a data set commonly involves a linear or nonlinear projection step in order to fulfill the sparseness constraints. In addition to the proximity measure used for the projection, the result of course is also intimately connected with the actual definition of the sparseness criterion. In this work, we introduce a novel sparseness measure and apply it to the problem of finding a sparse projection of a given signal. Here, sparseness is defined as the fixed ratio of p-over 2-norm, and existence and uniqueness of the projection holds. This framework extends previous work by Hoyer in the case of p = 1, where it is easy to give a deterministic, more or less closed-form solution. This is not possible for p ≠ 1, so we introduce an algorithm based on alternating projections onto spheres (POSH), which is similar to the projection onto convex sets (OCS). Although the assumption of convexity does not hold in our setting, we observe not only convergence of the algorithm, but also convergence to the correct minimal distance solution. Indications for a proof of this surprising property are given. Simulations confirm these results. Fabian J. Theis, Toshihisa Tanaka 0001 |
ICASSP (5) | 1 |
| 2006 | A Fast Predictive Lossless Coder for fMRI Data SetsabstractWe present a novel lossless compression algorithm which compresses sequences of three-dimensional (3D) volumes collected during a functional magnetic resonance imaging (fMRI) experiment. The large data sets involved in this popular biomedical application necessitate fast and efficient compression methods. We propose to use 3D prediction, temporal decorrelation and entropy coding with context modeling for encoding the fMRI scans after preprocessing with the region-of-interest (ROI) masking. The proposed algorithm is conceptually simple and can achieve fast implementation and efficient coding performance. We illustrate computer simulations to show advantages over conventional coding methods. Toshihisa Tanaka 0001, Yumi Murakami, Fabian J. Theis |
ICIP | 3 |
| 2006 | Region of Interest Based Independent Component Analysis
Ingo R. Keck, Jan Churan, Fabian J. Theis, Peter Gruber 0002, Elmar Wolfgang Lang, Carlos García Puntonet |
ICONIP (1) | 3 |
| 2006 | Stability Analysis of an Unsupervised Competitive Neural NetworkabstractUnsupervised competitive neural networks (UCNN) are an established technique in pattern recognition for feature extraction and cluster analysis. A novel model of an unsupervised competitive neural network implementing a multi—time scale dynamics is proposed in this paper. The global asymptotic stability of the equilibrium points of this continuous—time recurrent system whose weights are adapted based on a competitive learning law is mathematically analyzed. The proposed neural network and the derived results are compared with those obtained from other multi—time scale architectures. Anke Meyer-Bäse, Vera Thümmler, Fabian J. Theis |
IJCNN | 3 |
| 2006 | New Riemannian metrics for speeding-up the convergence of over- and underdetermined ICAabstractIn this paper some alternative Riemannian metrics are defined on the parameter space of non-square matrices, corresponding to various translations defined therein. Such metrics allow the authors to derive novel learning rules for two ICA based algorithms for over-determined blind source separation (BSS), which tries to separate less sources from more sensors. Computer simulations show a significant improvement of the convergence speed when second-order translations are employed in contrast to their first-order counterparts, extending known results for complete BSS Stefano Squartini, Francesco Piazza, Fabian J. Theis |
ISCAS | 3 |
| 2006 | On the use of joint diagonalization in blind signal processingabstractBlind source separation (BSS) tries to decompose a given multivariate data set into the product of a mixing matrix and a source vector, both of which are unknown. The sources can be recovered if we pose additional constraints to this model. One class of BSS algorithms is given by algebraic BSS, which recovers the mixing structure by jointly diagonalizing various source condition matrices corresponding to different source models. We review classical BSS algorithms such as FOBI, JADE, AMUSE, SOBI, TDSEP and SONS within this framework; combination of the respective source conditions can then yield additional algorithms as implemented e.g. by JADETD. Extensions to dependent component analysis models such as spatiotemporal or multidimensional BSS are discussed Fabian J. Theis, Yujiro Inouye |
ISCAS | 1 |
| 2006 | Towards a general independent subspace analysisabstractThe increasingly popular independent component analysis (ICA) may only be applied to data following the generative ICA model in order to guarantee algorithmindependent and theoretically valid results. Subspace ICA models generalize the assumption of component independence to independence between groups of components. They are attractive candidates for dimensionality reduction methods, however are currently limited by the assumption of equal group sizes or less general semi-parametric models. By introducing the concept of irreducible independent subspaces or components, we present a generalization to a parameter-free mixture model. Moreover, we relieve the condition of at-most-one-Gaussian by including previous results on non-Gaussian component analysis. After introducing this general model, we discuss joint block diagonalization with unknown block sizes, on which we base a simple extension of JADE to algorithmically perform the subspace analysis. Simulations confirm the feasibility of the algorithm. 1 Independent subspace analysis A random vector Y is called an independent component of the random vector X, if there exists an invertible matrix A and a decomposition X = A(Y, Z) such that Y and Z are stochastically independent. The goal of a general independent subspace analysis (ISA) or multidimensional independent component analysis is the decomposition of an arbitrary random vector X into independent components. If X is to be decomposed into one-dimensional components, this coincides with ordinary independent component analysis (ICA). Similarly, if the independent components are required to be of the same dimension k , then this is denoted by multidimensional ICA of fixed group size k or simply k -ISA. So 1-ISA is equivalent to ICA. 1.1 Why extend ICA? An important structural aspect in the search for decompositions is the knowledge of the number of solutions i.e. the indeterminacies of the problem. Without it, the result of any ICA or ISA algorithm cannot be compared with other solutions, so for instance blind source separation (BSS) would be impossible. Clearly, given an ISA solution, invertible transforms in each component (scaling matrices L) as well as permutations of components of the same dimension (permutation matrices P) give again an ISA of X. And indeed, in the special case of ICA, scaling and permutation are already all indeterminacies given that at most one Gaussian is contained in X [6]. This is one of the key theoretical results in ICA, allowing the usage of ICA for solving BSS problems and hence stimulating many applications. It has been shown that also for k -ISA, scalings and permutations as above are the only indeterminacies [11], given some additional rather weak restrictions to the model. However, a serious drawback of k -ISA (and hence of ICA) lies in the fact that the requirement fixed group-size k does not allow us to apply this analysis to an arbitrary random vector. Indeed, 4 crosstalking error Fabian J. Theis |
NIPS | 1 |
| 2006 | Blind source separation based on self-organizing neural network
Anke Meyer-Bäse, Peter Gruber 0002, Fabian J. Theis, Simon Y. Foo |
Eng. Appl. Artif. Intell. | 3 |
| 2006 | Denoising using local projective subspace methods
Peter Gruber 0002, Kurt Stadlthanner, Matthias Böhm 0002, Fabian J. Theis, Elmar Wolfgang Lang, Ana Maria Tomé, Ana R. Teixeira, Carlos García Puntonet, Juan Manuel Górriz |
Neurocomputing | 4 |
| 2006 | Separation of water artifacts in 2D NOESY protein spectra using congruent matrix pencils
Kurt Stadlthanner, Ana Maria Tomé, Fabian J. Theis, Elmar Wolfgang Lang, Wolfram Gronwald, Hans Robert Kalbitzer |
Neurocomputing | 3 |
| 2006 | On the use of sparse signal decomposition in the analysis of multi-channel surface electromyograms
Fabian J. Theis, Gonzalo A. García |
Signal Process. | 1 |
| 2006 | Median-based clustering for underdetermined blind signal processingabstractIn underdetermined blind source separation, more sources are to be extracted from less observed mixtures without knowing both sources and mixing matrix. k-means-style clustering algorithms are commonly used to do this algorithmically given sufficiently sparse sources, but in any case other than deterministic sources, this lacks theoretical justification. After establishing that mean-based algorithms converge to wrong solutions in practice, we propose a median-based clustering scheme. Theoretical justification as well as algorithmic realizations (both online and batch) are given and illustrated by some examples. Fabian J. Theis, Carlos García Puntonet, Elmar Wolfgang Lang |
IEEE Signal Process. Lett. | 1 |
| 2005 | Bispectrum-Based Statistical Tests for VAD
Juan Manuel Górriz, Javier Ramírez 0001, Carlos García Puntonet, Fabian J. Theis, Elmar Wolfgang Lang |
ICANN (2) | 4 |
| 2005 | Functional MRI Analysis by a Novel Spatiotemporal ICA Algorithm
Fabian J. Theis, Peter Gruber 0002, Ingo R. Keck, Elmar Wolfgang Lang |
ICANN (1) | 1 |
| 2005 | A Fast and Efficient Method for Compressing fMRI Data Sets
Fabian J. Theis, Toshihisa Tanaka 0001 |
ICANN (2) | 1 |
| 2005 | An algorithm for automatic assignment of artifact-related independent components in biomedical signal analysisabstractIn this work an automatic assignment tool for estimated independent components within an independent component analysis is presented. The tool is applied to the problem of removing the water resonance and related artifacts from multi-dimensional proton NMR spectra. The algorithm uses local PCA to approximate the water artifact and defines a suitable cost function which is optimized using simulated annealing. The blind extraction of artifact-related source signals is effected by a recently developed algorithm called dAMUSE. Matthias Böhm 0002, Kurt Stadlthanner, Elmar Wolfgang Lang, Fabian J. Theis, Peter Gruber 0002, Ana Maria Tomé, Ana R. Teixeira, Carlos García Puntonet |
IJCNN | 4 |
| 2005 | On model identifiability in analytic postnonlinear ICA
Fabian J. Theis, Peter Gruber 0002 |
Neurocomputing | 1 |
| 2005 | Sparse component analysis and blind source separation of underdetermined mixturesabstractIn this letter, we solve the problem of identifying matrices S is an element of R(n x N) and A is an element of R(m x n) knowing only their multiplication X = AS, under some conditions, expressed either in terms of A and sparsity of S (identifiability conditions), or in terms of X (sparse component analysis (SCA) conditions). We present algorithms for such identification and illustrate them by examples. Pando G. Georgiev, Fabian J. Theis, Andrzej Cichocki |
IEEE Trans. Neural Networks | 2 |
| 2004 | Separability of analytic postnonlinear blind source separation with bounded sources
Fabian J. Theis, Peter Gruber 0002 |
ESANN | 1 |
| 2004 | Robust overcomplete matrix recovery for sparse sources using a generalized Hough transform
Fabian J. Theis, Pando G. Georgiev, Andrzej Cichocki |
ESANN | 1 |
| 2004 | Linearization identification and an application to BSS using a SOM
Fabian J. Theis, Elmar Wolfgang Lang |
ESANN | 1 |
| 2004 | Blind source separation and sparse component analysis of overcomplete mixturesabstractWe formulate conditions (k-SCA-conditions) under which we can represent a given (m/spl times/N)-matrix, X, (data set) uniquely (up to scaling and permutation) as a multiplication of m/spl times/n and n/spl times/N matrices, A and S, (often called mixing matrix or dictionary and source matrix, respectively), such that S is sparse of level n-m+k in the sense that each column of S has at least n-m+k zero elements. We call this the k-sparse component analysis problem (k-SCA). Conditions on a matrix, S, are presented such that the k-SCA-conditions are satisfied for the matrix X=AS, where A is an arbitrary matrix from some class. This is the blind source separation problem and the above conditions are called identifiability conditions. We present new algorithms for matrix identification (under k-SCA-conditions), and for source recovery (under identifiability conditions). The methods are illustrated with examples, showing good separation of the high-frequency part of mixtures of images after appropriate sparsification. Pando G. Georgiev, Fabian J. Theis, Andrzej Cichocki |
ICASSP (5) | 2 |
| 2004 | Denoising using local ICA and kernel-PCAabstractWe present a denoising algorithm for enhancing noisy signals based on local independent component analysis (ICA). This is done by applying ICA to the signal in localized delayed coordinates. The components resembling the signals can be detected by various criteria depending on the nature of the signal. Estimators of kurtosis or the variance of the autocorrelation have been considered. The algorithm proposed can favorably be applied to the problem of denoising multidimensional data like images or fMRI data sets. In comparison to denoising algorithms using wavelets, Wiener filters and kernel PCA the local PCA and ICA algorithms perform considerably better. We provide applications of the algorithm to images and the analysis of protein NMR spectra. Peter Gruber 0002, Fabian J. Theis, Kurt Stadlthanner, Elmar Wolfgang Lang, Ana Maria Tomé, Ana R. Teixeira |
IJCNN | 2 |
| 2004 | 3D spatial analysis of fMRI data: a comparison of ICA and GLM analysis on a word perception taskabstractWe discuss a comparative 3D spatial analysis of fMRI data taken during a combined word perception and motor task. We show that a classical GLM analysis using SPM does not yield reasonable results. Only with BSS techniques using the fastICA algorithm can we get meaningful and interesting results. The event-based experiment was part of a study to investigate the network of neurons involved in the perception of speech and the decoding of auditory speech stimuli. Corresponding to 4 different stimuli different independent components (IC) could be identified in the auditory cortex and, most interesting, an IC representing a network of 3 simultaneously active areas in the inferior frontal gyrus could be detected. Ingo R. Keck, Fabian J. Theis, Peter Gruber 0002, Elmar Wolfgang Lang, Karsten Specht, Carlos García Puntonet |
IJCNN | 2 |
| 2004 | Clustering of dependent components: a new paradigm for fMRI signal detectionabstractAbs. Anke Meyer-Bäse, Fabian J. Theis, Oliver Lange, Axel Wismüller |
IJCNN | 2 |
| 2004 | Kernel-PCA denoising of artifact-free protein NMR spectraabstractMultidimensional /sup 1/H NMR spectra of biomolecules dissolved in light water are contaminated by an intense water artifact. Generalized eigenvalue decomposition methods using congruent matrix pencils are used to separate the water artefact from the protein spectra. Due to the statistical separation process, however, noise is introduced into the reconstructed spectra. Hence Kernel-based denoising techniques are discussed to obtain noise- and artifact-free 2D NOESY NMR spectra of proteins. Kurt Stadlthanner, Elmar Wolfgang Lang, Peter Gruber 0002, Fabian J. Theis, Ana Maria Tomé, Ana R. Teixeira, Carlos García Puntonet |
IJCNN | 4 |
| 2004 | Postnonlinear blind source separation via linearization identificationabstractIn the first part of the paper, the one-dimensional functional equation g(y(t))=cg(z(t)) with known functions y and z and constant c is studied. Its indeterminacies are calculated, and an algorithm for approximating g is proposed. Then, this linearization identification algorithm is applied to the postnonlinear blind source separation (BSS) problem. In the case of bounded sources, a self-organizing map is used to approximate the boundary, and the postnonlinearity estimation is reduced to the one-dimensional equation from above. For super Gaussian sources, the density maxima are interpolated by performing linear BSS within concentric rings. Postnonlinearity estimation using ring approximation separates the mixtures. Fabian J. Theis, Elmar Wolfgang Lang |
IJCNN | 1 |
| 2004 | A geometric algorithm for overcomplete linear ICA
Fabian J. Theis, Elmar Wolfgang Lang, Carlos García Puntonet |
Neurocomputing | 1 |
| 2004 | A New Concept for Separability Problems in Blind Source SeparationabstractThe goal of blind source separation (BSS) lies in recovering the original independent sources of a mixed random vector without knowing the mixing structure. A key ingredient for performing BSS successfully is to know the indeterminacies of the problem-that is, to know how the separating model relates to the original mixing model (separability). For linear BSS, Comon (1994) showed using the Darmois-Skitovitch theorem that the linear mixing matrix can be found except for permutation and scaling. In this work, a much simpler, direct proof for linear separability is given. The idea is based on the fact that a random vector is independent if and only if the Hessian of its logarithmic density (resp. characteristic function) is diagonal everywhere. This property is then exploited to propose a new algorithm for performing BSS. Furthermore, first ideas of how to generalize separability results based on Hessian diagonalization to more complicated nonlinear models are studied in the setting of postnonlinear BSS. Fabian J. Theis |
Neural Comput. | 1 |
| 2004 | Uniqueness of complex and multidimensional independent component analysis
Fabian J. Theis |
Signal Process. | 1 |
| 2003 | Local features in biomedical image clusters extracted with independent component analysisabstractA neural network model for the identification and classification of malign and benign skin lesions from ALA-induced fluorescence images is presented. A self-organizing feature map or generative topographic mapping is used to cluster images patches according to their inherent local features, which then can be extracted with ICA. These components are used to distinguish skin cancer from benign lesions achieving an average classification rate of 70% so far. Christoph Bauer, Fabian J. Theis, Wolfgang Bäumler, Elmar Wolfgang Lang |
IJCNN | 2 |
| 2003 | SOMICA - an application of self-organizing maps to geometric independent component analysisabstractGuided by the principles of geometric independent component analysis (ICA), we present a new approach (SOM-ICA) to linear geometric ICA using self-organizing map (SOM). We observe a considerable improvement in separation quality of different distributions, albeit at high computational costs. The SOMICA algorithm is therefore primarily interesting from a theoretical point of view bringing together ICA and SOMs; this intersection could lead to new proofs in geometric ICA based on similar theorems in the SOM theory. Fabian J. Theis, Carlos García Puntonet, Elmar Wolfgang Lang |
IJCNN | 1 |
| 2003 | Linear Geometric ICA: Fundamentals and AlgorithmsabstractGeometric algorithms for linear independent component analysis (ICA) have recently received some attention due to their pictorial description and their relative ease of implementation. The geometric approach to ICA was proposed first by Puntonet and Prieto (1995). We will reconsider geometric ICA in a theoretic framework showing that fixed points of geometric ICA fulfill a geometric convergence condition (GCC), which the mixed images of the unit vectors satisfy too. This leads to a conjecture claiming that in the nongaussian unimodal symmetric case, there is only one stable fixed point, implying the uniqueness of geometric ICA after convergence. Guided by the principles of ordinary geometric ICA, we then present a new approach to linear geometric ICA based on histograms observing a considerable improvement in separation quality of different distributions and a sizable reduction in computational cost, by a factor of 100, compared to the ordinary geometric approach. Furthermore, we explore the accuracy of the algorithm depending on the number of samples and the choice of the mixing matrix, and compare geometric algorithms with classical ICA algorithms, namely, Extended Infomax and FastICA. Finally, we discuss the problem of high-dimensional data sets within the realm of geometrical ICA algorithms. Fabian J. Theis, Andreas Jung, Carlos García Puntonet, Elmar Wolfgang Lang |
Neural Comput. | 1 |
| 2002 | How to generalize geometric ICA to higher dimensions
Fabian J. Theis, Elmar Wolfgang Lang |
ESANN | 1 |
| 2002 | Geometric overcomplete ICA
Fabian J. Theis, Elmar Wolfgang Lang |
ESANN | 1 |
| 2002 | Overcomplete ICA with a Geometric Algorithm
Fabian J. Theis, Elmar Wolfgang Lang, Tobias Westenhuber, Carlos García Puntonet |
ICANN | 1 |
| 2002 | Comparison of maximum entropy and minimal mutual information in a nonlinear setting
Fabian J. Theis, Christoph Bauer, Elmar Wolfgang Lang |
Signal Process. | 1 |