EDBT 2026 Demo / reviewers in the wild / expert
Roland Eils
dblp:63/1724
· DBLP profile ↗
57ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-0034-4036ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 44 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9Artificial intelligence and machine learning · 6 · 3 since 2021Security and privacy · 3Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Managing workflow executions with WESkitabstractSUMMARY: In biomedical research, managing computational workflows across numerous projects-with varying parameters, tools, and environments-creates major challenges in scalability, reproducibility, and collaboration. Here, we present WESkit, an implementation of the Global Alliance for Genomics and Health (GA4GH) Workflow Execution Service (WES) interface, designed to streamline the execution, monitoring, and documentation of data processing workflows. It addresses the complexities involved in managing numerous executions with varying parameters across diverse research projects. Supporting both Snakemake and Nextflow, the system enables consistent automation and centralized monitoring, which benefits research groups aiming for long-term reproducibility and scalable collaboration. Its suitability for larger teams and service units is further enhanced by seamless integration into cloud environments, contributing to the GA4GH cloud framework. AVAILABILITY AND IMPLEMENTATION: The software WESkit is available under MIT license at the GitLab repository (https://gitlab.com/one-touch-pipeline/weskit). The WESkit main repository is archived at Software Heritage (https://archive.softwareheritage.org/browse/origin/directory/?origin_url=https://gitlab.com/one-touch-pipeline/weskit/api.git) and can be found using "one-touch-pipeline/weskit" term in the search section. Valentin Schneider-Lunitz, Philip R. Kensche, Landfried Kraatz, Philipp Strubel, Stefan Borufka, Gurudeep Parala, Alexander Kanitz, Roland Eils, Ivo Buchhalter, Sven Twardziok |
Bioinform. | 8 |
| 2025 | JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation ModelabstractLarge language models (LLMs) have revolutionized natural language processing and are increasingly applied to other sequential data types, including genetic sequences. However, adapting LLMs to genetics presents significant challenges. Capturing complex genomic interactions requires modeling long-range global dependencies within DNA sequences, where interactions often span over 10,000 base pairs, even within a single gene. This poses substantial computational demands under conventional model architectures and training paradigms. Additionally, traditional LLM training approaches are suboptimal for DNA sequences: autoregressive training, while efficient for training, only supports unidirectional sequence understanding. However, DNA is inherently bidirectional. For instance, bidirectional promoters regulate gene expression in both directions and govern approximately 11% of human gene expression. Masked language models (MLMs) enable bidirectional understanding. However, they are inefficient since only masked tokens contribute to loss calculations at each training step. To address these limitations, we introduce JanusDNA, the first bidirectional DNA foundation model built upon a novel pretraining paradigm, integrating the optimization efficiency of autoregressive modeling with the bidirectional comprehension capability of masked modeling. JanusDNA's architecture leverages a Mamba-Attention Mixture-of-Experts (MoE) design, combining the global, high-resolution context awareness of attention mechanisms with the efficient sequential representation learning capabilities of Mamba. The MoE layers further enhance the model's capacity through sparse parameter scaling, while maintaining manageable computational costs. Notably, JanusDNA can process up to 1 million base pairs at single-nucleotide resolution on a single 80GB GPU using its hybrid architecture. Extensive experiments and ablation studies demonstrate that JanusDNA achieves new state-of-the-art performance on three genomic representation benchmarks. Remarkably, JanusDNA surpasses models with 250x more activated parameters, underscoring its efficiency and effectiveness. Code available at https://anonymous.4open.science/r/JanusDNA/. Qihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann, Roland Eils, Benjamin Wild |
NeurIPS | 6 |
| 2025 | A review of breast cancer histopathology image analysis with deep learning: Challenges, innovations, and clinical integrationabstractBreast cancer (BC) is the most frequently diagnosed cancer among women and a leading cause of cancer-related mortality globally. Accurate and timely diagnosis is essential for improving patient outcomes. However, traditional histopathological assessments are labor-intensive and subjective, leading to inter-observer variability and diagnostic inconsistencies, especially in resource-limited settings. Furthermore, variability in tissue staining, limited availability of standardized annotated datasets, and subtle morphological patterns complicate the consistent characterization of tumors. Deep learning (DL) has recently emerged as a transformative technology in breast cancer pathology, providing automated and objective solutions for cancer detection, classification, and segmentation from histopathological images. This review systematically evaluates advanced deep learning (DL) architectures, including convolutional neural networks (CNNs), generative adversarial networks (GANs), autoencoders, deep belief networks (DBNs), extreme learning machines (ELMs), and transformer-based models such as Vision Transformers (ViTs) as well as transfer learning, attention-based explainable AI techniques, and multimodal integration to address these diagnostic challenges. Analyzing 199 references, including 182 peer-reviewed studies published between 2014 and 2025 and 17 reputable online sources (websites, databases, etc.), we identify key innovations, limitations, and opportunities for future research. Furthermore, we explore the critical roles of synthetic data augmentation, explainable AI (XAI), and multimodal integration to enhance clinical trust, model interpretability, and diagnostic precision, ultimately facilitating personalized and efficient patient care. Inayatul Haq, Haomin Liang, Rashid Khan, Roland Eils, Bingding Huang |
Image Vis. Comput. | 7 |
| 2023 | Differentiable sorting for censored time-to-event dataabstractSurvival analysis is a crucial semi-supervised task in machine learning with significant real-world applications, especially in healthcare. The most common approach to survival analysis, Cox’s partial likelihood, can be interpreted as a ranking model optimized on a lower bound of the concordance index. We follow these connections further, with listwise ranking losses that allow for a relaxation of the pairwise independence assumption. Given the inherent transitivity of ranking, we explore differentiable sorting networks as a means to introduce a stronger transitive inductive bias during optimization. Despite their potential, current differentiable sorting methods cannot account for censoring, a crucial aspect of many real-world datasets. We propose a novel method, Diffsurv, to overcome this limitation by extending differentiable sorting methods to handle censored tasks. Diffsurv predicts matrices of possible permutations that accommodate the label uncertainty introduced by censored samples. Our experiments reveal that Diffsurv outperforms established baselines in various simulated and real-world risk prediction scenarios. Furthermore, we demonstrate the algorithmic advantages of Diffsurv by presenting a novel method for top-k risk prediction that surpasses current methods. Andre Vauvelle, Benjamin Wild, Roland Eils, Spiros C. Denaxas |
NeurIPS | 3 |
| 2023 | Eos and OMOCL: Towards a seamless integration of openEHR records into the OMOP Common Data ModelabstractBACKGROUND: The reuse of data from electronic health records (EHRs) for research purposes promises to improve the data foundation for clinical trials and may even support to enable them. Nevertheless, EHRs are characterized by both, heterogeneous structure and semantics. To standardize this data for research, the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) standard has recently seen an increase in use. However, the conversion of these EHRs into the OMOP CDM requires complex and resource intensive Extract Transform and Load (ETL) processes. This hampers the reuse of clinical data for research. To solve the issues of heterogeneity of EHRs and the lack of semantic precision on the care site, the openEHR standard has recently seen wider adoption. A standardized process to integrate openEHR records into the CDM potentially lowers the barriers of making EHRs accessible for research. Yet, a comprehensive approach about the integration of openEHR records into the OMOP CDM has not yet been made. METHODS: We analyzed both standards and compared their models to identify possible mappings. Based on this, we defined the necessary processes to transform openEHR records into CDM tables. We also discuss the limitation of openEHR with its unspecific demographics model and propose two possible solutions. RESULTS: We developed the OMOP Conversion Language (OMOCL) which enabled us to define a declarative openEHR archetype-to-CDM mapping language. Using OMOCL, it was possible to define a set of mappings. As a proof-of-concept, we implemented the Eos tool, which uses the OMOCL-files to successfully automatize the ETL from real-world and sample EHRs into the OMOP CDM. DISCUSSION: Both Eos and OMOCL provide a way to define generic mappings for an integration of openEHR records into OMOP. Thus, it represents a significant step towards achieving interoperability between the clinical and the research data domains. However, the transformation of openEHR data into the less expressive OMOP CDM leads to a loss of semantics. Severin Kohler, Diego Boscá, Florian Kärcher, Birger Haarbrandt, Manuel Prinz, Michael Marschollek, Roland Eils |
J. Biomed. Informatics | 7 |
| 2021 | Knowledge bases and software support for variant interpretation in precision oncologyabstract[14 June 2021]: notice amended to state that the funding information has been corrected to remove a duplicate reference to two funding organizations introduced in error during the initial correction. In the originally published version of this manuscript, requested author amendments to the Funding section were inadvertently omitted prior to publishing. The section should read: ``German Federal Ministry of Research and Education (01ZZ1802); Physician-Scientist Program of the University of Heidelberg, Faculty of Medicine, DKTK (German Cancer Consortium) School of Oncology and the Cancer Core Europe TRYTRAC program (to A.M.); MTB-Report project (VolkswagenStiftung ZN3424) (to J.H.).'' instead of ``German Federal Ministry of Research and Education (01ZZ1802); University of Heidelberg, Faculty of Medicine (to A.M.); Volkswagen Foundation (to J.H.).'' This has now been corrected. Florian Borchert, Andreas Mock, Aurelie Tomczak, Jonas Hügel, Samer Alkarkoukly, Alexander Knurr, Anna-Lena Volckmar, Albrecht Stenzinger, Peter Schirmacher, Jürgen Debus, Dirk Jäger, Thomas Longerich, Stefan Fröhling, Roland Eils, Nina Bougatf, Ulrich Sax, Matthieu-P. Schapranow |
Briefings Bioinform. | 14 |
| 2021 | Knowledge bases and software support for variant interpretation in precision oncologyabstractPrecision oncology is a rapidly evolving interdisciplinary medical specialty. Comprehensive cancer panels are becoming increasingly available at pathology departments worldwide, creating the urgent need for scalable cancer variant annotation and molecularly informed treatment recommendations. A wealth of mainly academia-driven knowledge bases calls for software tools supporting the multi-step diagnostic process. We derive a comprehensive list of knowledge bases relevant for variant interpretation by a review of existing literature followed by a survey among medical experts from university hospitals in Germany. In addition, we review cancer variant interpretation tools, which integrate multiple knowledge bases. We categorize the knowledge bases along the diagnostic process in precision oncology and analyze programmatic access options as well as the integration of knowledge bases into software tools. The most commonly used knowledge bases provide good programmatic access options and have been integrated into a range of software tools. For the wider set of knowledge bases, access options vary across different parts of the diagnostic process. Programmatic access is limited for information regarding clinical classifications of variants and for therapy recommendations. The main issue for databases used for biological classification of pathogenic variants and pathway context information is the lack of standardized interfaces. There is no single cancer variant interpretation tool that integrates all identified knowledge bases. Specialized tools are available and need to be further developed for different steps in the diagnostic process. Florian Borchert, Andreas Mock, Aurelie Tomczak, Jonas Hügel, Samer Alkarkoukly, Alexander Knurr, Anna-Lena Volckmar, Albrecht Stenzinger, Peter Schirmacher, Jürgen Debus, Dirk Jäger, Thomas Longerich, Stefan Fröhling, Roland Eils, Nina Bougatf, Ulrich Sax, Matthieu-P. Schapranow |
Briefings Bioinform. | 14 |
| 2021 | Implementing FAIR data management within the German Network for Bioinformatics Infrastructure (de.NBI) exemplified by selected use casesabstractThis article describes some use case studies and self-assessments of FAIR status of de.NBI services to illustrate the challenges and requirements for the definition of the needs of adhering to the FAIR (findable, accessible, interoperable and reusable) data principles in a large distributed bioinformatics infrastructure. We address the challenge of heterogeneity of wet lab technologies, data, metadata, software, computational workflows and the levels of implementation and monitoring of FAIR principles within the different bioinformatics sub-disciplines joint in de.NBI. On the one hand, this broad service landscape and the excellent network of experts are a strong basis for the development of useful research data management plans. On the other hand, the large number of tools and techniques maintained by distributed teams renders FAIR compliance challenging. Gerhard Mayer, Wolfgang Müller 0001, Karin Schork, Julian Uszkoreit, Andreas Weidemann, Ulrike Wittig, Maja Rey, Christian Quast, Janine Felden, Frank Oliver Glöckner, Matthias Lange 0001, Daniel Arend, Sebastian Beier, Astrid Junker, Uwe Scholz, Danuta Schüler, Hans A. Kestler, Daniel Wibberg, Alfred Pühler, Sven Twardziok, Roland Eils, Steve Hoffmann, Martin Eisenacher, Michael Turewicz |
Briefings Bioinform. | 22 |
| 2020 | Membership Inference Against DNA Methylation DatabasesabstractBiomedical data sharing is one of the key elements fostering the advancement of biomedical research but poses severe risks towards the privacy of individuals contributing their data, as already demonstrated for genomic data. In this paper, we study whether and to which extent DNA methylation data, one of the most important epigenetic elements regulating human health, is prone to membership inference attacks, a critical type of attack that reveals an individual's participation in a given database. We design and evaluate three different attacks exploiting published summary statistics, among which one is based on machine learning and another is exploiting the dependencies between genome and methylation data. Our extensive evaluation on six datasets containing a diverse set of tissues and diseases collected from more than 1,300 individuals in total shows that such membership inference attacks are effective, even when the target's methylation profile is not accessible. It further shows that the machine-learning approach outperforms the statistical attacks, and that learned models are transferable across different datasets. Inken Hagestedt, Mathias Humbert, Pascal Berrang, Irina Lehmann, Roland Eils, Michael Backes 0001, Yang Zhang 0016 |
EuroS&P | 5 |
| 2018 | Dissecting Privacy Risks in Biomedical DataabstractThe decreasing costs of molecular profiling has fueled the biomedical research community with a plethora of new types of biomedical data, enabling a breakthrough towards a more precise and personalized medicine. However, the release of these intrinsically highly sensitive data poses a new severe privacy threat. While biomedical data is largely associated with our health, there also exist various correlations between different types of biomedical data, along the temporal dimension, and also in-between family members. However, so far, the security community has focused on privacy risks stemming from genomic data, largely overlooking the manifold interdependencies between other biomedical data. In this paper, we present a generic framework for quantifying the privacy risks in biomedical data taking into account the various interdependencies between data (i) of different types, (ii) from different individuals, and (iii) at different time. To this end, we rely on a Bayesian network model that allows us to take all aforementioned dependencies into account and run exact probabilistic inference attacks very efficiently. Furthermore, we introduce a generic algorithm for building the Bayesian network, which encompasses expert knowledge for known dependencies, such as genetic inheritance laws, and learns previously unknown dependencies from the data. Then, we conduct a thorough inference risk evaluation with a very rich dataset containing genomic and epigenomic data of mothers and children over multiple years. Besides effective probabilistic inference, we further demonstrate that our Bayesian network model can also serve as a building block for other attacks. We show that, with our framework, an adversary can efficiently identify the parent-child relationships based on methylation data with a success rate of 95%. Pascal Berrang, Mathias Humbert, Yang Zhang 0016, Irina Lehmann, Roland Eils, Michael Backes 0001 |
EuroS&P | 5 |
| 2018 | Web-based design and analysis tools for CRISPR base editingabstractBACKGROUND: As a result of its simplicity and high efficiency, the CRISPR-Cas system has been widely used as a genome editing tool. Recently, CRISPR base editors, which consist of deactivated Cas9 (dCas9) or Cas9 nickase (nCas9) linked with a cytidine or a guanine deaminase, have been developed. Base editing tools will be very useful for gene correction because they can produce highly specific DNA substitutions without the introduction of any donor DNA, but dedicated web-based tools to facilitate the use of such tools have not yet been developed. RESULTS: We present two web tools for base editors, named BE-Designer and BE-Analyzer. BE-Designer provides all possible base editor target sequences in a given input DNA sequence with useful information including potential off-target sites. BE-Analyzer, a tool for assessing base editing outcomes from next generation sequencing (NGS) data, provides information about mutations in a table and interactive graphs. Furthermore, because the tool runs client-side, large amounts of targeted deep sequencing data (< 1 GB) do not need to be uploaded to a server, substantially reducing running time and increasing data security. BE-Designer and BE-Analyzer can be freely accessed at http://www.rgenome.net/be-designer/ and http://www.rgenome.net/be-analyzer /, respectively. CONCLUSION: We develop two useful web tools to design target sequence (BE-Designer) and to analyze NGS data from experimental results (BE-Analyzer) for CRISPR base editors. Gue-Ho Hwang, Jeongbin Park, Kayeong Lim, Jihyeon Yu, Eunchong Yu, Sang-Tae Kim, Roland Eils, Sangsu Bae |
BMC Bioinform. | 8 |
| 2017 | Identifying Personal DNA Methylation Profiles by Genotype InferenceabstractSince the first whole-genome sequencing, the biomedical research community has made significant steps towards a more precise, predictive and personalized medicine. Genomic data is nowadays widely considered privacy-sensitive and consequently protected by strict regulations and released only after careful consideration. Various additional types of biomedical data, however, are not shielded by any dedicated legal means and consequently disseminated much less thoughtfully. This in particular holds true for DNA methylation data as one of the most important and well-understood epigenetic element influencing human health. In this paper, we show that, in contrast to the aforementioned belief, releasing one's DNA methylation data causes privacy issues akin to releasing one's actual genome. We show that already a small subset of methylation regions influenced by genomic variants are sufficient to infer parts of someone's genome, and to further map this DNA methylation profile to the corresponding genome. Notably, we show that such re-identification is possible with 97.5% accuracy, relying on a dataset of more than 2500 genomes, and that we can reject all wrongly matched genomes using an appropriate statistical test. We provide means for countering this threat by proposing a novel cryptographic scheme for privately classifying tumors that enables a privacy-respecting medical diagnosis in a common clinical setting. The scheme relies on a combination of random forests and homomorphic encryption, and it is proven secure in the honest-but-curious model. We evaluate this scheme on real DNA methylation data, and show that we can keep the computational overhead to acceptable values for our application scenario. Michael Backes 0001, Pascal Berrang, Matthias Bieg, Roland Eils, Carl Herrmann, Mathias Humbert, Irina Lehmann |
IEEE Symposium on Security and Privacy | 4 |
| 2017 | Correlated receptor transport processes buffer single-cell heterogeneityabstractCells typically vary in their response to extracellular ligands. Receptor transport processes modulate ligand-receptor induced signal transduction and impact the variability in cellular responses. Here, we quantitatively characterized cellular variability in erythropoietin receptor (EpoR) trafficking at the single-cell level based on live-cell imaging and mathematical modeling. Using ensembles of single-cell mathematical models reduced parameter uncertainties and showed that rapid EpoR turnover, transport of internalized EpoR back to the plasma membrane, and degradation of Epo-EpoR complexes were essential for receptor trafficking. EpoR trafficking dynamics in adherent H838 lung cancer cells closely resembled the dynamics previously characterized by mathematical modeling in suspension cells, indicating that dynamic properties of the EpoR system are widely conserved. Receptor transport processes differed by one order of magnitude between individual cells. However, the concentration of activated Epo-EpoR complexes was less variable due to the correlated kinetics of opposing transport processes acting as a buffering system. Stefan M. Kallenberger, Anne L. Unger, Stefan Legewie, Konstantinos Lymperopoulos, Ursula Klingmüller, Roland Eils, Dirk-Peter Herten |
PLoS Comput. Biol. | 6 |
| 2016 | A comprehensive comparison of tools for differential ChIP-seq analysisabstractChIP-seq has become a widely adopted genomic assay in recent years to determine binding sites for transcription factors or enrichments for specific histone modifications. Beside detection of enriched or bound regions, an important question is to determine differences between conditions. While this is a common analysis for gene expression, for which a large number of computational approaches have been validated, the same question for ChIP-seq is particularly challenging owing to the complexity of ChIP-seq data in terms of noisiness and variability. Many different tools have been developed and published in recent years. However, a comprehensive comparison and review of these tools is still missing. Here, we have reviewed 14 tools, which have been developed to determine differential enrichment between two conditions. They differ in their algorithmic setups, and also in the range of applicability. Hence, we have benchmarked these tools on real data sets for transcription factors and histone modifications, as well as on simulated data sets to quantitatively evaluate their performance. Overall, there is a great variety in the type of signal detected by these tools with a surprisingly low level of agreement. Depending on the type of analysis performed, the choice of method will crucially impact the outcome. Sebastian Steinhauser, Nils Kurzawa, Roland Eils, Carl Herrmann |
Briefings Bioinform. | 3 |
| 2016 | HilbertCurve: an R/Bioconductor package for high-resolution visualization of genomic dataabstractUNLABELLED: : Hilbert curves enable high-resolution visualization of genomic data on a chromosome- or genome-wide scale. Here we present the HilbertCurve package that provides an easy-to-use interface for mapping genomic data to Hilbert curves. The package transforms the curve as a virtual axis, thereby hiding the details of the curve construction from the user. HilbertCurve supports multiple-layer overlay that makes it a powerful tool to correlate the spatial distribution of multiple feature types. AVAILABILITY AND IMPLEMENTATION: The HilbertCurve package and documentation are freely available from the Bioconductor project: http://www.bioconductor.org/packages/devel/bioc/html/HilbertCurve.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zuguang Gu, Roland Eils, Matthias Schlesner |
Bioinform. | 2 |
| 2016 | Complex heatmaps reveal patterns and correlations in multidimensional genomic dataabstractUNLABELLED: Parallel heatmaps with carefully designed annotation graphics are powerful for efficient visualization of patterns and relationships among high dimensional genomic data. Here we present the ComplexHeatmap package that provides rich functionalities for customizing heatmaps, arranging multiple parallel heatmaps and including user-defined annotation graphics. We demonstrate the power of ComplexHeatmap to easily reveal patterns and correlations among multiple sources of information with four real-world datasets. AVAILABILITY AND IMPLEMENTATION: The ComplexHeatmap package and documentation are freely available from the Bioconductor project: http://www.bioconductor.org/packages/devel/bioc/html/ComplexHeatmap.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zuguang Gu, Roland Eils, Matthias Schlesner |
Bioinform. | 2 |
| 2016 | gtrellis: an R/Bioconductor package for making genome-level Trellis graphicsabstractBACKGROUND: Trellis graphics are a visualization method that splits data by one or more categorical variables and displays subsets of the data in a grid of panels. Trellis graphics are broadly used in genomic data analysis to compare statistics over different categories in parallel and reveal multivariate relationships. However, current software packages to produce Trellis graphics have not been designed with genomic data in mind and lack some functionality that is required for effective visualization of genomic data. RESULTS: Here we introduce the gtrellis package which provides an efficient and extensible way to visualize genomic data in a Trellis layout. gtrellis provides highly flexible Trellis layouts which allow efficient arrangement of genomic categories on the plot. It supports multiple-track visualization, which makes it straightforward to visualize several properties of genomic data in parallel to explain complex relationships. In addition, gtrellis provides an extensible framework that allows adding user-defined graphics. CONCLUSIONS: The gtrellis package provides an easy and effective way to visualize genomic data and reveal high dimensional relationships on a genome-wide scale. gtrellis can be flexibly extended and thus can also serve as a base package for highly specific purposes. gtrellis makes it easy to produce novel visualizations, which can lead to the discovery of previously unrecognized patterns in genomic data. Zuguang Gu, Roland Eils, Matthias Schlesner |
BMC Bioinform. | 2 |
| 2015 | Non-rigid multi-frame registration of cell nuclei in live cell fluorescence microscopy image data
Marco Tektonidis, Il-Han Kim, Roland Eils, David L. Spector, Karl Rohr |
Medical Image Anal. | 4 |
| 2015 | Tracking Virus Particles in Fluorescence Microscopy Images Using Multi-Scale Detection and Multi-Frame AssociationabstractAutomatic fluorescent particle tracking is an essential task to study the dynamics of a large number of biological structures at a sub-cellular level. We have developed a probabilistic particle tracking approach based on multi-scale detection and two-step multi-frame association. The multi-scale detection scheme allows coping with particles in close proximity. For finding associations, we have developed a two-step multi-frame algorithm, which is based on a temporally semiglobal formulation as well as spatially local and global optimization. In the first step, reliable associations are determined for each particle individually in local neighborhoods. In the second step, the global spatial information over multiple frames is exploited jointly to determine optimal associations. The multi-scale detection scheme and the multi-frame association finding algorithm have been combined with a probabilistic tracking approach based on the Kalman filter. We have successfully applied our probabilistic tracking approach to synthetic as well as real microscopy image sequences of virus particles and quantified the performance. We found that the proposed approach outperforms previous approaches. Astha Jaiswal, William J. Godinez, Roland Eils, Maik Jörg Lehmann, Karl Rohr |
IEEE Trans. Image Process. | 3 |
| 2014 | circlize implements and enhances circular visualization in RabstractSUMMARY: Circular layout is an efficient way for the visualization of huge amounts of genomic information. Here we present the circlize package, which provides an implementation of circular layout generation in R as well as an enhancement of available software. The flexibility of this package is based on the usage of low-level graphics functions such that self-defined high-level graphics can be easily implemented by users for specific purposes. Together with the seamless connection between the powerful computational and visual environment in R, circlize gives users more convenience and freedom to design figures for better understanding genomic patterns behind multi-dimensional data. AVAILABILITY AND IMPLEMENTATION: circlize is available at the Comprehensive R Archive Network (CRAN): http://cran.r-project.org/web/packages/circlize/ Zuguang Gu, Roland Eils, Matthias Schlesner, Benedikt Brors |
Bioinform. | 3 |
| 2014 | Estimating the activity of transcription factors by the effect on their target genesabstractMOTIVATION: Understanding regulation of transcription is central for elucidating cellular regulation. Several statistical and mechanistic models have come up the last couple of years explaining gene transcription levels using information of potential transcriptional regulators as transcription factors (TFs) and information from epigenetic modifications. The activity of TFs is often inferred by their transcription levels, promoter binding and epigenetic effects. However, in principle, these methods do not take hard-to-measure influences such as post-transcriptional modifications into account. RESULTS: For TFs, we present a novel concept circumventing this problem. We estimate the regulatory activity of TFs using their cumulative effects on their target genes. We established our model using expression data of 59 cell lines from the National Cancer Institute. The trained model was applied to an independent expression dataset of melanoma cells yielding excellent expression predictions and elucidated regulation of melanogenesis. AVAILABILITY AND IMPLEMENTATION: Using mixed-integer linear programming, we implemented a switch-like optimization enabling a constrained but optimal selection of TFs and optimal model selection estimating their effects. The method is generic and can also be applied to further regulators of transcription. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Theresa Schacht, Marcus Oswald, Roland Eils, Stefan B. Eichmüller, Rainer König |
Bioinform. | 3 |
| 2014 | Data-Derived Modeling Characterizes Plasticity of MAPK Signaling in MelanomaabstractThe majority of melanomas have been shown to harbor somatic mutations in the RAS-RAF-MEK-MAPK and PI3K-AKT pathways, which play a major role in regulation of proliferation and survival. The prevalence of these mutations makes these kinase signal transduction pathways an attractive target for cancer therapy. However, tumors have generally shown adaptive resistance to treatment. This adaptation is achieved in melanoma through its ability to undergo neovascularization, migration and rearrangement of signaling pathways. To understand the dynamic, nonlinear behavior of signaling pathways in cancer, several computational modeling approaches have been suggested. Most of those models require that the pathway topology remains constant over the entire observation period. However, changes in topology might underlie adaptive behavior to drug treatment. To study signaling rearrangements, here we present a new approach based on Fuzzy Logic (FL) that predicts changes in network architecture over time. This adaptive modeling approach was used to investigate pathway dynamics in a newly acquired experimental dataset describing total and phosphorylated protein signaling over four days in A375 melanoma cell line exposed to different kinase inhibitors. First, a generalized strategy was established to implement a parameter-reduced FL model encoding non-linear activity of a signaling network in response to perturbation. Next, a literature-based topology was generated and parameters of the FL model were derived from the full experimental dataset. Subsequently, the temporal evolution of model performance was evaluated by leaving time-defined data points out of training. Emerging discrepancies between model predictions and experimental data at specific time points allowed the characterization of potential network rearrangement. We demonstrate that this adaptive FL modeling approach helps to enhance our mechanistic understanding of the molecular plasticity of melanoma. Marti Bernardo-Faura, Stefan Massen, Christine S. Falk, Nathan R. Brady, Roland Eils |
PLoS Comput. Biol. | 5 |
| 2014 | Characterizing Protein Interactions Employing a Genome-Wide siRNA Cellular Phenotyping ScreenabstractCharacterizing the activating and inhibiting effect of protein-protein interactions (PPI) is fundamental to gain insight into the complex signaling system of a human cell. A plethora of methods has been suggested to infer PPI from data on a large scale, but none of them is able to characterize the effect of this interaction. Here, we present a novel computational development that employs mitotic phenotypes of a genome-wide RNAi knockdown screen and enables identifying the activating and inhibiting effects of PPIs. Exemplarily, we applied our technique to a knockdown screen of HeLa cells cultivated at standard conditions. Using a machine learning approach, we obtained high accuracy (82% AUC of the receiver operating characteristics) by cross-validation using 6,870 known activating and inhibiting PPIs as gold standard. We predicted de novo unknown activating and inhibiting effects for 1,954 PPIs in HeLa cells covering the ten major signaling pathways of the Kyoto Encyclopedia of Genes and Genomes, and made these predictions publicly available in a database. We finally demonstrate that the predicted effects can be used to cluster knockdown genes of similar biological processes in coherent subgroups. The characterization of the activating or inhibiting effect of individual PPIs opens up new perspectives for the interpretation of large datasets of PPIs and thus considerably increases the value of PPIs as an integrated resource for studying the detailed function of signaling pathways of the cellular system of interest. Apichat Suratanee, Martin H. Schaefer 0001, Matthew J. Betts, Zita Soons, Heiko A. Mannsperger, Nathalie Harder, Marcus Oswald, Markus Gipp, Ellen Ramminger, Guillermo Marcus Martinez, Reinhard Männer, Karl Rohr, Erich E. Wanker, Robert B. Russell, Miguel A. Andrade-Navarro, Roland Eils, Rainer König |
PLoS Comput. Biol. | 16 |
| 2013 | SplicingCompass: differential splicing detection using RNA-Seq dataabstractMOTIVATION: Alternative splicing is central for cellular processes and substantially increases transcriptome and proteome diversity. Aberrant splicing events often have pathological consequences and are associated with various diseases and cancer types. The emergence of next-generation RNA sequencing (RNA-seq) provides an exciting new technology to analyse alternative splicing on a large scale. However, algorithms that enable the analysis of alternative splicing from short-read sequencing are not fully established yet and there are still no standard solutions available for a variety of data analysis tasks. RESULTS: We present a new method and software to predict genes that are differentially spliced between two different conditions using RNA-seq data. Our method uses geometric angles between the high dimensional vectors of exon read counts. With this, differential splicing can be detected even if the splicing events are composed of higher complexity and involve previously unknown splicing patterns. We applied our approach to two case studies including neuroblastoma tumour data with favourable and unfavourable clinical courses. We show the validity of our predictions as well as the applicability of our method in the context of patient clustering. We verified our predictions by several methods including simulated experiments and complementary in silico analyses. We found a significant number of exons with specific regulatory splicing factor motifs for predicted genes and a substantial number of publications linking those genes to alternative splicing. Furthermore, we could successfully exploit splicing information to cluster tissues and patients. Finally, we found additional evidence of splicing diversity for many predicted genes in normalized read coverage plots and in reads that span exon-exon junctions. AVAILABILITY: SplicingCompass is licensed under the GNU GPL and freely available as a package in the statistical language R at http://www.ichip.de/software/SplicingCompass.html Moritz Aschoff, Agnes Hotz-Wagenblatt, Karl-Heinz Glatting, Matthias Fischer 0003, Roland Eils, Rainer König |
Bioinform. | 5 |
| 2013 | Disease-gene discovery by integration of 3D gene expression and transcription factor binding affinitiesabstractMOTIVATION: The computational evaluation of candidate genes for hereditary disorders is a non-trivial task. Several excellent methods for disease-gene prediction have been developed in the past 2 decades, exploiting widely differing data sources to infer disease-relevant functional relationships between candidate genes and disorders. We have shown recently that spatially mapped, i.e. 3D, gene expression data from the mouse brain can be successfully used to prioritize candidate genes for human Mendelian disorders of the central nervous system. RESULTS: We improved our previous work 2-fold: (i) we demonstrate that condition-independent transcription factor binding affinities of the candidate genes' promoters are relevant for disease-gene prediction and can be integrated with our previous approach to significantly enhance its predictive power; and (ii) we define a novel similarity measure-termed Relative Intensity Overlap-for both 3D gene expression patterns and binding affinity profiles that better exploits their disease-relevant information content. Finally, we present novel disease-gene predictions for eight loci associated with different syndromes of unknown molecular basis that are characterized by mental retardation. Rosario M. Piro, Ivan Molineris, Ferdinando Di Cunto, Roland Eils, Rainer König |
Bioinform. | 4 |
| 2013 | Mining Quasi-Bicliques from HIV-1-Human Protein Interaction Network: A Multiobjective Biclustering ApproachabstractIn this work, we model the problem of mining quasi-bicliques from weighted viral-host protein-protein interaction network as a biclustering problem for identifying strong interaction modules. In this regard, a multiobjective genetic algorithm-based biclustering technique is proposed that simultaneously optimizes three objective functions to obtain dense biclusters having high mean interaction strengths. The performance of the proposed technique has been compared with that of other existing biclustering methods on an artificial data. Subsequently, the proposed biclustering method is applied on the records of biologically validated and predicted interactions between a set of HIV-1 proteins and a set of human proteins to identify strong interaction modules. For this, the entire interaction information is realized as a bipartite graph. We have further investigated the biological significance of the obtained biclusters. The human proteins involved in the strong interaction module have been found to share common biological properties and they are identified as the gateways of viral infection leading to various diseases. These human proteins can be potential drug targets for developing anti-HIV drugs. Ujjwal Maulik, Anirban Mukhopadhyay 0001, Malay Bhattacharyya 0001, Lars Kaderali, Benedikt Brors, Sanghamitra Bandyopadhyay, Roland Eils |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2012 | Identifying Virus-Cell Fusion in Two-Channel Fluorescence Microscopy Image Sequences Based on a Layered Probabilistic ApproachabstractThe entry process of virus particles into cells is decisive for infection. In this work, we investigate fusion of virus particles with the cell membrane via time-lapse fluorescence microscopy. To automatically identify fusion for single particles based on their intensity over time, we have developed a layered probabilistic approach. The approach decomposes the action of a single particle into three abstractions: the intensity over time, the underlying temporal intensity model, as well as a high level behavior. Each abstraction corresponds to a layer and these layers are represented via stochastic hybrid systems and hidden Markov models. We use a maxbelief strategy to efficiently combine both representations. To compute estimates for the abstractions we use a hybrid particle filter and the Viterbi algorithm. Based on synthetic image sequences, we characterize the performance of the approach as a function of the image noise. We also characterize the performance as a function of the tracking error. We have also successfully applied the approach to real image sequences displaying pseudotyped HIV-1 particles in contact with host cells and compared the experimental results with ground truth obtained by manual analysis. William J. Godinez, Marko Lampe, Peter Koch 0003, Roland Eils, Barbara Müller, Karl Rohr |
IEEE Trans. Medical Imaging | 4 |
| 2011 | Reconstructing the Stochastic Evolution Diagram of Dynamic Complex SystemsabstractThe behavior and dynamics of complex systems are in focus of many research fields. The complexity of such systems comes not only from the number of their elements, but also from the unavoidable emergence of new properties of the system, which are not just a simple summation of the properties of its elements. The behavior of complex systems can be fitted with a number of well developed models, which, however, do not incorporate the modularity and the evolution of a system simultaneously. In this work, we propose a generalized model that addresses this issue. Our model is developed within the Random Set Theory’s framework and allows for reconstructing the stochastic evolution diagrams of complex systems. Navid Bazzazzadeh, Benedikt Brors, Roland Eils |
AAAI | 3 |
| 2011 | Reconstructing evolutionary modular networks from time series data
Navid Bazzazzadeh, Benedikt Brors, Roland Eils |
FUSION | 3 |
| 2011 | RIP: the regulatory interaction predictor - a machine learning-based approach for predicting target genes of transcription factorsabstractMOTIVATION: Understanding transcriptional gene regulation is essential for studying cellular systems. Identifying genome-wide targets of transcription factors (TFs) provides the basis to discover the involvement of TFs and TF cooperativeness in cellular systems and pathogenesis. RESULTS: We present the regulatory interaction predictor (RIP), a machine learning approach that inferred 73 923 regulatory interactions (RIs) for 301 human TFs and 11 263 target genes with considerably good quality and 4516 RIs with very high quality. The inference of RIs is independent of any specific condition. Our approach employs support vector machines (SVMs) trained on a set of experimentally proven RIs from a public repository (TRANSFAC). Features of RIs for the learning process are based on a correlation meta-analysis of 4064 gene expression profiles from 76 studies, in silico predictions of transcription factor binding sites (TFBSs) and combinations of these employing knowledge about co-regulation of genes by a common TF (TF-module). The trained SVMs were applied to infer new RIs for a large set of TFs and genes. In a case study, we employed the inferred RIs to analyze an independent microarray dataset. We identified key TFs regulating the transcriptional response upon interferon alpha stimulation of monocytes, most prominently interferon-stimulated gene factor 3 (ISGF3). Furthermore, predicted TF-modules were highly associated to their functionally related pathways. CONCLUSION: Descriptors of gene expression, TFBS predictions, experimentally verified binding information and statistical combination of this enabled inferring RIs on a genome-wide scale for human genes with considerably good precision serving as a good basis for expression profiling studies. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tobias Bauer, Roland Eils, Rainer König |
Bioinform. | 2 |
| 2011 | Normalizing for individual cell population context in the analysis of high-content cellular screensabstractBACKGROUND: High-content, high-throughput RNA interference (RNAi) offers unprecedented possibilities to elucidate gene function and involvement in biological processes. Microscopy based screening allows phenotypic observations at the level of individual cells. It was recently shown that a cell's population context significantly influences results. However, standard analysis methods for cellular screens do not currently take individual cell data into account unless this is important for the phenotype of interest, i.e. when studying cell morphology. RESULTS: We present a method that normalizes and statistically scores microscopy based RNAi screens, exploiting individual cell information of hundreds of cells per knockdown. Each cell's individual population context is employed in normalization. We present results on two infection screens for hepatitis C and dengue virus, both showing considerable effects on observed phenotypes due to population context. In addition, we show on a non-virus screen that these effects can be found also in RNAi data in the absence of any virus. Using our approach to normalize against these effects we achieve improved performance in comparison to an analysis without this normalization and hit scoring strategy. Furthermore, our approach results in the identification of considerably more significantly enriched pathways in hepatitis C virus replication than using a standard analysis approach. CONCLUSIONS: Using a cell-based analysis and normalization for population context, we achieve improved sensitivity and specificity not only on a individual protein level, but especially also on a pathway level. This leads to the identification of new host dependency factors of the hepatitis C and dengue viruses and higher reproducibility of results. Bettina Knapp, Ilka Rebhan, Anil Kumar 0006, Petr Matula, Narsis A. Kiani, Marco Binder, Holger Erfle, Karl Rohr, Roland Eils, Ralf Bartenschlager, Lars Kaderali |
BMC Bioinform. | 9 |
| 2011 | Nonrigid Registration of 2-D and 3-D Dynamic Cell Nuclei Images for Improved Classification of Subcellular Particle MotionabstractThe observed motion of subcellular particles in fluorescence microscopy image sequences of live cells is generally a superposition of the motion and deformation of the cell and the motion of the particles. Decoupling the two types of movements to enable accurate classification of the particle motion requires the application of registration algorithms. We have developed an intensity-based approach for nonrigid registration of multichannel microscopy image sequences of cell nuclei. First, based on 3-D synthetic images we demonstrate that cell nucleus deformations change the observed motion types of particles and that our approach allows to recover the original motion. Second, we have successfully applied our approach to register 2-D and 3-D real microscopy image sequences. A quantitative experimental comparison with previous approaches for nonrigid registration of cell microscopy has also been performed. Il-Han Kim, David L. Spector, Roland Eils, Karl Rohr |
IEEE Trans. Image Process. | 4 |
| 2010 | PathWave: discovering patterns of differentially regulated enzymes in metabolic pathwaysabstractMOTIVATION: Gene expression profiling by microarrays or transcript sequencing enables observing the pathogenic function of tumors on a mesoscopic level. RESULTS: We investigated neuroblastoma tumors that clinically exhibit a very heterogeneous course ranging from rapid growth with fatal outcome to spontaneous regression and detected regulatory oncogenetic shifts in their metabolic networks. In contrast to common enrichment tests, we took network topology into account by applying adjusted wavelet transforms on an elaborated and new 2D grid representation of curated pathway maps from the Kyoto Enzyclopedia of Genes and Genomes. The aggressive form of the tumors showed regulatory shifts for purine and pyrimidine biosynthesis as well as folate-mediated metabolism of the one-carbon pool in respect to increased nucleotide production. We spotted an oncogentic regulatory switch in glutamate metabolism for which we provided experimental validation, being the first steps towards new possible drug therapy. The pattern recognition method we used complements normal enrichment tests to detect such functionally related regulation patterns. AVAILABILITY AND IMPLEMENTATION: PathWave is implemented in a package for R (www.r-project.org) version 2.6.0 or higher. It is freely available from http://www.ichip.de/software/pathwave.html. Gunnar Schramm, Stefan Wiesberg, Nicolle Diessl, Anna-Lena Kranz, Vitalia Sagulenko, Marcus Oswald, Gerhard Reinelt, Frank Westermann, Roland Eils, Rainer König |
Bioinform. | 9 |
| 2010 | Detecting host factors involved in virus infection by observing the clustering of infected cells in siRNA screening imagesabstractMOTIVATION: Detecting human proteins that are involved in virus entry and replication is facilitated by modern high-throughput RNAi screening technology. However, hit lists from different laboratories have shown only little consistency. This may be caused by not only experimental discrepancies, but also not fully explored possibilities of the data analysis. We wanted to improve reliability of such screens by combining a population analysis of infected cells with an established dye intensity readout. RESULTS: Viral infection is mainly spread by cell-cell contacts and clustering of infected cells can be observed during spreading of the infection in situ and in vivo. We employed this clustering feature to define knockdowns which harm viral infection efficiency of human Hepatitis C Virus. Images of knocked down cells for 719 human kinase genes were analyzed with an established point pattern analysis method (Ripley's K-function) to detect knockdowns in which virally infected cells did not show any clustering and therefore were hindered to spread their infection to their neighboring cells. The results were compared with a statistical analysis using a common intensity readout of the GFP-expressing viruses and a luciferase-based secondary screen yielding five promising host factors which may suit as potential targets for drug therapy. CONCLUSION: We report of an alternative method for high-throughput imaging methods to detect host factors being relevant for the infection efficiency of viruses. The method is generic and has the potential to be used for a large variety of different viruses and treatments being screened by imaging techniques. Apichat Suratanee, Ilka Rebhan, Petr Matula, Anil Kumar 0006, Lars Kaderali, Karl Rohr, Ralf Bartenschlager, Roland Eils, Rainer König |
Bioinform. | 8 |
| 2010 | Bayesian statistical modelling of human protein interaction network incorporating protein disorder informationabstractBACKGROUND: We present a statistical method of analysis of biological networks based on the exponential random graph model, namely p2-model, as opposed to previous descriptive approaches. The model is capable to capture generic and structural properties of a network as emergent from local interdependencies and uses a limited number of parameters. Here, we consider one global parameter capturing the density of edges in the network, and local parameters representing each node's contribution to the formation of edges in the network. The modelling suggests a novel definition of important nodes in the network, namely social, as revealed based on the local sociality parameters of the model. Moreover, the sociality parameters help to reveal organizational principles of the network. An inherent advantage of our approach is the possibility of hypotheses testing: a priori knowledge about biological properties of the nodes can be incorporated into the statistical model to investigate its influence on the structure of the network. RESULTS: We applied the statistical modelling to the human protein interaction network obtained with Y2H experiments. Bayesian approach for the estimation of the parameters was employed. We deduced social proteins, essential for the formation of the network, while incorporating into the model information on protein disorder. Intrinsically disordered are proteins which lack a well-defined three-dimensional structure under physiological conditions. We predicted the fold group (ordered or disordered) of proteins in the network from their primary sequences. The network analysis indicated that protein disorder has a positive effect on the connectivity of proteins in the network, but do not fully explains the interactivity. CONCLUSIONS: The approach opens a perspective to study effects of biological properties of individual entities on the structure of biological networks. Svetlana Bulashevska, Alla Bulashevska, Roland Eils |
BMC Bioinform. | 3 |
| 2009 | RNAither, an automated pipeline for the statistical analysis of high-throughput RNAi screensabstractSUMMARY: We present RNAither, a package for the free statistical environment R which performs an analysis of high-throughput RNA interference (RNAi) knock-down experiments, generating lists of relevant genes and pathways out of raw experimental data. The library provides a quality assessment of the signal intensities, as well as a broad range of options for data normalization, different statistical tests for the identification of significant siRNAs, and a significance analysis of the biological processes involving corresponding genes. The results of the analysis are presented as a set of HTML pages. Additionally, all values and plots are available as either text files or pdf and png files. AVAILABILITY: http://bioconductor.org/ Nora Rieber, Bettina Knapp, Roland Eils, Lars Kaderali |
Bioinform. | 3 |
| 2009 | Deterministic and probabilistic approaches for tracking virus particles in time-lapse fluorescence microscopy image sequences
William J. Godinez, Marko Lampe, Stefan Wörz, Barbara Müller, Roland Eils, Karl Rohr |
Medical Image Anal. | 5 |
| 2009 | Optimal Experimental Design for Parameter Estimation of a Cell Signaling ModelabstractDifferential equation models that describe the dynamic changes of biochemical signaling states are important tools to understand cellular behavior. An essential task in building such representations is to infer the affinities, rate constants, and other parameters of a model from actual measurement data. However, intuitive measurement protocols often fail to generate data that restrict the range of possible parameter values. Here we utilized a numerical method to iteratively design optimal live-cell fluorescence microscopy experiments in order to reveal pharmacological and kinetic parameters of a phosphatidylinositol 3,4,5-trisphosphate (PIP(3)) second messenger signaling process that is deregulated in many tumors. The experimental approach included the activation of endogenous phosphoinositide 3-kinase (PI3K) by chemically induced recruitment of a regulatory peptide, reversible inhibition of PI3K using a kinase inhibitor, and monitoring of the PI3K-mediated production of PIP(3) lipids using the pleckstrin homology (PH) domain of Akt. We found that an intuitively planned and established experimental protocol did not yield data from which relevant parameters could be inferred. Starting from a set of poorly defined model parameters derived from the intuitively planned experiment, we calculated concentration-time profiles for both the inducing and the inhibitory compound that would minimize the predicted uncertainty of parameter estimates. Two cycles of optimization and experimentation were sufficient to narrowly confine the model parameters, with the mean variance of estimates dropping more than sixty-fold. Thus, optimal experimental design proved to be a powerful strategy to minimize the number of experiments needed to infer biological parameters from a cell signaling assay. Samuel Bandara, Johannes P. Schlöder, Roland Eils, Hans Georg Bock |
PLoS Comput. Biol. | 3 |
| 2008 | Systems Biology and Artificial Life: Towards Predictive Modeling of Biological SystemsabstractSystems biology has been defined as ‘‘the search for the syntax of biological information, that is, the study of the dynamic networks of interacting biological elements’’ [1]. Based on large, quantitative data sets, systems biology aims to generate models that facilitate an integrative understanding of biological systems that goes beyond increasingly detailed knowledge of the system’s elements. According to a classical definition, ‘‘Artificial Life is a field of study devoted to understanding life by attempting to abstract the fundamental dynamical principles underlying biological phenomena, and recreating these dynamics in other physical media—such as computers—making them accessible to new kinds of experimental manipulation and testing’’ [2]. These definitions show that systems biology and artificial life share the same objective: a principled and comprehensive understanding of living systems. Both interdisciplinary fields employ computational and other formal models of biological systems, and apply mathematical tools to analyze models and complex systems. The use of models is partially complementary: Systems biology focuses on analyzing and understanding experimental data using fairly generic modeling techniques, while artificial life considers rather elaborate and specific computational and other formal models as objects of experimentation, aiming to understand general biological features that are not necessarily represented by quantitative data. Using computational models to generate synthetic data is a recurring element in the contributions to this special issue. Ohno et al. present a model-based investigation of ammonia detoxification by the liver, and Ogawa et al. implemented multiple models of the regulatory networks realizing the Drosophila circadian clock for comparative study. Both these studies use the E-cell simulation framework. Van Leemput et al. use the SynTReN simulator to conduct a comparative study of regulatory network inference algorithms. Methods to analyze data and to infer networks complement synthetic data generation. A novel network inference algorithm based on variational Bayes expectation maximization (VBEM) is introduced by Tienda-Luna et al., and Matsubara et al. present extreme signal flows (ESFs) as a technique to analyze signaling networks. Models of biological systems will increase in complexity in the future. The overlap of systems biology and artificial life can be expected to grow in this process, as formal and computational techniques to model levels of biological organization such as tissues, growth, ecology, and evolution are combined with advanced techniques for data-driven model inference. As model complexity grows, the relative amount of molecular biology data and other prior knowledge to validate model inference methods decreases. Therefore, models that incorporate multiple key levels of biological organization will become increasingly important as a source of realistic synthetic test data. It would be desirable to develop a standard for specifying processes of synthetic data generation, as this would facilitate comparable validation of model inference algorithms as well as comparative studies of models. Biological systems typically interact with other levels of organization, and considering such interactions in modeling and analysis is often crucial for understanding a given biological system. In Jan T. Kim, Roland Eils |
Artif. Life | 2 |
| 2008 | Nonrigid Registration of 3-D Multichannel Microscopy Images of Cell NucleiabstractWe present an intensity-based nonrigid registration approach for the normalization of 3-D multichannel microscopy images of cell nuclei. A main problem with cell nuclei images is that the intensity structure of different nuclei differs very much; thus, an intensity-based registration scheme cannot be used directly. Instead, we first perform a segmentation of the images from the cell nucleus channel, smooth the resulting images by a Gaussian filter, and then apply an intensity-based registration algorithm. The obtained transformation is applied to the images from the nucleus channel as well as to the images from the other channels. To improve the convergence rate of the algorithm, we propose an adaptive step length optimization scheme and also employ a multiresolution scheme. Our approach has been successfully applied using 2-D cell-like synthetic images, 3-D phantom images as well as 3-D multichannel microscopy images representing different chromosome territories and gene regions. We also describe an extension of our approach, which is applied for the registration of 3D + t (4-D) image series of moving cell nuclei. Siwei Yang, Daniela Köhler, Kathrin Teller, Thomas Cremer, Patricia Le Baccon, Edith Heard, Roland Eils, Karl Rohr |
IEEE Trans. Image Process. | 7 |
| 2007 | Geometrical probability approach for analysis of 3D chromatin structure in interphase cell nucleiabstractInvestigation of 3D chromatin structure in interphase cell nuclei is important for the understanding of genome function. For a reconstruction of the 3D architecture of the human genome, systematic fluorescent in situ hybridization in combination with 3D confocal laser scanning microscopy is applied. The position of two or three genomic loci plus the overall nuclear shape were simultaneously recorded, resulting in statistical series of pair and triple loci combinations probed along the human chromosome 1 q-arm. For interpretation of statistical distributions of geometrical features (e.g. distances, angles, etc.) resulting from finite point sampling experiments, a Monte-Carlo-based approach to numerical computation of geometrical probability density functions (PDFs) for arbitrarily-shaped confined spatial domains is developed. Simulated PDFs are used as bench marks for evaluation of experimental PDFs and quantitative analysis of dimension and shape of probed 3D chromatin regions. Preliminary results of our numerical simulations show that the proposed numerical model is capable to reproduce experimental observations, and support the assumption of confined random folding of 3D chromatin fiber in interphase cell nuclei Evgeny Gladilin, Sandra Götze, Jose Mateos-Langerak, Roel van Driel, Karl Rohr, Roland Eils |
CIBCB | 6 |
| 2007 | Using gene expression data and network topology to detect substantial pathways, clusters and switches during oxygen deprivation of Escherichia coliabstractBACKGROUND: Biochemical investigations over the last decades have elucidated an increasingly complete image of the cellular metabolism. To derive a systems view for the regulation of the metabolism when cells adapt to environmental changes, whole genome gene expression profiles can be analysed. Moreover, utilising a network topology based on gene relationships may facilitate interpreting this vast amount of information, and extracting significant patterns within the networks. RESULTS: Interpreting expression levels as pixels with grey value intensities and network topology as relationships between pixels, allows for an image-like representation of cellular metabolism. While the topology of a regular image is a lattice grid, biological networks demonstrate scale-free architecture and thus advanced image processing methods such as wavelet transforms cannot directly be applied. In the study reported here, one-dimensional enzyme-enzyme pairs were tracked to reveal sub-graphs of a biological interaction network which showed significant adaptations to a changing environment. As a case study, the response of the hetero-fermentative bacterium E. coli to oxygen deprivation was investigated. With our novel method, we detected, as expected, an up-regulation in the pathways of hexose nutrients up-take and metabolism and formate fermentation. Furthermore, our approach revealed a down-regulation in iron processing as well as the up-regulation of the histidine biosynthesis pathway. The latter may reflect an adaptive response of E. coli against an increasingly acidic environment due to the excretion of acidic products during anaerobic growth in a batch culture. CONCLUSION: Based on microarray expression profiling data of prokaryotic cells exposed to fundamental treatment changes, our novel technique proved to extract system changes for a rather broad spectrum of the biochemical network. Gunnar Schramm, Marc Zapatka, Roland Eils, Rainer König |
BMC Bioinform. | 3 |
| 2006 | Feature Selection for Evaluating Fluorescence Microscopy Images in Genome-Wide Cell ScreensabstractWe investigate different approaches for efficient feature space reduction and compare different methods for cell classification. The application context is the development of automatic methods for analysing fluorescence microscopy images with the goal to identify those genes that are involved in the mitosis of human cells (cell division). We distinguish four cell classes comprising interphase cells, mitotic cells, apoptotic cells, and cells with clustered nuclei. Feature space reduction was performed using the Principal Component Analysis and Independent Component Analysis methods. Six classification methods were examined including unsupervised clustering algorithms such as K-means, Hard Competitive Learning, and Neural Gas as well as Hierarchical Clustering, Support Vector Machines, and Random Forests classifiers. Detailed results on the cell image classification accuracy and computational efficiency achieved using different feature sets and different classification methods are reported. Vassili Kovalev, Nathalie Harder, Beate Neumann, Michael Held, Urban Liebel, Holger Erfle, Jan Ellenberg, Roland Eils, Karl Rohr |
CVPR (1) | 8 |
| 2006 | Automated Analysis of the Mitotic Phases of Human Cells in 3D Fluorescence Microscopy Image Sequences
Nathalie Harder, Felipe Mora-Bermúdez, William J. Godinez, Jan Ellenberg, Roland Eils, Karl Rohr |
MICCAI (1) | 5 |
| 2006 | Non-rigid Registration of 3D Multi-channel Microscopy Images of Cell Nuclei
Siwei Yang, Daniela Köhler, Kathrin Teller, Thomas Cremer, Patricia Le Baccon, Edith Heard, Roland Eils, Karl Rohr |
MICCAI (1) | 7 |
| 2006 | Group testing for pathway analysis improves comparability of different microarray datasetsabstractMOTIVATION: The wide use of DNA microarrays for the investigation of the cell transcriptome triggered the invention of numerous methods for the processing of microarray data and lead to a growing number of microarray studies that examine the same biological conditions. However, comparisons made on the level of gene lists obtained by different statistical methods or from different datasets hardly converge. We aimed at examining such discrepancies on the level of apparently affected biologically related groups of genes, e.g. metabolic or signalling pathways. This can be achieved by group testing procedures, e.g. over-representation analysis, functional class scoring (FCS), or global tests. RESULTS: Three public prostate cancer datasets obtained with the same microarray platform (HGU95A/HGU95Av2) were analyzed. Each dataset was subjected to normalization by either variance stabilizing normalization (vsn) or mixed model normalization (MMN). Then, statistical analysis of microarrays was applied to the vsn-normalized data and mixed model analysis to the data normalized by MMN. For multiple testing adjustment the false discovery rate was calculated and the threshold was set to 0.05. Gene lists from the same method applied to different datasets showed overlaps between 42 and 52%, while lists from different methods applied to the same dataset had between 63 and 85% of genes in common. A number of six gene lists obtained by the two statistical methods applied to the three datasets was then subjected to group testing by Fisher's exact test. Group testing by GSEA and global test was applied to the three datasets, as well. Fisher's exact test followed by global test showed more consistent results with respect to the concordance between analyses on gene lists obtained by different methods and different datasets than the GSEA. However, all group testing methods identified pathways that had already been described to be involved in the pathogenesis of prostate cancer. Moreover, pathways recurrently identified in these analyses are more likely to be reliable than those from a single analysis on a single dataset. Theodora Manoli, Norbert Gretz, Hermann-Josef Gröne, Marc Kenzelmann, Roland Eils, Benedikt Brors |
Bioinform. | 5 |
| 2006 | Tropical - parameter estimation and simulation of reaction-diffusion models based on spatio-temporal microscopy imagesabstractUNLABELLED: Tropical is a software for simulation and parameter estimation of reaction-diffusion models. Based on spatio-temporal microscopy images, Tropical estimates reaction and diffusion coefficients for user-defined models. Tropical allows the investigation of systems with an inhomogeneous distribution of molecules, making it well suited for quantitative analyses of microscopy experiments such as fluorescence recovery after photobleaching (FRAP). AVAILABILITY: Tropical is available free of charge for academic use at http://www.dkfz.de/tbi/projects/modellingAndSimulationOfCelluarSystems/tropical.jsp after signing a material transfer agreement. Markus Ulrich, Constantin Kappel, Joël Beaudouin, Stefan Hezel, Jochen Ulrich, Roland Eils |
Bioinform. | 6 |
| 2006 | Predicting protein subcellular locations using hierarchical ensemble of Bayesian classifiers based on Markov chainsabstractBACKGROUND: The subcellular location of a protein is closely related to its function. It would be worthwhile to develop a method to predict the subcellular location for a given protein when only the amino acid sequence of the protein is known. Although many efforts have been made to predict subcellular location from sequence information only, there is the need for further research to improve the accuracy of prediction. RESULTS: A novel method called HensBC is introduced to predict protein subcellular location. HensBC is a recursive algorithm which constructs a hierarchical ensemble of classifiers. The classifiers used are Bayesian classifiers based on Markov chain models. We tested our method on six various datasets; among them are Gram-negative bacteria dataset, data for discriminating outer membrane proteins and apoptosis proteins dataset. We observed that our method can predict the subcellular location with high accuracy. Another advantage of the proposed method is that it can improve the accuracy of the prediction of some classes with few sequences in training and is therefore useful for datasets with imbalanced distribution of classes. CONCLUSION: This study introduces an algorithm which uses only the primary sequence of a protein to predict its subcellular location. The proposed recursive scheme represents an interesting methodology for learning and combining classifiers. The method is computationally efficient and competitive with the previously reported approaches in terms of prediction accuracies as empirical results indicate. The code for the software is available upon request. Alla Bulashevska, Roland Eils |
BMC Bioinform. | 2 |
| 2006 | Discovering functional gene expression patterns in the metabolic network of Escherichia coli with wavelets transformsabstractBACKGROUND: Microarray technology produces gene expression data on a genomic scale for an endless variety of organisms and conditions. However, this vast amount of information needs to be extracted in a reasonable way and funneled into manageable and functionally meaningful patterns. Genes may be reasonably combined using knowledge about their interaction behaviour. On a proteomic level, biochemical research has elucidated an increasingly complete image of the metabolic architecture, especially for less complex organisms like the well studied bacterium Escherichia coli. RESULTS: We sought to discover central components of the metabolic network, regulated by the expression of associated genes under changing conditions. We mapped gene expression data from E. coli under aerobic and anaerobic conditions onto the enzymatic reaction nodes of its metabolic network. An adjacency matrix of the metabolites was created from this graph. A consecutive ones clustering method was used to obtain network clusters in the matrix. The wavelet method was applied on the adjacency matrices of these clusters to collect features for the classifier. With a feature extraction method the most discriminating features were selected. We yielded network sub-graphs from these top ranking features representing formate fermentation, in good agreement with the anaerobic response of hetero-fermentative bacteria. Furthermore, we found a switch in the starting point for NAD biosynthesis, and an adaptation of the l-aspartate metabolism, in accordance with its higher abundance under anaerobic conditions. CONCLUSION: We developed and tested a novel method, based on a combination of rationally chosen machine learning methods, to analyse gene expression data on the basis of interaction data, using a metabolic network of enzymes. As a case study, we applied our method to E. coli under oxygen deprived conditions and extracted physiologically relevant patterns that represent an adaptation of the cells to changing environmental conditions. In general, our concept may be transferred to network analyses on biological interaction data, when data for two comparable states of the associated nodes are made available. Rainer König, Gunnar Schramm, Marcus Oswald, Hanna Seitz, Sebastian Sager, Marc Zapatka, Gerhard Reinelt, Roland Eils |
BMC Bioinform. | 8 |
| 2006 | GOPET: A tool for automated predictions of Gene Ontology termsabstractBACKGROUND: Vast progress in sequencing projects has called for annotation on a large scale. A Number of methods have been developed to address this challenging task. These methods, however, either apply to specific subsets, or their predictions are not formalised, or they do not provide precise confidence values for their predictions. DESCRIPTION: We recently established a learning system for automated annotation, trained with a broad variety of different organisms to predict the standardised annotation terms from Gene Ontology (GO). Now, this method has been made available to the public via our web-service GOPET (Gene Ontology term Prediction and Evaluation Tool). It supplies annotation for sequences of any organism. For each predicted term an appropriate confidence value is provided. The basic method had been developed for predicting molecular function GO-terms. It is now expanded to predict biological process terms. This web service is available via http://genius.embnet.dkfz-heidelberg.de/menu/biounit/open-husar CONCLUSION: Our web service gives experimental researchers as well as the bioinformatics community a valuable sequence annotation device. Additionally, GOPET also provides less significant annotation data which may serve as an extended discovery platform for the user. Arunachalam Vinayagam, Coral del Val, Falk Schubert, Roland Eils, Karl-Heinz Glatting, Sándor Suhai, Rainer König |
BMC Bioinform. | 4 |
| 2005 | Inferring genetic regulatory logic from expression dataabstractMotivation: High-throughput molecular genetics methods allow the collection of data about the expression of genes at different time points and under different conditions. The challenge is to infer gene regulatory interactions from these data and to get an insight into the mechanisms of genetic regulation. Results: We propose a model for genetic regulatory interactions, which has a biologically motivated Boolean logic semantics, but is of a probabilistic nature, and is hence able to confront noisy biological processes and data. We propose a method for learning the model from data based on the Bayesian approach and utilizing Gibbs sampling. We tested our method with previously published data of the Saccharomyces cerevisiae cell cycle and found relations between genes consistent with biological knowledge. Availability: The code for the software BUGS is available upon request. Contact: [email protected] Supplementary information: http://oslo.inet.dkfz-heidelberg.de/ibios_old/people/bulashev/Supplement/ Svetlana Bulashevska, Roland Eils |
Bioinform. | 2 |
| 2005 | CGH-Profiler: Data mining based on genomic aberration profilesabstractBACKGROUND: CGH-Profiler is a program that supports the analysis of genomic aberrations measured by Comparative Genomic Hybridisation (CGH). Comparative genomic hybridisation (CGH) is a well-established, molecular cytogenetic method that allows the detection of chromosomal imbalances in entire genomes. This technique is widely used in routine molecular diagnostics. Typically, chromosomal imbalances are described in a complex syntax based on the International Standard for Cytogenetic Nomenclature (ISCN). This semantic description of chromosomal imbalances hinders a large-scale statistical analysis across different experiments, e.g. for finding aberration patterns associated with a particular disease type or state. RESULTS: CGH-Profiler circumvents the semantic ISCN description by importing data from different CGH system vendors and by directly transferring the data into a table format that is readily accessible for subsequent statistical analysis. CGH-profiler comes with different consistency checks, calculates various statistics and automatically assigns a median copy number ratio to each chromosomal band. Import of CGH profiles from different CGH system vendors is already supported; its extension to other systems can be readily achieved through Perl scripts.CGH profiler can also be used to analyse comparative expressed sequence hybridisation (CESH) data. CESH reveals gene expression patterns according to chromosomal locations in a similar manner as CGH detects chromosomal imbalances. CONCLUSION: CGH-Profiler is a useful tool for processing of CGH and CESH data. Falk Schubert, Bernhard Tausch, Stefan Joos, Roland Eils |
BMC Bioinform. | 4 |
| 2005 | Cross-platform analysis of cancer microarray data improves gene expression based classification of phenotypesabstractBACKGROUND: The extensive use of DNA microarray technology in the characterization of the cell transcriptome is leading to an ever increasing amount of microarray data from cancer studies. Although similar questions for the same type of cancer are addressed in these different studies, a comparative analysis of their results is hampered by the use of heterogeneous microarray platforms and analysis methods. RESULTS: In contrast to a meta-analysis approach where results of different studies are combined on an interpretative level, we investigate here how to directly integrate raw microarray data from different studies for the purpose of supervised classification analysis. We use median rank scores and quantile discretization to derive numerically comparable measures of gene expression from different platforms. These transformed data are then used for training of classifiers based on support vector machines. We apply this approach to six publicly available cancer microarray gene expression data sets, which consist of three pairs of studies, each examining the same type of cancer, i.e. breast cancer, prostate cancer or acute myeloid leukemia. For each pair, one study was performed by means of cDNA microarrays and the other by means of oligonucleotide microarrays. In each pair, high classification accuracies (> 85%) were achieved with training and testing on data instances randomly chosen from both data sets in a cross-validation analysis. To exemplify the potential of this cross-platform classification analysis, we use two leukemia microarray data sets to show that important genes with regard to the biology of leukemia are selected in an integrated analysis, which are missed in either single-set analysis. CONCLUSION: Cross-platform classification of multiple cancer microarray data sets yields discriminative gene expression signatures that are found and validated on a large number of microarray samples, generated by different laboratories and microarray technologies. Predictive models generated by this approach are better validated than those generated on a single data set, while showing high predictive power and improved generalization performance. Patrick Warnat, Roland Eils, Benedikt Brors |
BMC Bioinform. | 2 |
| 2004 | Gene expression analysis on biochemical networks using the Potts spin modelabstractMOTIVATION: Microarray technology allows us to profile the expression of a large subset or all genes of a cell. Biochemical research over the last three decades has elucidated an increasingly complete image of the metabolic architecture. For less complex organisms, such as Escherichia coli, the biochemical network has been described in much detail. Here, we investigate the clustering of such networks by applying gene expression data that define edge lengths in the network. RESULTS: The Potts spin model is used as a nearest neighbour based clustering algorithm to discover fragmentation of the network in mutants or in biological samples when treated with drugs. As an example, we tested our method with gene expression data from E.coli treated with tryptophan excess, starvation and trpyptophan repressor mutants. We observed fragmentation of the tryptophan biosynthesis pathway, which corresponds well to the commonly known regulatory response of the cells. Rainer König, Roland Eils |
Bioinform. | 2 |
| 2004 | Applying Support Vector Machines for Gene ontology based gene function predictionabstractBACKGROUND: The current progress in sequencing projects calls for rapid, reliable and accurate function assignments of gene products. A variety of methods has been designed to annotate sequences on a large scale. However, these methods can either only be applied for specific subsets, or their results are not formalised, or they do not provide precise confidence estimates for their predictions. RESULTS: We have developed a large-scale annotation system that tackles all of these shortcomings. In our approach, annotation was provided through Gene Ontology terms by applying multiple Support Vector Machines (SVM) for the classification of correct and false predictions. The general performance of the system was benchmarked with a large dataset. An organism-wise cross-validation was performed to define confidence estimates, resulting in an average precision of 80% for 74% of all test sequences. The validation results show that the prediction performance was organism-independent and could reproduce the annotation of other automated systems as well as high-quality manual annotations. We applied our trained classification system to Xenopus laevis sequences, yielding functional annotation for more than half of the known expressed genome. Compared to the currently available annotation, we provided more than twice the number of contigs with good quality annotation, and additionally we assigned a confidence value to each predicted GO term. CONCLUSIONS: We present a complete automated annotation system that overcomes many of the usual problems by applying a controlled vocabulary of Gene Ontology and an established classification method on large and well-described sequence data sets. In a case study, the function for Xenopus laevis contig sequences was predicted and the results are publicly available at ftp://genome.dkfz-heidelberg.de/pub/agd/gene_association.agd_Xenopus. Arunachalam Vinayagam, Rainer König, Jutta Moormann, Falk Schubert, Roland Eils, Karl-Heinz Glatting, Sándor Suhai |
BMC Bioinform. | 5 |
| 2001 | An Active Contour Model for Segmentation Based on Cubic B-splines and Gradient Vector Flow
Matthias Gebhard, Julian Mattes, Roland Eils |
MICCAI | 3 |
| 2001 | New Tools for Visualization and Quantification in Dynamic Processes: Application to the Nuclear Envelope Dynamics During Mitosis
Julian Mattes, Johannes Fieres, Joël Beaudouin, Daniel Gerlich, Jan Ellenberg, Roland Eils |
MICCAI | 6 |