Peter K. Sorger

dblp:43/6532 · DBLP profile ↗
← Back
27ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0002-3364-1838ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Virtual Multiplex Staining for Histological Images Using a Marker-Wise Conditioned Diffusion Model
abstract
Multiplex imaging is revolutionizing pathology by enabling the simultaneous visualization of multiple biomarkers within tissue samples, providing molecular-level insights that traditional hematoxylin and eosin (H&E) staining cannot provide. However, the complexity and cost of multiplex data acquisition have hindered its widespread adoption. Additionally, most existing large repositories of H&E images lack corresponding multiplex images, limiting opportunities for multi-modal analysis. To address these challenges, we leverage recent advances in latent diffusion models (LDMs), which excel at modeling complex data distributions by utilizing their powerful priors for fine-tuning to a target domain. In this paper, we introduce a novel framework for virtual multiplex staining that utilizes pretrained LDM parameters to generate multiplex images from H&E images using a conditional diffusion model. Our approach enables marker-by-marker generation by conditioning the diffusion model on each marker, while sharing the same architecture across all markers. To tackle the challenge of varying pixel value distributions across different marker stains and to improve inference speed, we fine-tune the model for single-step sampling, enhancing both color contrast fidelity and inference efficiency through pixel-level loss functions. We validate our framework on two publicly available datasets, notably demonstrating its effectiveness in generating up to 18 different marker types with improved accuracy, a substantial increase over the 2-3 marker types achieved in previous approaches. This validation highlights the potential of our framework, pioneering virtual multiplex staining. Finally, this paper bridges the gap between H&E and multiplex imaging, potentially enabling retrospective studies and large-scale analyses of existing H&E image repositories.
Hyun-Jic Oh, Junsik Kim 0001, Zhiyi Shi, Yu-An Chen, Peter K. Sorger, Hanspeter Pfister, Won-Ki Jeong
AAAI6
2026 A survey on computational pathology foundation models: datasets, adaptation strategies, and evaluation tasks
abstract
Abstract Computational pathology foundation models (CPathFMs) have emerged as a powerful approach for analyzing histopathological data, leveraging self-supervised learning to extract robust feature representations from unlabeled whole-slide images. These models, categorized into uni-modal and multi-modal frameworks, have demonstrated promise in automating complex pathology tasks such as segmentation, classification, and biomarker discovery. However, the development of CPathFMs presents significant challenges, such as limited data accessibility, high variability across datasets, the necessity for domain-specific adaptation, and the lack of standardized evaluation benchmarks. This survey provides a comprehensive review of CPathFMs in computational pathology, focusing on pre-training datasets, adaptation strategies, and evaluation tasks. We analyze key techniques, such as contrastive learning, masked image modeling and multi-modal integration, and highlight existing gaps in current research. Finally, we explore future directions from four perspectives for advancing CPathFMs. This survey serves as a valuable resource for researchers, clinicians, and AI practitioners, guiding the advancement of CPathFMs toward robust and clinically applicable AI-driven pathology solutions.
Dong Li 0034, Guihong Wan, Xintao Wu, Yi He 0007, Zhong Chen 0003, Ajit Johnson Nirmal, Christine G. Lian, Peter K. Sorger, Yevgeniy R. Semenov, Chen Zhao 0010
Knowl. Inf. Syst.9
2026 SEAL: Spatially-resolved Embedding Analysis with Linked Imaging Data
abstract
Dimensionality reduction techniques help analysts make sense of complex, high-dimensional spatial datasets, such as multiplexed tissue imaging, satellite imagery, and astronomical observations, by projecting data attributes into a two-dimensional space. However, these techniques typically abstract away crucial spatial, positional, and morphological contexts, complicating interpretation and limiting insights. To address these limitations, we present SEAL, an interactive visual analytics system designed to bridge the gap between abstract 2D embeddings and their rich spatial imaging context. SEAL introduces a novel hybrid-embedding visualization that preserves image and morphological information while integrating critical high-dimensional feature data. By adapting set visualization methods, SEAL allows analysts to identify, visualize, and compare selections-defined manually or algorithmically-in both the embedding and original spatial views, facilitating a deeper understanding of the spatial arrangement and morphological characteristics of entities of interest. To elucidate differences between selected sets of items, SEAL employs a scalable surrogate model to calculate feature importance scores, identifying the most influential features governing the position of objects within embeddings. These importance scores are visually summarized across selections, with mathematical set operations enabling detailed comparative analyses. We demonstrate SEAL's effectiveness and versatility through three case studies: colorectal cancer tissue analysis with a pharmacologist, melanoma investigation with a cell biologist, and exploration of sky survey data with an astronomer. These studies underscore the importance of integrating image context into embedding spaces when interpreting complex imaging datasets. Implemented as a standalone tool while also integrating seamlessly with computational notebooks, SEAL provides an interactive platform for spatially informed exploration of high-dimensional datasets, significantly enhancing interpretability and insight generation.
Simon Warchol, Grace Guo 0001, Johannes Knittel, Dan Freeman, Usha Shalla, Jeremy Muhlich, Peter K. Sorger, Hanspeter Pfister
IEEE Trans. Vis. Comput. Graph.7
2025 Dynamic modelling of cell cycle arrest through integrated single-cell and mathematical modelling approaches
abstract
Highly multiplexed imaging assays allow simultaneous quantification of multiple protein and phosphorylation markers, providing a static snapshots of cell types and states. Pseudo-time techniques can transform these static snapshots of unsynchronized cells into dynamic trajectories, enabling the study of dynamic processes such as development trajectories and the cell cycle. Such ordering also enables training of mathematical models on these data, but technical challenges have hitherto made it difficult to integrate multiple experimental conditions, limiting the predictive power and insights these models can generate. In this work, we propose data processing and model training approaches for integrating multiplexed, multi-condition immunofluorescence data with mathematical modelling. We devise training strategies for mathematical models that are applicable to datasets where cells exhibit oscillatory as well as arrested dynamics and use them to train a cell cycle model on a dataset of MCF-10A mammary epithelial cells exposed to cell-cycle arresting small molecules. We validate the model by investigating predicted growth factor sensitivities and responses to inhibitors of cells at different initial conditions. We anticipate that our framework will generalise to other highly multiplexed measurement techniques such as mass-cytometry, rendering larger bodies of data accessible to dynamic modelling and paving the way to deeper biological insights.
Javiera Cortés-Ríos, Maria Rodriguez-Fernandez, Peter K. Sorger, Fabian Fröhlich
PLoS Comput. Biol.3
2025 Cell2Cell: Explorative Cell Interaction Analysis in Multi-Volumetric Tissue Data
abstract
We present Cell2Cell, a novel visual analytics approach for quantifying and visualizing networks of cell-cell interactions in three-dimensional (3D) multi-channel cancerous tissue data. By analyzing cellular interactions, biomedical experts can gain a more accurate understanding of the intricate relationships between cancer and immune cells. Recent methods have focused on inferring interaction based on the proximity of cells in low-resolution 2D multi-channel imaging data. By contrast, we analyze cell interactions by quantifying the presence and levels of specific proteins within a tissue sample (protein expressions) extracted from high-resolution 3D multi-channel volume data. Such analyses have a strong exploratory nature and require a tight integration of domain experts in the analysis loop to leverage their deep knowledge. We propose two complementary semi-automated approaches to cope with the increasing size and complexity of the data interactively: On the one hand, we interpret cell-to-cell interactions as edges in a cell graph and analyze the image signal (protein expressions) along those edges, using spatial as well as abstract visualizations. Complementary, we propose a cell-centered approach, enabling scientists to visually analyze polarized distributions of proteins in three dimensions, which also captures neighboring cells with biochemical and cell biological consequences. We evaluate our application in three case studies, where biologists and medical experts use Cell2Cell to investigate tumor micro-environments to identify and quantify T-cell activation in human tissue data. We confirmed that our tool can fully solve the use cases and enables a streamlined and detailed analysis of cell-cell interactions.
Eric Mörth, Kevin Sidak, Zoltan Maliga, Torsten Möller, Nils Gehlenborg, Peter K. Sorger, Hanspeter Pfister, Johanna Beyer, Robert Krüger
IEEE Trans. Vis. Comput. Graph.6
2024 Pass-Efficient Algorithms for Graph Spectral Clustering (Student Abstract)
abstract
Graph spectral clustering is a fundamental technique in data analysis, which utilizes eigenpairs of the Laplacian matrix to partition graph vertices into clusters. However, classical spectral clustering algorithms require eigendecomposition of the Laplacian matrix, which has cubic time complexity. In this work, we describe pass-efficient spectral clustering algorithms that leverage recent advances in randomized eigendecomposition and the structure of the graph vertex-edge matrix. Furthermore, we derive formulas for their efficient implementation. The resulting algorithms have a linear time complexity with respect to the number of vertices and edges and pass over the graph constant times, making them suitable for processing large graphs stored on slow memory. Experiments validate the accuracy and efficiency of the algorithms.
Boshen Yan, Guihong Wan, Haim Schweitzer, Zoltan Maliga, Sara Khattab, Kun-Hsing Yu, Peter K. Sorger, Yevgeniy R. Semenov
AAAI7
2024 SpatialCells: automated profiling of tumor microenvironments with spatially resolved multiplexed single-cell data
abstract
Cancer is a complex cellular ecosystem where malignant cells coexist and interact with immune, stromal and other cells within the tumor microenvironment (TME). Recent technological advancements in spatially resolved multiplexed imaging at single-cell resolution have led to the generation of large-scale and high-dimensional datasets from biological specimens. This underscores the necessity for automated methodologies that can effectively characterize molecular, cellular and spatial properties of TMEs for various malignancies. This study introduces SpatialCells, an open-source software package designed for region-based exploratory analysis and comprehensive characterization of TMEs using multiplexed single-cell data. The source code and tutorials are available at https://semenovlab.github.io/SpatialCells. SpatialCells efficiently streamlines the automated extraction of features from multiplexed single-cell data and can process samples containing millions of cells. Thus, SpatialCells facilitates subsequent association analyses and machine learning predictions, making it an essential tool in advancing our understanding of tumor growth, invasion and metastasis.
Guihong Wan, Zoltan Maliga, Boshen Yan, Tuulia Vallius, Yingxiao Shi, Sara Khattab, Crystal T. Chang, Ajit Johnson Nirmal, Kun-Hsing Yu, David S. L. Wei, Christine G. Lian, Mia S. Desimone, Peter K. Sorger, Yevgeniy R. Semenov
Briefings Bioinform.13
2024 psudo: Exploring Multi-Channel Biomedical Image Data with Spatially and Perceptually Optimized Pseudocoloring
abstract
Over the past century, multichannel fluorescence imaging has been pivotal in myriad scientific breakthroughs by enabling the spatial visualization of proteins within a biological sample. With the shift to digital methods and visualization software, experts can now flexibly pseudocolor and combine image channels, each corresponding to a different protein, to explore their spatial relationships. We thus propose psudo, an interactive system that allows users to create optimal color palettes for multichannel spatial data. In psudo, a novel optimization method generates palettes that maximize the perceptual differences between channels while mitigating confusing color blending in overlapping channels. We integrate this method into a system that allows users to explore multi-channel image data and compare and evaluate color palettes for their data. An interactive lensing approach provides on-demand feedback on channel overlap and a color confusion metric while giving context to the underlying channel values. Color palettes can be applied globally or, using the lens, to local regions of interest. We evaluate our palette optimization approach using three graphical perception tasks in a crowdsourced user study with 150 participants, showing that users are more accurate at discerning and comparing the underlying data using our approach. Additionally, we showcase psudo in a case study exploring the complex immune responses in cancer tissue data with a biologist.
Simon Warchol, Jakob Troidl, Jeremy Muhlich, Robert Krüger, John Hoffer, Tica Lin, Johanna Beyer, Elena L. Glassman, Peter K. Sorger, Hanspeter Pfister
Comput. Graph. Forum9
2024 Residency Octree: A Hybrid Approach for Scalable Web-Based Multi-Volume Rendering
abstract
We present a hybrid multi-volume rendering approach based on a novel Residency Octree that combines the advantages of out-of-core volume rendering using page tables with those of standard octrees. Octree approaches work by performing hierarchical tree traversal. However, in octree volume rendering, tree traversal and the selection of data resolution are intrinsically coupled. This makes fine-grained empty-space skipping costly. Page tables, on the other hand, allow access to any cached brick from any resolution. However, they do not offer a clear and efficient strategy for substituting missing high-resolution data with lower-resolution data. We enable flexible mixed-resolution out-of-core multi-volume rendering by decoupling the cache residency of multi-resolution data from a resolution-independent spatial subdivision determined by the tree. Instead of one-to-one node-to-brick correspondences, each residency octree node is mapped to a set of bricks from different resolution levels. This makes it possible to efficiently and adaptively choose and mix resolutions, adapt sampling rates, and compensate for cache misses. At the same time, residency octrees support fine-grained empty-space skipping, independent of the data subdivision used for caching. Finally, to facilitate collaboration and outreach, and to eliminate local data storage, our implementation is a web-based, pure client-side renderer using WebGPU and WebAssembly. Our method is faster than prior approaches and efficient for many data channels with a flexible and adaptive choice of data resolution.
Lukas Herzberger, Markus Hadwiger, Robert Krüger, Peter K. Sorger, Hanspeter Pfister, M. Eduard Gröller, Johanna Beyer
IEEE Trans. Vis. Comput. Graph.4
2023 Visinity: Visual Spatial Neighborhood Analysis for Multiplexed Tissue Imaging Data
abstract
New highly-multiplexed imaging technologies have enabled the study of tissues in unprecedented detail. These methods are increasingly being applied to understand how cancer cells and immune response change during tumor development, progression, and metastasis, as well as following treatment. Yet, existing analysis approaches focus on investigating small tissue samples on a per-cell basis, not taking into account the spatial proximity of cells, which indicates cell-cell interaction and specific biological processes in the larger cancer microenvironment. We present Visinity, a scalable visual analytics system to analyze cell interaction patterns across cohorts of whole-slide multiplexed tissue images. Our approach is based on a fast regional neighborhood computation, leveraging unsupervised learning to quantify, compare, and group cells by their surrounding cellular neighborhood. These neighborhoods can be visually analyzed in an exploratory and confirmatory workflow. Users can explore spatial patterns present across tissues through a scalable image viewer and coordinated views highlighting the neighborhood composition and spatial arrangements of cells. To verify or refine existing hypotheses, users can query for specific patterns to determine their presence and statistical significance. Findings can be interactively annotated, ranked, and compared in the form of small multiples. In two case studies with biomedical experts, we demonstrate that Visinity can identify common biological processes within a human tonsil and uncover novel white-blood cell networks and immune-tumor interactions.
Simon Warchol, Robert Krüger, Ajit Johnson Nirmal, Giorgio Gaglia, Jared Jessup, Cecily C. Ritch, John Hoffer, Jeremy Muhlich, Megan L. Burger, Tyler Jacks, Sandro Santagata, Peter K. Sorger, Hanspeter Pfister
IEEE Trans. Vis. Comput. Graph.12
2022 Stitching and registering highly multiplexed whole-slide images of tissues and tumors using ASHLAR
abstract
MOTIVATION: Stitching microscope images into a mosaic is an essential step in the analysis and visualization of large biological specimens, particularly human and animal tissues. Recent approaches to highly multiplexed imaging generate high-plex data from sequential rounds of lower-plex imaging. These multiplexed imaging methods promise to yield precise molecular single-cell data and information on cellular neighborhoods and tissue architecture. However, attaining mosaic images with single-cell accuracy requires robust image stitching and image registration capabilities that are not met by existing methods. RESULTS: We describe the development and testing of ASHLAR, a Python tool for coordinated stitching and registration of 103 or more individual multiplexed images to generate accurate whole-slide mosaics. ASHLAR reads image formats from most commercial microscopes and slide scanners, and we show that it performs better than existing open-source and commercial software. ASHLAR outputs standard OME-TIFF images that are ready for analysis by other open-source tools and recently developed image analysis pipelines. AVAILABILITY AND IMPLEMENTATION: ASHLAR is written in Python and is available under the MIT license at https://github.com/labsyspharm/ashlar. The newly published data underlying this article are available in Sage Synapse at https://dx.doi.org/10.7303/syn25826362; the availability of other previously published data re-analyzed in this article is described in Supplementary Table S4. An informational website with user guides and test data is available at https://labsyspharm.github.io/ashlar/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jeremy Muhlich, Yu-An Chen, Clarence Han-Wei Yapp, Douglas Russell, Sandro Santagata, Peter K. Sorger
Bioinform.6
2022 Fides: Reliable trust-region optimization for parameter estimation of ordinary differential equation models
abstract
Ordinary differential equation (ODE) models are widely used to study biochemical reactions in cellular networks since they effectively describe the temporal evolution of these networks using mass action kinetics. The parameters of these models are rarely known a priori and must instead be estimated by calibration using experimental data. Optimization-based calibration of ODE models on is often challenging, even for low-dimensional problems. Multiple hypotheses have been advanced to explain why biochemical model calibration is challenging, including non-identifiability of model parameters, but there are few comprehensive studies that test these hypotheses, likely because tools for performing such studies are also lacking. Nonetheless, reliable model calibration is essential for uncertainty analysis, model comparison, and biological interpretation. We implemented an established trust-region method as a modular Python framework (fides) to enable systematic comparison of different approaches to ODE model calibration involving a variety of Hessian approximation schemes. We evaluated fides on a recently developed corpus of biologically realistic benchmark problems for which real experimental data are available. Unexpectedly, we observed high variability in optimizer performance among different implementations of the same mathematical instructions (algorithms). Analysis of possible sources of poor optimizer performance identified limitations in the widely used Gauss-Newton, BFGS and SR1 Hessian approximation schemes. We addressed these drawbacks with a novel hybrid Hessian approximation scheme that enhances optimizer performance and outperforms existing hybrid approaches. When applied to the corpus of test models, we found that fides was on average more reliable and efficient than existing methods using a variety of criteria. We expect fides to be broadly useful for ODE constrained optimization problems in biochemical models and to be a foundation for future methods development.
Fabian Fröhlich, Peter K. Sorger
PLoS Comput. Biol.2
2022 Scope2Screen: Focus+Context Techniques for Pathology Tumor Assessment in Multivariate Image Data
abstract
Inspection of tissues using a light microscope is the primary method of diagnosing many diseases, notably cancer. Highly multiplexed tissue imaging builds on this foundation, enabling the collection of up to 60 channels of molecular information plus cell and tissue morphology using antibody staining. This provides unique insight into disease biology and promises to help with the design of patient-specific therapies. However, a substantial gap remains with respect to visualizing the resulting multivariate image data and effectively supporting pathology workflows in digital environments on screen. We, therefore, developed Scope2Screen, a scalable software system for focus+context exploration and annotation of whole-slide, high-plex, tissue images. Our approach scales to analyzing 100GB images of 109or more pixels per channel, containing millions of individual cells. A multidisciplinary team of visualization experts, microscopists, and pathologists identified key image exploration and annotation tasks involving finding, magnifying, quantifying, and organizing regions of interest (ROIs) in an intuitive and cohesive manner. Building on a scope-to-screen metaphor, we present interactive lensing techniques that operate at single-cell and tissue levels. Lenses are equipped with task-specific functionality and descriptive statistics, making it possible to analyze image features, cell types, and spatial arrangements (neighborhoods) across image channels and scales. A fast sliding-window search guides users to regions similar to those under the lens; these regions can be analyzed and considered either separately or as part of a larger image collection. A novel snapshot method enables linked lens configurations and image statistics to be saved, restored, and shared with these regions. We validate our designs with domain experts and apply Scope2Screen in two case studies involving lung and colorectal cancers to discover cancer-relevant image features.
Jared Jessup, Robert Krüger, Simon Warchol, John Hoffer, Jeremy Muhlich, Cecily C. Ritch, Giorgio Gaglia, Shannon Coy, Yu-An Chen, Jia-Ren Lin, Sandro Santagata, Peter K. Sorger, Hanspeter Pfister
IEEE Trans. Vis. Comput. Graph.12
2020 Channel Embedding for Informative Protein Identification from Highly Multiplexed Images
Salma Abdel Magid, Won-Dong Jang, Denis Schapiro, Donglai Wei 0001, James Tompkin 0001, Peter K. Sorger, Hanspeter Pfister
MICCAI (5)6
2020 Facetto: Combining Unsupervised and Supervised Learning for Hierarchical Phenotype Analysis in Multi-Channel Image Data
abstract
Facetto is a scalable visual analytics application that is used to discover single-cell phenotypes in high-dimensional multi-channel microscopy images of human tumors and tissues. Such images represent the cutting edge of digital histology and promise to revolutionize how diseases such as cancer are studied, diagnosed, and treated. Highly multiplexed tissue images are complex, comprising 109 or more pixels, 60-plus channels, and millions of individual cells. This makes manual analysis challenging and error-prone. Existing automated approaches are also inadequate, in large part, because they are unable to effectively exploit the deep knowledge of human tissue biology available to anatomic pathologists. To overcome these challenges, Facetto enables a semi-automated analysis of cell types and states. It integrates unsupervised and supervised learning into the image and feature exploration process and offers tools for analytical provenance. Experts can cluster the data to discover new types of cancer and immune cells and use clustering results to train a convolutional neural network that classifies new cells accordingly. Likewise, the output of classifiers can be clustered to discover aggregate patterns and phenotype subsets. We also introduce a new hierarchical approach to keep track of analysis steps and data subsets created by users; this assists in the identification of cell types. Users can build phenotype trees and interact with the resulting hierarchical structures of both high-dimensional feature and image spaces. We report on use-cases in which domain scientists explore various large-scale fluorescence imaging datasets. We demonstrate how Facetto assists users in steering the clustering and classification process, inspecting analysis results, and gaining new scientific insights into cancer biology.
Robert Krüger, Johanna Beyer, Won-Dong Jang, Artem Sokolov 0003, Peter K. Sorger, Hanspeter Pfister
IEEE Trans. Vis. Comput. Graph.6
2019 INDRA-IPM: interactive pathway modeling using natural language with automated assembly
abstract
SUMMARY: INDRA-IPM (Interactive Pathway Map) is a web-based pathway map modeling tool that combines natural language processing with automated model assembly and visualization. INDRA-IPM contextualizes models with expression data and exports them to standard formats. AVAILABILITY AND IMPLEMENTATION: INDRA-IPM is available at: http://pathwaymap.indra.bio. Source code is available at http://github.com/sorgerlab/indra_pathway_map. The underlying web service API is available at http://api.indra.bio:8000. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Petar V. Todorov, Benjamin M. Gyori, John A. Bachman, Peter K. Sorger
Bioinform.4
2019 Inferring reaction network structure from single-cell, multiplex data, using toric systems theory
abstract
The goal of many single-cell studies on eukaryotic cells is to gain insight into the biochemical reactions that control cell fate and state. In this paper we introduce the concept of Effective Stoichiometric Spaces (ESS) to guide the reconstruction of biochemical networks from multiplexed, fixed time-point, single-cell data. In contrast to methods based solely on statistical models of data, the ESS method leverages the power of the geometric theory of toric varieties to begin unraveling the structure of chemical reaction networks (CRN). This application of toric theory enables a data-driven mapping of covariance relationships in single-cell measurements into stoichiometric information, one in which each cell subpopulation has its associated ESS interpreted in terms of CRN theory. In the development of ESS we reframe certain aspects of the theory of CRN to better match data analysis. As an application of our approach we process cytomery- and image-based single-cell datasets and identify differences in cells treated with kinase inhibitors. Our approach is directly applicable to data acquired using readily accessible experimental methods such as Fluorescence Activated Cell Sorting (FACS) and multiplex immunofluorescence.
Jia-Ren Lin, Eduardo D. Sontag, Peter K. Sorger
PLoS Comput. Biol.4
2018 FamPlex: a resource for entity recognition and relationship resolution of human protein families and complexes in biomedical text mining
abstract
BACKGROUND: For automated reading of scientific publications to extract useful information about molecular mechanisms it is critical that genes, proteins and other entities be correctly associated with uniform identifiers, a process known as named entity linking or "grounding." Correct grounding is essential for resolving relationships among mined information, curated interaction databases, and biological datasets. The accuracy of this process is largely dependent on the availability of machine-readable resources associating synonyms and abbreviations commonly found in biomedical literature with uniform identifiers. RESULTS: In a task involving automated reading of ∼215,000 articles using the REACH event extraction software we found that grounding was disproportionately inaccurate for multi-protein families (e.g., "AKT") and complexes with multiple subunits (e.g."NF- κB"). To address this problem we constructed FamPlex, a manually curated resource defining protein families and complexes as they are commonly encountered in biomedical text. In FamPlex the gene-level constituents of families and complexes are defined in a flexible format allowing for multi-level, hierarchical membership. To create FamPlex, text strings corresponding to entities were identified empirically from literature and linked manually to uniform identifiers; these identifiers were also mapped to equivalent entries in multiple related databases. FamPlex also includes curated prefix and suffix patterns that improve named entity recognition and event extraction. Evaluation of REACH extractions on a test corpus of ∼54,000 articles showed that FamPlex significantly increased grounding accuracy for families and complexes (from 15 to 71%). The hierarchical organization of entities in FamPlex also made it possible to integrate otherwise unconnected mechanistic information across families, subfamilies, and individual proteins. Applications of FamPlex to the TRIPS/DRUM reading system and the Biocreative VI Bioentity Normalization Task dataset demonstrated the utility of FamPlex in other settings. CONCLUSION: FamPlex is an effective resource for improving named entity recognition, grounding, and relationship resolution in automated reading of biomedical text. The content in FamPlex is available in both tabular and Open Biomedical Ontology formats at https://github.com/sorgerlab/famplex under the Creative Commons CC0 license and has been integrated into the TRIPS/DRUM and REACH reading systems.
John A. Bachman, Benjamin M. Gyori, Peter K. Sorger
BMC Bioinform.3
2013 Discovering causal pathways linking genomic events to transcriptional states using Tied Diffusion Through Interacting Events (TieDIE)
abstract
MOTIVATION: Identifying the cellular wiring that connects genomic perturbations to transcriptional changes in cancer is essential to gain a mechanistic understanding of disease initiation, progression and ultimately to predict drug response. We have developed a method called Tied Diffusion Through Interacting Events (TieDIE) that uses a network diffusion approach to connect genomic perturbations to gene expression changes characteristic of cancer subtypes. The method computes a subnetwork of protein-protein interactions, predicted transcription factor-to-target connections and curated interactions from literature that connects genomic and transcriptomic perturbations. RESULTS: Application of TieDIE to The Cancer Genome Atlas and a breast cancer cell line dataset identified key signaling pathways, with examples impinging on MYC activity. Interlinking genes are predicted to correspond to essential components of cancer signaling and may provide a mechanistic explanation of tumor character and suggest subtype-specific drug targets. AVAILABILITY: Software is available from the Stuart lab's wiki: https://sysbiowiki.soe.ucsc.edu/tiedie. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Evan O. Paull, Daniel E. Carlin, Mario Niepel, Peter K. Sorger, David Haussler, Joshua M. Stuart
Bioinform.4
2012 Exploring the Contextual Sensitivity of Factors that Determine Cell-to-Cell Variability in Receptor-Mediated Apoptosis
abstract
Stochastic fluctuations in gene expression give rise to cell-to-cell variability in protein levels which can potentially cause variability in cellular phenotype. For TRAIL (TNF-related apoptosis-inducing ligand) variability manifests itself as dramatic differences in the time between ligand exposure and the sudden activation of the effector caspases that kill cells. However, the contribution of individual proteins to phenotypic variability has not been explored in detail. In this paper we use feature-based sensitivity analysis as a means to estimate the impact of variation in key apoptosis regulators on variability in the dynamics of cell death. We use Monte Carlo sampling from measured protein concentration distributions in combination with a previously validated ordinary differential equation model of apoptosis to simulate the dynamics of receptor-mediated apoptosis. We find that variation in the concentrations of some proteins matters much more than variation in others and that precisely which proteins matter depends both on the concentrations of other proteins and on whether correlations in protein levels are taken into account. A prediction from simulation that we confirm experimentally is that variability in fate is sensitive to even small increases in the levels of Bcl-2. We also show that sensitivity to Bcl-2 levels is itself sensitive to the levels of interacting proteins. The contextual dependency is implicit in the mathematical formulation of sensitivity, but our data show that it is also important for biologically relevant parameter values. Our work provides a conceptual and practical means to study and understand the impact of cell-to-cell variability in protein expression levels on cell fate using deterministic models and sampling from parameter distributions.
Suzanne Gaudet, Sabrina L. Spencer, William W. Chen, Peter K. Sorger
PLoS Comput. Biol.4
2011 Training Signaling Pathway Maps to Biochemical Data with Constrained Fuzzy Logic: Quantitative Analysis of Liver Cell Responses to Inflammatory Stimuli
abstract
Predictive understanding of cell signaling network operation based on general prior knowledge but consistent with empirical data in a specific environmental context is a current challenge in computational biology. Recent work has demonstrated that Boolean logic can be used to create context-specific network models by training proteomic pathway maps to dedicated biochemical data; however, the Boolean formalism is restricted to characterizing protein species as either fully active or inactive. To advance beyond this limitation, we propose a novel form of fuzzy logic sufficiently flexible to model quantitative data but also sufficiently simple to efficiently construct models by training pathway maps on dedicated experimental measurements. Our new approach, termed constrained fuzzy logic (cFL), converts a prior knowledge network (obtained from literature or interactome databases) into a computable model that describes graded values of protein activation across multiple pathways. We train a cFL-converted network to experimental data describing hepatocytic protein activation by inflammatory cytokines and demonstrate the application of the resultant trained models for three important purposes: (a) generating experimentally testable biological hypotheses concerning pathway crosstalk, (b) establishing capability for quantitative prediction of protein activity, and (c) prediction and understanding of the cytokine release phenotypic response. Our methodology systematically and quantitatively trains a protein pathway map summarizing curated literature to context-specific biochemical data. This process generates a computable model yielding successful prediction of new test data and offering biological insight into complex datasets that are difficult to fully analyze by intuition alone.
Melody K. Morris, Julio Saez-Rodriguez, David C. Clarke, Peter K. Sorger, Douglas A. Lauffenburger
PLoS Comput. Biol.4
2010 Systematic calibration of a cell signaling network model
abstract
BACKGROUND: Mathematical modeling is being applied to increasingly complex biological systems and datasets; however, the process of analyzing and calibrating against experimental data is often challenging and a rate limiting step in model development. To address this problem, we developed a systematic methodology for calibrating quantitative models of dynamic biological processes and illustrate its utility by validating a model of TRAIL (Tumor necrosis factor Related Apoptosis-Inducing Ligand)-induced cell death. RESULTS: We propose a serial framework integrating analysis and calibration modules and we compare various methods for global sensitivity analysis and global parameter estimation. First, adequacy of the network structure is checked by global sensitivity analysis to changes in concentrations of molecular species, validating that the model can reproduce qualitative features of the system behavior derived from experiments or literature surveys. Second, rate parameters are ranked by importance using gradient-based and variance-based sensitivity indices, and we systematically determine the optimal number of parameters to include in model calibration. Third, deterministic, stochastic and hybrid algorithms for global optimization are applied to estimate the values of the most important parameters by fitting to time series data. We compare the performance of these three optimization algorithms. CONCLUSIONS: Our proposed framework covers the entire process from validating a proto-model to establishing a realistic model for in silico experiments and thereby provides a generalized workflow for the construction of predictive models of complex network systems.
Kyoung Ae Kim, Sabrina L. Spencer, John G. Albeck, John M. Burke, Peter K. Sorger, Suzanne Gaudet
BMC Bioinform.5
2009 Fuzzy Logic Analysis of Kinase Pathway Crosstalk in TNF/EGF/Insulin-Induced Signaling
abstract
When modeling cell signaling networks, a balance must be struck between mechanistic detail and ease of interpretation. In this paper we apply a fuzzy logic framework to the analysis of a large, systematic dataset describing the dynamics of cell signaling downstream of TNF, EGF, and insulin receptors in human colon carcinoma cells. Simulations based on fuzzy logic recapitulate most features of the data and generate several predictions involving pathway crosstalk and regulation. We uncover a relationship between MK2 and ERK pathways that might account for the previously identified pro-survival influence of MK2. We also find unexpected inhibition of IKK following EGF treatment, possibly due to down-regulation of autocrine signaling. More generally, fuzzy logic models are flexible, able to incorporate qualitative and noisy data, and powerful enough to produce quantitative predictions and new biological insights about the operation of signaling networks.
Bree B. Aldridge, Julio Saez-Rodriguez, Jeremy Muhlich, Peter K. Sorger, Douglas A. Lauffenburger
PLoS Comput. Biol.4
2009 The Logic of EGFR/ErbB Signaling: Theoretical Properties and Analysis of High-Throughput Data
abstract
The epidermal growth factor receptor (EGFR) signaling pathway is probably the best-studied receptor system in mammalian cells, and it also has become a popular example for employing mathematical modeling to cellular signaling networks. Dynamic models have the highest explanatory and predictive potential; however, the lack of kinetic information restricts current models of EGFR signaling to smaller sub-networks. This work aims to provide a large-scale qualitative model that comprises the main and also the side routes of EGFR/ErbB signaling and that still enables one to derive important functional properties and predictions. Using a recently introduced logical modeling framework, we first examined general topological properties and the qualitative stimulus-response behavior of the network. With species equivalence classes, we introduce a new technique for logical networks that reveals sets of nodes strongly coupled in their behavior. We also analyzed a model variant which explicitly accounts for uncertainties regarding the logical combination of signals in the model. The predictive power of this model is still high, indicating highly redundant sub-structures in the network. Finally, one key advance of this work is the introduction of new techniques for assessing high-throughput data with logical models (and their underlying interaction graph). By employing these techniques for phospho-proteomic data from primary hepatocytes and the HepG2 cell line, we demonstrate that our approach enables one to uncover inconsistencies between experimental results and our current qualitative knowledge and to generate new hypotheses and conclusions. Our results strongly suggest that the Rac/Cdc42 induced p38 and JNK cascades are independent of PI3K in both primary hepatocytes and HepG2. Furthermore, we detected that the activation of JNK in response to neuregulin follows a PI3K-dependent signaling pathway.
Regina Samaga, Julio Saez-Rodriguez, Leonidas G. Alexopoulos, Peter K. Sorger, Steffen Klamt
PLoS Comput. Biol.4
2008 Modeling pro-death signaling pathways in cancer hepatocytes using multi-combinatorial treatments of inhibitors and stimuli
abstract
Cell death in hepatocellular carcinoma (HCC) cells as well as in other cell types is driven by a complex signaling transduction network comprised of several different pathways. In this study, we measured cell death of the HCC cell line C3A and we quantify the contribution of several signaling pathways using a reverse approach; namely instead of measuring the activity of the intracellular signal, we correlated the inhibition of that signal to the cell death. To achieve that, cells were treated with 6 stimuli and 7 inhibitors in a multi-combinatorial manner. Approximately 400 observations were made with the simultaneous treatment of up to 2 inhibitors with up to 3 stimuli. A modified linear regression model was developed to predict cell death as indicated by lactate dehydrogenase (LDH) activity. The correlation coefficients of this model were used to quantify the role of each pathway on HCC cell death. Our results are in good agreement with the literature; caspase 8 was revealed as the strongest pro-death mediator whereas the Akt pathway was shown to be the most pro-survival signal. MEK/ERK pathway had a dual role depending on the applied stimuli. Experimentally, the present method is fast, cost efficient, and easy. The coupling of the experiments to a mathematical model allowed us to quantify the contributions of a broad spectrum of pathways on cellular behavior. Additionally, such approach has a significant advantage in cases where measurements of signaling activities (i.e. from cell lysates) are experimentally impractical. For example, using the current methodology we might be able to monitor chondrocyte signaling transduction within its native environment: the articular cartilage extracellular matrix which presents a major challenge in cartilage biology.
Leonidas G. Alexopoulos, Douglas A. Lauffenburger, Peter K. Sorger
BIBE3
2008 Flexible informatics for linking experimental data to mathematical models via DataRail
abstract
MOTIVATION: Linking experimental data to mathematical models in biology is impeded by the lack of suitable software to manage and transform data. Model calibration would be facilitated and models would increase in value were it possible to preserve links to training data along with a record of all normalization, scaling, and fusion routines used to assemble the training data from primary results. RESULTS: We describe the implementation of DataRail, an open source MATLAB-based toolbox that stores experimental data in flexible multi-dimensional arrays, transforms arrays so as to maximize information content, and then constructs models using internal or external tools. Data integrity is maintained via a containment hierarchy for arrays, imposition of a metadata standard based on a newly proposed MIDAS format, assignment of semantically typed universal identifiers, and implementation of a procedure for storing the history of all transformations with the array. We illustrate the utility of DataRail by processing a newly collected set of approximately 22 000 measurements of protein activities obtained from cytokine-stimulated primary and transformed human liver cells. AVAILABILITY: DataRail is distributed under the GNU General Public License and available at http://code.google.com/p/sbpipeline/
Julio Saez-Rodriguez, Arthur Goldsipe, Jeremy Muhlich, Leonidas G. Alexopoulos, Bjorn Millard, Douglas A. Lauffenburger, Peter K. Sorger
Bioinform.7
2007 Phenotypic clustering of yeast mutants based on kinetochore microtubule dynamics
abstract
MOTIVATION: Kinetochores are multiprotein complexes which mediate chromosome attachment to microtubules (MTs) of the mitotic spindle. They regulate MT dynamics during chromosome segregation. Our goal is to identify groups of kinetochore proteins with similar effects on MT dynamics, revealing pathways through which kinetochore proteins transform chemical and mechanical input signals into cues of MT regulation. RESULTS: We have developed a hierarchical, agglomerative clustering algorithm that groups Saccharomyces cerevisiae strains based on MT-mediated chromosome dynamics measured by high-resolution live cell microscopy. Clustering is based on parameters of autoregressive moving average (ARMA) models of the probed dynamics. We have found that the regulation of wildtype MT dynamics varies with cell cycle and temperature, but not with the chromosome an MT is attached to. By clustering the dynamics of mutants, we discovered that the three genes IPL1, DAM1 and KIP3 co-regulate MT dynamics. Our study establishes the clustering of chromosome and MT dynamics by ARMA descriptors as a sensitive framework for the systematic identification of kinetochore protein subcomplexes and pathways for the regulation of MT dynamics. AVAILABILITY: The clustering code, written in Matlab, can be downloaded from http://lccb.scripps.edu. ('download' hyperlink at bottom of website). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Khuloud Jaqaman, Jonas F. Dorn, Eugenio Marco, Peter K. Sorger, Gaudenz Danuser
Bioinform.4