EDBT 2026 Demo / reviewers in the wild / expert
Robert F. Murphy
dblp:88/220
· DBLP profile ↗
44ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-0358-901XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Artificial intelligence and machine learning · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CytoSpatio: Learning cell type spatial relationships using multirange, multitype point process modelsabstractRecent advances in multiplexed fluorescence imaging have provided new opportunities for deciphering the complex spatial relationships among various cell types across diverse tissues. We introduce CytoSpatio, open-source software that constructs generative, multirange, and multitype point process models that capture interactions among multiple cell types at various distances simultaneously. On analyzing five cell types across five tissues, our software showed consistent spatial relationships within the same tissue type, with certain cell types like proliferating T cells consistently clustering across tissue types. It also revealed that the attraction-repulsion relationships between cell types like B cells and CD4-positive T cells vary with tissue type. Models for a published dataset demonstrated consistency with prior findings. CytoSpatio can also generate synthetic tissue patterns from learned models, a capability not provided by previous descriptive, motif-based approaches. This potentially allows spatially realistic simulations of how cell relationships affect tissue biochemistry. Yangyuan Zhang, Robert F. Murphy |
PLoS Comput. Biol. | 3 |
| 2024 | Expanding the coverage of spatial proteomics: a machine learning approachabstractMOTIVATION: Multiplexed protein imaging methods use a chosen set of markers and provide valuable information about complex tissue structure and cellular heterogeneity. However, the number of markers that can be measured in the same tissue sample is inherently limited. RESULTS: In this paper, we present an efficient method to choose a minimal predictive subset of markers that for the first time allows the prediction of full images for a much larger set of markers. We demonstrate that our approach also outperforms previous methods for predicting cell-level protein composition. Most importantly, we demonstrate that our approach can be used to select a marker set that enables prediction of a much larger set than could be measured concurrently. AVAILABILITY AND IMPLEMENTATION: All code and intermediate results are available in a Reproducible Research Archive at https://github.com/murphygroup/CODEXPanelOptimization. Huangqingbo Sun, Robert F. Murphy |
Bioinform. | 3 |
| 2023 | Implementing and Evaluating ASSISTments Online Math Homework Support At large Scale over Two Years: Findings and Lessons Learned
Mingyu Feng, Neil T. Heffernan, Kelly Collins, Cristina Heffernan, Robert F. Murphy |
AIED | 5 |
| 2022 | Improving and evaluating deep learning models of cellular organizationabstractMOTIVATION: Cells contain dozens of major organelles and thousands of other structures, many of which vary extensively in their number, size, shape and spatial distribution. This complexity and variation dramatically complicates the use of both traditional and deep learning methods to build accurate models of cell organization. Most cellular organelles are distinct objects with defined boundaries that do not overlap, while the pixel resolution of most imaging methods is n sufficient to resolve these boundaries. Thus while cell organization is conceptually object-based, most current methods are pixel-based. Using extensive image collections in which particular organelles were fluorescently labeled, deep learning methods can be used to build conditional autoencoder models for particular organelles. A major advance occurred with the use of a U-net approach to make multiple models all conditional upon a common reference, unlabeled image, allowing the relationships between different organelles to be at least partially inferred. RESULTS: We have developed improved Generative Adversarial Networks-based approaches for learning these models and have also developed novel criteria for evaluating how well synthetic cell images reflect the properties of real images. The first set of criteria measure how well models preserve the expected property that organelles do not overlap. We also developed a modified loss function that allows retraining of the models to minimize that overlap. The second set of criteria uses object-based modeling to compare object shape and spatial distribution between synthetic and real images. Our work provides the first demonstration that, at least for some organelles, deep learning models can capture object-level properties of cell images. AVAILABILITY AND IMPLEMENTATION: http://murphylab.cbd.cmu.edu/Software/2022_insilico. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huangqingbo Sun, Xuecong Fu, Serena Abraham, Shen Jin, Robert F. Murphy |
Bioinform. | 5 |
| 2021 | Evaluation of categorical matrix completion algorithms: toward improved active learning for drug discoveryabstractMOTIVATION: High throughput and high content screening are extensively used to determine the effect of small molecule compounds and other potential therapeutics upon particular targets as part of the early drug development process. However, screening is typically used to find compounds that have a desired effect but not to identify potential undesirable side effects. This is because the size of the search space precludes measuring the potential effect of all compounds on all targets. Active machine learning has been proposed as a solution to this problem. RESULTS: In this article, we describe an improved imputation method, Impute by Committee, for completion of matrices containing categorical values. We compare this method to existing approaches in the context of modeling the effects of many compounds on many targets using latent similarities between compounds and conditions. We also compare these methods for the task of driving active learning in well-characterized settings for synthetic and real datasets. Our new approach performed the best overall both in the accuracy of matrix completion itself and in the number of experiments needed to train an accurate predictive model compared to random selection of experiments. We further improved upon the performance of our new method by developing an adaptive switching strategy for active learning that iteratively chooses between different matrix completion methods. AVAILABILITY AND IMPLEMENTATION: A Reproducible Research Archive containing all data and code is available at http://murphylab.cbd.cmu.edu/software. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huangqingbo Sun, Robert F. Murphy |
Bioinform. | 2 |
| 2020 | Learning complex subcellular distribution patterns of proteins via analysis of immunohistochemistry imagesabstractMOTIVATION: Systematic and comprehensive analysis of protein subcellular location as a critical part of proteomics ('location proteomics') has been studied for many years, but annotating protein subcellular locations and understanding variation of the location patterns across various cell types and states is still challenging. RESULTS: In this work, we used immunohistochemistry images from the Human Protein Atlas as the source of subcellular location information, and built classification models for the complex protein spatial distribution in normal and cancerous tissues. The models can automatically estimate the fractions of protein in different subcellular locations, and can help to quantify the changes of protein distribution from normal to cancer tissues. In addition, we examined the extent to which different annotated protein pathways and complexes showed similarity in the locations of their member proteins, and then predicted new potential proteins for these networks. AVAILABILITY AND IMPLEMENTATION: The dataset and code are available at: www.csbio.sjtu.edu.cn/bioinf/complexsubcellularpatterns. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ying-Ying Xu, Hong-Bin Shen, Robert F. Murphy |
Bioinform. | 3 |
| 2019 | Evaluation of methods for generative modeling of cell and nuclear shapeabstractMOTIVATION: Cell shape provides both geometry for, and a reflection of, cell function. Numerous methods for describing and modeling cell shape have been described, but previous evaluation of these methods in terms of the accuracy of generative models has been limited. RESULTS: Here we compare traditional methods and deep autoencoders to build generative models for cell shapes in terms of the accuracy with which shapes can be reconstructed from models. We evaluated the methods on different collections of 2D and 3D cell images, and found that none of the methods gave accurate reconstructions using low dimensional encodings. As expected, much higher accuracies were observed using high dimensional encodings, with outline-based methods significantly outperforming image-based autoencoders. The latter tended to encode all cells as having smooth shapes, even for high dimensions. For complex 3D cell shapes, we developed a significant improvement of a method based on the spherical harmonic transform that performs significantly better than other methods. We obtained similar results for the joint modeling of cell and nuclear shape. Finally, we evaluated the modeling of shape dynamics by interpolation in the shape space. We found that our modified method provided lower deformation energies along linear interpolation paths than other methods. This allows practical shape evolution in high dimensional shape spaces. We conclude that our improved spherical harmonic based methods are preferable for cell and nuclear shape modeling, providing better representations, higher computational efficiency and requiring fewer training images than deep learning methods. AVAILABILITY AND IMPLEMENTATION: All software and data is available at http://murphylab.cbd.cmu.edu/software. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiongtao Ruan, Robert F. Murphy |
Bioinform. | 2 |
| 2019 | Learning the sequence of influenza A genome assembly during viral replication using point process models and fluorescence in situ hybridizationabstractWithin influenza virus infected cells, viral genomic RNA are selectively packed into progeny virions, which predominantly contain a single copy of 8 viral RNA segments. Intersegmental RNA-RNA interactions are thought to mediate selective packaging of each viral ribonucleoprotein complex (vRNP). Clear evidence of a specific interaction network culminating in the full genomic set has yet to be identified. Using multi-color fluorescence in situ hybridization to visualize four vRNP segments within a single cell, we developed image-based models of vRNP-vRNP spatial dependence. These models were used to construct likely sequences of vRNP associations resulting in the full genomic set. Our results support the notion that selective packaging occurs during cytoplasmic transport and identifies the formation of multiple distinct vRNP sub-complexes that likely form as intermediate steps toward full genomic inclusion into a progeny virion. The methods employed demonstrate a statistically driven, model based approach applicable to other interaction and assembly problems. Timothy Majarian, Robert F. Murphy, Seema S. Lakdawala |
PLoS Comput. Biol. | 2 |
| 2018 | Learning Generative Models of Tissue Organization with Supervised GANsabstractA key step in understanding the spatial organization of cells and tissues is the ability to construct generative models that accurately reflect that organization. In this paper, we focus on building generative models of electron microscope (EM) images in which the positions of cell membranes and mitochondria have been densely annotated, and propose a two-stage procedure that produces realistic images using Generative Adversarial Networks (or GANs) in a supervised way. In the first stage, we synthesize a label "image" given a noise "image" as input, which then provides supervision for EM image synthesis in the second stage. The full model naturally generates label-image pairs. We show that accurate synthetic EM images are produced using assessment via (1) shape features and global statistics, (2) segmentation accuracies, and (3) user studies. We also demonstrate further improvements by enforcing a reconstruction loss on intermediate synthetic labels and thus unifying the two stages into one single end-to-end framework. Ligong Han, Robert F. Murphy, Deva Ramanan |
WACV | 2 |
| 2017 | Image-based spatiotemporal causality inference for protein signaling networksabstractMOTIVATION: Efforts to model how signaling and regulatory networks work in cells have largely either not considered spatial organization or have used compartmental models with minimal spatial resolution. Fluorescence microscopy provides the ability to monitor the spatiotemporal distribution of many molecules during signaling events, but as of yet no methods have been described for large scale image analysis to learn a complex protein regulatory network. Here we present and evaluate methods for identifying how changes in concentration in one cell region influence concentration of other proteins in other regions. RESULTS: Using 3D confocal microscope movies of GFP-tagged T cells undergoing costimulation, we learned models containing putative causal relationships among 12 proteins involved in T cell signaling. The models included both relationships consistent with current knowledge and novel predictions deserving further exploration. Further, when these models were applied to the initial frames of movies of T cells that had been only partially stimulated, they predicted the localization of proteins at later times with statistically significant accuracy. The methods, consisting of spatiotemporal alignment, automated region identification, and causal inference, are anticipated to be applicable to a number of biological systems. AVAILABILITY AND IMPLEMENTATION: The source code and data are available as a Reproducible Research Archive at http://murphylab.cbd.cmu.edu/software/2017_TcellCausalModels/. CONTACT: [email protected]. Xiongtao Ruan, Christoph Wülfing, Robert F. Murphy |
Bioinform. | 3 |
| 2016 | Unbiased Rare Event Sampling in Spatial Stochastic Systems Biology Models Using a Weighted Ensemble of TrajectoriesabstractThe long-term goal of connecting scales in biological simulation can be facilitated by scale-agnostic methods. We demonstrate that the weighted ensemble (WE) strategy, initially developed for molecular simulations, applies effectively to spatially resolved cell-scale simulations. The WE approach runs an ensemble of parallel trajectories with assigned weights and uses a statistical resampling strategy of replicating and pruning trajectories to focus computational effort on difficult-to-sample regions. The method can also generate unbiased estimates of non-equilibrium and equilibrium observables, sometimes with significantly less aggregate computing time than would be possible using standard parallelization. Here, we use WE to orchestrate particle-based kinetic Monte Carlo simulations, which include spatial geometry (e.g., of organelles, plasma membrane) and biochemical interactions among mobile molecular species. We study a series of models exhibiting spatial, temporal and biochemical complexity and show that although WE has important limitations, it can achieve performance significantly exceeding standard parallel simulation--by orders of magnitude for some observables. Rory M. Donovan, José Juan Tapia, Devin P. Sullivan, James R. Faeder, Robert F. Murphy, Markus Dittrich, Daniel M. Zuckerman |
PLoS Comput. Biol. | 5 |
| 2015 | Design Automation for Biological Models: A Pipeline that Incorporates Spatial and Molecular ComplexityabstractUnderstanding the dynamics of biochemical networks is a major goal of systems biology. Due to the heterogeneity of cells and the low copy numbers of key molecules, spatially resolved approaches are required to fully understand and model these systems. Until recently, most spatial modeling was performed using geometries obtained either through manual segmentation or manual fabrication both of which are time-consuming and tedious. Similarly, the system of reactions associated with the model had to be manually defined, a process that is both tedious and error-prone for large networks. As a result, spatially resolved simulations have typically only been performed in a limited number of geometries, which are often highly simplified, and with small reaction networks. Devin P. Sullivan, Rohan Arepally, Robert F. Murphy, José Juan Tapia, James R. Faeder, Markus Dittrich, Jacob Czech |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | Deciding When to Stop: Efficient Experimentation to Learn to Predict Drug-Target Interactions (Extended Abstract)
Maja Temerinac-Ott, Armaghan W. Naik, Robert F. Murphy |
RECOMB | 3 |
| 2015 | Deciding when to stop: efficient experimentation to learn to predict drug-target interactionsabstractBACKGROUND: Active learning is a powerful tool for guiding an experimentation process. Instead of doing all possible experiments in a given domain, active learning can be used to pick the experiments that will add the most knowledge to the current model. Especially, for drug discovery and development, active learning has been shown to reduce the number of experiments needed to obtain high-confidence predictions. However, in practice, it is crucial to have a method to evaluate the quality of the current predictions and decide when to stop the experimentation process. Only by applying reliable stopping criteria to active learning can time and costs in the experimental process actually be saved. RESULTS: We compute active learning traces on simulated drug-target matrices in order to determine a regression model for the accuracy of the active learner. By analyzing the performance of the regression model on simulated data, we design stopping criteria for previously unseen experimental matrices. We demonstrate on four previously characterized drug effect data sets that applying the stopping criteria can result in upto 40 % savings of the total experiments for highly accurate predictions. CONCLUSIONS: We show that active learning accuracy can be predicted using simulated data and results in substantial savings in the number of experiments required to make accurate drug-target predictions. Maja Temerinac-Ott, Armaghan W. Naik, Robert F. Murphy |
BMC Bioinform. | 3 |
| 2015 | Automated Learning of Subcellular Variation among Punctate Protein Patterns and a Generative Model of Their Relation to MicrotubulesabstractCharacterizing the spatial distribution of proteins directly from microscopy images is a difficult problem with numerous applications in cell biology (e.g. identifying motor-related proteins) and clinical research (e.g. identification of cancer biomarkers). Here we describe the design of a system that provides automated analysis of punctate protein patterns in microscope images, including quantification of their relationships to microtubules. We constructed the system using confocal immunofluorescence microscopy images from the Human Protein Atlas project for 11 punctate proteins in three cultured cell lines. These proteins have previously been characterized as being primarily located in punctate structures, but their images had all been annotated by visual examination as being simply "vesicular". We were able to show that these patterns could be distinguished from each other with high accuracy, and we were able to assign to one of these subclasses hundreds of proteins whose subcellular localization had not previously been well defined. In addition to providing these novel annotations, we built a generative approach to modeling of punctate distributions that captures the essential characteristics of the distinct patterns. Such models are expected to be valuable for representing and summarizing each pattern and for constructing systems biology simulations of cell behaviors. Gregory R. Johnson, Jieyue Li, Aabid Shariff, Gustavo K. Rohde, Robert F. Murphy |
PLoS Comput. Biol. | 5 |
| 2014 | Implementation of an Intelligent Tutoring System for Online Homework Support in an Efficacy Trial
Mingyu Feng, Jeremy Roschelle, Neil T. Heffernan, Janet Fairman, Robert F. Murphy |
Intelligent Tutoring Systems | 5 |
| 2014 | A new era in bioimage informaticsabstractBioimage informatics arose from efforts to automate pathology and cytology tasks (Eaves, 1967). With few exceptions, much of the software developed during these early days, whether in academic or commercial institutions, was proprietary. The primary paradigm was production of hand-tuned engineered systems that could reproduce human performance, and visualization was emphasized for interpreting results or providing assistance to clinicians (Bartels and Wied, 1977;,Kaman, et al., 1984;,van Driel-Kulker and Ploem, 1982). The computational resources available at the time were frequently limiting. Essentially, no successful commercial systems came from these efforts for many years, until the US Food and Drug Administration’s approval of automated Pap smear analysis in the mid 1990s (Patten et al., 1996). Beginning around this time, a new era began with automation of tasks associated with more basic measurement of cellular and subcellular phenomena, such as detection of drug effects (Giuliano et al., 1997) and recognition of subcellular patterns (Boland et al., 1998). The field gained significant exposure, and a new name, with the founding of two National Science Foundation-sponsored Centers for Bioimage Informatics at the University of California, Santa Barbara and Carnegie Mellon University in 2003. This led to the first Bioimage Informatics conference held in Santa Barbara in 2006. The primary paradigm in this era became supervised and semisupervised machine learning, and automated systems were reported that were able to outperform humans at recognizing cell positions in tissues (Nattkemper et al., 2003) and recognizing subcellular patterns (Murphy et al., 2003). Cutting-edge systems during this period increasingly avoided visualization, hand-tuning or user intervention (with the exception of marking things like cell or nuclear boundaries for training purposes), with a primary goal to produce a simple actionable answer. Examples include identifying which compounds or inhibitory RNAs produce particular effects, and answers typically require further follow-up work, usually not involving imaging, to produce a final biological result. Another shift has been the increased prevalence of open-source software (Eliceiri et al., 2012), although the frequency with which investigator-provided software has been usable by others has varied extensively. Papers have often emphasized the further refinement of methods for benchmark datasets, such as the original 2D HeLa dataset, with the risk that systems become overly tuned to the benchmark. Bioimage informatics is entering its third era, in which the goal is the fully automated production of models of biological systems. This includes a shift toward unsupervised machine learning, and especially toward structure learning methods. Examples of initial steps include building models of signaling networks from multiprobe images (Welch, et al., 2011), identification of regulatory modules from gene expression images of embryos (Puniyani and Xing, 2013), initial attempts at building image-derived generative models of cells (Buck et al., 2012) and identification of proteins involved in cell shape regulation (Sailem et al., 2014). Provision of data and software in reproducible research archives, begun already in the early 2000s, is growing. Other significant trends are the use of image datasets from different sources, and integration with non-imaging data such as from genomics, transcriptomics, proteomics and metabolomics studies. In this third era, image analysis papers published in Bioinformatics should reflect these trends and follow the journal’s focus on analysis of systems at a molecular level. Incremental refinements to algorithms, methods requiring hand-labeling and -tuning and studies in which the primary outputs are visualizations are all expected to give way to those producing verifiable models that can be combined to produce multiscale, multicomponent representations of the molecular basis and dynamics of cell and tissue organization and behavior. Robert F. Murphy |
Bioinform. | 1 |
| 2014 | Efficient discovery of responses of proteins to compounds using active learningabstractBACKGROUND: Drug discovery and development has been aided by high throughput screening methods that detect compound effects on a single target. However, when using focused initial screening, undesirable secondary effects are often detected late in the development process after significant investment has been made. An alternative approach would be to screen against undesired effects early in the process, but the number of possible secondary targets makes this prohibitively expensive. RESULTS: This paper describes methods for making this global approach practical by constructing predictive models for many target responses to many compounds and using them to guide experimentation. We demonstrate for the first time that by jointly modeling targets and compounds using descriptive features and using active machine learning methods, accurate models can be built by doing only a small fraction of possible experiments. The methods were evaluated by computational experiments using a dataset of 177 assays and 20,000 compounds constructed from the PubChem database. CONCLUSIONS: An average of nearly 60% of all hits in the dataset were found after exploring only 3% of the experimental space which suggests that active learning can be used to enable more complete characterization of compound effects than otherwise affordable. The methods described are also likely to find widespread application outside drug discovery, such as for characterizing the effects of a large number of compounds or inhibitory RNAs on a large number of cell or tissue phenotypes. Joshua D. Kangas, Armaghan W. Naik, Robert F. Murphy |
BMC Bioinform. | 3 |
| 2013 | Determining the subcellular location of new proteins from microscope images using local featuresabstractMOTIVATION: Evaluation of previous systems for automated determination of subcellular location from microscope images has been done using datasets in which each location class consisted of multiple images of the same representative protein. Here, we frame a more challenging and useful problem where previously unseen proteins are to be classified. RESULTS: Using CD-tagging, we generated two new image datasets for evaluation of this problem, which contain several different proteins for each location class. Evaluation of previous methods on these new datasets showed that it is much harder to train a classifier that generalizes across different proteins than one that simply recognizes a protein it was trained on. We therefore developed and evaluated additional approaches, incorporating novel modifications of local features techniques. These extended the notion of local features to exploit both the protein image and any reference markers that were imaged in parallel. With these, we obtained a large accuracy improvement in our new datasets over existing methods. Additionally, these features help achieve classification improvements for other previously studied datasets. AVAILABILITY: The datasets are available for download at http://murphylab.web.cmu.edu/data/. The software was written in Python and C++ and is available under an open-source license at http://murphylab.web.cmu.edu/software/. The code is split into a library, which can be easily reused for other data and a small driver script for reproducing all results presented here. A step-by-step tutorial on applying the methods to new datasets is also available at that address. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Luís Pedro Coelho, Joshua D. Kangas, Armaghan W. Naik, Elvira Osuna-Highley, Estelle Glory-Afshar, Margaret Fuhrman, Ramanuja Simha, Peter B. Berget, Jonathan W. Jarvik, Robert F. Murphy |
Bioinform. | 10 |
| 2012 | (3) The CellOrganizer project: An open source system to learn image-derived models of subcellular organization over time and spaceabstractThe CellOrganizer project (http://cellorganizer.org) provides open source tools for learning generative models of cell organization directly from images and for synthesizing cell images (or other representations) from one or more of those models. Model learning captures variation among cells in a collection of images. Images used for model learning and instances synthesized from models can be two- or three-dimensional static images or movies. Current components of CellOrganizer can learn models of cell shape, nuclear shape, chromatin texture, vesicular organe lie number, size, shape and position, and microtubule distribution. These models can be conditional upon each other: for example, for a given synthesized cell instance, organelle position will be dependent upon the cell and nuclear shape of that instance. The models can be parametric, in which a choice is made about an explicit form to represent a particular structure, or non-parametric, in which distributions are learned empirically. One of the main uses of the system is in support of cell simulations: models learned from separate experiments can be combined into one or more synthetic cell instances that are output in a form compatible with cell simulation engines such as MCell, Virtual Cell and Smoldyn. Another important application of the system is in comparison of target patterns and perturbagen effects in high content screening and analysis. This is currently done using numerical features, but these are difficult to compare across different microscope systems or cell types since features can be affected by changes in more than one aspect of cell organization. More robust comparisons can be made using generative model parameters, since these can distinguish effects on cell size or shape from effects on organelle pattern. Ultimately, it is anticipated that collaborative efforts by many groups will enable creation of image-derived generative models that permit accurate modeling of cell behaviors, and that can be used to drive experimentation to improve them through active learning. [replace "perturbation" for the word "perturbagen"] Robert F. Murphy |
BIBM | 1 |
| 2012 | Protein subcellular location pattern classification in cellular images using latent discriminative modelsabstractMOTIVATION: Knowledge of the subcellular location of a protein is crucial for understanding its functions. The subcellular pattern of a protein is typically represented as the set of cellular components in which it is located, and an important task is to determine this set from microscope images. In this article, we address this classification problem using confocal immunofluorescence images from the Human Protein Atlas (HPA) project. The HPA contains images of cells stained for many proteins; each is also stained for three reference components, but there are many other components that are invisible. Given one such cell, the task is to classify the pattern type of the stained protein. We first randomly select local image regions within the cells, and then extract various carefully designed features from these regions. This region-based approach enables us to explicitly study the relationship between proteins and different cell components, as well as the interactions between these components. To achieve these two goals, we propose two discriminative models that extend logistic regression with structured latent variables. The first model allows the same protein pattern class to be expressed differently according to the underlying components in different regions. The second model further captures the spatial dependencies between the components within the same cell so that we can better infer these components. To learn these models, we propose a fast approximate algorithm for inference, and then use gradient-based methods to maximize the data likelihood. RESULTS: In the experiments, we show that the proposed models help improve the classification accuracies on synthetic data and real cellular images. The best overall accuracy we report in this article for classifying 942 proteins into 13 classes of patterns is about 84.6%, which to our knowledge is the best so far. In addition, the dependencies learned are consistent with prior knowledge of cell organization. AVAILABILITY: http://murphylab.web.cmu.edu/software/. Jieyue Li, Liang Xiong, Jeff G. Schneider, Robert F. Murphy |
Bioinform. | 4 |
| 2011 | Learning Cellular Sorting Pathways Using Protein Interactions and Sequence Motifs
Tien-ho Lin, Ziv Bar-Joseph, Robert F. Murphy |
RECOMB | 3 |
| 2011 | Model building and intelligent acquisition with application to protein subcellular location classificationabstractMOTIVATION: We present a framework and algorithms to intelligently acquire movies of protein subcellular location patterns by learning their models as they are being acquired, and simultaneously determining how many cells to acquire as well as how many frames to acquire per cell. This is motivated by the desire to minimize acquisition time and photobleaching, given the need to build such models for all proteins, in all cell types, under all conditions. Our key innovation is to build models during acquisition rather than as a post-processing step, thus allowing us to intelligently and automatically adapt the acquisition process given the model acquired. RESULTS: We validate our framework on protein subcellular location classification, and show that the combination of model building and intelligent acquisition results in time and storage savings without loss of classification accuracy, or alternatively, higher classification accuracy for the same total acquisition time. AVAILABILITY AND IMPLEMENTATION: The data and software used for this study will be made available upon publication at http://murphylab.web.cmu.edu/software and http://www.andrew.cmu.edu/user/jelenak/Software. CONTACT: [email protected]. Charles Jackson, Estelle Glory-Afshar, Robert F. Murphy, Jelena Kovacevic |
Bioinform. | 3 |
| 2011 | Discriminative Motif Finding for Predicting Protein Subcellular LocalizationabstractMany methods have been described to predict the subcellular location of proteins from sequence information. However, most of these methods either rely on global sequence properties or use a set of known protein targeting motifs to predict protein localization. Here, we develop and test a novel method that identifies potential targeting motifs using a discriminative approach based on hidden Markov models (discriminative HMMs). These models search for motifs that are present in a compartment but absent in other, nearby, compartments by utilizing an hierarchical structure that mimics the protein sorting mechanism. We show that both discriminative motif finding and the hierarchical structure improve localization prediction on a benchmark data set of yeast proteins. The motifs identified can be mapped to known targeting motifs and they are more conserved than the average protein sequence. Using our motif-based predictions, we can identify potential annotation errors in public databases for the location of some of the proteins. A software implementation and the data set described in this paper are available from http://murphylab.web.cmu.edu/software/2009_TCBB_motif/. Tien-ho Lin, Robert F. Murphy, Ziv Bar-Joseph |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Quantifying the distribution of probes between subcellular locations using unsupervised pattern unmixingabstractMOTIVATION: Proteins exhibit complex subcellular distributions, which may include localizing in more than one organelle and varying in location depending on the cell physiology. Estimating the amount of protein distributed in each subcellular location is essential for quantitative understanding and modeling of protein dynamics and how they affect cell behaviors. We have previously described automated methods using fluorescent microscope images to determine the fractions of protein fluorescence in various subcellular locations when the basic locations in which a protein can be present are known. As this set of basic locations may be unknown (especially for studies on a proteome-wide scale), we here describe unsupervised methods to identify the fundamental patterns from images of mixed patterns and estimate the fractional composition of them. METHODS: We developed two approaches to the problem, both based on identifying types of objects present in images and representing patterns by frequencies of those object types. One is a basis pursuit method (which is based on a linear mixture model), and the other is based on latent Dirichlet allocation (LDA). For testing both approaches, we used images previously acquired for testing supervised unmixing methods. These images were of cells labeled with various combinations of two organelle-specific probes that had the same fluorescent properties to simulate mixed patterns of subcellular location. RESULTS: We achieved 0.80 and 0.91 correlation between estimated and underlying fractions of the two probes (fundamental patterns) with basis pursuit and LDA approaches, respectively, indicating that our methods can unmix the complex subcellular distribution with reasonably high accuracy. AVAILABILITY: http://murphylab.web.cmu.edu/software. Luís Pedro Coelho, Tao Peng 0004, Robert F. Murphy |
Bioinform. | 3 |
| 2010 | Automated analysis of protein subcellular location in time series imagesabstractMOTIVATION: Image analysis, machine learning and statistical modeling have become well established for the automatic recognition and comparison of the subcellular locations of proteins in microscope images. By using a comprehensive set of features describing static images, major subcellular patterns can be distinguished with near perfect accuracy. We now extend this work to time series images, which contain both spatial and temporal information. The goal is to use temporal features to improve recognition of protein patterns that are not fully distinguishable by their static features alone. RESULTS: We have adopted and designed five sets of features for capturing temporal behavior in 2D time series images, based on object tracking, temporal texture, normal flow, Fourier transforms and autoregression. Classification accuracy on an image collection for 12 fluorescently tagged proteins was increased when temporal features were used in addition to static features. Temporal texture, normal flow and Fourier transform features were most effective at increasing classification accuracy. We therefore extended these three feature sets to 3D time series images, but observed no significant improvement over results for 2D images. The methods for 2D and 3D temporal pattern analysis do not require segmentation of images into single cell regions, and are suitable for automated high-throughput microscopy applications. AVAILABILITY: Images, source code and results will be available upon publication at http://murphylab.web.cmu.edu/software CONTACT: [email protected]. Yanhua Hu, Elvira Osuna-Highley, Juchang Hua, Theodore Scott Nowicki, Robert Stolz, Camille McKayle, Robert F. Murphy |
Bioinform. | 7 |
| 2010 | Structured literature image finder: Parsing text and figures in biomedical literature
Amr Ahmed 0001, Andrew Arnold, Luís Pedro Coelho, Joshua D. Kangas, Abdul-Saboor Sheikh, Eric P. Xing, William W. Cohen, Robert F. Murphy |
J. Web Semant. | 8 |
| 2009 | Workshop summary: Automated interpretation and modelling of cell imagesabstractNo abstract available. Robert F. Murphy, Chun-Nan Hsu, Loris Nanni |
ICML | 1 |
| 2009 | Structured correspondence topic models for mining captioned figures in biological literatureabstractA major source of information (often the most crucial and informative part) in scholarly articles from scientific journals, proceedings and books are the figures that directly provide images and other graphical illustrations of key experimental results and other scientific contents. In biological articles, a typical figure often comprises multiple panels, accompanied by either scoped or global captioned text. Moreover, the text in the caption contains important semantic entities such as protein names, gene ontology, tissues labels, etc., relevant to the images in the figure. Due to the avalanche of biological literature in recent years, and increasing popularity of various bio-imaging techniques, automatic retrieval and summarization of biological information from literature figures has emerged as a major unsolved challenge in computational knowledge extraction and management in the life science. We present a new structured probabilistic topic model built on a realistic figure generation scheme to model the structurally annotated biological figures, and we derive an efficient inference algorithm based on collapsed Gibbs sampling for information retrieval and visualization. The resulting program constitutes one of the key IR engines in our SLIF system that has recently entered the final round (4 out 70 competing systems) of the Elsevier Grand Challenge on Knowledge Enhancement in the Life Science. Here we present various evaluations on a number of data mining tasks to illustrate our method. Amr Ahmed 0001, Eric P. Xing, William W. Cohen, Robert F. Murphy |
KDD | 4 |
| 2009 | Intelligent Acquisition and Learning of Fluorescence Microscope Data ModelsabstractWe propose a mathematical framework and algorithms both to build accurate models of fluorescence microscope time series, as well as to design intelligent acquisition systems based on these models. Model building allows the information contained in the 2-D and 3-D time series to be presented in a more useful and concise form than the raw image data. This is particularly relevant as the trend in biology tends more and more towards high-throughput applications, and the resulting increase in the amount of acquired image data makes visual inspection impractical. The intelligent acquisition system uses an active learning approach to choose the acquisition regions that let us build our model most efficiently, resulting in a shorter acquisition time, as well as a reduction of the amount of photobleaching and phototoxicity incurred during acquisition. We validate our methodology by modeling object motion within a cell. For intelligent acquisition, we propose a set of algorithms to evaluate the information contained in a given acquisition region, as well as the costs associated with acquiring this region in terms of the resulting photobleaching and phototoxicity and the amount of time taken for acquisition. We use these algorithms to determine an acquisition strategy: where and when to acquire, as well as when to stop acquiring. Results, both on synthetic as well as real data, demonstrate accurate model building and large efficiency gains during acquisition. Charles Jackson, Robert F. Murphy, Jelena Kovacevic |
IEEE Trans. Image Process. | 2 |
| 2008 | Improved recognition of figures containing fluorescence microscope images in online journal articles using graphical modelsabstractMOTIVATION: There is extensive interest in automating the collection, organization and analysis of biological data. Data in the form of images in online literature present special challenges for such efforts. The first steps in understanding the contents of a figure are decomposing it into panels and determining the type of each panel. In biological literature, panel types include many kinds of images collected by different techniques, such as photographs of gels or images from microscopes. We have previously described the SLIF system (http://slif.cbi.cmu.edu) that identifies panels containing fluorescence microscope images among figures in online journal articles as a prelude to further analysis of the subcellular patterns in such images. This system contains a pretrained classifier that uses image features to assign a type (class) to each separate panel. However, the types of panels in a figure are often correlated, so that we can consider the class of a panel to be dependent not only on its own features but also on the types of the other panels in a figure. RESULTS: In this article, we introduce the use of a type of probabilistic graphical model, a factor graph, to represent the structured information about the images in a figure, and permit more robust and accurate inference about their types. We obtain significant improvement over results for considering panels separately. AVAILABILITY: The code and data used for the experiments described here are available from http://murphylab.web.cmu.edu/software. Yuntao Qian, Robert F. Murphy |
Bioinform. | 2 |
| 2008 | Graphical Models for Structured Classification, with an Application to Interpreting Images of Protein Subcellular Location Patterns
Shann-Ching Chen, Geoffrey J. Gordon, Robert F. Murphy |
J. Mach. Learn. Res. | 3 |
| 2007 | Efficient Acquisition and Learning of Fluorescence Microscope Data ModelsabstractWe present a method for efficient acquisition of fluorescence microscope datasets, to allow for higher spatial and temporal resolution, and with less damage from photobleaching. Our proposal is to restrict acquisition to regions where we expect to find an object. Given that the objects are continuously moving, we must have an accurate model to describe objects' motion to predict their future locations. We outline a system for learning and applying this motion model, provide details from some simple simulations, and summarize results from more complex applications. Charles Jackson, Robert F. Murphy, Jelena Kovacevic |
ICIP (6) | 2 |
| 2007 | A multiresolution approach to automated classification of protein subcellular location imagesabstractBACKGROUND: Fluorescence microscopy is widely used to determine the subcellular location of proteins. Efforts to determine location on a proteome-wide basis create a need for automated methods to analyze the resulting images. Over the past ten years, the feasibility of using machine learning methods to recognize all major subcellular location patterns has been convincingly demonstrated, using diverse feature sets and classifiers. On a well-studied data set of 2D HeLa single-cell images, the best performance to date, 91.5%, was obtained by including a set of multiresolution features. This demonstrates the value of multiresolution approaches to this important problem. RESULTS: We report here a novel approach for the classification of subcellular location patterns by classifying in multiresolution subspaces. Our system is able to work with any feature set and any classifier. It consists of multiresolution (MR) decomposition, followed by feature computation and classification in each MR subspace, yielding local decisions that are then combined into a global decision. With 26 texture features alone and a neural network classifier, we obtained an increase in accuracy on the 2D HeLa data set to 95.3%. CONCLUSION: We demonstrate that the space-frequency localized information in the multiresolution subspaces adds significantly to the discriminative power of the system. Moreover, we show that a vastly reduced set of features is sufficient, consisting of our novel modified Haralick texture features. Our proposed system is general, allowing for any combinations of sets of features and any combination of classifiers. Amina Chebira, Yann Barbotin, Charles Jackson, Thomas E. Merryman, Gowri Srinivasa, Robert F. Murphy, Jelena Kovacevic |
BMC Bioinform. | 6 |
| 2006 | A Novel Graphical Model Approach to Segmenting Cell ImagesabstractSuccessful biological image analysis usually requires satisfactory segmentations to identify regions of interest as an intermediate step. Here we present a novel graphical model approach for segmentation of multi-cell yeast images acquired by fluorescence microscopy. Yeast cells are often clustered together, so they are hard to segment by conventional techniques. Our approach assumes that two parallel images are available for each field: an image containing information about the nuclear positions (such as an image of a DNA probe) and an image containing information about the cell boundaries (such as a differential interference contrast, or DIC, image). The nuclear information provides an initial assignment of whether each pixel belongs to the background or one of the cells. The boundary information is used to estimate the probability that any two pixels in the graph are separated by a cell boundary. From these two kinds of information, we construct a graph that links nearby pairs of pixels, and seek to infer a good segmentation from this graph. We pose this problem as inference in a Bayes network, and use a fast approximation approach to iteratively improve the estimated probability of each class for each pixel. The resulting algorithm can efficiently generate segmentation masks which are highly consistent with hand-labeled data, and results suggest that the work will be of particular use for large scale determination of protein location patterns by automated microscopy Shann-Ching Chen, Geoffrey J. Gordon, Robert F. Murphy |
CIBCB | 4 |
| 2006 | A graphical model approach to automated classification of protein subcellular location patterns in multi-cell imagesabstractBACKGROUND: Knowledge of the subcellular location of a protein is critical to understanding how that protein works in a cell. This location is frequently determined by the interpretation of fluorescence microscope images. In recent years, automated systems have been developed for consistent and objective interpretation of such images so that the protein pattern in a single cell can be assigned to a known location category. While these systems perform with nearly perfect accuracy for single cell images of all major subcellular structures, their ability to distinguish subpatterns of an organelle (such as two Golgi proteins) is not perfect. Our goal in the work described here was to improve the ability of an automated system to decide which of two similar patterns is present in a field of cells by considering more than one cell at a time. Since cells displaying the same location pattern are often clustered together, considering multiple cells may be expected to improve discrimination between similar patterns. RESULTS: We describe how to take advantage of information on experimental conditions to construct a graphical representation for multiple cells in a field. Assuming that a field is composed of a small number of classes, the classification accuracy can be improved by allowing the computed probability of each pattern for each cell to be influenced by the probabilities of its neighboring cells in the model. We describe a novel way to allow this influence to occur, in which we adjust the prior probabilities of each class to reflect the patterns that are present. When this graphical model approach is used on synthetic multi-cell images in which the true class of each cell is known, we observe that the ability to distinguish similar classes is improved without suffering any degradation in ability to distinguish dissimilar classes. The computational complexity of the method is sufficiently low that improved assignments of classes can be obtained for fields of twelve cells in under 0.04 second on a 1600 megahertz processor. CONCLUSION: We demonstrate that graphical models can be used to improve the accuracy of classification of subcellular patterns in multi-cell fluorescence microscope images. We also describe a novel algorithm for inferring classes from a graphical model. The performance and speed suggest that the method will be particularly valuable for analysis of images from high-throughput microscopy. We also anticipate that it will be useful for analyzing the mixtures of cell types typically present in images of tissues. Lastly, we anticipate that the method can be generalized to other problems. Shann-Ching Chen, Robert F. Murphy |
BMC Bioinform. | 2 |
| 2005 | Adaptive Multirate Data Acquisition of 3D Cell ImagesabstractWe present an algorithm for efficient acquisition of fluorescence microscopy data sets, a problem not addressed until now in the literature. We do this as part of a larger system for protein classification based on their subcellular location patterns, and thus strive to maintain the achieved level of classification accuracy as much as possible. This problem is similar to image compression but unique due to additional restrictions, namely causality; we have access only to the information that has been scanned up to that point. While we do want to acquire fewer samples with as low distortion as possible to achieve compression, our goal is to do so while affecting the overall classification accuracy as little as possible. We achieve this by using an adaptive multiresolution scanning scheme which samples the regions of the image area that hold the most pertinent information. Our results show that we can achieve significant compression which we can then use to increase either time or space resolution of our data set, all while minimally affecting the classification accuracy of the entire system. Thomas E. Merryman, Jelena Kovacevic, Elvira Garcia Osuna, Robert F. Murphy |
ICASSP (2) | 4 |
| 2005 | Research issues in protein location image databasesabstractWhich proteins have similar locations within cells? How many distinct location patters do cells display? How do we answer these questions quickly, from a large collection of microscope images such as in on-line journals? Robert F. Murphy, Christos Faloutsos |
SIGMOD Conference | 1 |
| 2005 | Object Type Recognition for Automated Analysis of Protein Subcellular LocationabstractThe new field of location proteomics seeks to provide a comprehensive, objective characterization of the subcellular locations of all proteins expressed in a given cell type. Previous work has demonstrated that automated classifiers can recognize the patterns of all major subcellular organelles and structures in fluorescence microscope images with high accuracy. However, since some proteins may be present in more than one organelle, this paper addresses a more difficult task: recognizing a pattern that is a mixture of two or more fundamental patterns. The approach utilizes an object-based image model, in which each image of a location pattern is represented by a set of objects of distinct, learned types. Using a two-stage approach in which object types are learned and then cell-level features are calculated based on the object types, the basic location patterns were well recognized. Given the object types, a multinomial mixture model was built to recognize mixture patterns. Under appropriate conditions, synthetic mixture patterns can be decomposed with over 80% accuracy, which, for the first time, shows that the problem of computationally decomposing subcellular patterns into fundamental organelle patterns can be solved. Meel Velliste, Michael V. Boland, Robert F. Murphy |
IEEE Trans. Image Process. | 4 |
| 2004 | Boosting accuracy of automated classification of fluorescence microscope images for location proteomicsabstractBACKGROUND: Detailed knowledge of the subcellular location of each expressed protein is critical to a full understanding of its function. Fluorescence microscopy, in combination with methods for fluorescent tagging, is the most suitable current method for proteome-wide determination of subcellular location. Previous work has shown that neural network classifiers can distinguish all major protein subcellular location patterns in both 2D and 3D fluorescence microscope images. Building on these results, we evaluate here new classifiers and features to improve the recognition of protein subcellular location patterns in both 2D and 3D fluorescence microscope images. RESULTS: We report here a thorough comparison of the performance on this problem of eight different state-of-the-art classification methods, including neural networks, support vector machines with linear, polynomial, radial basis, and exponential radial basis kernel functions, and ensemble methods such as AdaBoost, Bagging, and Mixtures-of-Experts. Ten-fold cross validation was used to evaluate each classifier with various parameters on different Subcellular Location Feature sets representing both 2D and 3D fluorescence microscope images, including new feature sets incorporating features derived from Gabor and Daubechies wavelet transforms. After optimal parameters were chosen for each of the eight classifiers, optimal majority-voting ensemble classifiers were formed for each feature set. Comparison of results for each image for all eight classifiers permits estimation of the lower bound classification error rate for each subcellular pattern, which we interpret to reflect the fraction of cells whose patterns are distorted by mitosis, cell death or acquisition errors. Overall, we obtained statistically significant improvements in classification accuracy over the best previously published results, with the overall error rate being reduced by one-third to one-half and with the average accuracy for single 2D images being higher than 90% for the first time. In particular, the classification accuracy for the easily confused endomembrane compartments (endoplasmic reticulum, Golgi, endosomes, lysosomes) was improved by 5-15%. We achieved further improvements when classification was conducted on image sets rather than on individual cell images. CONCLUSIONS: The availability of accurate, fast, automated classification systems for protein location patterns in conjunction with high throughput fluorescence microscope imaging techniques enables a new subfield of proteomics, location proteomics. The accuracy and sensitivity of this approach represents an important alternative to low-resolution assignments by curation or sequence-based prediction. Robert F. Murphy |
BMC Bioinform. | 2 |
| 2003 | Understanding captions in biomedical publicationsabstractFrom the standpoint of the automated extraction of scientific knowledge, an important but little-studied part of scientific publications are the figures and accompanying captions. Captions are dense in information, but also contain many extra-grammatical constructs, making them awkward to process with standard information extraction methods. We propose a scheme for "understanding" captions in biomedical publications by extracting and classifying "image pointers" (references to the accompanying image). We evaluate a number of automated methods for this task, including hand-coded methods, methods based on existing learning techniques, and methods based on novel learning techniques. The best of these methods leads to a usefully accurate tool for caption-understanding, with both recall and precision in excess of 94% on the most important single class in a combined extraction/classification task. William W. Cohen, Richard C. Wang, Robert F. Murphy |
KDD | 3 |
| 2001 | Searching Online Journals for Fluorescence Microscope Images Depicting Protein Subcellular Location PatternsabstractThere is extensive interest in automating the collection, organization and analysis of biological data. Data in the form of images present special challenges for such efforts. Since fluorescence microscope images are a primary source of information about the location of proteins within cells, we have set as a long-term goal the building of a knowledge base system that can interpret such images in online journals. To this end, we first developed a robot that searches online journals and finds fluorescence microscope images of individual cells. We then characterized the applicability of pattern classification methods we have previously used on images obtained under controlled conditions to images from different sources and to images subjected to manipulations commonly performed during publication. The results indicate the feasibility of developing search engines to find fluorescence microscope images depicting particular subcellular patterns. Robert F. Murphy, Meel Velliste, Gregory Porreca |
BIBE | 1 |
| 2001 | A neural network classifier capable of recognizing the patterns of all major subcellular structures in fluorescence microscope images of HeLa cellsabstractMOTIVATION: Assessment of protein subcellular location is crucial to proteomics efforts since localization information provides a context for a protein's sequence, structure, and function. The work described below is the first to address the subcellular localization of proteins in a quantitative, comprehensive manner. RESULTS: Images for ten different subcellular patterns (including all major organelles) were collected using fluorescence microscopy. The patterns were described using a variety of numeric features, including Zernike moments, Haralick texture features, and a set of new features developed specifically for this purpose. To test the usefulness of these features, they were used to train a neural network classifier. The classifier was able to correctly recognize an average of 83% of previously unseen cells showing one of the ten patterns. The same classifier was then used to recognize previously unseen sets of homogeneously prepared cells with 98% accuracy. AVAILABILITY: Algorithms were implemented using the commercial products Matlab, S-Plus, and SAS, as well as some functions written in C. The scripts and source code generated for this work are available at http://murphylab.web.cmu.edu/software. CONTACT: [email protected] Michael V. Boland, Robert F. Murphy |
Bioinform. | 2 |
| 2000 | Towards a Systematics for Protein Subcellular Location: Quantitative Description of Protein Localization Patterns and Automated Analysis of Fluorescence Microscope Images
Robert F. Murphy, Michael V. Boland, Meel Velliste |
ISMB | 1 |