Richard H. Scheuermann

dblp:71/4512 · DBLP profile ↗
← Back
23ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0003-1355-892XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2023 A cell-level discriminative neural network model for diagnosis of blood cancers
abstract
MOTIVATION: Precise identification of cancer cells in patient samples is essential for accurate diagnosis and clinical monitoring but has been a significant challenge in machine learning approaches for cancer precision medicine. In most scenarios, training data are only available with disease annotation at the subject or sample level. Traditional approaches separate the classification process into multiple steps that are optimized independently. Recent methods either focus on predicting sample-level diagnosis without identifying individual pathologic cells or are less effective for identifying heterogeneous cancer cell phenotypes. RESULTS: We developed a generalized end-to-end differentiable model, the Cell Scoring Neural Network (CSNN), which takes sample-level training data and predicts the diagnosis of the testing samples and the identity of the diagnostic cells in the sample, simultaneously. The cell-level density differences between samples are linked to the sample diagnosis, which allows the probabilities of individual cells being diagnostic to be calculated using backpropagation. We applied CSNN to two independent clinical flow cytometry datasets for leukemia diagnosis. In both qualitative and quantitative assessments, CSNN outperformed preexisting neural network modeling approaches for both cancer diagnosis and cell-level classification. Post hoc decision trees and 2D dot plots were generated for interpretation of the identified cancer cells, showing that the identified cell phenotypes match the cancer endotypes observed clinically in patient cohorts. Independent data clustering analysis confirmed the identified cancer cell populations. AVAILABILITY AND IMPLEMENTATION: The source code of CSNN and datasets used in the experiments are publicly available on GitHub (http://github.com/erobl/csnn). Raw FCS files can be downloaded from FlowRepository (ID: FR-FCM-Z6YK).
Edgar E. Robles, Padhraic Smyth, Richard H. Scheuermann, Jack D. Bui, Huan-You Wang, Jean Oak
Bioinform.4
2022 FastMix: a versatile data integration pipeline for cell type-specific biomarker inference
abstract
MOTIVATION: Flow cytometry (FCM) and transcription profiling are the two widely used assays in translational immunology research. However, there is no data integration pipeline for analyzing these two types of assays together with experiment variables for biomarker inference. Current FCM data analysis mainly relies on subjective manual gating analysis, which is difficult to be directly integrated with other automated computational methods. Existing deconvolutional analysis of bulk transcriptomics relies on predefined marker genes in the transcriptomics data, which are unavailable for novel cell types and does not utilize the FCM data that provide canonical phenotypic definitions of the cell types. RESULTS: We developed a novel analytics pipeline-FastMix-for computational immunology, which integrates flow cytometry, bulk transcriptomics and clinical covariates for identifying cell type-specific gene expression signatures and biomarker genes. FastMix addresses the 'large p, small n' problem in the gene expression and flow cytometry integration analysis via a linear mixed effects model (LMER) for both cross-sectional and longitudinal studies. Its novel moment-based estimator not only reduces bias in parameter estimation but also is more efficient than iterative optimization. The FastMix pipeline also includes a cutting-edge flow cytometry data analysis method-DAFi-for identifying cell populations of interest and their characteristics. Simulation studies showed that FastMix produced smaller type I/II errors than competing methods. Validation using real data of two vaccine studies showed that FastMix identified a consistent set of signature genes as in independent single-cell RNA-seq analysis, producing additional interesting findings. AVAILABILITY AND IMPLEMENTATION: Source code of FastMix is publicly available at https://github.com/terrysun0302/FastMix. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yun Zhang 0021, Aishwarya Mandava, Brian D. Aevermann, Tobias R. Kollmann, Richard H. Scheuermann, Xing Qiu
Bioinform.6
2021 Abstract: Explainable artificial intelligence and single cell genomics to understand the cellular complexity of human brain
abstract
Elucidating the cellular architecture of the human cerebral cortex is central to understanding our cognitive abilities and susceptibility to disease. Using single-nucleus RNA-sequencing analysis to perform a comprehensive study of transcriptomic cell types in the human middle temporal gyrus and primary motor cortex, we have identified a highly diverse set of excitatory and inhibitory neuron types in healthy brain specimens (e.g., PMID: 31435019]. Using these data, we have developed an explainable machine learning method, NS-Forest [PMID: 34088715], to identify a minimum sets of marker genes that define each cell type identified and used these marker gene sets for the semantic representation of this knowledge in biomedical ontologies. NS-Forest marker gene identification also serves as a feature selection step that has been incorporated into a statistical graph-based framework using the FR-Match algorithm[PMID: 332494 for comparing and matching transcriptomic data clusters for novel cell type discovery. The complete characterization of the cell type diversity of healthy human brain will serve as a foundation for understanding the cellular determinants of neurodegenerative, psychiatric, and other brain disorders.
Richard H. Scheuermann
BIBM1
2021 FR-Match: robust matching of cell type clusters from single cell RNA sequencing data using the Friedman-Rafsky non-parametric test
abstract
Single cell/nucleus RNA sequencing (scRNAseq) is emerging as an essential tool to unravel the phenotypic heterogeneity of cells in complex biological systems. While computational methods for scRNAseq cell type clustering have advanced, the ability to integrate datasets to identify common and novel cell types across experiments remains a challenge. Here, we introduce a cluster-to-cluster cell type matching method-FR-Match-that utilizes supervised feature selection for dimensionality reduction and incorporates shared information among cells to determine whether two cell type clusters share the same underlying multivariate gene expression distribution. FR-Match is benchmarked with existing cell-to-cell and cell-to-cluster cell type matching methods using both simulated and real scRNAseq data. FR-Match proved to be a stringent method that produced fewer erroneous matches of distinct cell subtypes and had the unique ability to identify novel cell phenotypes in new datasets. In silico validation demonstrated that the proposed workflow is the only self-contained algorithm that was robust to increasing numbers of true negatives (i.e. non-represented cell types). FR-Match was applied to two human brain scRNAseq datasets sampled from cortical layer 1 and full thickness middle temporal gyrus. When mapping cell types identified in specimens isolated from these overlapping human brain regions, FR-Match precisely recapitulated the laminar characteristics of matched cell type clusters, reflecting their distinct neuroanatomical distributions. An R package and Shiny application are provided at https://github.com/JCVenterInstitute/FRmatch for users to interactively explore and match scRNAseq cell type clusters with complementary visualization tools.
Yun Zhang 0021, Brian D. Aevermann, Trygve E. Bakken, Jeremy A. Miller, Rebecca D. Hodge, Ed S. Lein, Richard H. Scheuermann
Briefings Bioinform.7
2017 Cell type discovery and representation in the era of high-content single cell phenotyping
abstract
BACKGROUND: A fundamental characteristic of multicellular organisms is the specialization of functional cell types through the process of differentiation. These specialized cell types not only characterize the normal functioning of different organs and tissues, they can also be used as cellular biomarkers of a variety of different disease states and therapeutic/vaccine responses. In order to serve as a reference for cell type representation, the Cell Ontology has been developed to provide a standard nomenclature of defined cell types for comparative analysis and biomarker discovery. Historically, these cell types have been defined based on unique cellular shapes and structures, anatomic locations, and marker protein expression. However, we are now experiencing a revolution in cellular characterization resulting from the application of new high-throughput, high-content cytometry and sequencing technologies. The resulting explosion in the number of distinct cell types being identified is challenging the current paradigm for cell type definition in the Cell Ontology. RESULTS: In this paper, we provide examples of state-of-the-art cellular biomarker characterization using high-content cytometry and single cell RNA sequencing, and present strategies for standardized cell type representations based on the data outputs from these cutting-edge technologies, including "context annotations" in the form of standardized experiment metadata about the specimen source analyzed and marker genes that serve as the most useful features in machine learning-based cell type classification models. We also propose a statistical strategy for comparing new experiment data to these standardized cell type representations. CONCLUSION: The advent of high-throughput/high-content single cell technologies is leading to an explosion in the number of distinct cell types being identified. It will be critical for the bioinformatics community to develop and adopt data standard conventions that will be compatible with these new technologies and support the data representation needs of the research community. The proposals enumerated here will serve as a useful starting point to address these challenges.
Trygve E. Bakken, Lindsay G. Cowell, Brian D. Aevermann, Mark Novotny, Rebecca D. Hodge, Jeremy A. Miller, Alexandra Lee, Ivan Chang, Jamison M. McCorrison, Bali Pulendran, Nicholas J. Schork, Roger S. Lasken, Ed S. Lein, Richard H. Scheuermann
BMC Bioinform.15
2017 VDJPipe: a pipelined tool for pre-processing immune repertoire sequencing data
abstract
BACKGROUND: Pre-processing of high-throughput sequencing data for immune repertoire profiling is essential to insure high quality input for downstream analysis. VDJPipe is a flexible, high-performance tool that can perform multiple pre-processing tasks with just a single pass over the data files. RESULTS: Processing tasks provided by VDJPipe include base composition statistics calculation, read quality statistics calculation, quality filtering, homopolymer filtering, length and nucleotide filtering, paired-read merging, barcode demultiplexing, 5' and 3' PCR primer matching, and duplicate reads collapsing. VDJPipe utilizes a pipeline approach whereby multiple processing steps are performed in a sequential workflow, with the output of each step passed as input to the next step automatically. The workflow is flexible enough to handle the complex barcoding schemes used in many immunosequencing experiments. Because VDJPipe is designed for computational efficiency, we evaluated this by comparing execution times with those of pRESTO, a widely-used pre-processing tool for immune repertoire sequencing data. We found that VDJPipe requires <10% of the run time required by pRESTO. CONCLUSIONS: VDJPipe is a high-performance tool that is optimized for pre-processing large immune repertoire sequencing data sets.
Scott Christley, Mikhail K. Levin, Inimary T. Toby, John M. Fonner, Nancy Monson, William Rounds, Florian Rubelt, Walter Scarborough, Richard H. Scheuermann, Lindsay G. Cowell
BMC Bioinform.9
2016 VDJML: a file format with tools for capturing the results of inferring immune receptor rearrangements
abstract
BACKGROUND: The genes that produce antibodies and the immune receptors expressed on lymphocytes are not germline encoded; rather, they are somatically generated in each developing lymphocyte by a process called V(D)J recombination, which assembles specific, independent gene segments into mature composite genes. The full set of composite genes in an individual at a single point in time is referred to as the immune repertoire. V(D)J recombination is the distinguishing feature of adaptive immunity and enables effective immune responses against an essentially infinite array of antigens. Characterization of immune repertoires is critical in both basic research and clinical contexts. Recent technological advances in repertoire profiling via high-throughput sequencing have resulted in an explosion of research activity in the field. This has been accompanied by a proliferation of software tools for analysis of repertoire sequencing data. Despite the widespread use of immune repertoire profiling and analysis software, there is currently no standardized format for output files from V(D)J analysis. Researchers utilize software such as IgBLAST and IMGT/High V-QUEST to perform V(D)J analysis and infer the structure of germline rearrangements. However, each of these software tools produces results in a different file format, and can annotate the same result using different labels. These differences make it challenging for users to perform additional downstream analyses. RESULTS: To help address this problem, we propose a standardized file format for representing V(D)J analysis results. The proposed format, VDJML, provides a common standardized format for different V(D)J analysis applications to facilitate downstream processing of the results in an application-agnostic manner. The VDJML file format specification is accompanied by a support library, written in C++ and Python, for reading and writing the VDJML file format. CONCLUSIONS: The VDJML suite will allow users to streamline their V(D)J analysis and facilitate the sharing of scientific knowledge within the community. The VDJML suite and documentation are available from https://vdjserver.org/vdjml/ . We welcome participation from the community in developing the file format standard, as well as code contributions.
Inimary T. Toby, Mikhail K. Levin, Edward Salinas, Scott Christley, Sanchita Bhattacharya, Felix Breden, Adam Buntzman, Brian Corrie, John M. Fonner, Namita T. Gupta, Uri Hershberg, Nishanth Marthandan, Aaron M. Rosenfeld, William Rounds, Florian Rubelt, Walter Scarborough, Jamie K. Scott, Mohamed Uduman, Jason A. Vander Heiden, Richard H. Scheuermann, Nancy Monson, Steven H. Kleinstein, Lindsay G. Cowell
BMC Bioinform.20
2015 flowCL: ontology-based cell population labelling in flow cytometry
abstract
MOTIVATION: Finding one or more cell populations of interest, such as those correlating to a specific disease, is critical when analysing flow cytometry data. However, labelling of cell populations is not well defined, making it difficult to integrate the output of algorithms to external knowledge sources. RESULTS: We developed flowCL, a software package that performs semantic labelling of cell populations based on their surface markers and applied it to labelling of the Federation of Clinical Immunology Societies Human Immunology Project Consortium lyoplate populations as a use case. CONCLUSION: By providing automated labelling of cell populations based on their immunophenotype, flowCL allows for unambiguous and reproducible identification of standardized cell types. AVAILABILITY AND IMPLEMENTATION: Code, R script and documentation are available under the Artistic 2.0 license through Bioconductor (http://www.bioconductor.org/packages/devel/bioc/html/flowCL.html). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mélanie Courtot, Justin Meskas, Alexander D. Diehl, Radina Droumeva, Raphael Gottardo, Adrin Jalali, Mohammad Jafar Taghiyar, Holden T. Maecker, J. Philip McCoy, Alan Ruttenberg, Richard H. Scheuermann, Ryan Remy Brinkman
Bioinform.11
2012 A Machine Learning Approach for Identifying Anatomical Locations of Actionable Findings in Radiology Reports
Kirk Roberts, Bryan Rink, Sanda M. Harabagiu, Richard H. Scheuermann, Seth M. Toomay, Travis Browning, Teresa Bosler, Ronald M. Peshock
AMIA4
2011 Hematopoietic cell types: Prototype for a revised cell ontology
Alexander D. Diehl, Alison Deckhut Augustine, Judith A. Blake, Lindsay G. Cowell, Elizabeth S. Gold, Timothy A. Gondré-Lewis, Anna Maria Masci, Terrence F. Meehan, Penelope A. Morel, Anastasia Nijnik, Björn Peters, Bali Pulendran, Richard H. Scheuermann, Q. Alison Yao, Martin S. Zand, Chris Mungall
J. Biomed. Informatics13
2011 Toward an ontology-based framework for clinical research databases
Y. Megan Kong, Carl Dahlke, Qun Xiang, David R. Karp, Richard H. Scheuermann
J. Biomed. Informatics6
2011 Ontologies for clinical and translational research: Introduction
Barry Smith 0001, Richard H. Scheuermann
J. Biomed. Informatics2
2010 GO-Bayes: Gene Ontology-based overrepresentation analysis using a Bayesian approach
abstract
MOTIVATION: A typical approach for the interpretation of high-throughput experiments, such as gene expression microarrays, is to produce groups of genes based on certain criteria (e.g. genes that are differentially expressed). To gain more mechanistic insights into the underlying biology, overrepresentation analysis (ORA) is often conducted to investigate whether gene sets associated with particular biological functions, for example, as represented by Gene Ontology (GO) annotations, are statistically overrepresented in the identified gene groups. However, the standard ORA, which is based on the hypergeometric test, analyzes each GO term in isolation and does not take into account the dependence structure of the GO-term hierarchy. RESULTS: We have developed a Bayesian approach (GO-Bayes) to measure overrepresentation of GO terms that incorporates the GO dependence structure by taking into account evidence not only from individual GO terms, but also from their related terms (i.e. parents, children, siblings, etc.). The Bayesian framework borrows information across related GO terms to strengthen the detection of overrepresentation signals. As a result, this method tends to identify sets of closely related GO terms rather than individual isolated GO terms. The advantage of the GO-Bayes approach is demonstrated with a simulation study and an application example.
Song Zhang 0001, Y. Megan Kong, Richard H. Scheuermann
Bioinform.4
2009 Development and Evaluation of a Study Design Typology for Human Research
Simona Carini, Brad Pollock, Harold P. Lehmann, Suzanne Bakken, Edward M. Barbour, Davera Gabriel, Herbert K. Hagler, Caryn R. Harper, Shamim Mollah, Meredith Nahm, Hien H. Nguyen, Richard H. Scheuermann, Ida Sim
AMIA12
2009 Core and periphery structures in protein interaction networks
abstract
BACKGROUND: Characterizing the structural properties of protein interaction networks will help illuminate the organizational and functional relationships among elements in biological systems. RESULTS: In this paper, we present a systematic exploration of the core/periphery structures in protein interaction networks (PINs). First, the concepts of cores and peripheries in PINs are defined. Then, computational methods are proposed to identify two types of cores, k-plex cores and star cores, from PINs. Application of these methods to a yeast protein interaction network has identified 110 k-plex cores and 109 star cores. We find that the k-plex cores consist of either "party" proteins, "date" proteins, or both. We also reveal that there are two classes of 1-peripheral proteins: "party" peripheries, which are more likely to be part of protein complex, and "connector" peripheries, which are more likely connected to different proteins or protein complexes. Our results also show that, besides connectivity, other variations in structural properties are related to the variation in biological properties. Furthermore, the negative correlation between evolutionary rate and connectivity are shown toysis. Moreover, the core/periphery structures help to reveal the existence of multiple levels of protein expression dynamics. CONCLUSION: Our results show that both the structure and connectivity can be used to characterize topological properties in protein interaction networks, illuminating the functional organization of cellular systems.
Feng Luo 0001, Bo Li 0036, Xiu-Feng Wan, Richard H. Scheuermann
BMC Bioinform.4
2009 An improved ontological representation of dendritic cells as a paradigm for all cell types
abstract
BACKGROUND: Recent increases in the volume and diversity of life science data and information and an increasing emphasis on data sharing and interoperability have resulted in the creation of a large number of biological ontologies, including the Cell Ontology (CL), designed to provide a standardized representation of cell types for data annotation. Ontologies have been shown to have significant benefits for computational analyses of large data sets and for automated reasoning applications, leading to organized attempts to improve the structure and formal rigor of ontologies to better support computation. Currently, the CL employs multiple is_a relations, defining cell types in terms of histological, functional, and lineage properties, and the majority of definitions are written with sufficient generality to hold across multiple species. This approach limits the CL's utility for computation and for cross-species data integration. RESULTS: To enhance the CL's utility for computational analyses, we developed a method for the ontological representation of cells and applied this method to develop a dendritic cell ontology (DC-CL). DC-CL subtypes are delineated on the basis of surface protein expression, systematically including both species-general and species-specific types and optimizing DC-CL for the analysis of flow cytometry data. We avoid multiple uses of is_a by linking DC-CL terms to terms in other ontologies via additional, formally defined relations such as has_function. CONCLUSION: This approach brings benefits in the form of increased accuracy, support for reasoning, and interoperability with other ontology resources. Accordingly, we propose our method as a general strategy for the ontological representation of cells. DC-CL is available from http://www.obofoundry.org.
Anna Maria Masci, Cecilia N. Arighi, Alexander D. Diehl, Anne E. Lieberman, Chris Mungall, Richard H. Scheuermann, Barry Smith 0001, Lindsay G. Cowell
BMC Bioinform.6
2009 FuGEFlow: data model and markup language for flow cytometry
abstract
BACKGROUND: Flow cytometry technology is widely used in both health care and research. The rapid expansion of flow cytometry applications has outpaced the development of data storage and analysis tools. Collaborative efforts being taken to eliminate this gap include building common vocabularies and ontologies, designing generic data models, and defining data exchange formats. The Minimum Information about a Flow Cytometry Experiment (MIFlowCyt) standard was recently adopted by the International Society for Advancement of Cytometry. This standard guides researchers on the information that should be included in peer reviewed publications, but it is insufficient for data exchange and integration between computational systems. The Functional Genomics Experiment (FuGE) formalizes common aspects of comprehensive and high throughput experiments across different biological technologies. We have extended FuGE object model to accommodate flow cytometry data and metadata. METHODS: We used the MagicDraw modelling tool to design a UML model (Flow-OM) according to the FuGE extension guidelines and the AndroMDA toolkit to transform the model to a markup language (Flow-ML). We mapped each MIFlowCyt term to either an existing FuGE class or to a new FuGEFlow class. The development environment was validated by comparing the official FuGE XSD to the schema we generated from the FuGE object model using our configuration. After the Flow-OM model was completed, the final version of the Flow-ML was generated and validated against an example MIFlowCyt compliant experiment description. RESULTS: The extension of FuGE for flow cytometry has resulted in a generic FuGE-compliant data model (FuGEFlow), which accommodates and links together all information required by MIFlowCyt. The FuGEFlow model can be used to build software and databases using FuGE software toolkits to facilitate automated exchange and manipulation of potentially large flow cytometry experimental data sets. Additional project documentation, including reusable design patterns and a guide for setting up a development environment, was contributed back to the FuGE project. CONCLUSION: We have shown that an extension of FuGE can be used to transform minimum information requirements in natural language to markup language in XML. Extending FuGE required significant effort, but in our experiences the benefits outweighed the costs. The FuGEFlow is expected to play a central role in describing flow cytometry experiments and ultimately facilitating data exchange including public flow cytometry repositories currently under development.
Olga Tchuvatkina, Josef Spidlen, Peter Wilkinson, Maura Gasparetto, Andrew R. Jones, Frank J. Manion, Richard H. Scheuermann, Rafick-Pierre Sekaly, Ryan Remy Brinkman
BMC Bioinform.8
2008 Exploring Core/Periphery Structures in Protein Interaction Networks Provides Structure-Property Relation Insights
abstract
In this study, we present a systematic exploration of the core/periphery structures in protein interaction networks (PINs). First, the concepts of cores and peripheries in PINs are defined. Then, computational methods are proposed to identify cores from PINs. Application of these methods to a combined yeast PIN has identified 110 k-plex cores and 138 star cores. Based on more precise structural characteristics, our studies reveal new prospects of principles and roles of proteins. Our results show that, aside from connectivity, the structural variations between different types of proteins are also related to the variation in biological properties. Two classes of 1-peripheral proteins have been identified: party peripheries, which are more likely to be part of protein complex, and connector peripheries, which are more likely connected to different complexes or individual proteins. This study may facilitate the understanding of the topological characteristics of proteins in interaction networks and thus help elucidate the organization of cellular systems.
Thomas Grindinger, Feng Luo 0001, Xiu-Feng Wan, Richard H. Scheuermann
BIBM4
2007 A distribution free summarization method for Affymetrix GeneChip® arrays
abstract
MOTIVATION: Affymetrix GeneChip arrays require summarization in order to combine the probe-level intensities into one value representing the expression level of a gene. However, probe intensity measurements are expected to be affected by different levels of non-specific- and cross-hybridization to non-specific transcripts. Here, we present a new summarization technique, the Distribution Free Weighted method (DFW), which uses information about the variability in probe behavior to estimate the extent of non-specific and cross-hybridization for each probe. The contribution of the probe is weighted accordingly during summarization, without making any distributional assumptions for the probe-level data. RESULTS: We compare DFW with several popular summarization methods on spike-in datasets, via both our own calculations and the 'Affycomp II' competition. The results show that DFW outperforms other methods when sensitivity and specificity are considered simultaneously. With the Affycomp spike-in datasets, the area under the receiver operating characteristic curve for DFW is nearly 1.0 (a perfect value), indicating that DFW can identify all differentially expressed genes with a few false positives. The approach used is also computationally faster than most other methods in current use. AVAILABILITY: The R code for DFW is available upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhongxue Chen, Monnie McGee, Qingzhong Liu, Richard H. Scheuermann
Bioinform.4
2007 Ontology development for biological systems: immunology
abstract
UNLABELLED: We recently implemented improvements to the representation of immunology content of the biological process branch of the Gene Ontology (GO). The aims of the revision were to provide a comprehensive representation of immunological processes and to improve the organization of immunology related terms in the GO to match current concepts in the field of immunology. With these improvements, the GO will better reflect current understanding in the field of immunology and thus prove to be a more valuable resource for knowledge representation in gene annotation and analysis in the areas of immunology related to genomics and bioinformatics. AVAILABILITY: http://www.geneontology.org.
Alexander D. Diehl, Jamie A. Lee, Richard H. Scheuermann, Judith A. Blake
Bioinform.3
2007 Modular organization of protein interaction networks
abstract
MOTIVATION: Accumulating evidence suggests that biological systems are composed of interacting, separable, functional modules. Identifying these modules is essential to understand the organization of biological systems. RESULT: In this paper, we present a framework to identify modules within biological networks. In this approach, the concept of degree is extended from the single vertex to the sub-graph, and a formal definition of module in a network is used. A new agglomerative algorithm was developed to identify modules from the network by combining the new module definition with the relative edge order generated by the Girvan-Newman (G-N) algorithm. A JAVA program, MoNet, was developed to implement the algorithm. Applying MoNet to the yeast core protein interaction network from the database of interacting proteins (DIP) identified 86 simple modules with sizes larger than three proteins. The modules obtained are significantly enriched in proteins with related biological process Gene Ontology terms. A comparison between the MoNet modules and modules defined by Radicchi et al. (2004) indicates that MoNet modules show stronger co-clustering of related genes and are more robust to ties in betweenness values. Further, the MoNet output retains the adjacent relationships between modules and allows the construction of an interaction web of modules providing insight regarding the relationships between different functional modules. Thus, MoNet provides an objective approach to understand the organization and interactions of biological processes in cellular systems. AVAILABILITY: MoNet is available upon request from the authors.
Feng Luo 0001, Yunfeng Yang, Chin-Fu Chen, Roger L. Chang, Jizhong Zhou, Richard H. Scheuermann
Bioinform.6
2007 Modular organization of protein interaction networks
Feng Luo 0001, Yunfeng Yang, Chin-Fu Chen, Roger L. Chang, Jizhong Zhou, Richard H. Scheuermann
Bioinform.6
2006 Components of the antigen processing and presentation pathway revealed by gene expression microarray analysis following B cell antigen receptor (BCR) stimulation
abstract
BACKGROUND: Activation of naïve B lymphocytes by extracellular ligands, e.g. antigen, lipopolysaccharide (LPS) and CD40 ligand, induces a combination of common and ligand-specific phenotypic changes through complex signal transduction pathways. For example, although all three of these ligands induce proliferation, only stimulation through the B cell antigen receptor (BCR) induces apoptosis in resting splenic B cells. In order to define the common and unique biological responses to ligand stimulation, we compared the gene expression changes induced in normal primary B cells by a panel of ligands using cDNA microarrays and a statistical approach, CLASSIFI (Cluster Assignment for Biological Inference), which identifies significant co-clustering of genes with similar Gene Ontology annotation. RESULTS: CLASSIFI analysis revealed an overrepresentation of genes involved in ion and vesicle transport, including multiple components of the proton pump, in the BCR-specific gene cluster, suggesting that activation of antigen processing and presentation pathways is a major biological response to antigen receptor stimulation. Proton pump components that were not included in the initial microarray data set were also upregulated in response to BCR stimulation in follow up experiments. MHC Class II expression was found to be maintained specifically in response to BCR stimulation. Furthermore, ligand-specific internalization of the BCR, a first step in B cell antigen processing and presentation, was demonstrated. CONCLUSION: These observations provide experimental validation of the computational approach implemented in CLASSIFI, demonstrating that CLASSIFI-based gene expression cluster analysis is an effective data mining tool to identify biological processes that correlate with the experimental conditional variables. Furthermore, this analysis has identified at least thirty-eight candidate components of the B cell antigen processing and presentation pathway and sets the stage for future studies focused on a better understanding of the components involved in and unique to B cell antigen processing and presentation.
Jamie A. Lee, Robert S. Sinkovits, Dennis Mock, Eva L. Rab, Jennifer Cai, Brian Saunders, Robert C. Hsueh, Sangdun Choi, Shankar Subramaniam, Richard H. Scheuermann
BMC Bioinform.11