EDBT 2026 Demo / reviewers in the wild / expert
Steven H. Kleinstein
dblp:80/2008
· DBLP profile ↗
24ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-4957-1544ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | <tt>nipalsMCIA</tt> : flexible multi-block dimensionality reduction in R via nonlinear iterative partial least squaresabstractSUMMARY: With the increased reliance on multi-omics data for bulk and single-cell analyses, the availability of robust approaches to perform unsupervised learning for clustering, visualization, and feature selection is imperative. We introduce nipalsMCIA, an implementation of multiple co-inertia analysis (MCIA) for joint dimensionality reduction that solves the objective function using an extension to Nonlinear Iterative Partial Least Squares. We applied nipalsMCIA to both bulk and single-cell datasets and observed significant speed-up over other implementations for data with a large sample size and/or feature dimension. AVAILABILITY AND IMPLEMENTATION: nipalsMCIA is available as a Bioconductor package at https://bioconductor.org/packages/release/bioc/html/nipalsMCIA.html, and includes detailed documentation and application vignettes. Max Mattessich, Joaquin Reyna, Edel Aron, Ferhat Ay, Misha Elena Kilmer, Steven H. Kleinstein, Anna Konstorum |
Bioinform. | 6 |
| 2025 | Putting computational models of immunity to the test - An invited challenge to predict B.pertussis vaccination responsesabstractSystems vaccinology studies have been used to build computational models that predict individual vaccine responses and identify the factors contributing to differences in outcome. Comparing such models is challenging due to variability in study designs. To address this, we established a community resource to compare models predicting B. pertussis booster responses and generate experimental data for the explicit purpose of model evaluation. We here describe our second computational prediction challenge using this resource, where we benchmarked 49 algorithms from 53 scientists. We found that the most successful models stood out in their handling of nonlinearities, reducing large feature sets to representative subsets, and advanced data preprocessing. In contrast, we found that models adopted from literature that were developed to predict vaccine antibody responses in other settings performed poorly, reinforcing the need for purpose-built models. Overall, this demonstrates the value of purpose-generated datasets for rigorous and open model evaluations to identify features that improve the reliability and applicability of computational models in vaccine response prediction. Pramod Shinde, Lisa Willemsen, Minori Aoki, Saonli Basu, Julie G. Burel, Souradipto Ghosh Dastidar, Aidan Dunleavy, Tal Einav, Jamie Forschmiedt, Slim Fourati, William Gibson, Jason Greenbaum, Leying Guan, Weikang Guan, Jeremy P. Gygi, Brendan Ha, Joe Hou, Jason Hsiao, Yunda Huang, Rick Jansen, Bhargob Kakoty, Zhiyu Kang, James J. Kobie, Mari Kojima, Anna Konstorum, Jiyeun Lee, Sloan A. Lewis, Aixin Li, Eric F. Lock, Jarjapu Mahita, Marcus Mendes, Hailong Meng, Aidan Neher, Somayeh Nili, Lars Rønn Olsen, Shelby Orfield, James A. Overton, Nidhi Pai, Cokie Parker, Brian Qian, Mikkel Rasmussen, Joaquin Reyna, Eve Richardson, Sandra Safo, Josey Sorenson, Aparna Srinivasan, Nicola Thrupp, Rashmi Tippalagama, Raphael Trevizani, Steffen Ventz, Jiuzhou Wang, Cheng-Chang Wu, Ferhat Ay, Barry Grant, Steven H. Kleinstein, Björn Peters |
PLoS Comput. Biol. | 58 |
| 2025 | Supervised fine-tuning of pre-trained antibody language models improves antigen specificity predictionabstractAntibodies play a crucial role in the adaptive immune response, with their specificity to antigens being a fundamental determinant of immune function. Accurate prediction of antibody-antigen specificity is vital for understanding immune responses, guiding vaccine design, and developing antibody-based therapeutics. In this study, we present a method of supervised fine-tuning for antibody language models, which improves on pre-trained antibody language model embeddings in binding specificity prediction to SARS-CoV-2 spike protein and influenza hemagglutinin. We perform supervised fine-tuning on four pre-trained antibody language models to predict specificity to these antigens and demonstrate that fine-tuned language model classifiers exhibit enhanced predictive accuracy compared to classifiers trained on pre-trained model embeddings. Additionally, we investigate the change of model attention activations after supervised fine-tuning to gain insights into the molecular basis of antigen recognition by antibodies. Furthermore, we apply the supervised fine-tuned models to BCR repertoire data related to influenza and SARS-CoV-2 vaccination, demonstrating their ability to capture changes in repertoire following vaccination. Overall, our study highlights the effect of supervised fine-tuning on pre-trained antibody language models as valuable tools to improve antigen specificity prediction. Jonathan Patsenker, Henry Li, Yuval Kluger, Steven H. Kleinstein |
PLoS Comput. Biol. | 5 |
| 2024 | A supervised Bayesian factor model for the identification of multi-omics signaturesabstractMOTIVATION: Predictive biological signatures provide utility as biomarkers for disease diagnosis and prognosis, as well as prediction of responses to vaccination or therapy. These signatures are identified from high-throughput profiling assays through a combination of dimensionality reduction and machine learning techniques. The genes, proteins, metabolites, and other biological analytes that compose signatures also generate hypotheses on the underlying mechanisms driving biological responses, thus improving biological understanding. Dimensionality reduction is a critical step in signature discovery to address the large number of analytes in omics datasets, especially for multi-omics profiling studies with tens of thousands of measurements. Latent factor models, which can account for the structural heterogeneity across diverse assays, effectively integrate multi-omics data and reduce dimensionality to a small number of factors that capture correlations and associations among measurements. These factors provide biologically interpretable features for predictive modeling. However, multi-omics integration and predictive modeling are generally performed independently in sequential steps, leading to suboptimal factor construction. Combining these steps can yield better multi-omics signatures that are more predictive while still being biologically meaningful. RESULTS: We developed a supervised variational Bayesian factor model that extracts multi-omics signatures from high-throughput profiling datasets that can span multiple data types. Signature-based multiPle-omics intEgration via lAtent factoRs (SPEAR) adaptively determines factor rank, emphasis on factor structure, data relevance and feature sparsity. The method improves the reconstruction of underlying factors in synthetic examples and prediction accuracy of coronavirus disease 2019 severity and breast cancer tumor subtypes. AVAILABILITY AND IMPLEMENTATION: SPEAR is a publicly available R-package hosted at https://bitbucket.org/kleinstein/SPEAR. Jeremy P. Gygi, Anna Konstorum, Shrikant Pawar, Edel Aron, Steven H. Kleinstein, Leying Guan |
Bioinform. | 5 |
| 2024 | nf-core/airrflow: An adaptive immune receptor repertoire analysis workflow employing the Immcantation frameworkabstractAdaptive Immune Receptor Repertoire sequencing (AIRR-seq) is a valuable experimental tool to study the immune state in health and following immune challenges such as infectious diseases, (auto)immune diseases, and cancer. Several tools have been developed to reconstruct B cell and T cell receptor sequences from AIRR-seq data and infer B and T cell clonal relationships. However, currently available tools offer limited parallelization across samples, scalability or portability to high-performance computing infrastructures. To address this need, we developed nf-core/airrflow, an end-to-end bulk and single-cell AIRR-seq processing workflow which integrates the Immcantation Framework following BCR and TCR sequencing data analysis best practices. The Immcantation Framework is a comprehensive toolset, which allows the processing of bulk and single-cell AIRR-seq data from raw read processing to clonal inference. nf-core/airrflow is written in Nextflow and is part of the nf-core project, which collects community contributed and curated Nextflow workflows for a wide variety of analysis tasks. We assessed the performance of nf-core/airrflow on simulated sequencing data with sequencing errors and show example results with real datasets. To demonstrate the applicability of nf-core/airrflow to the high-throughput processing of large AIRR-seq datasets, we validated and extended previously reported findings of convergent antibody responses to SARS-CoV-2 by analyzing 97 COVID-19 infected individuals and 99 healthy controls, including a mixture of bulk and single-cell sequencing datasets. Using this dataset, we extended the convergence findings to 20 additional subjects, highlighting the applicability of nf-core/airrflow to validate findings in small in-house cohorts with reanalysis of large publicly available AIRR datasets. Gisela Gabernet, Susanna Marquez, Robert Bjornson, Alexander Peltzer, Hailong Meng, Edel Aron, Noah Y. Lee, Cole G. Jensen, David Ladd, Mark Polster, Friederike Hanssen, Simon Heumos, Gur Yaari, Markus C. Kowarik, Sven Nahnsen, Steven H. Kleinstein |
PLoS Comput. Biol. | 16 |
| 2023 | A pipeline for the retrieval and extraction of domain-specific information with application to COVID-19 immune signaturesabstractBACKGROUND: The accelerating pace of biomedical publication has made it impractical to manually, systematically identify papers containing specific information and extract this information. This is especially challenging when the information itself resides beyond titles or abstracts. For emerging science, with a limited set of known papers of interest and an incomplete information model, this is of pressing concern. A timely example in retrospect is the identification of immune signatures (coherent sets of biomarkers) driving differential SARS-CoV-2 infection outcomes. IMPLEMENTATION: We built a classifier to identify papers containing domain-specific information from the document embeddings of the title and abstract. To train this classifier with limited data, we developed an iterative process leveraging pre-trained SPECTER document embeddings, SVM classifiers and web-enabled expert review to iteratively augment the training set. This training set was then used to create a classifier to identify papers containing domain-specific information. Finally, information was extracted from these papers through a semi-automated system that directly solicited the paper authors to respond via a web-based form. RESULTS: We demonstrate a classifier that retrieves papers with human COVID-19 immune signatures with a positive predictive value of 86%. The type of immune signature (e.g., gene expression vs. other types of profiling) was also identified with a positive predictive value of 74%. Semi-automated queries to the corresponding authors of these publications requesting signature information achieved a 31% response rate. CONCLUSIONS: Our results demonstrate the efficacy of using a SVM classifier with document embeddings of the title and abstract, to retrieve papers with domain-specific information, even when that information is rarely present in the abstract. Targeted author engagement based on classifier predictions offers a promising pathway to build a semi-structured representation of such information. Through this approach, partially automated literature mining can help rapidly create semi-structured knowledge repositories for automatic analysis of emerging health threats. Adam J. H. Newton, David Chartash, Steven H. Kleinstein, Robert A. McDougal |
BMC Bioinform. | 3 |
| 2022 | Phylogenetic analysis of migration, differentiation, and class switching in B cellsabstractB cells undergo rapid mutation and selection for antibody binding affinity when producing antibodies capable of neutralizing pathogens. This evolutionary process can be intermixed with migration between tissues, differentiation between cellular subsets, and switching between functional isotypes. B cell receptor (BCR) sequence data has the potential to elucidate important information about these processes. However, there is currently no robust, generalizable framework for making such inferences from BCR sequence data. To address this, we develop three parsimony-based summary statistics to characterize migration, differentiation, and isotype switching along B cell phylogenetic trees. We use simulations to demonstrate the effectiveness of this approach. We then use this framework to infer patterns of cellular differentiation and isotype switching from high throughput BCR sequence datasets obtained from patients in a study of HIV infection and a study of food allergy. These methods are implemented in the R package dowser, available at https://dowser.readthedocs.io. Kenneth B. Hoehn, Oliver G. Pybus, Steven H. Kleinstein |
PLoS Comput. Biol. | 3 |
| 2021 | LinkedImm: a linked data graph database for integrating immunological dataabstractBACKGROUND: Many systems biology studies leverage the integration of multiple data types (across different data sources) to offer a more comprehensive view of the biological system being studied. While SQL (Structured Query Language) databases are popular in the biomedical domain, NoSQL database technologies have been used as a more relationship-based, flexible and scalable method of data integration. RESULTS: We have created a graph database integrating data from multiple sources. In addition to using a graph-based query language (Cypher) for data retrieval, we have developed a web-based dashboard that allows users to easily browse and plot data without the need to learn Cypher. We have also implemented a visual graph query interface for users to browse graph data. Finally, we have built a prototype to allow the user to query the graph database in natural language. CONCLUSION: We have demonstrated the feasibility and flexibility of using a graph database for storing and querying immunological data with complex biological relationships. Querying a graph database through such relationships has the potential to discover novel relationships among heterogeneous biological data and metadata. Syed Ahmad Chan Bukhari, Shrikant Pawar, Jeff Mandell, Steven H. Kleinstein, Kei-Hoi Cheung |
BMC Bioinform. | 4 |
| 2020 | Somatic hypermutation analysis for improved identification of B cell clonal families from next-generation sequencing dataabstractAdaptive immune receptor repertoire sequencing (AIRR-Seq) offers the possibility of identifying and tracking B cell clonal expansions during adaptive immune responses. Members of a B cell clone are descended from a common ancestor and share the same initial V(D)J rearrangement, but their B cell receptor (BCR) sequence may differ due to the accumulation of somatic hypermutations (SHMs). Clonal relationships are learned from AIRR-seq data by analyzing the BCR sequence, with the most common methods focused on the highly diverse junction region. However, clonally related cells often share SHMs which have been accumulated during affinity maturation. Here, we investigate whether shared SHMs in the V and J segments of the BCR can be leveraged along with the junction sequence to improve the ability to identify clonally related sequences. We develop independent distance functions that capture junction similarity and shared mutations, and combine these in a spectral clustering framework to infer the BCR clonal relationships. Using both simulated and experimental data, we show that this model improves both the sensitivity and specificity for identifying B cell clones. Source code for this method is freely available in the SCOPer (Spectral Clustering for clOne Partitioning) R package (version 0.2 or newer) in the Immcantation framework: www.immcantation.org under the AGPLv3 license. Nima Nouri, Steven H. Kleinstein |
PLoS Comput. Biol. | 2 |
| 2019 | A linked data graph approach to integration of immunological dataabstractSystems biology involves the integration of multiple data types (across different data sources) to offer a more complete picture of the biological system being studied. While many existing biological databases are implemented using the traditional SQL (Structured Query Language) database technology, NoSQL database technologies have been explored as a more relationship-based, flexible and scalable method of data integration. In this paper, we describe how to use the Neo4J graph database to integrate a variety of types of data sets in the context of systems vaccinology. Specifically, we have converted into a common graph model diverse types of vaccine response measurement data from the NIH/NIAID ImmPort data repository, pathway data from Reactome, influenza virus strains from WHO, and taxonomic data from NCBI Taxon. While Neo4J provides a graph-based query language (Cypher) for data retrieval, we develop a web-based dashboard for users to easily browse and visualize data without the need to learn Cypher. In addition, we have prototyped a natural language query interface for users to interact with our system. In conclusion, we demonstrate the feasibility of using a graph-based database for storing and querying immunological data with complex biological relationships. Querying a graph database through such relationships has the potential to reveal novel relationships among heterogeneous biological data. Syed Ahmad Chan Bukhari, Jeff Mandell, Steven H. Kleinstein, Kei-Hoi Cheung |
BIBM | 3 |
| 2019 | Reporting and connecting cell type names and gating definitions through ontologiesabstractBACKGROUND: Human immunology studies often rely on the isolation and quantification of cell populations from an input sample based on flow cytometry and related techniques. Such techniques classify cells into populations based on the detection of a pattern of markers. The description of the cell populations targeted in such experiments typically have two complementary components: the description of the cell type targeted (e.g. 'T cells'), and the description of the marker pattern utilized (e.g. CD14-, CD3+). RESULTS: We here describe our attempts to use ontologies to cross-compare cell types and marker patterns (also referred to as gating definitions). We used a large set of such gating definitions and corresponding cell types submitted by different investigators into ImmPort, a central database for immunology studies, to examine the ability to parse gating definitions using terms from the Protein Ontology (PRO) and cell type descriptions, using the Cell Ontology (CL). We then used logical axioms from CL to detect discrepancies between the two. CONCLUSIONS: We suggest adoption of our proposed format for describing gating and cell type definitions to make comparisons easier. We also suggest a number of new terms to describe gating definitions in flow cytometry that are not based on molecular markers captured in PRO, but on forward- and side-scatter of light during data acquisition, which is more appropriate to capture in the Ontology for Biomedical Investigations (OBI). Finally, our approach results in suggestions on what logical axioms and new cell types could be considered for addition to the Cell Ontology. James A. Overton, Randi Vita, Patrick Dunn, Julie G. Burel, Syed Ahmad Chan Bukhari, Kei-Hoi Cheung, Steven H. Kleinstein, Alexander D. Diehl, Björn Peters |
BMC Bioinform. | 7 |
| 2019 | Gene set meta-analysis with Quantitative Set Analysis for Gene Expression (QuSAGE)abstractSmall sample sizes combined with high person-to-person variability can make it difficult to detect significant gene expression changes from transcriptional profiling studies. Subtle, but coordinated, gene expression changes may be detected using gene set analysis approaches. Meta-analysis is another approach to increase the power to detect biologically relevant changes by integrating information from multiple studies. Here, we present a framework that combines both approaches and allows for meta-analysis of gene sets. QuSAGE meta-analysis extends our previously published QuSAGE framework, which offers several advantages for gene set analysis, including fully accounting for gene-gene correlations and quantifying gene set activity as a full probability density function. Application of QuSAGE meta-analysis to influenza vaccination response shows it can detect significant activity that is not apparent in individual studies. Hailong Meng, Gur Yaari, Christopher R. Bolen, Stefan Avey, Steven H. Kleinstein |
PLoS Comput. Biol. | 5 |
| 2018 | A spectral clustering-based method for identifying clones from high-throughput B cell repertoire sequencing dataabstractMotivation: B cells derive their antigen-specificity through the expression of Immunoglobulin (Ig) receptors on their surface. These receptors are initially generated stochastically by somatic re-arrangement of the DNA and further diversified following antigen-activation by a process of somatic hypermutation, which introduces mainly point substitutions into the receptor DNA at a high rate. Recent advances in next-generation sequencing have enabled large-scale profiling of the B cell Ig repertoire from blood and tissue samples. A key computational challenge in the analysis of these data is partitioning the sequences to identify descendants of a common B cell (i.e. a clone). Current methods group sequences using a fixed distance threshold, or a likelihood calculation that is computationally-intensive. Here, we propose a new method based on spectral clustering with an adaptive threshold to determine the local sequence neighborhood. Validation using simulated and experimental datasets demonstrates that this method has high sensitivity and specificity compared to a fixed threshold that is optimized for these measures. In addition, this method works on datasets where choosing an optimal fixed threshold is difficult and is more computationally efficient in all cases. The ability to quickly and accurately identify members of a clone from repertoire sequencing data will greatly improve downstream analyses. Clonally-related sequences cannot be treated independently in statistical models, and clonal partitions are used as the basis for the calculation of diversity metrics, lineage reconstruction and selection analysis. Thus, the spectral clustering-based method here represents an important contribution to repertoire analysis. Availability and implementation: Source code for this method is freely available in the SCOPe (Spectral Clustering for clOne Partitioning) R package in the Immcantation framework: www.immcantation.org under the CC BY-SA 4.0 license. Supplementary information: Supplementary data are available at Bioinformatics online. Nima Nouri, Steven H. Kleinstein |
Bioinform. | 2 |
| 2018 | CEDAR OnDemand: a browser extension to generate ontology-based scientific metadataabstractBACKGROUND: Public biomedical data repositories often provide web-based interfaces to collect experimental metadata. However, these interfaces typically reflect the ad hoc metadata specification practices of the associated repositories, leading to a lack of standardization in the collected metadata. This lack of standardization limits the ability of the source datasets to be broadly discovered, reused, and integrated with other datasets. To increase reuse, discoverability, and reproducibility of the described experiments, datasets should be appropriately annotated by using agreed-upon terms, ideally from ontologies or other controlled term sources. RESULTS: This work presents "CEDAR OnDemand", a browser extension powered by the NCBO (National Center for Biomedical Ontology) BioPortal that enables users to seamlessly enter ontology-based metadata through existing web forms native to individual repositories. CEDAR OnDemand analyzes the web page contents to identify the text input fields and associate them with relevant ontologies which are recommended automatically based upon input fields' labels (using the NCBO ontology recommender) and a pre-defined list of ontologies. These field-specific ontologies are used for controlling metadata entry. CEDAR OnDemand works for any web form designed in the HTML format. We demonstrate how CEDAR OnDemand works through the NCBI (National Center for Biotechnology Information) BioSample web-based metadata entry. CONCLUSION: CEDAR OnDemand helps lower the barrier of incorporating ontologies into standardized metadata entry for public data repositories. CEDAR OnDemand is available freely on the Google Chrome store https://chrome.google.com/webstore/search/CEDAROnDemand. Syed Ahmad Chan Bukhari, Marcos Martínez Romero, Martin J. O'Connor, Attila L. Egyedi, Debra Willrett, John B. Graybeal, Mark A. Musen, Kei-Hoi Cheung, Steven H. Kleinstein |
BMC Bioinform. | 9 |
| 2017 | Multiple network-constrained regressions expand insights into influenza vaccination responsesabstractMOTIVATION: Systems immunology leverages recent technological advancements that enable broad profiling of the immune system to better understand the response to infection and vaccination, as well as the dysregulation that occurs in disease. An increasingly common approach to gain insights from these large-scale profiling experiments involves the application of statistical learning methods to predict disease states or the immune response to perturbations. However, the goal of many systems studies is not to maximize accuracy, but rather to gain biological insights. The predictors identified using current approaches can be biologically uninterpretable or present only one of many equally predictive models, leading to a narrow understanding of the underlying biology. RESULTS: Here we show that incorporating prior biological knowledge within a logistic modeling framework by using network-level constraints on transcriptional profiling data significantly improves interpretability. Moreover, incorporating different types of biological knowledge produces models that highlight distinct aspects of the underlying biology, while maintaining predictive accuracy. We propose a new framework, Logistic Multiple Network-constrained Regression (LogMiNeR), and apply it to understand the mechanisms underlying differential responses to influenza vaccination. Although standard logistic regression approaches were predictive, they were minimally interpretable. Incorporating prior knowledge using LogMiNeR led to models that were equally predictive yet highly interpretable. In this context, B cell-specific genes and mTOR signaling were associated with an effective vaccination response in young adults. Overall, our results demonstrate a new paradigm for analyzing high-dimensional immune profiling data in which multiple networks encoding prior knowledge are incorporated to improve model interpretability. AVAILABILITY AND IMPLEMENTATION: The R source code described in this article is publicly available at https://bitbucket.org/kleinstein/logminer . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Stefan Avey, Subhasis Mohanty, Jean Wilson, Heidi Zapata, Samit R. Joshi, Barbara Siconolfi, Sui Tsang, Albert C. Shaw, Steven H. Kleinstein |
Bioinform. | 9 |
| 2016 | VDJML: a file format with tools for capturing the results of inferring immune receptor rearrangementsabstractBACKGROUND: The genes that produce antibodies and the immune receptors expressed on lymphocytes are not germline encoded; rather, they are somatically generated in each developing lymphocyte by a process called V(D)J recombination, which assembles specific, independent gene segments into mature composite genes. The full set of composite genes in an individual at a single point in time is referred to as the immune repertoire. V(D)J recombination is the distinguishing feature of adaptive immunity and enables effective immune responses against an essentially infinite array of antigens. Characterization of immune repertoires is critical in both basic research and clinical contexts. Recent technological advances in repertoire profiling via high-throughput sequencing have resulted in an explosion of research activity in the field. This has been accompanied by a proliferation of software tools for analysis of repertoire sequencing data. Despite the widespread use of immune repertoire profiling and analysis software, there is currently no standardized format for output files from V(D)J analysis. Researchers utilize software such as IgBLAST and IMGT/High V-QUEST to perform V(D)J analysis and infer the structure of germline rearrangements. However, each of these software tools produces results in a different file format, and can annotate the same result using different labels. These differences make it challenging for users to perform additional downstream analyses. RESULTS: To help address this problem, we propose a standardized file format for representing V(D)J analysis results. The proposed format, VDJML, provides a common standardized format for different V(D)J analysis applications to facilitate downstream processing of the results in an application-agnostic manner. The VDJML file format specification is accompanied by a support library, written in C++ and Python, for reading and writing the VDJML file format. CONCLUSIONS: The VDJML suite will allow users to streamline their V(D)J analysis and facilitate the sharing of scientific knowledge within the community. The VDJML suite and documentation are available from https://vdjserver.org/vdjml/ . We welcome participation from the community in developing the file format standard, as well as code contributions. Inimary T. Toby, Mikhail K. Levin, Edward Salinas, Scott Christley, Sanchita Bhattacharya, Felix Breden, Adam Buntzman, Brian Corrie, John M. Fonner, Namita T. Gupta, Uri Hershberg, Nishanth Marthandan, Aaron M. Rosenfeld, William Rounds, Florian Rubelt, Walter Scarborough, Jamie K. Scott, Mohamed Uduman, Jason A. Vander Heiden, Richard H. Scheuermann, Nancy Monson, Steven H. Kleinstein, Lindsay G. Cowell |
BMC Bioinform. | 22 |
| 2015 | Improvement of cytokine annotation using ontology synonym mapping
Demetrios Sarantis, Steven H. Kleinstein, Kei Cheung |
AMIA | 2 |
| 2015 | Change-O: a toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing dataabstractUNLABELLED: Advances in high-throughput sequencing technologies now allow for large-scale characterization of B cell immunoglobulin (Ig) repertoires. The high germline and somatic diversity of the Ig repertoire presents challenges for biologically meaningful analysis, which requires specialized computational methods. We have developed a suite of utilities, Change-O, which provides tools for advanced analyses of large-scale Ig repertoire sequencing data. Change-O includes tools for determining the complete set of Ig variable region gene segment alleles carried by an individual (including novel alleles), partitioning of Ig sequences into clonal populations, creating lineage trees, inferring somatic hypermutation targeting models, measuring repertoire diversity, quantifying selection pressure, and calculating sequence chemical properties. All Change-O tools utilize a common data format, which enables the seamless integration of multiple analyses into a single workflow. AVAILABILITY AND IMPLEMENTATION: Change-O is freely available for non-commercial use and may be downloaded from http://clip.med.yale.edu/changeo. CONTACT: [email protected]. Namita T. Gupta, Jason A. Vander Heiden, Mohamed Uduman, Daniel Gadala-Maria, Gur Yaari, Steven H. Kleinstein |
Bioinform. | 6 |
| 2015 | The center for expanded data annotation and retrievalabstractThe Center for Expanded Data Annotation and Retrieval is studying the creation of comprehensive and expressive metadata for biomedical datasets to facilitate data discovery, data interpretation, and data reuse. We take advantage of emerging community-based standard templates for describing different kinds of biomedical datasets, and we investigate the use of computational techniques to help investigators to assemble templates and to fill in their values. We are creating a repository of metadata from which we plan to identify metadata patterns that will drive predictive data entry when filling in metadata templates. The metadata repository not only will capture annotations specified when experimental datasets are initially created, but also will incorporate links to the published literature, including secondary analyses and possible refinements or retractions of experimental interpretations. By working initially with the Human Immunology Project Consortium and the developers of the ImmPort data repository, we are developing and evaluating an end-to-end solution to the problems of metadata authoring and management that will generalize to other data-management environments. Mark A. Musen, Carol A. Bean, Kei-Hoi Cheung, Michel Dumontier, Kim A. Durante, Olivier Gevaert, Alejandra N. González-Beltrán, Purvesh Khatri, Steven H. Kleinstein, Martin J. O'Connor, Yannick Pouliot, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jeffrey A. Wiser |
J. Am. Medical Informatics Assoc. | 9 |
| 2014 | pRESTO: a toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoiresabstractUNLABELLED: Driven by dramatic technological improvements, large-scale characterization of lymphocyte receptor repertoires via high-throughput sequencing is now feasible. Although promising, the high germline and somatic diversity, especially of B-cell immunoglobulin repertoires, presents challenges for analysis requiring the development of specialized computational pipelines. We developed the REpertoire Sequencing TOolkit (pRESTO) for processing reads from high-throughput lymphocyte receptor studies. pRESTO processes raw sequences to produce error-corrected, sorted and annotated sequence sets, along with a wealth of metrics at each step. The toolkit supports multiplexed primer pools, single- or paired-end reads and emerging technologies that use single-molecule identifiers. pRESTO has been tested on data generated from Roche and Illumina platforms. It has a built-in capacity to parallelize the work between available processors and is able to efficiently process millions of sequences generated by typical high-throughput projects. AVAILABILITY AND IMPLEMENTATION: pRESTO is freely available for academic use. The software package and detailed tutorials may be downloaded from http://clip.med.yale.edu/presto. Jason A. Vander Heiden, Gur Yaari, Mohamed Uduman, Joel N. H. Stern, Kevin C. O'Connor, David A. Hafler, Francois Vigneault, Steven H. Kleinstein |
Bioinform. | 8 |
| 2013 | Reconstruction of regulatory networks through temporal enrichment profiling and its application to H1N1 influenza viral infectionabstractBACKGROUND: H1N1 influenza viruses were responsible for the 1918 pandemic that caused millions of deaths worldwide and the 2009 pandemic that caused approximately twenty thousand deaths. The cellular response to such virus infections involves extensive genetic reprogramming resulting in an antiviral state that is critical to infection control. Identifying the underlying transcriptional network driving these changes, and how this program is altered by virally-encoded immune antagonists, is a fundamental challenge in systems immunology. RESULTS: Genome-wide gene expression patterns were measured in human monocyte-derived dendritic cells (DCs) infected in vitro with seasonal H1N1 influenza A/New Caledonia/20/1999. To provide a mechanistic explanation for the timing of gene expression changes over the first 12 hours post-infection, we developed a statistically rigorous enrichment approach integrating genome-wide expression kinetics and time-dependent promoter analysis. Our approach, TIme-Dependent Activity Linker (TIDAL), generates a regulatory network that connects transcription factors associated with each temporal phase of the response into a coherent linked cascade. TIDAL infers 12 transcription factors and 32 regulatory connections that drive the antiviral response to influenza. To demonstrate the generality of this approach, TIDAL was also used to generate a network for the DC response to measles infection. The software implementation of TIDAL is freely available at http://tsb.mssm.edu/primeportal/?q=tidal_prog. CONCLUSIONS: We apply TIDAL to reconstruct the transcriptional programs activated in monocyte-derived human dendritic cells in response to influenza and measles infections. The application of this time-centric network reconstruction method in each case produces a single transcriptional cascade that recapitulates the known biology of the response with high precision and recall, in addition to identifying potentially novel antiviral factors. The ability to reconstruct antiviral networks with TIDAL enables comparative analysis of antiviral responses, such as the differences between pandemic and seasonal influenza infections. Elena Zaslavsky, German Nudelman, Susanna Marquez, Uri Hershberg, Boris M. Hartmann, Juilee Thakar, Stuart C. Sealfon, Steven H. Kleinstein |
BMC Bioinform. | 8 |
| 2011 | Cell Subset Prediction for Blood Genomic StudiesabstractBACKGROUND: Genome-wide transcriptional profiling of patient blood samples offers a powerful tool to investigate underlying disease mechanisms and personalized treatment decisions. Most studies are based on analysis of total peripheral blood mononuclear cells (PBMCs), a mixed population. In this case, accuracy is inherently limited since cell subset-specific differential expression of gene signatures will be diluted by RNA from other cells. While using specific PBMC subsets for transcriptional profiling would improve our ability to extract knowledge from these data, it is rarely obvious which cell subset(s) will be the most informative. RESULTS: We have developed a computational method (Subset Prediction from Enrichment Correlation, SPEC) to predict the cellular source for a pre-defined list of genes (i.e. a gene signature) using only data from total PBMCs. SPEC does not rely on the occurrence of cell subset-specific genes in the signature, but rather takes advantage of correlations with subset-specific genes across a set of samples. Validation using multiple experimental datasets demonstrates that SPEC can accurately identify the source of a gene signature as myeloid or lymphoid, as well as differentiate between B cells, T cells, NK cells and monocytes. Using SPEC, we predict that myeloid cells are the source of the interferon-therapy response gene signature associated with HCV patients who are non-responsive to standard therapy. CONCLUSIONS: SPEC is a powerful technique for blood genomic studies. It can help identify specific cell subsets that are important for understanding disease and therapy response. SPEC is widely applicable since only gene expression profiles from total PBMCs are required, and thus it can easily be used to mine the massive amount of existing microarray or RNA-seq data. Christopher R. Bolen, Mohamed Uduman, Steven H. Kleinstein |
BMC Bioinform. | 3 |
| 2009 | Activated Germinal-Center B Cells Undergo Directed MigrationabstractAffinity maturation, the fundamental basis for adaptive immunity, is accomplished through somatic hypermutation of B-cell receptors followed by the expansion of rare mutants with higher affinity for the immunizing antigen. This process occurs over a period of weeks in unique microanatomic sites known as germinal centers. Two-photon microscopy has recently made it possible to track individual B cells moving within germinal centers in living animals. Characterizing the migration patterns of B cells within germinal centers is critical for understanding the mechanisms underlying affinity maturation. Here we present the results of two statistical approaches designed to test the hypothesis that the motion of B cells within germinal centers is random. Analysis of four different experiments shows that activated B cells move in a directed manner that sharply contrasts with the behavior of naive B cells. Mark J. O'Connor, Anja E. Hauser, Ann M. Haberman, Steven H. Kleinstein |
BIBM | 4 |
| 2008 | Getting Started in Computational ImmunologyabstractThe immune system acts across multiple scales involving complex interactions and feedback, from somatic modifications of DNA to the systemic inflammatory reaction. Computational modeling provides a framework to integrate observational data collected from multiple modes of experimentation and insight into the immune response in health and disease. This Message attempts to illustrate how different computational methods have been integrated with experimental observations to study an immunological question from multiple perspectives by focusing on a very particular, though fundamental, component of adaptive immunity: B cells and affinity maturation (Figure 1). B cells bind foreign antigens through their Immunoglobulin (Ig) receptor. Affinity maturation is the process by which B cell receptors that initially bind antigen with low affinity are modified through cycles of somatic mutation and affinity-dependent selection to produce high-affinity memory and plasma cells. How this process can reliably generate orders of magnitude increases in affinity over a period of weeks is one of the many questions where computational modeling has made important contributions (for example, the cyclic re-entry model [1]). Yet, even the seemingly straightforward matter of detecting antigen-driven selection remains controversial, and such fundamental questions as whether increased proliferation or decreased death drives the preferential expansion of higher-affinity B cell mutants remain unanswered. A good biological introduction to the immune system is available on the NIH website [2], while more detailed information can be found in any number of textbooks [3]. An animation by Julian Kirk-Elleker provides a visual introduction to the affinity maturation process (http://web.mac.com/patrickwlee/Antibody-affinity_maturation/Movie.html). The kinds of computational techniques described here have been widely applied in other areas of immunology, including the innate response [4],[5], viral dynamics [6], and immune memory [7]. A classic introduction to computational immunology geared to the more mathematically inclined was written by Perelson and Weisbuch [8]. The rapidly expanding area of immunoinformatics was covered in a recent issue of PLoS Computational Biology [9], and several other applications were explored in a 2007 volume of Immunological Reviews (216) devoted to quantitative modeling of immune responses.
Figure 1
A wide range of experimental techniques are used in combination with computational modeling to probe the process of affinity maturation at multiple scales (from DNA to tissue). Steven H. Kleinstein |
PLoS Comput. Biol. | 1 |