EDBT 2026 Demo / reviewers in the wild / expert
Gur Yaari
dblp:70/7208
· DBLP profile ↗
11ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0001-9311-9884ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An unbiased comparison of immunoglobulin sequence alignersabstractAdaptive Immune Receptor Repertoire sequencing (AIRR-seq) is critical for our understanding of the adaptive immune system's dynamics in health and disease. Reliable analysis of AIRR-seq data depends on accurate rearranged immunoglobulin (Ig) sequence alignment. Various Ig sequence aligners exist, but there is no unified benchmarking standard representing the complexities of AIRR-seq data, obscuring objective comparisons of aligners across tasks. Here, we introduce GenAIRR, a modular simulation framework for generating Ig sequences alongside their ground truths. GenAIRR realistically simulates the intricacies of V(D)J recombination, somatic hypermutation, and an array of sequence corruptions. We comprehensively assessed prominent Ig sequence aligners across various metrics, unveiling unique performance characteristics for each aligner. The GenAIRR-produced datasets, combined with the proposed rigorous evaluation criteria, establish a solid basis for unbiased benchmarking of immunogenetics computational tools. It sets up the ground for further improving the crucial task of Ig sequence alignment, ultimately enhancing our understanding of adaptive immunity. Thomas Konstantinovsky, Ayelet Peres, Pazit Polak, Gur Yaari |
Briefings Bioinform. | 4 |
| 2024 | Guidelines for reproducible analysis of adaptive immune receptor repertoire sequencing dataabstractEnhancing the reproducibility and comprehension of adaptive immune receptor repertoire sequencing (AIRR-seq) data analysis is critical for scientific progress. This study presents guidelines for reproducible AIRR-seq data analysis, and a collection of ready-to-use pipelines with comprehensive documentation. To this end, ten common pipelines were implemented using ViaFoundry, a user-friendly interface for pipeline management and automation. This is accompanied by versioned containers, documentation and archiving capabilities. The automation of pre-processing analysis steps and the ability to modify pipeline parameters according to specific research needs are emphasized. AIRR-seq data analysis is highly sensitive to varying parameters and setups; using the guidelines presented here, the ability to reproduce previously published results is demonstrated. This work promotes transparency, reproducibility, and collaboration in AIRR-seq data analysis, serving as a model for handling and documenting bioinformatics pipelines in other research domains. Ayelet Peres, Vered Klein, Boaz Frankel, William D. Lees, Pazit Polak, Mark Meehan, Artur Rocha, João Correia Lopes, Gur Yaari |
Briefings Bioinform. | 9 |
| 2024 | Digger: directed annotation of immunoglobulin and T cell receptor V, D, and J gene sequences and assembliesabstractSUMMARY: Knowledge of immunoglobulin and T cell receptor encoding genes is derived from high-quality genomic sequencing. High-throughput sequencing is delivering large volumes of data, and precise, high-throughput approaches to annotation are needed. Digger is an automated tool that identifies coding and regulatory regions of these genes, with results comparable to those obtained by current expert curational methods. AVAILABILITY AND IMPLEMENTATION: Digger is published under open source license at https://github.com/williamdlees/Digger and is available as a Python package and a Docker container. William D. Lees, Swati Saha, Gur Yaari, Corey T. Watson |
Bioinform. | 3 |
| 2024 | nf-core/airrflow: An adaptive immune receptor repertoire analysis workflow employing the Immcantation frameworkabstractAdaptive Immune Receptor Repertoire sequencing (AIRR-seq) is a valuable experimental tool to study the immune state in health and following immune challenges such as infectious diseases, (auto)immune diseases, and cancer. Several tools have been developed to reconstruct B cell and T cell receptor sequences from AIRR-seq data and infer B and T cell clonal relationships. However, currently available tools offer limited parallelization across samples, scalability or portability to high-performance computing infrastructures. To address this need, we developed nf-core/airrflow, an end-to-end bulk and single-cell AIRR-seq processing workflow which integrates the Immcantation Framework following BCR and TCR sequencing data analysis best practices. The Immcantation Framework is a comprehensive toolset, which allows the processing of bulk and single-cell AIRR-seq data from raw read processing to clonal inference. nf-core/airrflow is written in Nextflow and is part of the nf-core project, which collects community contributed and curated Nextflow workflows for a wide variety of analysis tasks. We assessed the performance of nf-core/airrflow on simulated sequencing data with sequencing errors and show example results with real datasets. To demonstrate the applicability of nf-core/airrflow to the high-throughput processing of large AIRR-seq datasets, we validated and extended previously reported findings of convergent antibody responses to SARS-CoV-2 by analyzing 97 COVID-19 infected individuals and 99 healthy controls, including a mixture of bulk and single-cell sequencing datasets. Using this dataset, we extended the convergence findings to 20 additional subjects, highlighting the applicability of nf-core/airrflow to validate findings in small in-house cohorts with reanalysis of large publicly available AIRR datasets. Gisela Gabernet, Susanna Marquez, Robert Bjornson, Alexander Peltzer, Hailong Meng, Edel Aron, Noah Y. Lee, Cole G. Jensen, David Ladd, Mark Polster, Friederike Hanssen, Simon Heumos, Gur Yaari, Markus C. Kowarik, Sven Nahnsen, Steven H. Kleinstein |
PLoS Comput. Biol. | 13 |
| 2023 | A novel approach to T-cell receptor beta chain (TCRB) repertoire encoding using lossless string compressionabstractMOTIVATION: T-cell receptor beta chain (TCRB) repertoires are crucial for understanding immune responses. However, their high diversity and complexity present significant challenges in representation and analysis. The main motivation of this study is to develop a unified and compact representation of a TCRB repertoire that can efficiently capture its inherent complexity and diversity and allow for direct inference. RESULTS: We introduce a novel approach to TCRB repertoire encoding and analysis, leveraging the Lempel-Ziv 76 algorithm. This approach allows us to create a graph-like model, identify-specific sequence features, and produce a new encoding approach for an individual's repertoire. The proposed representation enables various applications, including generation probability inference, informative feature vector derivation, sequence generation, a new measure for diversity estimation, and a new sequence centrality measure. The approach was applied to four large-scale public TCRB sequencing datasets, demonstrating its potential for a wide range of applications in big biological sequencing data. AVAILABILITY AND IMPLEMENTATION: Python package for implementation is available https://github.com/MuteJester/LZGraphs. Thomas Konstantinovsky, Gur Yaari |
Bioinform. | 2 |
| 2021 | Breast cancer is marked by specific, Public T-cell receptor CDR3 regions shared by mice and humansabstractThe partial success of tumor immunotherapy induced by checkpoint blockade, which is not antigen-specific, suggests that the immune system of some patients contain antigen receptors able to specifically identify tumor cells. Here we focused on T-cell receptor (TCR) repertoires associated with spontaneous breast cancer. We studied the alpha and beta chain CDR3 domains of TCR repertoires of CD4 T cells using deep sequencing of cell populations in mice and applied the results to published TCR sequence data obtained from human patients. We screened peripheral blood T cells obtained monthly from individual mice spontaneously developing breast tumors by 5 months. We then looked at identical TCR sequences in published human studies; we used TCGA data from tumors and healthy tissues of 1,256 breast cancer resections and from 4 focused studies including sequences from tumors, lymph nodes, blood and healthy tissues, and from single cell dataset of 3 breast cancer subjects. We now report that mice spontaneously developing breast cancer manifest shared, Public CDR3 regions in both their alpha and beta and that a significant number of women with early breast cancer manifest identical CDR3 sequences. These findings suggest that the development of breast cancer is associated, across species, with biomarker, exclusive TCR repertoires. Miri Gordin, Hagit Philip, Alona Zilberberg, Moriah Gidoni, Raanan Margalit, Christopher Clouser, Kristofor Adams, Francois Vigneault, Irun R. Cohen, Gur Yaari, Sol Efroni |
PLoS Comput. Biol. | 10 |
| 2019 | RAbHIT: R Antibody Haplotype Inference ToolabstractSUMMARY: Antibody haplotype inference (chromosomal phasing) may have clinical implications for the identification of genetic predispositions to diseases. Yet, our knowledge of the genomic loci encoding for the variable regions of the antibody is only partial, mostly due to the challenge of aligning short reads from genome sequencing to these highly repetitive loci. A powerful approach to infer the content of these loci relies on analyzing repertoires of rearranged V(D)J sequences. We present here RAbHIT, an R Haplotype Antibody Inference Tool, that implements a novel algorithm to infer V(D)J haplotypes by adapting a Bayesian framework. RAbHIT offers inference of haplotype and gene deletions. It may be applied to sequences from naïve and non-naïve B-cells, sequenced by different library preparation protocols. AVAILABILITY AND IMPLEMENTATION: RAbHIT is freely available for academic use from comprehensive R archive network (CRAN) (https://cran.r-project.org/web/packages/rabhit/) under CC BY-SA 4.0 license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ayelet Peres, Moriah Gidoni, Pazit Polak, Gur Yaari |
Bioinform. | 4 |
| 2019 | Gene set meta-analysis with Quantitative Set Analysis for Gene Expression (QuSAGE)abstractSmall sample sizes combined with high person-to-person variability can make it difficult to detect significant gene expression changes from transcriptional profiling studies. Subtle, but coordinated, gene expression changes may be detected using gene set analysis approaches. Meta-analysis is another approach to increase the power to detect biologically relevant changes by integrating information from multiple studies. Here, we present a framework that combines both approaches and allows for meta-analysis of gene sets. QuSAGE meta-analysis extends our previously published QuSAGE framework, which offers several advantages for gene set analysis, including fully accounting for gene-gene correlations and quantifying gene set activity as a full probability density function. Application of QuSAGE meta-analysis to influenza vaccination response shows it can detect significant activity that is not apparent in individual studies. Hailong Meng, Gur Yaari, Christopher R. Bolen, Stefan Avey, Steven H. Kleinstein |
PLoS Comput. Biol. | 2 |
| 2015 | Change-O: a toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing dataabstractUNLABELLED: Advances in high-throughput sequencing technologies now allow for large-scale characterization of B cell immunoglobulin (Ig) repertoires. The high germline and somatic diversity of the Ig repertoire presents challenges for biologically meaningful analysis, which requires specialized computational methods. We have developed a suite of utilities, Change-O, which provides tools for advanced analyses of large-scale Ig repertoire sequencing data. Change-O includes tools for determining the complete set of Ig variable region gene segment alleles carried by an individual (including novel alleles), partitioning of Ig sequences into clonal populations, creating lineage trees, inferring somatic hypermutation targeting models, measuring repertoire diversity, quantifying selection pressure, and calculating sequence chemical properties. All Change-O tools utilize a common data format, which enables the seamless integration of multiple analyses into a single workflow. AVAILABILITY AND IMPLEMENTATION: Change-O is freely available for non-commercial use and may be downloaded from http://clip.med.yale.edu/changeo. CONTACT: [email protected]. Namita T. Gupta, Jason A. Vander Heiden, Mohamed Uduman, Daniel Gadala-Maria, Gur Yaari, Steven H. Kleinstein |
Bioinform. | 5 |
| 2014 | pRESTO: a toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoiresabstractUNLABELLED: Driven by dramatic technological improvements, large-scale characterization of lymphocyte receptor repertoires via high-throughput sequencing is now feasible. Although promising, the high germline and somatic diversity, especially of B-cell immunoglobulin repertoires, presents challenges for analysis requiring the development of specialized computational pipelines. We developed the REpertoire Sequencing TOolkit (pRESTO) for processing reads from high-throughput lymphocyte receptor studies. pRESTO processes raw sequences to produce error-corrected, sorted and annotated sequence sets, along with a wealth of metrics at each step. The toolkit supports multiplexed primer pools, single- or paired-end reads and emerging technologies that use single-molecule identifiers. pRESTO has been tested on data generated from Roche and Illumina platforms. It has a built-in capacity to parallelize the work between available processors and is able to efficiently process millions of sequences generated by typical high-throughput projects. AVAILABILITY AND IMPLEMENTATION: pRESTO is freely available for academic use. The software package and detailed tutorials may be downloaded from http://clip.med.yale.edu/presto. Jason A. Vander Heiden, Gur Yaari, Mohamed Uduman, Joel N. H. Stern, Kevin C. O'Connor, David A. Hafler, Francois Vigneault, Steven H. Kleinstein |
Bioinform. | 2 |
| 2010 | Optimizing Metapopulation Sustainability through a Checkerboard StrategyabstractThe persistence of a spatially structured population is determined by the rate of dispersal among habitat patches. If the local dynamic at the subpopulation level is extinction-prone, the system viability is maximal at intermediate connectivity where recolonization is allowed, but full synchronization that enables correlated extinction is forbidden. Here we developed and used an algorithm for agent-based simulations in order to study the persistence of a stochastic metapopulation. The effect of noise is shown to be dramatic, and the dynamics of the spatial population differs substantially from the predictions of deterministic models. This has been validated for the stochastic versions of the logistic map, the Ricker map and the Nicholson-Bailey host-parasitoid system. To analyze the possibility of extinction, previous studies were focused on the attractiveness (Lyapunov exponent) of stable solutions and the structure of their basin of attraction (dependence on initial population size). Our results suggest that these features are of secondary importance in the presence of stochasticity. Instead, optimal sustainability is achieved when decoherence is maximal. Individual-based simulations of metapopulations of different sizes, dimensions and noise types, show that the system's lifetime peaks when it displays checkerboard spatial patterns. This conclusion is supported by the results of a recently published Drosophila experiment. The checkerboard strategy provides a technique for the manipulation of migration rates (e.g., by constructing corridors) in order to affect the persistence of a metapopulation. It may be used in order to minimize the risk of extinction of an endangered species, or to maximize the efficiency of an eradication campaign. Yossi Ben Zion, Gur Yaari, Nadav M. Shnerb |
PLoS Comput. Biol. | 2 |