Victor Greiff

dblp:211/6517 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-2622-5032ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 7 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
YearPublicationVenuePosition
2026 PEPE: scalable extraction of multi-modal protein language model representations
abstract
SUMMARY: Protein language models (PLMs) capture intricate amino-acid dependencies, producing embeddings that encode rich structural, functional, and evolutionary information. Despite their potential, current extraction workflows rely on arbitrary choices, with respect to embedding layer, pooling, and padding, that frequently yield suboptimal representations for feature extraction and downstream analyses. Large-scale embedding generation is further limited by inefficiencies in computation and memory: (i) accumulating all model outputs in memory before writing to disk causes severe bottlenecks, and (ii) repeatedly embedding identical sequences to extract different modes introduces redundant computation and drastically reduces throughput and scalability. We introduce PEPE (Parallel Extraction for Protein Embeddings), a command-line tool and Python library that enables efficient, high-throughput, and multimodal extraction from protein language models. PEPE's parallelized and streaming-based architecture achieves runtimes several orders of magnitude faster than sequential approaches. Unlike conventional methods-whose peak memory usage scales linearly with output size and fails when memory capacity is exceeded-PEPE maintains stable, low memory consumption, enabling multimodal embedding extraction even beyond available RAM. PEPE supports a wide range of state-of-the-art and custom PLMs through a simple, flexible interface. By combining scalability, robustness, and ease of use, PEPE allows researchers to generate massive, information-rich embedding datasets efficiently, and facilitate the discovery of optimal representations for structural, functional, and evolutionary downstream tasks. By streamlining the generation of diverse embedding configurations, PEPE provides researchers with the necessary data to identify high-performing latent states for specific biological contexts without requiring additional computational resources. AVAILABILITY AND IMPLEMENTATION: PEPE is a command-line tool written in Python and published under MIT license. The source code and documentation are available at https://github.com/csi-greifflab/pepe-cli. PEPE is also available for installation from PyPI under https://pypi.org/project/pepe-cli and deposited on Zenodo at https://zenodo.org/records/20268104.
Jahn Zhong, Niccolò Cardente, Geir Kjetil Sandve, Habib Bashour, Maria Francesca Abbate, Victor Greiff
Bioinform.6
2026 TCR2HLA: Calibrated inference of HLA genotypes from TCR repertoires enables identification of immunologically relevant metaclonotypes
abstract
T cell receptors (TCRs) recognize peptides presented by polymorphic human leukocyte antigen (HLA) molecules, but HLA genotype data are often missing from TCR repertoire sequencing studies. To address this, we developed TCR2HLA, an open-source tool that infers HLA genotypes from TCRβ repertoires. Expanding on work linking public TRBV-CDR3 sequences to HLA genotypes, we incorporated "quasi-public" metaclonotypes - composed of rarer TCRβ sequences with shared amino acid features - enriched by HLA genotypes. Using four TCRβseq datasets from 3,150 individuals, we applied TRBV gene partitioning and locality-sensitive hashing to identify ~96,000 TCRβ features strongly associated with specific HLA alleles from 71M input TCRs. Binary HLA classifiers built with these features achieved high balanced accuracy (>0.9) across common HLA-A (9/12), B (9/12), C (6/13), DRB1 (11/11) alleles and prevalent DPA1/DPB1 (6/10), DQA1/DQB1 (8/17) heterodimers. We also introduced a high-sensitivity calibration to support predictions in samples with as few as 5,000 unique clonotypes. Calibrated predictions with confidence filtering improved reliability. Beyond genotype imputation, TCR2HLA enables the discovery of novel HLA- and exposure-associated TCRs, as shown by the identification of SARS-CoV-2 related TCRs in a large COVID-19 dataset lacking HLA data. TCR2HLA provides a scalable framework for bridging the gap between TCRseq data and HLA genotype for biomarker discovery.
Koshlan Mayer-Blackwell, Anastasia A. Minervina, Mikhail Pogorelyy, Puneet Rawat, Melanie R. Shapiro, Leeana D. Peters, Emily S. Ford, Amanda Posgai, Kasi Vegesana, Samuel Minot, David M. Koelle, Victor Greiff, Philip Bradley, Todd M. Brusko, Paul G. Thomas, Andrew Fiore-Gartland
PLoS Comput. Biol.12
2024 Incorporating probabilistic domain knowledge into deep multiple instance learning
abstract
Deep learning methods, including deep multiple instance learning methods, have been criticized for their limited ability to incorporate domain knowledge. A reason that knowledge incorporation is challenging in deep learning is that the models usually lack a mapping between their model components and the entities of the domain, making it a non-trivial task to incorporate probabilistic prior information. In this work, we show that such a mapping between domain entities and model components can be defined for a multiple instance learning setting and propose a framework DeeMILIP that encompasses multiple strategies to exploit this mapping for prior knowledge incorporation. We motivate and formalize these strategies from a probabilistic perspective. Experiments on an immune-based diagnostics case show that our proposed strategies allow to learn generalizable models even in settings with weak signals, limited dataset size, and limited compute.
Ghadi S. Al Hajj, Aliaksandr Hubin, Chakravarthi Kanduri, Milena Pavlovic, Knut D. Rand, Michael Widrich, Anne H. Schistad Solberg, Victor Greiff, Johan Pensar, Günter Klambauer, Geir Kjetil Sandve
ICML8
2024 Predictability of antigen binding based on short motifs in the antibody CDRH3
abstract
Adaptive immune receptors, such as antibodies and T-cell receptors, recognize foreign threats with exquisite specificity. A major challenge in adaptive immunology is discovering the rules governing immune receptor-antigen binding in order to predict the antigen binding status of previously unseen immune receptors. Many studies assume that the antigen binding status of an immune receptor may be determined by the presence of a short motif in the complementarity determining region 3 (CDR3), disregarding other amino acids. To test this assumption, we present a method to discover short motifs which show high precision in predicting antigen binding and generalize well to unseen simulated and experimental data. Our analysis of a mutagenesis-based antibody dataset reveals 11 336 position-specific, mostly gapped motifs of 3-5 amino acids that retain high precision on independently generated experimental data. Using a subset of only 178 motifs, a simple classifier was made that on the independently generated dataset outperformed a deep learning model proposed specifically for such datasets. In conclusion, our findings support the notion that for some antibodies, antigen binding may be largely determined by a short CDR3 motif. As more experimental data emerge, our methodology could serve as a foundation for in-depth investigations into antigen binding signals.
Lonneke Scheffer, Eric Emanuel Reber, Brij Bhushan Mehta, Milena Pavlovic, Maria Chernigovskaya, Eve Richardson, Rahmad Akbar, Fridtjof Lund-Johansen, Victor Greiff, Ingrid Hobæk Haff, Geir Kjetil Sandve
Briefings Bioinform.9
2022 TCRpower: quantifying the detection power of T-cell receptor sequencing with a novel computational pipeline calibrated by spike-in sequences
abstract
T-cell receptor (TCR) sequencing has enabled the development of innovative diagnostic tests for cancers, autoimmune diseases and other applications. However, the rarity of many T-cell clonotypes presents a detection challenge, which may lead to misdiagnosis if diagnostically relevant TCRs remain undetected. To address this issue, we developed TCRpower, a novel computational pipeline for quantifying the statistical detection power of TCR sequencing methods. TCRpower calculates the probability of detecting a TCR sequence as a function of several key parameters: in-vivo TCR frequency, T-cell sample count, read sequencing depth and read cutoff. To calibrate TCRpower, we selected unique TCRs of 45 T-cell clones (TCCs) as spike-in TCRs. We sequenced the spike-in TCRs from TCCs, together with TCRs from peripheral blood, using a 5' RACE protocol. The 45 spike-in TCRs covered a wide range of sample frequencies, ranging from 5 per 100 to 1 per 1 million. The resulting spike-in TCR read counts and ground truth frequencies allowed us to calibrate TCRpower. In our TCR sequencing data, we observed a consistent linear relationship between sample and sequencing read frequencies. We were also able to reliably detect spike-in TCRs with frequencies as low as one per million. By implementing an optimized read cutoff, we eliminated most of the falsely detected sequences in our data (TCR α-chain 99.0% and TCR β-chain 92.4%), thereby improving diagnostic specificity. TCRpower is publicly available and can be used to optimize future TCR sequencing experiments, and thereby enable reliable detection of disease-relevant TCRs for diagnostic applications.
Shiva Dahal-Koirala, Gabriel Balaban, Ralf Stefan Neumann, Lonneke Scheffer, Knut Erik Aslaksen Lundin, Victor Greiff, Ludvig Magne Sollid, Shuo-Wang Qiao, Geir Kjetil Sandve
Briefings Bioinform.6
2022 Machine-designed biotherapeutics: opportunities, feasibility and advantages of deep learning in computational antibody discovery
abstract
Antibodies are versatile molecular binders with an established and growing role as therapeutics. Computational approaches to developing and designing these molecules are being increasingly used to complement traditional lab-based processes. Nowadays, in silico methods fill multiple elements of the discovery stage, such as characterizing antibody-antigen interactions and identifying developability liabilities. Recently, computational methods tackling such problems have begun to follow machine learning paradigms, in many cases deep learning specifically. This paradigm shift offers improvements in established areas such as structure or binding prediction and opens up new possibilities such as language-based modeling of antibody repertoires or machine-learning-based generation of novel sequences. In this review, we critically examine the recent developments in (deep) machine learning approaches to therapeutic antibody design with implications for fully computational antibody design.
Wiktoria Wilman, Sonia Wróbel, Weronika Bielska, Piotr Deszynski, Pawel Dudzic, Igor Jaszczyszyn, Jedrzej Kaniewski, Jakub Mlokosiewicz, Anahita Rouyan, Tadeusz Satlawa, Victor Greiff, Konrad Krawczyk
Briefings Bioinform.12
2022 CompAIRR: ultra-fast comparison of adaptive immune receptor repertoires by exact and approximate sequence matching
abstract
MOTIVATION: Adaptive immune receptor (AIR) repertoires (AIRRs) record past immune encounters with exquisite specificity. Therefore, identifying identical or similar AIR sequences across individuals is a key step in AIRR analysis for revealing convergent immune response patterns that may be exploited for diagnostics and therapy. Existing methods for quantifying AIRR overlap scale poorly with increasing dataset numbers and sizes. To address this limitation, we developed CompAIRR, which enables ultra-fast computation of AIRR overlap, based on either exact or approximate sequence matching. RESULTS: CompAIRR improves computational speed 1000-fold relative to the state of the art and uses only one-third of the memory: on the same machine, the exact pairwise AIRR overlap of 104 AIRRs with 105 sequences is found in ∼17 min, while the fastest alternative tool requires 10 days. CompAIRR has been integrated with the machine learning ecosystem immuneML to speed up commonly used AIRR-based machine learning applications. AVAILABILITY AND IMPLEMENTATION: CompAIRR code and documentation are available at https://github.com/uio-bmi/compairr. Docker images are available at https://hub.docker.com/r/torognes/compairr. The code to replicate the synthetic datasets, scripts for benchmarking and creating figures, and all raw data underlying the figures are available at https://github.com/uio-bmi/compairr-benchmarking. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Torbjørn Rognes, Lonneke Scheffer, Victor Greiff, Geir Kjetil Sandve
Bioinform.3
2022 Access to ground truth at unconstrained size makes simulated data as indispensable as experimental data for bioinformatics methods development and benchmarking
abstract
The author instructions of this journal (OUP Bioinformatics) include several detailed requirements for manuscripts to be considered for publication (https://academic.oup.com/bioinformatics/pages/instructions_for_authors). Such requirements clarify expectations for authors and establish important peer-review standards. Importantly, it also allows for an open discussion of these requirements, which for a leading journal such as Bioinformatics contributes to shaping the fields of computational biology and bioinformatics. One of the requirements is that manuscripts presenting new methodology must include ‘actual biological data’ as opposed to simulated data. We find this requirement potentially counterproductive for several reasons and argue for emphasizing the complementarity of simulated and experimental data. Before we outline our argument, we would like to remark that terminology has connotations that may influence how different types of scientific evidence are valued. Specifically, we find the term ‘actual biological data’ to be problematic because primary data in the biological research domain is generated either from wet-lab experiments or computational simulations, both of which have their idiosyncrasies relative to the underlying biology (Leek et al., 2010). The aim, and often the very purpose, of computational methodology in bioinformatics is to model biological phenomena at a resolution and scale that transcends the current experimental state-of-the-art, to prepare for the arrival of biological data generation at a scale that allows better coverage of biological phenomena. Thus, a term such as ‘experimental calibration’ may be more appropriate to describe the use and purpose of experimental data in the majority of the bioinformatics literature, as opposed to the currently predominant denomination of ‘experimental validation’ (Jafari et al., 2021). Here, we argue that in the majority of bioinformatics settings, available experimental data do not have the size, resolution and sufficient set of controls (hereafter referred to as ‘limited data’) that would allow for rigorous method assessment. With limited data, performance estimates may be uncertain and sensitive to external factors such as parameter choices. This makes it challenging to judge whether observed improvements over previous methods are substantial, that is, biologically relevant, or merely the result of deliberate tuning of a method to perform particularly well on the experimental dataset(s) at hand (Castaldi et al., 2011; Salzberg, 1997). As a reviewer or critical reader, it is usually unfeasible to generate corroborating (or falsifying) experimental data with similar properties and thus not possible to rule out chance or tuning. Therefore, we argue that increased emphasis on experimental data may lead to insufficient and potentially misleading method evaluation. In contrast, simulation enables the generation of datasets of virtually unconstrained size, with precise control over introduced signals (ground truth) (Morris et al., 2019). This confers a critical reader the competence to challenge a reported assessment (Meyer and Birney, 2018) and rule out chance results and inappropriate tuning by simply generating new data from the same simulation process, thereby ensuring that conclusions can be meaningfully reproduced and cover a biologically relevant parameter range. Additionally, the specification of a simulation algorithm makes data assumptions for a method explicit and thus contributes to the transparency of a method both in terms of advances over the state of the art as well as its limitations. For example, in immunoinformatics of adaptive immunity, natural immune receptor sequence diversity, which is of the order of >1013, is routinely modeled using simulation frameworks for testing biological assumptions and the benchmarking of novel methods (Davidsen et al., 2019; Marcou et al., 2018; Pavlovic et al., 2021; Safonova et al., 2015; Weber et al., 2020). We stress that simulated data are only meaningful for bioinformatics method development and assessment if it reflects method-relevant underlying biology. The same criterion should be applied to experimental data. We agree that, if available at a sufficient scale, resolution and quality, experimental data are unsurpassed for assessing the capacity of bioinformatics methods to handle the types of signal complexities and data distributions that distinguishes bioinformatics method development from general informatics. Thus, novel methods should indeed be required to be evaluated on experimental data in domains where the available data are sufficiently robust to admit a rigorous assessment. However, we have several concerns with the quality of typically available experimental data for assessment purposes. A first concern is one of size. The performance of a method on a given dataset is always an uncertain estimate of its true performance on the underlying distribution that the observed data reflects. For many biological problems, available experimental datasets are so small that estimate uncertainties can easily be larger than any performance differences observed between competing methods. We fear that a strict journal requirement of employing experimental data may push authors to draw unwarranted conclusions from too small datasets and that reviewers may allow this to pass through due to the lack of good alternatives for the authors. An author’s requirement to include at least rudimentary measures of uncertainty for any reported performance measurement could alleviate these concerns (Walsh et al., 2021). A second concern is that experimental data are often available only for one particular problem setup. This leaves no opportunities to test the sensitivity of a method to variation in problem configuration (to assess how broadly it generalizes) or to test how it performs on data outside the training distribution (whether it is robust to domain shift). A third concern is that there is no possibility to know whether the patterns that a method extracts from experimental data reflect underlying causal relations or not. Furthermore, suboptimal study designs may introduce spurious correlations in datasets, and there is a risk that the best-performing methods at least partly exploit such artificial data patterns. Also, we fear that a strict, general requirement to assess novel methodology on experimental data may impede progress in method development in the many domains where available data are scarce. In data-scarce domains, we hold that authors should instead be urged to provide a rigorous assessment on simulated data, where authors should explicitly argue for the biological relevance based on either underlying mechanistic knowledge or by calibrating their simulation procedure with experimental data. In particular, we consider such experimentally calibrated simulation to often provide a better assessment of the capabilities of a method than the (in our opinion) too common reliance on anecdotal findings on small experimental datasets, which may reveal more about the ingenuity of the authors than of the proposed method. The discovery of novel biological knowledge does not in itself establish the usefulness or novelty of a new bioinformatics method—it is merely a corollary of, for example, a new method’s greater sensitivity, applicability to a wider parameter range or scalability. While it may be tempting to boost impact by combining novel methodology and novel biological findings in the same paper, this interferes with the assessment of methods on their own merit and thus undermines the selection pressures for the evolutionary process of method improvement in the field. Our view on the complementarity of simulated and experimental data is in line with the approach to method assessment taken in the machine learning field. Here, the evaluation of simulated data has always had a prominent role. Nowadays, the availability of large and well-curated databases such as ImageNet (Deng et al., 2009) and MNIST (Deng, 2012) make it natural to expect that novel methodologies are also assessed in such real-world data collections. However, when for instance the long short-term memory model was introduced in 1997 (Hochreiter and Schmidhuber, 1997), the authors explicitly asked in their paper ‘which tasks are appropriate to demonstrate the quality of a novel long-time-lag algorithm’ and answered their question based on a collection of exclusively synthetic datasets. Years later, improved data availability revealed that the model is indeed able to learn relevant patterns in a wide variety of real-world domains. The top-cited paper of the present journal (according to ISI web of science) (Li and Durbin, 2009) contains two sections in the Results section entitled ‘Evaluation on simulated data’ and ‘Evaluation on real data’. While we suggest that well-argued exceptions to the inclusion of experimental data assessment should be allowed for bioinformatics methods research, we can hardly think of any circumstance with compelling reasons for not including any assessment on simulated data. Since method developers should always have a conscious relationship to the data assumptions that they build their models and algorithms on, it should usually be straightforward to implement a simulation of data according to these same assumptions. This allows method developers to confirm that their method behaves as expected (e.g. identification of any bugs), it reveals to developers and readers the range of data parameters within which the method provides sensible results (method transparency) and allows developers or readers to reproduce assessments under identical or modified data assumptions (method reproducibility). We thus encourage basic assessment on simulated data to be considered an integral part of good bioinformatics method craftsmanship. Once simulated data have shown that a method works as intended, experimental data may be used to show that the software works on the field-specific experimental data formats and, ideally, recovers orthogonally validated biological or technological signals. We have argued that simulated and experimental data should be considered complementary and of equal importance for assessing methods in typical bioinformatics settings. They should both be strongly encouraged as part of a rigorous review process of novel methodology, where reviewers should ensure that a given paper exploits the best available data sources for assessment (be it experimental or well-established simulated datasets) and when necessary combines data sources for a comprehensive assessment. When available in high quantity, fidelity and generality, experimental data may ensure assessment validity—that a method handles relevant signals and noise profiles from the biological domain. But for many bioinformatics application areas, experimental data are not available at sufficient scale or annotation quality to allow conclusive assessment. Through full control over ground truth and unconstrained data size, simulated data may ensure assessment reliability—that the reported performance of a given method is representative and can be reproduced under the same or modified assumptions of the underlying data generating process. Importantly, sophisticated simulation processes, where signals and noise are calibrated by experimental data or knowledge of underlying mechanisms (Cao et al., 2021; Prakash et al., 2021; Schuler et al., 2017), allows methodology to be developed, assessed and improved early in a field so as to reach a good level of maturity at the time large-scale experimental data starts to become available. Well-calibrated simulation schemes may even be used to explore targeted hypotheses relating to complex biological systems in a way that can guide future experimental data collection (Azencott et al., 2017). In addition, simulation makes explicit the assumptions and layers of biological complexity understood so far and helps identify methodological errors or software bugs. In summary, we suggest that new bioinformatics methods should be shown to perform comparatively well on ground truth data of a size that allows reliable assessment, be it experimental or simulated. Method developers should be encouraged to make use of both simulated and experimental data, in complementary ways, to cover the multiple purposes of method assessment. When certain roles of assessments are not fully covered, method developers should be expected to provide compelling, explicit reasons—be it reasons for not including assessments involving simulated or experimental data. We would like to thank Michael Widrich and Günter Klambauer for their helpful suggestions. This work was supported by the Research Council of Norway [IKTPLUSS project (#311341 to G.K.S. and V.G.)]. Conflict of Interest: V.G. declares advisory board positions in aiNET GmbH, Enpicom B.V, Specifica Inc, Adaptyv Biosystems and EVQLV. V.G. is a consultant for Roche/Genentech. No new data were generated or analyzed in support of this research.
Geir Kjetil Sandve, Victor Greiff
Bioinform.2
2020 Modern Hopfield Networks and Attention for Immune Repertoire Classification
abstract
A central mechanism in machine learning is to identify, store, and recognize patterns. How to learn, access, and retrieve such patterns is crucial in Hopfield networks and the more recent transformer architectures. We show that the attention mechanism of transformer architectures is actually the update rule of modern Hopfield networks that can store exponentially many patterns. We exploit this high storage capacity of modern Hopfield networks to solve a challenging multiple instance learning (MIL) problem in computational biology: immune repertoire classification. In immune repertoire classification, a vast number of immune receptors are used to predict the immune status of an individual. This constitutes a MIL problem with an unprecedentedly massive number of instances, two orders of magnitude larger than currently considered problems, and with an extremely low witness rate. Accurate and interpretable machine learning methods solving this problem could pave the way towards new vaccines and therapies, which is currently a very relevant research topic intensified by the COVID-19 crisis. In this work, we present our novel method DeepRC that integrates transformer-like attention, or equivalently modern Hopfield networks, into deep learning architectures for massive MIL such as immune repertoire classification. We demonstrate that DeepRC outperforms all other methods with respect to predictive performance on large-scale experiments including simulated and real-world virus infection data and enables the extraction of sequence motifs that are connected to a given disease class. Source code and datasets: https://github.com/ml-jku/DeepRC
Michael Widrich, Bernhard Schäfl, Milena Pavlovic, Hubert Ramsauer, Lukas Gruber, Markus Holzleitner, Johannes Brandstetter, Geir Kjetil Sandve, Victor Greiff, Sepp Hochreiter, Günter Klambauer
NeurIPS9
2020 Benchmarking immunoinformatic tools for the analysis of antibody repertoire sequences
abstract
SUMMARY: Antibody repertoires reveal insights into the biology of the adaptive immune system and empower diagnostics and therapeutics. There are currently multiple tools available for the annotation of antibody sequences. All downstream analyses such as choosing lead drug candidates depend on the correct annotation of these sequences; however, a thorough comparison of the performance of these tools has not been investigated. Here, we benchmark the performance of commonly used immunoinformatic tools, i.e. IMGT/HighV-QUEST, IgBLAST and MiXCR, in terms of reproducibility of annotation output, accuracy and speed using simulated and experimental high-throughput sequencing datasets.We analyzed changes in IMGT reference germline database in the last 10 years in order to assess the reproducibility of the annotation output. We found that only 73/183 (40%) V, D and J human genes were shared between the reference germline sets used by the tools. We found that the annotation results differed between tools. In terms of alignment accuracy, MiXCR had the highest average frequency of gene mishits, 0.02 mishit frequency and IgBLAST the lowest, 0.004 mishit frequency. Reproducibility in the output of complementarity determining three regions (CDR3 amino acids) ranged from 4.3% to 77.6% with preprocessed data. In addition, run time of the tools was assessed: MiXCR was the fastest tool for number of sequences processed per unit of time. These results indicate that immunoinformatic analyses greatly depend on the choice of bioinformatics tool. Our results support informed decision-making to immunoinformaticians based on repertoire composition and sequencing platforms. AVAILABILITY AND IMPLEMENTATION: All tools utilized in the paper are free for academic use. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Erand Smakaj, Lmar Babrak, Mats Ohlin, Mikhail Shugay, Bryan S. Briney, Deniz Tosoni, Christopher Galli, Vendi Grobelsek, Igor D'Angelo, Branden J. Olson, Sai T. Reddy, Victor Greiff, Johannes Trück, Susanna Marquez, William D. Lees, Enkelejda Miho
Bioinform.12
2020 immuneSIM: tunable multi-feature simulation of B- and T-cell receptor repertoires for immunoinformatics benchmarking
abstract
SUMMARY: B- and T-cell receptor repertoires of the adaptive immune system have become a key target for diagnostics and therapeutics research. Consequently, there is a rapidly growing number of bioinformatics tools for immune repertoire analysis. Benchmarking of such tools is crucial for ensuring reproducible and generalizable computational analyses. Currently, however, it remains challenging to create standardized ground truth immune receptor repertoires for immunoinformatics tool benchmarking. Therefore, we developed immuneSIM, an R package that allows the simulation of native-like and aberrant synthetic full-length variable region immune receptor sequences by tuning the following immune receptor features: (i) species and chain type (BCR, TCR, single and paired), (ii) germline gene usage, (iii) occurrence of insertions and deletions, (iv) clonal abundance, (v) somatic hypermutation and (vi) sequence motifs. Each simulated sequence is annotated by the complete set of simulation events that contributed to its in silico generation. immuneSIM permits the benchmarking of key computational tools for immune receptor analysis, such as germline gene annotation, diversity and overlap estimation, sequence similarity, network architecture, clustering analysis and machine learning methods for motif detection. AVAILABILITY AND IMPLEMENTATION: The package is available via https://github.com/GreiffLab/immuneSIM and on CRAN at https://cran.r-project.org/web/packages/immuneSIM. The documentation is hosted at https://immuneSIM.readthedocs.io. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Cédric R. Weber, Rahmad Akbar, Alexander Yermanos, Milena Pavlovic, Igor Snapkov, Geir Kjetil Sandve, Sai T. Reddy, Victor Greiff
Bioinform.8
2017 Comparison of methods for phylogenetic B-cell lineage inference using time-resolved antibody repertoire simulations (AbSim)
abstract
MOTIVATION: The evolution of antibody repertoires represents a hallmark feature of adaptive B-cell immunity. Recent advancements in high-throughput sequencing have dramatically increased the resolution to which we can measure the molecular diversity of antibody repertoires, thereby offering for the first time the possibility to capture the antigen-driven evolution of B cells. However, there does not exist a repertoire simulation framework yet that enables the comparison of commonly utilized phylogenetic methods with regard to their accuracy in inferring antibody evolution. RESULTS: Here, we developed AbSim, a time-resolved antibody repertoire simulation framework, which we exploited for testing the accuracy of methods for the phylogenetic reconstruction of B-cell lineages and antibody molecular evolution. AbSim enables the (i) simulation of intermediate stages of antibody sequence evolution and (ii) the modeling of immunologically relevant parameters such as duration of repertoire evolution, and the method and frequency of mutations. First, we validated that our repertoire simulation framework recreates replicates topological similarities observed in experimental sequencing data. Second, we leveraged Absim to show that current methods fail to a certain extent to predict the true phylogenetic tree correctly. Finally, we formulated simulation-validated guidelines for antibody evolution, which in the future will enable the development of accurate phylogenetic methods. AVAILABILITY AND IMPLEMENTATION: https://cran.r-project.org/web/packages/AbSim/index.html. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexander Yermanos, Victor Greiff, Nike Julia Krautler, Ulrike Menzel, Andreas Dounas, Enkelejda Miho, Annette Oxenius, Tanja Stadler, Sai T. Reddy
Bioinform.2