Sai T. Reddy

dblp:211/6662 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021
YearPublicationVenuePosition
2025 PLMFit: benchmarking transfer learning with protein language models for protein engineering
abstract
Protein language models (PLMs) have emerged as a useful resource for protein engineering applications. Transfer learning (TL) leverages pre-trained parameters to extract features to train machine learning models or adjust the weights of PLMs for novel tasks via fine-tuning (FT) through back-propagation. TL methods have shown potential for enhancing protein predictions performance when paired with PLMs, however there is a notable lack of comparative analyses that benchmark TL methods applied to state-of-the-art PLMs, identify optimal strategies for transferring knowledge and determine the most suitable approach for specific tasks. Here, we report PLMFit, a benchmarking study that combines, three state-of-the-art PLMs (ESM2, ProGen2, ProteinBert), with three TL methods (feature extraction, low-rank adaptation, bottleneck adapters) for five protein engineering datasets. We conducted over >3150 in silico experiments, altering PLM sizes and layers, TL hyperparameters and different training procedures. Our experiments reveal three key findings: (i) utilizing a partial fraction of PLM for TL does not detrimentally impact performance, (ii) the choice between feature extraction (FE) and fine-tuning is primarily dictated by the amount and diversity of data, and (iii) FT is most effective when generalization is necessary and only limited data is available. We provide PLMFit as an open-source software package, serving as a valuable resource for the scientific community to facilitate the FE and FT of PLMs for various applications.
Thomas Bikias, Evangelos Stamkopoulos, Sai T. Reddy
Briefings Bioinform.3
2025 Protein language model pseudolikelihoods capture features of in vivo B cell selection and evolution
abstract
B cell selection and evolution play crucial roles in dictating successful immune responses. Recent advancements in sequencing technologies and deep-learning strategies have paved the way for generating and exploiting an ever-growing wealth of antibody repertoire data. The self-supervised nature of protein language models (PLMs) has demonstrated the ability to learn complex representations of antibody sequences and has been leveraged for a wide range of applications including diagnostics, structural modeling, and antigen-specificity predictions. PLM-derived likelihoods have been used to improve antibody affinities in vitro, raising the question of whether PLMs can capture and predict features of B cell selection in vivo. Here, we explore how general and antibody-specific PLM-generated sequence pseudolikelihoods (SPs) relate to features of in vivo B cell selection such as expansion, isotype usage, and somatic hypermutation (SHM) at single-cell resolution. Our results demonstrate that the type of PLM and the region of the antibody input sequence significantly affect the generated SP. Contrary to previous in vitro reports, we observe a negative correlation between SPs and binding affinity, whereas repertoire features such as SHM and isotype usage were strongly correlated with SPs. By constructing evolutionary lineage trees of B cell clones from human and mouse repertoires, we observe that SHMs are routinely among the most likely mutations suggested by PLMs and that mutating residues have lower absolute likelihoods than conserved residues. Our findings highlight the potential of PLMs to predict features of antibody selection and further suggest their potential to assist in antibody discovery and engineering.
Daphne van Ginneken, Anamay Samant, Karlis Daga-Krumins, Wiona Glänzer, Andreas Agrafiotis, Evgenios Kladis, Sai T. Reddy, Alexander Yermanos
Briefings Bioinform.7
2025 Delineating inter- and intra-antibody repertoire evolution with AntibodyForests
abstract
MOTIVATION: The rapid advancements in immune repertoire sequencing, powered by single-cell technologies and artificial intelligence, have created unprecedented opportunities to study B cell evolution at a novel scale and resolution. However, fully leveraging these data requires specialized software capable of performing inter- and intra-repertoire analyses to unravel the complex dynamics of B cell repertoire evolution during immune responses. RESULTS: Here, we present AntibodyForests, software to infer B cell lineages, quantify inter- and intra-antibody repertoire evolution, and analyze somatic hypermutation using protein language models and protein structure. AVAILABILITY AND IMPLEMENTATION: This R package is available on CRAN and Github at https://github.com/alexyermanos/AntibodyForests, a vignette is available at https://cran.case.edu/web/packages/AntibodyForests/vignettes/AntibodyForests_vignette.html.
Daphne van Ginneken, Valentijn Tromp, Lucas Stalder, Tudor-Stefan Cotet, Sophie Bakker, Anamay Samant, Sai T. Reddy, Alexander Yermanos
Bioinform.7
2024 TCR clustering by contrastive learning on antigen specificity
abstract
Effective clustering of T-cell receptor (TCR) sequences could be used to predict their antigen-specificities. TCRs with highly dissimilar sequences can bind to the same antigen, thus making their clustering into a common antigen group a central challenge. Here, we develop TouCAN, a method that relies on contrastive learning and pretrained protein language models to perform TCR sequence clustering and antigen-specificity predictions. Following training, TouCAN demonstrates the ability to cluster highly dissimilar TCRs into common antigen groups. Additionally, TouCAN demonstrates TCR clustering performance and antigen-specificity predictions comparable to other leading methods in the field.
Margarita Pertseva, Oceane Follonier, Daniele Scarcella, Sai T. Reddy
Briefings Bioinform.4
2023 ePlatypus: an ecosystem for computational analysis of immunogenomics data
abstract
MOTIVATION: The maturation of systems immunology methodologies requires novel and transparent computational frameworks capable of integrating diverse data modalities in a reproducible manner. RESULTS: Here, we present the ePlatypus computational immunology ecosystem for immunogenomics data analysis, with a focus on adaptive immune repertoires and single-cell sequencing. ePlatypus is an open-source web-based platform and provides programming tutorials and an integrative database that helps elucidate signatures of B and T cell clonal selection. Furthermore, the ecosystem links novel and established bioinformatics pipelines relevant for single-cell immune repertoires and other aspects of computational immunology such as predicting ligand-receptor interactions, structural modeling, simulations, machine learning, graph theory, pseudotime, spatial transcriptomics, and phylogenetics. The ePlatypus ecosystem helps extract deeper insight in computational immunology and immunogenomics and promote open science. AVAILABILITY AND IMPLEMENTATION: Platypus code used in this manuscript can be found at github.com/alexyermanos/Platypus.
Tudor-Stefan Cotet, Andreas Agrafiotis, Victor Kreiner, Raphael Kuhn, Danielle Shlesinger, Marcos Manero-Carranza, Keywan Khodaverdi, Evgenios Kladis, Aurora Desideri Perea, Dylan Maassen-Veeters, Wiona Glänzer, Solène Massery, Lorenzo Guerci, Kai-Lin Hong, Jiami Han, Kostas Stiklioraitis, Vittoria Martinolli D'arcy, Raphael Dizerens, Samuel Kilchenmann, Lucas Stalder, Leon Nissen, Basil Vogelsanger, Stine Anzböck, Daria Laslo, Sophie Bakker, Melinda Kondorosy, Marco Venerito, Alejandro Sanz García, Isabelle Feller, Annette Oxenius, Sai T. Reddy, Alexander Yermanos
Bioinform.31
2020 Benchmarking immunoinformatic tools for the analysis of antibody repertoire sequences
abstract
SUMMARY: Antibody repertoires reveal insights into the biology of the adaptive immune system and empower diagnostics and therapeutics. There are currently multiple tools available for the annotation of antibody sequences. All downstream analyses such as choosing lead drug candidates depend on the correct annotation of these sequences; however, a thorough comparison of the performance of these tools has not been investigated. Here, we benchmark the performance of commonly used immunoinformatic tools, i.e. IMGT/HighV-QUEST, IgBLAST and MiXCR, in terms of reproducibility of annotation output, accuracy and speed using simulated and experimental high-throughput sequencing datasets.We analyzed changes in IMGT reference germline database in the last 10 years in order to assess the reproducibility of the annotation output. We found that only 73/183 (40%) V, D and J human genes were shared between the reference germline sets used by the tools. We found that the annotation results differed between tools. In terms of alignment accuracy, MiXCR had the highest average frequency of gene mishits, 0.02 mishit frequency and IgBLAST the lowest, 0.004 mishit frequency. Reproducibility in the output of complementarity determining three regions (CDR3 amino acids) ranged from 4.3% to 77.6% with preprocessed data. In addition, run time of the tools was assessed: MiXCR was the fastest tool for number of sequences processed per unit of time. These results indicate that immunoinformatic analyses greatly depend on the choice of bioinformatics tool. Our results support informed decision-making to immunoinformaticians based on repertoire composition and sequencing platforms. AVAILABILITY AND IMPLEMENTATION: All tools utilized in the paper are free for academic use. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Erand Smakaj, Lmar Babrak, Mats Ohlin, Mikhail Shugay, Bryan S. Briney, Deniz Tosoni, Christopher Galli, Vendi Grobelsek, Igor D'Angelo, Branden J. Olson, Sai T. Reddy, Victor Greiff, Johannes Trück, Susanna Marquez, William D. Lees, Enkelejda Miho
Bioinform.11
2020 immuneSIM: tunable multi-feature simulation of B- and T-cell receptor repertoires for immunoinformatics benchmarking
abstract
SUMMARY: B- and T-cell receptor repertoires of the adaptive immune system have become a key target for diagnostics and therapeutics research. Consequently, there is a rapidly growing number of bioinformatics tools for immune repertoire analysis. Benchmarking of such tools is crucial for ensuring reproducible and generalizable computational analyses. Currently, however, it remains challenging to create standardized ground truth immune receptor repertoires for immunoinformatics tool benchmarking. Therefore, we developed immuneSIM, an R package that allows the simulation of native-like and aberrant synthetic full-length variable region immune receptor sequences by tuning the following immune receptor features: (i) species and chain type (BCR, TCR, single and paired), (ii) germline gene usage, (iii) occurrence of insertions and deletions, (iv) clonal abundance, (v) somatic hypermutation and (vi) sequence motifs. Each simulated sequence is annotated by the complete set of simulation events that contributed to its in silico generation. immuneSIM permits the benchmarking of key computational tools for immune receptor analysis, such as germline gene annotation, diversity and overlap estimation, sequence similarity, network architecture, clustering analysis and machine learning methods for motif detection. AVAILABILITY AND IMPLEMENTATION: The package is available via https://github.com/GreiffLab/immuneSIM and on CRAN at https://cran.r-project.org/web/packages/immuneSIM. The documentation is hosted at https://immuneSIM.readthedocs.io. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Cédric R. Weber, Rahmad Akbar, Alexander Yermanos, Milena Pavlovic, Igor Snapkov, Geir Kjetil Sandve, Sai T. Reddy, Victor Greiff
Bioinform.7
2017 Comparison of methods for phylogenetic B-cell lineage inference using time-resolved antibody repertoire simulations (AbSim)
abstract
MOTIVATION: The evolution of antibody repertoires represents a hallmark feature of adaptive B-cell immunity. Recent advancements in high-throughput sequencing have dramatically increased the resolution to which we can measure the molecular diversity of antibody repertoires, thereby offering for the first time the possibility to capture the antigen-driven evolution of B cells. However, there does not exist a repertoire simulation framework yet that enables the comparison of commonly utilized phylogenetic methods with regard to their accuracy in inferring antibody evolution. RESULTS: Here, we developed AbSim, a time-resolved antibody repertoire simulation framework, which we exploited for testing the accuracy of methods for the phylogenetic reconstruction of B-cell lineages and antibody molecular evolution. AbSim enables the (i) simulation of intermediate stages of antibody sequence evolution and (ii) the modeling of immunologically relevant parameters such as duration of repertoire evolution, and the method and frequency of mutations. First, we validated that our repertoire simulation framework recreates replicates topological similarities observed in experimental sequencing data. Second, we leveraged Absim to show that current methods fail to a certain extent to predict the true phylogenetic tree correctly. Finally, we formulated simulation-validated guidelines for antibody evolution, which in the future will enable the development of accurate phylogenetic methods. AVAILABILITY AND IMPLEMENTATION: https://cran.r-project.org/web/packages/AbSim/index.html. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexander Yermanos, Victor Greiff, Nike Julia Krautler, Ulrike Menzel, Andreas Dounas, Enkelejda Miho, Annette Oxenius, Tanja Stadler, Sai T. Reddy
Bioinform.9