VLDB 2026 Research / reviewers in the wild / expert
Sai T. Reddy
dblp:211/6662
· DBLP profile ↗
8ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PLMFit: benchmarking transfer learning with protein language models for protein engineeringabstractProtein language models (PLMs) have emerged as a useful resource for protein engineering applications. Transfer learning (TL) leverages pre-trained parameters to extract features to train machine learning models or adjust the weights of PLMs for novel tasks via fine-tuning (FT) through back-propagation. TL methods have shown potential for enhancing protein predictions performance when paired with PLMs, however there is a notable lack of comparative analyses that benchmark TL methods applied to state-of-the-art PLMs, identify optimal strategies for transferring knowledge and determine the most suitable approach for specific tasks. Here, we report PLMFit, a benchmarking study that combines, three state-of-the-art PLMs (ESM2, ProGen2, ProteinBert), with three TL methods (feature extraction, low-rank adaptation, bottleneck adapters) for five protein engineering datasets. We conducted over >3150 in silico experiments, altering PLM sizes and layers, TL hyperparameters and different training procedures. Our experiments reveal three key findings: (i) utilizing a partial fraction of PLM for TL does not detrimentally impact performance, (ii) the choice between feature extraction (FE) and fine-tuning is primarily dictated by the amount and diversity of data, and (iii) FT is most effective when generalization is necessary and only limited data is available. We provide PLMFit as an open-source software package, serving as a valuable resource for the scientific community to facilitate the FE and FT of PLMs for various applications. Thomas Bikias, Evangelos Stamkopoulos, Sai T. Reddy |
Briefings Bioinform. | 3 |
| 2025 | Protein language model pseudolikelihoods capture features of in vivo B cell selection and evolutionabstractB cell selection and evolution play crucial roles in dictating successful immune responses. Recent advancements in sequencing technologies and deep-learning strategies have paved the way for generating and exploiting an ever-growing wealth of antibody repertoire data. The self-supervised nature of protein language models (PLMs) has demonstrated the ability to learn complex representations of antibody sequences and has been leveraged for a wide range of applications including diagnostics, structural modeling, and antigen-specificity predictions. PLM-derived likelihoods have been used to improve antibody affinities in vitro, raising the question of whether PLMs can capture and predict features of B cell selection in vivo. Here, we explore how general and antibody-specific PLM-generated sequence pseudolikelihoods (SPs) relate to features of in vivo B cell selection such as expansion, isotype usage, and somatic hypermutation (SHM) at single-cell resolution. Our results demonstrate that the type of PLM and the region of the antibody input sequence significantly affect the generated SP. Contrary to previous in vitro reports, we observe a negative correlation between SPs and binding affinity, whereas repertoire features such as SHM and isotype usage were strongly correlated with SPs. By constructing evolutionary lineage trees of B cell clones from human and mouse repertoires, we observe that SHMs are routinely among the most likely mutations suggested by PLMs and that mutating residues have lower absolute likelihoods than conserved residues. Our findings highlight the potential of PLMs to predict features of antibody selection and further suggest their potential to assist in antibody discovery and engineering. Daphne van Ginneken, Anamay Samant, Karlis Daga-Krumins, Wiona Glänzer, Andreas Agrafiotis, Evgenios Kladis, Sai T. Reddy, Alexander Yermanos |
Briefings Bioinform. | 7 |
| 2025 | Delineating inter- and intra-antibody repertoire evolution with AntibodyForestsabstractMOTIVATION: The rapid advancements in immune repertoire sequencing, powered by single-cell technologies and artificial intelligence, have created unprecedented opportunities to study B cell evolution at a novel scale and resolution. However, fully leveraging these data requires specialized software capable of performing inter- and intra-repertoire analyses to unravel the complex dynamics of B cell repertoire evolution during immune responses. RESULTS: Here, we present AntibodyForests, software to infer B cell lineages, quantify inter- and intra-antibody repertoire evolution, and analyze somatic hypermutation using protein language models and protein structure. AVAILABILITY AND IMPLEMENTATION: This R package is available on CRAN and Github at https://github.com/alexyermanos/AntibodyForests, a vignette is available at https://cran.case.edu/web/packages/AntibodyForests/vignettes/AntibodyForests_vignette.html. Daphne van Ginneken, Valentijn Tromp, Lucas Stalder, Tudor-Stefan Cotet, Sophie Bakker, Anamay Samant, Sai T. Reddy, Alexander Yermanos |
Bioinform. | 7 |
| 2024 | TCR clustering by contrastive learning on antigen specificityabstractEffective clustering of T-cell receptor (TCR) sequences could be used to predict their antigen-specificities. TCRs with highly dissimilar sequences can bind to the same antigen, thus making their clustering into a common antigen group a central challenge. Here, we develop TouCAN, a method that relies on contrastive learning and pretrained protein language models to perform TCR sequence clustering and antigen-specificity predictions. Following training, TouCAN demonstrates the ability to cluster highly dissimilar TCRs into common antigen groups. Additionally, TouCAN demonstrates TCR clustering performance and antigen-specificity predictions comparable to other leading methods in the field. Margarita Pertseva, Oceane Follonier, Daniele Scarcella, Sai T. Reddy |
Briefings Bioinform. | 4 |
| 2023 | ePlatypus: an ecosystem for computational analysis of immunogenomics dataabstractMOTIVATION: The maturation of systems immunology methodologies requires novel and transparent computational frameworks capable of integrating diverse data modalities in a reproducible manner. RESULTS: Here, we present the ePlatypus computational immunology ecosystem for immunogenomics data analysis, with a focus on adaptive immune repertoires and single-cell sequencing. ePlatypus is an open-source web-based platform and provides programming tutorials and an integrative database that helps elucidate signatures of B and T cell clonal selection. Furthermore, the ecosystem links novel and established bioinformatics pipelines relevant for single-cell immune repertoires and other aspects of computational immunology such as predicting ligand-receptor interactions, structural modeling, simulations, machine learning, graph theory, pseudotime, spatial transcriptomics, and phylogenetics. The ePlatypus ecosystem helps extract deeper insight in computational immunology and immunogenomics and promote open science. AVAILABILITY AND IMPLEMENTATION: Platypus code used in this manuscript can be found at github.com/alexyermanos/Platypus. Tudor-Stefan Cotet, Andreas Agrafiotis, Victor Kreiner, Raphael Kuhn, Danielle Shlesinger, Marcos Manero-Carranza, Keywan Khodaverdi, Evgenios Kladis, Aurora Desideri Perea, Dylan Maassen-Veeters, Wiona Glänzer, Solène Massery, Lorenzo Guerci, Kai-Lin Hong, Jiami Han, Kostas Stiklioraitis, Vittoria Martinolli D'arcy, Raphael Dizerens, Samuel Kilchenmann, Lucas Stalder, Leon Nissen, Basil Vogelsanger, Stine Anzböck, Daria Laslo, Sophie Bakker, Melinda Kondorosy, Marco Venerito, Alejandro Sanz García, Isabelle Feller, Annette Oxenius, Sai T. Reddy, Alexander Yermanos |
Bioinform. | 31 |
| 2020 | Benchmarking immunoinformatic tools for the analysis of antibody repertoire sequencesabstractSUMMARY: Antibody repertoires reveal insights into the biology of the adaptive immune system and empower diagnostics and therapeutics. There are currently multiple tools available for the annotation of antibody sequences. All downstream analyses such as choosing lead drug candidates depend on the correct annotation of these sequences; however, a thorough comparison of the performance of these tools has not been investigated. Here, we benchmark the performance of commonly used immunoinformatic tools, i.e. IMGT/HighV-QUEST, IgBLAST and MiXCR, in terms of reproducibility of annotation output, accuracy and speed using simulated and experimental high-throughput sequencing datasets.We analyzed changes in IMGT reference germline database in the last 10 years in order to assess the reproducibility of the annotation output. We found that only 73/183 (40%) V, D and J human genes were shared between the reference germline sets used by the tools. We found that the annotation results differed between tools. In terms of alignment accuracy, MiXCR had the highest average frequency of gene mishits, 0.02 mishit frequency and IgBLAST the lowest, 0.004 mishit frequency. Reproducibility in the output of complementarity determining three regions (CDR3 amino acids) ranged from 4.3% to 77.6% with preprocessed data. In addition, run time of the tools was assessed: MiXCR was the fastest tool for number of sequences processed per unit of time. These results indicate that immunoinformatic analyses greatly depend on the choice of bioinformatics tool. Our results support informed decision-making to immunoinformaticians based on repertoire composition and sequencing platforms. AVAILABILITY AND IMPLEMENTATION: All tools utilized in the paper are free for academic use. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Erand Smakaj, Lmar Babrak, Mats Ohlin, Mikhail Shugay, Bryan S. Briney, Deniz Tosoni, Christopher Galli, Vendi Grobelsek, Igor D'Angelo, Branden J. Olson, Sai T. Reddy, Victor Greiff, Johannes Trück, Susanna Marquez, William D. Lees, Enkelejda Miho |
Bioinform. | 11 |
| 2020 | immuneSIM: tunable multi-feature simulation of B- and T-cell receptor repertoires for immunoinformatics benchmarkingabstractSUMMARY: B- and T-cell receptor repertoires of the adaptive immune system have become a key target for diagnostics and therapeutics research. Consequently, there is a rapidly growing number of bioinformatics tools for immune repertoire analysis. Benchmarking of such tools is crucial for ensuring reproducible and generalizable computational analyses. Currently, however, it remains challenging to create standardized ground truth immune receptor repertoires for immunoinformatics tool benchmarking. Therefore, we developed immuneSIM, an R package that allows the simulation of native-like and aberrant synthetic full-length variable region immune receptor sequences by tuning the following immune receptor features: (i) species and chain type (BCR, TCR, single and paired), (ii) germline gene usage, (iii) occurrence of insertions and deletions, (iv) clonal abundance, (v) somatic hypermutation and (vi) sequence motifs. Each simulated sequence is annotated by the complete set of simulation events that contributed to its in silico generation. immuneSIM permits the benchmarking of key computational tools for immune receptor analysis, such as germline gene annotation, diversity and overlap estimation, sequence similarity, network architecture, clustering analysis and machine learning methods for motif detection. AVAILABILITY AND IMPLEMENTATION: The package is available via https://github.com/GreiffLab/immuneSIM and on CRAN at https://cran.r-project.org/web/packages/immuneSIM. The documentation is hosted at https://immuneSIM.readthedocs.io. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cédric R. Weber, Rahmad Akbar, Alexander Yermanos, Milena Pavlovic, Igor Snapkov, Geir Kjetil Sandve, Sai T. Reddy, Victor Greiff |
Bioinform. | 7 |
| 2017 | Comparison of methods for phylogenetic B-cell lineage inference using time-resolved antibody repertoire simulations (AbSim)abstractMOTIVATION: The evolution of antibody repertoires represents a hallmark feature of adaptive B-cell immunity. Recent advancements in high-throughput sequencing have dramatically increased the resolution to which we can measure the molecular diversity of antibody repertoires, thereby offering for the first time the possibility to capture the antigen-driven evolution of B cells. However, there does not exist a repertoire simulation framework yet that enables the comparison of commonly utilized phylogenetic methods with regard to their accuracy in inferring antibody evolution. RESULTS: Here, we developed AbSim, a time-resolved antibody repertoire simulation framework, which we exploited for testing the accuracy of methods for the phylogenetic reconstruction of B-cell lineages and antibody molecular evolution. AbSim enables the (i) simulation of intermediate stages of antibody sequence evolution and (ii) the modeling of immunologically relevant parameters such as duration of repertoire evolution, and the method and frequency of mutations. First, we validated that our repertoire simulation framework recreates replicates topological similarities observed in experimental sequencing data. Second, we leveraged Absim to show that current methods fail to a certain extent to predict the true phylogenetic tree correctly. Finally, we formulated simulation-validated guidelines for antibody evolution, which in the future will enable the development of accurate phylogenetic methods. AVAILABILITY AND IMPLEMENTATION: https://cran.r-project.org/web/packages/AbSim/index.html. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander Yermanos, Victor Greiff, Nike Julia Krautler, Ulrike Menzel, Andreas Dounas, Enkelejda Miho, Annette Oxenius, Tanja Stadler, Sai T. Reddy |
Bioinform. | 9 |