VLDB 2026 Research / reviewers in the wild / expert
Björn Peters
dblp:68/4768 · also Bjoern Peters
· DBLP profile ↗
31ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 31 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Putting computational models of immunity to the test - An invited challenge to predict B.pertussis vaccination responsesabstractSystems vaccinology studies have been used to build computational models that predict individual vaccine responses and identify the factors contributing to differences in outcome. Comparing such models is challenging due to variability in study designs. To address this, we established a community resource to compare models predicting B. pertussis booster responses and generate experimental data for the explicit purpose of model evaluation. We here describe our second computational prediction challenge using this resource, where we benchmarked 49 algorithms from 53 scientists. We found that the most successful models stood out in their handling of nonlinearities, reducing large feature sets to representative subsets, and advanced data preprocessing. In contrast, we found that models adopted from literature that were developed to predict vaccine antibody responses in other settings performed poorly, reinforcing the need for purpose-built models. Overall, this demonstrates the value of purpose-generated datasets for rigorous and open model evaluations to identify features that improve the reliability and applicability of computational models in vaccine response prediction. Pramod Shinde, Lisa Willemsen, Minori Aoki, Saonli Basu, Julie G. Burel, Souradipto Ghosh Dastidar, Aidan Dunleavy, Tal Einav, Jamie Forschmiedt, Slim Fourati, William Gibson, Jason Greenbaum, Leying Guan, Weikang Guan, Jeremy P. Gygi, Brendan Ha, Joe Hou, Jason Hsiao, Yunda Huang, Rick Jansen, Bhargob Kakoty, Zhiyu Kang, James J. Kobie, Mari Kojima, Anna Konstorum, Jiyeun Lee, Sloan A. Lewis, Aixin Li, Eric F. Lock, Jarjapu Mahita, Marcus Mendes, Hailong Meng, Aidan Neher, Somayeh Nili, Lars Rønn Olsen, Shelby Orfield, James A. Overton, Nidhi Pai, Cokie Parker, Brian Qian, Mikkel Rasmussen, Joaquin Reyna, Eve Richardson, Sandra Safo, Josey Sorenson, Aparna Srinivasan, Nicola Thrupp, Rashmi Tippalagama, Raphael Trevizani, Steffen Ventz, Jiuzhou Wang, Cheng-Chang Wu, Ferhat Ay, Barry Grant, Steven H. Kleinstein, Björn Peters |
PLoS Comput. Biol. | 59 |
| 2023 | PEPMatch: a tool to identify short peptide sequence matches in large sets of proteinsabstractBACKGROUND: Numerous tools exist for biological sequence comparisons and search. One case of particular interest for immunologists is finding matches for linear peptide T cell epitopes, typically between 8 and 15 residues in length, in a large set of protein sequences. Both to find exact matches or matches that account for residue substitutions. The utility of such tools is critical in applications ranging from identifying conservation across viral epitopes, identifying putative epitope targets for allergens, and finding matches for cancer-associated neoepitopes to examine the role of tolerance in tumor recognition. RESULTS: We defined a set of benchmarks that reflect the different practical applications of short peptide sequence matching. We evaluated a suite of existing methods for speed and recall and developed a new tool, PEPMatch. The tool uses a deterministic k-mer mapping algorithm that preprocesses proteomes before searching, achieving a 50-fold increase in speed over methods such as the Basic Local Alignment Search Tool (BLAST) without compromising recall. PEPMatch's code and benchmark datasets are publicly available. CONCLUSIONS: PEPMatch offers significant speed and recall advantages for peptide sequence matching. While it is of immediate utility for immunologists, the developed benchmarking framework also provides a standard against which future tools can be evaluated for improvements. The tool is available at https://nextgen-tools.iedb.org , and the source code can be found at https://github.com/IEDB/PEPMatch . Daniel Marrama, William D. Chronister, Luise Westernberg, Randi Vita, Zeynep Kosaloglu-Yalçin, Alessandro Sette, Morten Nielsen 0001, Jason Greenbaum, Björn Peters |
BMC Bioinform. | 9 |
| 2022 | A comprehensive analysis of the IEDB MHC class-I automated benchmarkabstractIn 2014, the Immune Epitope Database automated benchmark was created to compare the performance of the MHC class I binding predictors. However, this is not a straightforward process due to the different and non-standardized outputs of the methods. Additionally, some methods are more restrictive regarding the HLA alleles and epitope sizes for which they predict binding affinities, while others are more comprehensive. To address how these problems impacted the ranking of the predictors, we developed an approach to assess the reliability of different metrics. We found that using percentile-ranked results improved the stability of the ranks and allowed the predictors to be reliably ranked despite not being evaluated on the same data. We also found that given the rate new data are incorporated into the benchmark, a new method must wait for at least 4 years to be ranked against the pre-existing methods. The best-performing tools with statistically indistinguishable scores in this benchmark were NetMHCcons, NetMHCpan4.0, ANN3.4, NetMHCpan3.0 and NetMHCpan2.8. The results of this study will be used to improve the evaluation and display of benchmark performance. We highly encourage anyone working on MHC binding predictions to participate in this benchmark to get an unbiased evaluation of their predictors. Raphael Trevizani, Jason Greenbaum, Alessandro Sette, Morten Nielsen 0001, Björn Peters |
Briefings Bioinform. | 6 |
| 2022 | Towards the prediction of non-peptidic epitopesabstractIn-silico methods for the prediction of epitopes can support and improve workflows for vaccine design, antibody production, and disease therapy. So far, the scope of B cell and T cell epitope prediction has been directed exclusively towards peptidic antigens. Nevertheless, various non-peptidic molecular classes can be recognized by immune cells. These compounds have not been systematically studied yet, and prediction approaches are lacking. The ability to predict the epitope activity of non-peptidic compounds could have vast implications; for example, for immunogenic risk assessment of the vast number of drugs and other xenobiotics. Here we present the first general attempt to predict the epitope activity of non-peptidic compounds using the Immune Epitope Database (IEDB) as a source for positive samples. The molecules stored in the Chemical Entities of Biological Interest (ChEBI) database were chosen as background samples. The molecules were clustered into eight homogeneous molecular groups, and classifiers were built for each cluster with the aim of separating the epitopes from the background. Different molecular feature encoding schemes and machine learning models were compared against each other. For those models where a high performance could be achieved based on simple decision rules, the molecular features were then further investigated. Additionally, the findings were used to build a web server that allows for the immunogenic investigation of non-peptidic molecules (http://tools-staging.iedb.org/np_epitope_predictor). The prediction quality was tested with samples from independent evaluation datasets, and the implemented method received noteworthy Receiver Operating Characteristic-Area Under Curve (ROC-AUC) values, ranging from 0.69-0.96 depending on the molecule cluster. Paul F. Zierep, Randi Vita, Nina Blazeska, Aurélien F. A. Moumbock, Jason Greenbaum, Björn Peters, Stefan Günther |
PLoS Comput. Biol. | 6 |
| 2021 | A metric for evaluating biological information in gene sets and its application to identify co-expressed gene clusters in PBMCabstractRecent technological advances have made the gathering of comprehensive gene expression datasets a commodity. This has shifted the limiting step of transcriptomic studies from the accumulation of data to their analyses and interpretation. The main problem in analyzing transcriptomics data is that the number of independent samples is typically much lower (<100) than the number of genes whose expression is quantified (typically >14,000). To address this, it would be desirable to reduce the gathered data's dimensionality without losing information. Clustering genes into discrete modules is one of the most commonly used tools to accomplish this task. While there are multiple clustering approaches, there is a lack of informative metrics available to evaluate the resultant clusters' biological quality. Here we present a metric that incorporates known ground truth gene sets to quantify gene clusters' biological quality derived from standard clustering techniques. The GECO (Ground truth Evaluation of Clustering Outcomes) metric demonstrates that quantitative and repeatable scoring of gene clusters is not only possible but computationally lightweight and robust. Unlike current methods, it allows direct comparison between gene clusters generated by different clustering techniques. It also reveals that current cluster analysis techniques often underestimate the number of clusters that should be formed from a dataset, which leads to fewer clusters of lower quality. As a test case, we applied GECO combined with k-means clustering to derive an optimal set of co-expressed gene modules derived from PBMC, which we show to be superior to previously generated modules generated on whole-blood. Overall, GECO provides a rational metric to test and compare different clustering approaches to analyze high-dimensional transcriptomic data. Jason Bennett, Mikhail Pomaznoy, Akul Singhania, Björn Peters |
PLoS Comput. Biol. | 4 |
| 2020 | Benchmarking predictions of MHC class I restricted T cell epitopes in a comprehensively studied model systemabstractT cell epitope candidates are commonly identified using computational prediction tools in order to enable applications such as vaccine design, cancer neoantigen identification, development of diagnostics and removal of unwanted immune responses against protein therapeutics. Most T cell epitope prediction tools are based on machine learning algorithms trained on MHC binding or naturally processed MHC ligand elution data. The ability of currently available tools to predict T cell epitopes has not been comprehensively evaluated. In this study, we used a recently published dataset that systematically defined T cell epitopes recognized in vaccinia virus (VACV) infected C57BL/6 mice (expressing H-2Db and H-2Kb), considering both peptides predicted to bind MHC or experimentally eluted from infected cells, making this the most comprehensive dataset of T cell epitopes mapped in a complex pathogen. We evaluated the performance of all currently publicly available computational T cell epitope prediction tools to identify these major epitopes from all peptides encoded in the VACV proteome. We found that all methods were able to improve epitope identification above random, with the best performance achieved by neural network-based predictions trained on both MHC binding and MHC ligand elution data (NetMHCPan-4.0 and MHCFlurry). Impressively, these methods were able to capture more than half of the major epitopes in the top N = 277 predictions within the N = 767,788 predictions made for distinct peptides of relevant lengths that can theoretically be encoded in the VACV proteome. These performance metrics provide guidance for immunologists as to which prediction methods to use, and what success rates are possible for epitope predictions when considering a highly controlled system of administered immunizations to inbred mice. In addition, this benchmark was implemented in an open and easy to reproduce format, providing developers with a framework for future comparisons against new tools. Sinu Paul, Nathan P. Croft, Anthony W. Purcell, David C. Tscharke, Alessandro Sette, Morten Nielsen 0001, Björn Peters |
PLoS Comput. Biol. | 7 |
| 2019 | Benchmark datasets of immune receptor-epitope structural complexesabstractBACKGROUND: The development of accurate epitope prediction tools is important in facilitating disease diagnostics, treatment and vaccine development. The advent of new approaches making use of antibody and TCR sequence information to predict receptor-specific epitopes have the potential to transform the epitope prediction field. Development and validation of these new generation of epitope prediction methods would benefit from regularly updated high-quality receptor-antigen complex datasets. RESULTS: To address the need for high-quality datasets to benchmark performance of these new generation of receptor-specific epitope prediction tools, a webserver called SCEptRe (Structural Complexes of Epitope-Receptor) was created. SCEptRe extracts weekly updated 3D complexes of antibody-antigen, TCR-pMHC and MHC-ligand from the Immune Epitope Database and clusters them based on antigen, receptor and epitope features to generate benchmark datasets. SCEptRe also provides annotated information such as CDR sequences and VDJ genes on the receptors. Users can generate custom datasets based by selecting thresholds for structural quality and clustering parameters (e.g. resolution, R-free factor, antigen or epitope sequence identity) based on their need. CONCLUSIONS: SCEptRe provides weekly updated, user-customized comprehensive benchmark datasets of immune receptor-epitope structural complexes. These datasets can be used to develop and benchmark performance of receptor-specific epitope prediction tools in the future. SCEptRe is freely accessible at http://tools.iedb.org/sceptre . Swapnil Mahajan, Martin Closter Jespersen, Kamilla Kjærgaard Jensen, Paolo Marcatili, Morten Nielsen 0001, Alessandro Sette, Björn Peters |
BMC Bioinform. | 8 |
| 2019 | Reporting and connecting cell type names and gating definitions through ontologiesabstractBACKGROUND: Human immunology studies often rely on the isolation and quantification of cell populations from an input sample based on flow cytometry and related techniques. Such techniques classify cells into populations based on the detection of a pattern of markers. The description of the cell populations targeted in such experiments typically have two complementary components: the description of the cell type targeted (e.g. 'T cells'), and the description of the marker pattern utilized (e.g. CD14-, CD3+). RESULTS: We here describe our attempts to use ontologies to cross-compare cell types and marker patterns (also referred to as gating definitions). We used a large set of such gating definitions and corresponding cell types submitted by different investigators into ImmPort, a central database for immunology studies, to examine the ability to parse gating definitions using terms from the Protein Ontology (PRO) and cell type descriptions, using the Cell Ontology (CL). We then used logical axioms from CL to detect discrepancies between the two. CONCLUSIONS: We suggest adoption of our proposed format for describing gating and cell type definitions to make comparisons easier. We also suggest a number of new terms to describe gating definitions in flow cytometry that are not based on molecular markers captured in PRO, but on forward- and side-scatter of light during data acquisition, which is more appropriate to capture in the Ontology for Biomedical Investigations (OBI). Finally, our approach results in suggestions on what logical axioms and new cell types could be considered for addition to the Cell Ontology. James A. Overton, Randi Vita, Patrick Dunn, Julie G. Burel, Syed Ahmad Chan Bukhari, Kei-Hoi Cheung, Steven H. Kleinstein, Alexander D. Diehl, Björn Peters |
BMC Bioinform. | 9 |
| 2018 | An automated benchmarking platform for MHC class II binding prediction methodsabstractMotivation: Computational methods for the prediction of peptide-MHC binding have become an integral and essential component for candidate selection in experimental T cell epitope discovery studies. The sheer amount of published prediction methods-and often discordant reports on their performance-poses a considerable quandary to the experimentalist who needs to choose the best tool for their research. Results: With the goal to provide an unbiased, transparent evaluation of the state-of-the-art in the field, we created an automated platform to benchmark peptide-MHC class II binding prediction tools. The platform evaluates the absolute and relative predictive performance of all participating tools on data newly entered into the Immune Epitope Database (IEDB) before they are made public, thereby providing a frequent, unbiased assessment of available prediction tools. The benchmark runs on a weekly basis, is fully automated, and displays up-to-date results on a publicly accessible website. The initial benchmark described here included six commonly used prediction servers, but other tools are encouraged to join with a simple sign-up procedure. Performance evaluation on 59 data sets composed of over 10 000 binding affinity measurements suggested that NetMHCIIpan is currently the most accurate tool, followed by NN-align and the IEDB consensus method. Availability and implementation: Weekly reports on the participating methods can be found online at: http://tools.iedb.org/auto_bench/mhcii/weekly/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Massimo Andreatta, Thomas Trolle, Jason Greenbaum, Björn Peters, Morten Nielsen 0001 |
Bioinform. | 5 |
| 2018 | ImmunomeBrowser: a tool to aggregate and visualize complex and heterogeneous epitopes in reference proteinsabstractMotivation: Datasets that are derived from different studies (e.g. MHC ligand elution, MHC binding, B/T cell epitope screening etc.) often vary in terms of experimental approaches, sizes of peptides tested, including partial and or nested overlapping peptides and in the number of donors tested. Results: We present a customized application of the Immune Epitope Database's ImmunomeBrowser tool, which can be used to effectively aggregate and visualize heterogeneous immunological data. User provided peptide sets and associated response data is mapped to a user-provided protein reference sequence. The output consists of tables and figures representing the aggregated data represented by a Response Frequency score and associated estimated confidence interval. This allows the user to visualizing regions associated with dominant responses and their boundaries. The results are presented both as a user interactive javascript based web interface and a tabular format in a selected reference sequence. Availability and implementation: The 'ImmunomeBrowser' has been a longstanding feature of the IEDB (http://www.iedb.org). The present application extends the use of this tool to work with user-provided datasets, rather than the output of IEDB queries. This new server version of the ImmunomeBrowser is freely accessible at http://tools.iedb.org/immunomebrowser/. Sandeep Kumar Dhanda, Randi Vita, Brendan Ha, Alba Grifoni, Björn Peters, Alessandro Sette |
Bioinform. | 5 |
| 2018 | GOnet: a tool for interactive Gene Ontology analysisabstractBACKGROUND: Biological interpretation of gene/protein lists resulting from -omics experiments can be a complex task. A common approach consists of reviewing Gene Ontology (GO) annotations for entries in such lists and searching for enrichment patterns. Unfortunately, there is a gap between machine-readable output of GO software and its human-interpretable form. This gap can be bridged by allowing users to simultaneously visualize and interact with term-term and gene-term relationships. RESULTS: We created the open-source GOnet web-application (available at http://tools.dice-database.org/GOnet/ ), which takes a list of gene or protein entries from human or mouse data and performs GO term annotation analysis (mapping of provided entries to GO subsets) or GO term enrichment analysis (scanning for GO categories overrepresented in the input list). The application is capable of producing parsable data formats and importantly, interactive visualizations of the GO analysis results. The interactive results allow exploration of genes and GO terms as a graph that depicts the natural hierarchy of the terms and retains relationships between terms and genes/proteins. As a result, GOnet provides insight into the functional interconnection of the submitted entries. CONCLUSIONS: The application can be used for GO analysis of any biological data sources resulting in gene/protein lists. It can be helpful for experimentalists as well as computational biologists working on biological interpretation of -omics data resulting in such lists. Mikhail Pomaznoy, Brendan Ha, Björn Peters |
BMC Bioinform. | 3 |
| 2018 | Putting benchmarks in their rightful place: The heart of computational biologyabstractResearch in computational biology has given rise to a vast number of methods developed to solve scientific problems.For areas in which many approaches exist, researchers have a hard time deciding which tool to select to address a scientific challenge, as essentially all publications introducing a new method will claim better performance than all others.Not all of these claims can be correct.Equally, for this same reason, developers struggle to demonstrate convincingly that they created a new and superior algorithm or implementation.Moreover, the developer community often has difficulty discerning which new approaches constitute true scientific advances for the field.The obvious answer to this conundrum is to develop benchmarks-meaning standard points of reference that facilitate evaluating the performance of different tools-allowing both users and developers to compare multiple tools in an unbiased fashion.Broadly speaking, benchmarks consist of input data that methods are meant to operate upon, expected output data against which tool output can be compared, a specification of metrics used to assess performance, and performance values of sets of tools that have been run through the benchmark.Developing good and comprehensive benchmarks, in which the performance metrics of each tool reflect its real-world utility, requires a significant effort.For highly competitive and established fields, such as protein structure predictions, community experiments evaluating the methods have been held periodically to provide blinded assessments of prediction performance.These blinded assessments are perhaps the gold standard on how benchmarks should be run.However, in most areas of computational biology, no such regular blinded contests are available.Instead, many tool developers end up generating their own benchmarks, which they publish alongside a newly developed tool to show its improved performance.The downside of this approach is that, if a new approach is developed in parallel to assembly of the benchmark on which it is evaluated, there is a strong selection bias encouraging the authors to report tool development approaches performing well against the benchmark compared to previous tools.This reporting bias makes most benchmarks that accompany newly developed tools questionable.Even if the authors are aware of this problem and take conscious steps to separate Björn Peters, Steven E. Brenner, Edwin Wang, Donna K. Slonim, Maricel G. Kann |
PLoS Comput. Biol. | 1 |
| 2015 | PEASE: predicting B-cell epitopes utilizing antibody sequenceabstractUNLABELLED: Antibody epitope mapping is a key step in understanding antibody-antigen recognition and is of particular interest for drug development, diagnostics and vaccine design. Most computational methods for epitope prediction are based on properties of the antigen sequence and/or structure, not taking into account the antibody for which the epitope is predicted. Here, we introduce PEASE, a web server predicting antibody-specific epitopes, utilizing the sequence of the antibody. The predictions are provided both at the residue level and as patches on the antigen structure. The tradeoff between recall and precision can be tuned by the user, by changing the default parameters. The results are provided as text and HTML files as well as a graph, and can be viewed on the antigen 3D structure. AVAILABILITY AND IMPLEMENTATION: PEASE is freely available on the web at www.ofranlab.org/PEASE. CONTACT: [email protected]. Inbal Sela-Culang, Shaul Ashkenazi, Björn Peters, Yanay Ofran |
Bioinform. | 3 |
| 2015 | Automated benchmarking of peptide-MHC class I binding predictionsabstractMOTIVATION: Numerous in silico methods predicting peptide binding to major histocompatibility complex (MHC) class I molecules have been developed over the last decades. However, the multitude of available prediction tools makes it non-trivial for the end-user to select which tool to use for a given task. To provide a solid basis on which to compare different prediction tools, we here describe a framework for the automated benchmarking of peptide-MHC class I binding prediction tools. The framework runs weekly benchmarks on data that are newly entered into the Immune Epitope Database (IEDB), giving the public access to frequent, up-to-date performance evaluations of all participating tools. To overcome potential selection bias in the data included in the IEDB, a strategy was implemented that suggests a set of peptides for which different prediction methods give divergent predictions as to their binding capability. Upon experimental binding validation, these peptides entered the benchmark study. RESULTS: The benchmark has run for 15 weeks and includes evaluation of 44 datasets covering 17 MHC alleles and more than 4000 peptide-MHC binding measurements. Inspection of the results allows the end-user to make educated selections between participating tools. Of the four participating servers, NetMHCpan performed the best, followed by ANN, SMM and finally ARB. AVAILABILITY AND IMPLEMENTATION: Up-to-date performance evaluations of each server can be found online at http://tools.iedb.org/auto_bench/mhci/weekly. All prediction tool developers are invited to participate in the benchmark. Sign-up instructions are available at http://tools.iedb.org/auto_bench/mhci/join. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas Trolle, Imir G. Metushi, Jason Greenbaum, John Sidney, Ole Lund, Alessandro Sette, Björn Peters, Morten Nielsen 0001 |
Bioinform. | 8 |
| 2014 | Dataset size and composition impact the reliability of performance benchmarks for peptide-MHC binding predictionsabstractBACKGROUND: It is important to accurately determine the performance of peptide:MHC binding predictions, as this enables users to compare and choose between different prediction methods and provides estimates of the expected error rate. Two common approaches to determine prediction performance are cross-validation, in which all available data are iteratively split into training and testing data, and the use of blind sets generated separately from the data used to construct the predictive method. In the present study, we have compared cross-validated prediction performances generated on our last benchmark dataset from 2009 with prediction performances generated on data subsequently added to the Immune Epitope Database (IEDB) which served as a blind set. RESULTS: We found that cross-validated performances systematically overestimated performance on the blind set. This was found not to be due to the presence of similar peptides in the cross-validation dataset. Rather, we found that small size and low sequence/affinity diversity of either training or blind datasets were associated with large differences in cross-validated vs. blind prediction performances. We use these findings to derive quantitative rules of how large and diverse datasets need to be to provide generalizable performance estimates. CONCLUSION: It has long been known that cross-validated prediction performance estimates often overestimate performance on independently generated blind set data. We here identify and quantify the specific factors contributing to this effect for MHC-I binding predictions. An increasing number of peptides for which MHC binding affinities are measured experimentally have been selected based on binding predictions and thus are less diverse than historic datasets sampling the entire sequence and affinity space, making them more difficult benchmark data sets. This has to be taken into account when comparing performance metrics between different benchmarks, and when deriving error estimates for predictions based on benchmark performance. John Sidney, Søren Buus, Alessandro Sette, Morten Nielsen 0001, Björn Peters |
BMC Bioinform. | 6 |
| 2013 | Properties of MHC Class I Presented Peptides That Enhance ImmunogenicityabstractT-cells have to recognize peptides presented on MHC molecules to be activated and elicit their effector functions. Several studies demonstrate that some peptides are more immunogenic than others and therefore more likely to be T-cell epitopes. We set out to determine which properties cause such differences in immunogenicity. To this end, we collected and analyzed a large set of data describing the immunogenicity of peptides presented on various MHC-I molecules. Two main conclusions could be drawn from this analysis: First, in line with previous observations, we showed that positions P4-6 of a presented peptide are more important for immunogenicity. Second, some amino acids, especially those with large and aromatic side chains, are associated with immunogenicity. This information was combined into a simple model that was used to demonstrate that immunogenicity is, to a certain extent, predictable. This model (made available at http://tools.iedb.org/immunogenicity/) was validated with data from two independent epitope discovery studies. Interestingly, with this model we could show that T-cells are equipped to better recognize viral than human (self) peptides. After the past successful elucidation of different steps in the MHC-I presentation pathway, the identification of variables that influence immunogenicity will be an important next step in the investigation of T-cell epitopes and our understanding of cellular immune responses. Jorg J. A. Calis, Matt Maybeno, Jason Greenbaum, Daniela Weiskopf, Aruna D. De Silva, Alessandro Sette, Can Kesmir, Björn Peters |
PLoS Comput. Biol. | 8 |
| 2013 | Positional Bias of MHC Class I Restricted T-Cell Epitopes in Viral Antigens Is Likely due to a Bias in ConservationabstractThe immune system rapidly responds to intracellular infections by detecting MHC class I restricted T-cell epitopes presented on infected cells. It was originally thought that viral peptides are liberated during constitutive protein turnover, but this conflicts with the observation that viral epitopes are detected within minutes of their synthesis even when their source proteins exhibit half-lives of days. The DRiPs hypothesis proposes that epitopes derive from Defective Ribosomal Products (DRiPs), rather than degradation of mature protein products. One potential source of DRiPs is premature translation termination. If this is a major source of DRiPs, this should be reflected in positional bias towards the N-terminus. By contrast, if downstream initiation is a major source of DRiPs, there should be positional bias towards the C-terminus. Here, we systematically assessed positional bias of epitopes in viral antigens, exploiting the large set of data available in the Immune Epitope Database and Analysis Resource. We show a statistically significant degree of positional skewing among epitopes; epitopes from both ends of antigens tend to be under-represented. Centric-skewing correlates with a bias towards class I binding peptides being over-represented in the middle, in parallel with a higher degree of evolutionary conservation. Jonathan W. Yewdell, Alessandro Sette, Björn Peters |
PLoS Comput. Biol. | 4 |
| 2012 | Structural Consensus among Antibodies Defines the Antigen Binding SiteabstractThe Complementarity Determining Regions (CDRs) of antibodies are assumed to account for the antigen recognition and binding and thus to contain also the antigen binding site. CDRs are typically discerned by searching for regions that are most different, in sequence or in structure, between different antibodies. Here, we show that ~20% of the antibody residues that actually bind the antigen fall outside the CDRs. However, virtually all antigen binding residues lie in regions of structural consensus across antibodies. Furthermore, we show that these regions of structural consensus which cover the antigen binding site are identifiable from the sequence of the antibody. Analyzing the predicted contribution of antigen binding residues to the stability of the antibody-antigen complex, we show that residues that fall outside of the traditionally defined CDRs are at least as important to antigen binding as residues within the CDRs, and in some cases, they are even more important energetically. Furthermore, antigen binding residues that fall outside of the structural consensus regions but within traditionally defined CDRs show a marginal energetic contribution to antigen binding. These findings allow for systematic and comprehensive identification of antigen binding sites, which can improve the understanding of antigenic interactions and may be useful in antibody engineering and B-cell epitope identification. Vered Kunik, Björn Peters, Yanay Ofran |
PLoS Comput. Biol. | 2 |
| 2011 | Cost sensitive hierarchical document classification to triage PubMed abstracts for manual curationabstractBACKGROUND: The Immune Epitope Database (IEDB) project manually curates information from published journal articles that describe immune epitopes derived from a wide variety of organisms and associated with different diseases. In the past, abstracts of scientific articles were retrieved by broad keyword queries of PubMed, and were classified as relevant (curatable) or irrelevant (not curatable) to the scope of the database by a Naïve Bayes classifier. The curatable abstracts were subsequently manually classified into categories corresponding to different disease domains. Over the past four years, we have examined how to further improve this approach in order to enhance classification performance and to reduce the need for manual intervention. RESULTS: Utilizing 89,884 abstracts classified by a domain expert as curatable or uncuratable, we found that a SVM classifier outperformed the previously used Naïve Bayes classifier for curatability predictions with an AUC of 0.899 and 0.854, respectively. Next, using a non-hierarchical and a hierarchical application of SVM classifiers trained on 22,833 curatable abstracts manually classified into three levels of disease specific categories we demonstrated that a hierarchical application of SVM classifiers outperformed non-hierarchical SVM classifiers for categorization. Finally, to optimize the hierarchical SVM classifiers' error profile for the curation process, cost sensitivity functions were developed to avoid serious misclassifications. We tested our design on a benchmark dataset of 1,388 references and achieved an overall category prediction accuracy of 94.4%, 93.9%, and 82.1% at the three levels of categorization, respectively. CONCLUSIONS: A hierarchical application of SVM algorithms with cost sensitive output weighting enabled high quality reference classification with few serious misclassifications. This enabled us to significantly reduce the manual component of abstract categorization. Our findings are relevant to other databases that are developing their own document classifier schema and the datasets we make available provide large scale real-life benchmark sets for method developers. Emily Seymour, Rohini Damle, Alessandro Sette, Björn Peters |
BMC Bioinform. | 4 |
| 2011 | Hematopoietic cell types: Prototype for a revised cell ontology
Alexander D. Diehl, Alison Deckhut Augustine, Judith A. Blake, Lindsay G. Cowell, Elizabeth S. Gold, Timothy A. Gondré-Lewis, Anna Maria Masci, Terrence F. Meehan, Penelope A. Morel, Anastasia Nijnik, Björn Peters, Bali Pulendran, Richard H. Scheuermann, Q. Alison Yao, Martin S. Zand, Chris Mungall |
J. Biomed. Informatics | 11 |
| 2010 | Peptide binding predictions for HLA DR, DP and DQ moleculesabstractBACKGROUND: MHC class II binding predictions are widely used to identify epitope candidates in infectious agents, allergens, cancer and autoantigens. The vast majority of prediction algorithms for human MHC class II to date have targeted HLA molecules encoded in the DR locus. This reflects a significant gap in knowledge as HLA DP and DQ molecules are presumably equally important, and have only been studied less because they are more difficult to handle experimentally. RESULTS: In this study, we aimed to narrow this gap by providing a large scale dataset of over 17,000 HLA-peptide binding affinities for a set of 11 HLA DP and DQ alleles. We also expanded our dataset for HLA DR alleles resulting in a total of 40,000 MHC class II binding affinities covering 26 allelic variants. Utilizing this dataset, we generated prediction tools utilizing several machine learning algorithms and evaluated their performance. CONCLUSION: We found that 1) prediction methodologies developed for HLA DR molecules perform equally well for DP or DQ molecules. 2) Prediction performances were significantly increased compared to previous reports due to the larger amounts of training data available. 3) The presence of homologous peptides between training and testing datasets should be avoided to give real-world estimates of prediction performance metrics, but the relative ranking of different predictors is largely unaffected by the presence of homologous peptides, and predictors intended for end-user applications should include all training data for maximum performance. 4) The recently developed NN-align prediction method significantly outperformed all other algorithms, including a naïve consensus based on all prediction methods. A new consensus method dropping the comparably weak ARB prediction method could outperform the NN-align method, but further research into how to best combine MHC class II binding predictions is required. Peng Wang 0014, John Sidney, Alessandro Sette, Ole Lund, Morten Nielsen 0001, Björn Peters |
BMC Bioinform. | 7 |
| 2010 | Towards the Automotive HMI of the Future: Overview of the AIDE-Integrated Project ResultsabstractThe Adaptive Integrated Driver-vehicle interfacE (AIDE) is an integrated project funded by the European Commission in the Sixth Framework Programme. The project, which involves 31 partners from the European automotive industry and academia, deals with behavioral and technical issues related to automotive human-machine interface (HMI) design, with a particular focus on integration and adaptation. The project involves tightly integrated empirical research, driver-behavior modeling, and methodological and technological development. This paper provides an overview of the AIDE Sub-Project 3 results dealing with the design, development, and integration of the AIDE system in three prototype vehicles, together with the evaluation results of the trials. Angelos Amditis, Luisa Andreone, Katia Pagle, Gustav Markkula, Enrica Deregibus, Maria Romera Rué, Francesco Bellotti, Andreas Engelsberg, Rino Brouwer, Björn Peters, Alessandro De Gloria |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2009 | Derivation of an amino acid similarity matrix for peptide: MHC binding and its application as a Bayesian priorabstractBACKGROUND: Experts in peptide:MHC binding studies are often able to estimate the impact of a single residue substitution based on a heuristic understanding of amino acid similarity in an experimental context. Our aim is to quantify this measure of similarity to improve peptide:MHC binding prediction methods. This should help compensate for holes and bias in the sequence space coverage of existing peptide binding datasets. RESULTS: Here, a novel amino acid similarity matrix (PMBEC) is directly derived from the binding affinity data of combinatorial peptide mixtures. Like BLOSUM62, this matrix captures well-known physicochemical properties of amino acid residues. However, PMBEC differs markedly from existing matrices in cases where residue substitution involves a reversal of electrostatic charge. To demonstrate its usefulness, we have developed a new peptide:MHC class I binding prediction method, using the matrix as a Bayesian prior. We show that the new method can compensate for missing information on specific residues in the training data. We also carried out a large-scale benchmark, and its results indicate that prediction performance of the new method is comparable to that of the best neural network based approaches for peptide:MHC class I binding. CONCLUSION: A novel amino acid similarity matrix has been derived for peptide:MHC binding interactions. One prominent feature of the matrix is that it disfavors substitution of residues with opposite charges. Given that the matrix was derived from experimentally determined peptide:MHC binding affinity measurements, this feature is likely shared by all peptide:protein interactions. In addition, we have demonstrated the usefulness of the matrix as a Bayesian prior in an improved scoring-matrix based peptide:MHC class I prediction method. A software implementation of the method is available at: http://www.mhc-pathway.net/smmpmbec. John Sidney, Clemencia Pinilla, Alessandro Sette, Björn Peters |
BMC Bioinform. | 5 |
| 2008 | ElliPro: a new structure-based tool for the prediction of antibody epitopesabstractBACKGROUND: Reliable prediction of antibody, or B-cell, epitopes remains challenging yet highly desirable for the design of vaccines and immunodiagnostics. A correlation between antigenicity, solvent accessibility, and flexibility in proteins was demonstrated. Subsequently, Thornton and colleagues proposed a method for identifying continuous epitopes in the protein regions protruding from the protein's globular surface. The aim of this work was to implement that method as a web-tool and evaluate its performance on discontinuous epitopes known from the structures of antibody-protein complexes. RESULTS: Here we present ElliPro, a web-tool that implements Thornton's method and, together with a residue clustering algorithm, the MODELLER program and the Jmol viewer, allows the prediction and visualization of antibody epitopes in a given protein sequence or structure. ElliPro has been tested on a benchmark dataset of discontinuous epitopes inferred from 3D structures of antibody-protein complexes. In comparison with six other structure-based methods that can be used for epitope prediction, ElliPro performed the best and gave an AUC value of 0.732, when the most significant prediction was considered for each protein. Since the rank of the best prediction was at most in the top three for more than 70% of proteins and never exceeded five, ElliPro is considered a useful research tool for identifying antibody epitopes in protein antigens. ElliPro is available at http://tools.immuneepitope.org/tools/ElliPro. CONCLUSION: The results from ElliPro suggest that further research on antibody epitopes considering more features that discriminate epitopes from non-epitopes may further improve predictions. As ElliPro is based on the geometrical properties of protein structure and does not require training, it might be more generally applied for predicting different types of protein-protein interactions. Julia V. Ponomarenko, Huynh-Hoa Bui, Nicholas Fusseder, Philip E. Bourne, Alessandro Sette, Björn Peters |
BMC Bioinform. | 7 |
| 2008 | Quantitative Predictions of Peptide Binding to Any HLA-DR Molecule of Known Sequence: NetMHCIIpanabstractCD4 positive T helper cells control many aspects of specific immunity. These cells are specific for peptides derived from protein antigens and presented by molecules of the extremely polymorphic major histocompatibility complex (MHC) class II system. The identification of peptides that bind to MHC class II molecules is therefore of pivotal importance for rational discovery of immune epitopes. HLA-DR is a prominent example of a human MHC class II. Here, we present a method, NetMHCIIpan, that allows for pan-specific predictions of peptide binding to any HLA-DR molecule of known sequence. The method is derived from a large compilation of quantitative HLA-DR binding events covering 14 of the more than 500 known HLA-DR alleles. Taking both peptide and HLA sequence information into account, the method can generalize and predict peptide binding also for HLA-DR molecules where experimental data is absent. Validation of the method includes identification of endogenously derived HLA class II ligands, cross-validation, leave-one-molecule-out, and binding motif identification for hitherto uncharacterized HLA-DR molecules. The validation shows that the method can successfully predict binding for HLA-DR molecules-even in the absence of specific data for the particular molecule in question. Moreover, when compared to TEPITOPE, currently the only other publicly available prediction method aiming at providing broad HLA-DR allelic coverage, NetMHCIIpan performs equivalently for alleles included in the training of TEPITOPE while outperforming TEPITOPE on novel alleles. We propose that the method can be used to identify those hitherto uncharacterized alleles, which should be addressed experimentally in future updates of the method to cover the polymorphism of HLA-DR most efficiently. We thus conclude that the presented method meets the challenge of keeping up with the MHC polymorphism discovery rate and that it can be used to sample the MHC "space," enabling a highly efficient iterative process for improving MHC class II binding predictions. Morten Nielsen 0001, Claus Lundegaard, Thomas Blicher, Björn Peters, Alessandro Sette, Sune Justesen, Søren Buus, Ole Lund |
PLoS Comput. Biol. | 4 |
| 2008 | A Systematic Assessment of MHC Class II Peptide Binding Predictions and Evaluation of a Consensus ApproachabstractThe identification of MHC class II restricted peptide epitopes is an important goal in immunological research. A number of computational tools have been developed for this purpose, but there is a lack of large-scale systematic evaluation of their performance. Herein, we used a comprehensive dataset consisting of more than 10,000 previously unpublished MHC-peptide binding affinities, 29 peptide/MHC crystal structures, and 664 peptides experimentally tested for CD4+ T cell responses to systematically evaluate the performances of publicly available MHC class II binding prediction tools. While in selected instances the best tools were associated with AUC values up to 0.86, in general, class II predictions did not perform as well as historically noted for class I predictions. It appears that the ability of MHC class II molecules to bind variable length peptides, which requires the correct assignment of peptide binding cores, is a critical factor limiting the performance of existing prediction tools. To improve performance, we implemented a consensus prediction approach that combines methods with top performances. We show that this consensus approach achieved best overall performance. Finally, we make the large datasets used publicly available as a benchmark to facilitate further development of MHC class II binding peptide prediction methods. Peng Wang 0014, John Sidney, Courtney Dow, Bianca Mothé, Alessandro Sette, Björn Peters |
PLoS Comput. Biol. | 6 |
| 2007 | Automating document classification for the Immune Epitope DatabaseabstractBACKGROUND: The Immune Epitope Database contains information on immune epitopes curated manually from the scientific literature. Like similar projects in other knowledge domains, significant effort is spent on identifying which articles are relevant for this purpose. RESULTS: We here report our experience in automating this process using Naïve Bayes classifiers trained on 20,910 abstracts classified by domain experts. Improvements on the basic classifier performance were made by a) utilizing information stored in PubMed beyond the abstract itself b) applying standard feature selection criteria and c) extracting domain specific feature patterns that e.g. identify peptides sequences. We have implemented the classifier into the curation process determining if abstracts are clearly relevant, clearly irrelevant, or if no certain classification can be made, in which case the abstracts are manually classified. Testing this classification scheme on an independent dataset, we achieve 95% sensitivity and specificity in the 51.1% of abstracts that were automatically classified. CONCLUSION: By implementing text classification, we have sped up the reference selection process without sacrificing sensitivity or specificity of the human expert classification. This study provides both practical recommendations for users of text classification tools, as well as a large dataset which can serve as a benchmark for tool developers. Peng Wang 0014, Alexander A. Morgan, Alessandro Sette, Björn Peters |
BMC Bioinform. | 5 |
| 2006 | Curation of complex, context-dependent immunological dataabstractBACKGROUND: The Immune Epitope Database and Analysis Resource (IEDB) is dedicated to capturing, housing and analyzing complex immune epitope related data http://www.immuneepitope.org. DESCRIPTION: To identify and extract relevant data from the scientific literature in an efficient and accurate manner, novel processes were developed for manual and semi-automated annotation. CONCLUSION: Formalized curation strategies enable the processing of a large volume of context-dependent data, which are now available to the scientific community in an accessible and transparent format. The experiences described herein are applicable to other databases housing complex biological data and requiring a high level of curation expertise. Randi Vita, Kerrie Vaughan, Laura Zarebski, Nima Salimi, Ward Fleri, Howard Grey, Muthu Sathiamurthy, John Mokili, Huynh-Hoa Bui, Philip E. Bourne, Julia V. Ponomarenko, Romulo de Castro Jr., Russell K. Chan, John Sidney, Stephen S. Wilson, Scott Stewart, Scott Way, Björn Peters, Alessandro Sette |
BMC Bioinform. | 18 |
| 2006 | A Community Resource Benchmarking Predictions of Peptide Binding to MHC-I MoleculesabstractRecognition of peptides bound to major histocompatibility complex (MHC) class I molecules by T lymphocytes is an essential part of immune surveillance. Each MHC allele has a characteristic peptide binding preference, which can be captured in prediction algorithms, allowing for the rapid scan of entire pathogen proteomes for peptide likely to bind MHC. Here we make public a large set of 48,828 quantitative peptide-binding affinity measurements relating to 48 different mouse, human, macaque, and chimpanzee MHC class I alleles. We use this data to establish a set of benchmark predictions with one neural network method and two matrix-based prediction methods extensively utilized in our groups. In general, the neural network outperforms the matrix-based predictions mainly due to its ability to generalize even on a small amount of data. We also retrieved predictions from tools publicly available on the internet. While differences in the data used to generate these predictions hamper direct comparisons, we do conclude that tools based on combinatorial peptide libraries perform remarkably well. The transparent prediction evaluation on this dataset provides tool developers with a benchmark for comparison of newly developed prediction methods. In addition, to generate and evaluate our own prediction methods, we have established an easily extensible web-based prediction framework that allows automated side-by-side comparisons of prediction methods implemented by experts. This is an advance over the current practice of tool developers having to generate reference predictions themselves, which can lead to underestimating the performance of prediction methods they are not as familiar with as their own. The overall goal of this effort is to provide a transparent prediction evaluation allowing bioinformaticians to identify promising features of prediction methods and providing guidance to immunologists regarding the reliability of prediction tools. Björn Peters, Huynh-Hoa Bui, Sune Pletscher-Frankild, Morten Nielsen 0001, Claus Lundegaard, Emrah Kostem, Derek Basch, Kasper Lamberth, Mikkel Harndahl, Ward Fleri, Stephen S. Wilson, John Sidney, Ole Lund, Søren Buus, Alessandro Sette |
PLoS Comput. Biol. | 1 |
| 2005 | Generating quantitative models describing the sequence specificity of biological processes with the stabilized matrix methodabstractBACKGROUND: Many processes in molecular biology involve the recognition of short sequences of nucleic-or amino acids, such as the binding of immunogenic peptides to major histocompatibility complex (MHC) molecules. From experimental data, a model of the sequence specificity of these processes can be constructed, such as a sequence motif, a scoring matrix or an artificial neural network. The purpose of these models is two-fold. First, they can provide a summary of experimental results, allowing for a deeper understanding of the mechanisms involved in sequence recognition. Second, such models can be used to predict the experimental outcome for yet untested sequences. In the past we reported the development of a method to generate such models called the Stabilized Matrix Method (SMM). This method has been successfully applied to predicting peptide binding to MHC molecules, peptide transport by the transporter associated with antigen presentation (TAP) and proteasomal cleavage of protein sequences. RESULTS: Herein we report the implementation of the SMM algorithm as a publicly available software package. Specific features determining the type of problems the method is most appropriate for are discussed. Advantageous features of the package are: (1) the output generated is easy to interpret, (2) input and output are both quantitative, (3) specific computational strategies to handle experimental noise are built in, (4) the algorithm is designed to effectively handle bounded experimental data, (5) experimental data from randomized peptide libraries and conventional peptides can easily be combined, and (6) it is possible to incorporate pair interactions between positions of a sequence. CONCLUSION: Making the SMM method publicly available enables bioinformaticians and experimental biologists to easily access it, to compare its performance to other prediction methods, and to extend it to other applications. Björn Peters, Alessandro Sette |
BMC Bioinform. | 1 |
| 2003 | Examining the independent binding assumption for binding of peptide epitopes to MHC-I moleculesabstractMOTIVATION: Various methods have been proposed to predict the binding affinities of peptides to Major Histocompatibility Complex class I (MHC-I) molecules based on experimental binding data. They can be classified into two groups: (1) AIB methods that assume independent contributions of all peptide positions to the binding to MHC-I molecule (e.g. scoring matrices) and (2) general methods which can take into account interactions between different positions (e.g. artificial neural networks). We aim to compare the prediction accuracies of these methods, and quantify the impact of interactions between peptide positions. RESULTS: We compared several previously published and widely used methods and discovered that the best AIB methods gave significantly better predictions than three previously published general methods, possibly due to the lack of a sufficient training data for the general methods. The best results, however, were achieved with our newly developed general method, which combined a matrix describing independent binding with pair coefficients describing pair-wise interactions between peptide positions. The pair coefficients consistently but only slightly improved prediction accuracy, and were much smaller than the matrix entries. This explains why neglecting them-as is done in AIB methods-can still lead to good predictions. AVAILABILITY: The new prediction model is implemented at http://zlab.bu.edu/SMM. The underlying matrix and pair coefficients are also available as supplementary materials. Björn Peters, Weiwei Tong, John Sidney, Alessandro Sette, Zhiping Weng |
Bioinform. | 1 |