Pietro Sormanni

dblp:213/6485 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0002-6228-2221ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Representation and self-supervised learning · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Improving Antibody Humanness Prediction using Patent Data · ICML 2024
Machine learning › Representation and self-supervised learning › contrastive learning
weakly-supervised contrastive learning
0.812024
Improving Antibody Humanness Prediction using Patent Data · ICML 2024
Bioinformatics and computational biology › protein design
antibody design
0.812024
Improving Antibody Humanness Prediction using Patent Data · ICML 2024
Bioinformatics and computational biology › immunoinformatics
paratope prediction
0.312018
Parapred: antibody paratope prediction using convolutional and recurrent neural networks · Bioinform. 2018
Bioinformatics and computational biology › structural biology
protein structure and function
0.312018
Parapred: antibody paratope prediction using convolutional and recurrent neural networks · Bioinform. 2018

Methods — techniques the papers use, named apart from their topics

multi-stage training · 1.5multi-loss training · 1.5cross-entropy loss · 1.5rigid docking · 0.3recurrent neural network · 0.3convolutional neural network · 0.3
YearPublicationVenuePosition
2024 Improving Antibody Humanness Prediction using Patent Data
abstract
We investigate the potential of patent data for improving the antibody humanness prediction using a multi-stage, multi-loss training process. Humanness serves as a proxy for the immunogenic response to antibody therapeutics, one of the major causes of attrition in drug discovery and a challenging obstacle for their use in clinical settings. We pose the initial learning stage as a weakly-supervised contrastive-learning problem, where each antibody sequence is associated with possibly multiple identifiers of function and the objective is to learn an encoder that groups them according to their patented properties. We then freeze a part of the contrastive encoder and continue training it on the patent data using the cross-entropy loss to predict the humanness score of a given antibody sequence. We illustrate the utility of the patent data and our approach by performing inference on three different immunogenicity datasets, unseen during training. Our empirical results demonstrate that the learned model consistently outperforms the alternative baselines and establishes new state-of-the-art on five out of six inference tasks, irrespective of the used metric.
Talip Ucar, Aubin Ramon, Dino Oglic, Rebecca Croasdale-Wood, Tom Diethe, Pietro Sormanni
ICML6
2023 Sequence-based prediction of pH-dependent protein solubility using CamSol
abstract
Solubility is a property of central importance for the use of proteins in research in molecular and cell biology and in applications in biotechnology and medicine. Since experimental methods for measuring protein solubility are material intensive and time consuming, computational methods have recently emerged to enable the rapid and inexpensive screening of solubility for large libraries of proteins, as it is routinely required in development pipelines. Here, we describe the development of one such method to include in the predictions the effect of the pH on solubility. We illustrate the resulting pH-dependent predictions on a variety of antibodies and other proteins to demonstrate that these predictions achieve an accuracy comparable with that of experimental methods. We make this method publicly available at https://www-cohsoftware.ch.cam.ac.uk/index.php/camsolph, as the version 3.0 of CamSol.
Marc Oeller, Ryan Kang, Rosie Bell, Hannes Ausserwöger, Pietro Sormanni, Michele Vendruscolo
Briefings Bioinform.5
2019 A chemical kinetic basis for measuring translation initiation and elongation rates from ribosome profiling data
abstract
Analysis methods based on simulations and optimization have been previously developed to estimate relative translation rates from next-generation sequencing data. Translation involves molecules and chemical reactions, hence bioinformatics methods consistent with the laws of chemistry and physics are more likely to produce accurate results. Here, we derive simple equations based on chemical kinetic principles to measure the translation-initiation rate, transcriptome-wide elongation rate, and individual codon translation rates from ribosome profiling experiments. Our methods reproduce the known rates from ribosome profiles generated from detailed simulations of translation. By applying our methods to data from S. cerevisiae and mouse embryonic stem cells, we find that the extracted rates reproduce expected correlations with various molecular properties, and we also find that mouse embryonic stem cells have a global translation speed of 5.2 AA/s, in agreement with previous reports that used other approaches. Our analysis further reveals that a codon can exhibit up to 26-fold variability in its translation rate depending upon its context within a transcript. This broad distribution means that the average translation rate of a codon is not representative of the rate at which most instances of that codon are translated, and it suggests that translational regulation might be used by cells to a greater degree than previously thought.
Ajeet K. Sharma, Pietro Sormanni, Nabeel Ahmed, Prajwal Ciryam, Ulrike A. Friedrich, Günter Kramer, Edward P. O'Brien
PLoS Comput. Biol.2
2018 Parapred: antibody paratope prediction using convolutional and recurrent neural networks
abstract
Motivation: Antibodies play essential roles in the immune system of vertebrates and are powerful tools in research and diagnostics. While hypervariable regions of antibodies, which are responsible for binding, can be readily identified from their amino acid sequence, it remains challenging to accurately pinpoint which amino acids will be in contact with the antigen (the paratope). Results: In this work, we present a sequence-based probabilistic machine learning algorithm for paratope prediction, named Parapred. Parapred uses a deep-learning architecture to leverage features from both local residue neighbourhoods and across the entire sequence. The method significantly improves on the current state-of-the-art methodology, and only requires a stretch of amino acid sequence corresponding to a hypervariable region as an input, without any information about the antigen. We further show that our predictions can be used to improve both speed and accuracy of a rigid docking algorithm. Availability and implementation: The Parapred method is freely available as a webserver at http://www-mvsoftware.ch.cam.ac.uk/and for download at https://github.com/eliberis/parapred. Supplementary information: Supplementary information is available at Bioinformatics online.
Edgar Liberis, Petar Velickovic, Pietro Sormanni, Michele Vendruscolo, Pietro Liò
Bioinform.3