VLDB 2026 Research / reviewers in the wild / expert
Gabriele Orlando
dblp:161/2883
· DBLP profile ↗
11ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-5935-5258ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FoldX force field revisited, an improved versionabstractMOTIVATION: The FoldX force field was originally validated with a database of 1000 mutants at a time when there were few high-resolution structures. Here, we have manually curated a database of 5556 mutants affecting protein stability, resulting in 2484 highly confident mutations denominated FoldX stability dataset (FSD), represented in non-redundant X-ray structures with <2.5 Å resolution, not involving duplicates, metals, or prosthetic groups. Using this database, we have created a new version of the FoldX force field by introducing pi stacking, pH dependency for all charged residues, improving aromatic-aromatic interactions, modifying the Ncap contribution and α-helix dipole, recalibrating the side-chain entropy of methionine, adjusting the H-bond parameters, and modifying the solvation contribution of tryptophan and others. RESULTS: These changes have led to significant improvements for the prediction of specific mutants involving the above residues/interactions and a statistically significant increase of FoldX predictions, as well as for the majority of the 20 aa. Removing all training sets data from FSD [Validation FoldX Stability Dataset (VFSD) dataset] resulted in improved predictions from R = 0.693 (RMSE = 1.277 kcal/mol) to R = 0.706 (RMSE = 1.252 kcal/mol) when compared with the previously released version. FoldX achieves 95% accuracy considering an error of ±0.85 kcal/mol in prediction and an area under the curve = 0.78 for the VFSD, predicting the sign of the energy change upon mutation. AVAILABILITY AND IMPLEMENTATION: FoldX versions 4.1 and 5.1 are freely available for academics at https://foldxsuite.crg.eu/. Javier Delgado, Raul Reche, Damiano Cianferoni, Gabriele Orlando, Rob van der Kant, Frederic Rousseau 0001, Joost Schymkowitz, Luis Serrano |
Bioinform. | 4 |
| 2025 | In silico identification of archaeal DNA-binding proteinsabstractMOTIVATION: The rapid advancement of next-generation sequencing technologies has generated an immense volume of genetic data. However, these data are unevenly distributed, with well-studied organisms being disproportionately represented, while other organisms, such as from archaea, remain significantly underexplored. The study of archaea is particularly challenging due to the extreme environments they inhabit and the difficulties associated with culturing them in the laboratory. Despite these challenges, archaea likely represent a crucial evolutionary link between eukaryotic and prokaryotic organisms, and their investigation could shed light on the early stages of life on Earth. Yet, a significant portion of archaeal proteins are annotated with limited or inaccurate information. Among the various classes of archaeal proteins, DNA-binding proteins are of particular importance. While they represent a large portion of every known proteome, their identification in archaea is complicated by the substantial evolutionary divergence between archaeal and the other better studied organisms. RESULTS: To address the challenges of identifying DNA-binding proteins in archaea, we developed Xenusia, a neural network-based tool capable of screening entire archaeal proteomes to identify DNA-binding proteins. Xenusia has proven effective across diverse datasets, including metagenomics data, successfully identifying novel DNA-binding proteins, with experimental validation of its predictions. AVAILABILITY AND IMPLEMENTATION: Xenusia is available as a PyPI package, with source code accessible at https://github.com/grogdrinker/xenusia, and as a Google Colab web server application at xenusia.ipynb. Linus Donvil, Joëlle A. J. Housmans, Eveline Peeters, Wim F. Vranken, Gabriele Orlando |
Bioinform. | 5 |
| 2025 | Charting the structure-sequence landscape of light chain amyloidsabstractMOTIVATION: Light chain amyloidosis is a disease where misfolded antibody light chains (LCs) form toxic amyloid fibrils, leading to organ damage. Although LC overproduction occurs in all cases, only certain individuals develop the disease, suggesting that specific LC sequences and properties drive amyloid formation. This process is complex, involving both protein sequence and environmental factors, but mutations that destabilize the LC fold are linked to amyloid aggregation. Despite the significance of the disease, our understanding of LC fibril formation remains limited due to the lack of extensive data and technical challenges in studying amyloid structures. To address this, a tool is needed to compare unknown LC sequences with known structures and predict which amyloids are likely to adopt new conformations, guiding experimental investigations. RESULTS: HMMSTUFF addresses this by using a Hidden Markov Model to generate similarity scores between LC sequences and existing PDB templates, eventually modeling the LC amyloid structures similar enough to known templates. HMMSTUFF on one side expands our understanding of LC amyloid fibril conformations, and on the other highlights the gaps in our current knowledge of LC structural space. AVAILABILITY AND IMPLEMENTATION: HMMSTUFF is available as pypi package and as source code at https://github.com/grogdrinker/hmmstuff. Gabriele Orlando, Rodrigo Gallardo, Alicia Colla, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 1 |
| 2024 | Integrating physics in deep learning algorithms: a force field as a PyTorch moduleabstractMOTIVATION: Deep learning algorithms applied to structural biology often struggle to converge to meaningful solutions when limited data is available, since they are required to learn complex physical rules from examples. State-of-the-art force-fields, however, cannot interface with deep learning algorithms due to their implementation. RESULTS: We present MadraX, a forcefield implemented as a differentiable PyTorch module, able to interact with deep learning algorithms in an end-to-end fashion. AVAILABILITY AND IMPLEMENTATION: MadraX documentation, together with tutorials and installation guide, is available at madrax.readthedocs.io. Gabriele Orlando, Luis Serrano, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 1 |
| 2021 | In silico prediction of in vitro protein liquid-liquid phase separation experiments outcomes with multi-head neural attentionabstractMOTIVATION: Proteins able to undergo liquid-liquid phase separation (LLPS) in vivo and in vitro are drawing a lot of interest, due to their functional relevance for cell life. Nevertheless, the proteome-scale experimental screening of these proteins seems unfeasible, because besides being expensive and time-consuming, LLPS is heavily influenced by multiple environmental conditions such as concentration, pH and temperature, thus requiring a combinatorial number of experiments for each protein. RESULTS: To overcome this problem, we propose a neural network model able to predict the LLPS behavior of proteins given specified experimental conditions, effectively predicting the outcome of in vitro experiments. Our model can be used to rapidly screen proteins and experimental conditions searching for LLPS, thus reducing the search space that needs to be covered experimentally. We experimentally validate Droppler's prediction on the TAR DNA-binding protein in different experimental conditions, showing the consistency of its predictions. AVAILABILITY AND IMPLEMENTATION: A python implementation of Droppler is available at https://bitbucket.org/grogdrinker/droppler. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniele Raimondi, Gabriele Orlando, Emiel Michiels, Donya Pakravan, Anna Bratek-Skicki, Ludo Van Den Bosch, Yves Moreau, Frederic Rousseau 0001, Joost Schymkowitz |
Bioinform. | 2 |
| 2020 | Accurate prediction of protein beta-aggregation with generalized statistical potentialsabstractMOTIVATION: Protein beta-aggregation is an important but poorly understood phenomena involved in diseases as well as in beneficial physiological processes. However, while this task has been investigated for over 50 years, very little is known about its mechanisms of action. Moreover, the identification of regions involved in aggregation is still an open problem and the state-of-the-art methods are often inadequate in real case applications. RESULTS: In this article we present AgMata, an unsupervised tool for the identification of such regions from amino acidic sequence based on a generalized definition of statistical potentials that includes biophysical information. The tool outperforms the state-of-the-art methods on two different benchmarks. As case-study, we applied our tool to human ataxin-3, a protein involved in Machado-Joseph disease. Interestingly, AgMata identifies aggregation-prone residues that share the very same structural environment. Additionally, it successfully predicts the outcome of in vitro mutagenesis experiments, identifying point mutations that lead to an alteration of the aggregation propensity of the wild-type ataxin-3. AVAILABILITY AND IMPLEMENTATION: A python implementation of the tool is available at https://bitbucket.org/bio2byte/agmata. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriele Orlando, Alexandra Silva 0004, Sandra Macedo-Ribeiro, Daniele Raimondi, Wim F. Vranken |
Bioinform. | 1 |
| 2020 | Insight into the protein solubility driving forces with neural attentionabstractProtein solubility is a key aspect for many biotechnological, biomedical and industrial processes, such as the production of active proteins and antibodies. In addition, understanding the molecular determinants of the solubility of proteins may be crucial to shed light on the molecular mechanisms of diseases caused by aggregation processes such as amyloidosis. Here we present SKADE, a novel Neural Network protein solubility predictor and we show how it can provide novel insight into the protein solubility mechanisms, thanks to its neural attention architecture. First, we show that SKADE positively compares with state of the art tools while using just the protein sequence as input. Then, thanks to the neural attention mechanism, we use SKADE to investigate the patterns learned during training and we analyse its decision process. We use this peculiarity to show that, while the attention profiles do not correlate with obvious sequence aspects such as biophysical properties of the aminoacids, they suggest that N- and C-termini are the most relevant regions for solubility prediction and are predictive for complex emergent properties such as aggregation-prone regions involved in beta-amyloidosis and contact density. Moreover, SKADE is able to identify mutations that increase or decrease the overall solubility of the protein, allowing it to be used to perform large scale in-silico mutagenesis of proteins in order to maximize their solubility. Daniele Raimondi, Gabriele Orlando, Piero Fariselli, Yves Moreau |
PLoS Comput. Biol. | 2 |
| 2019 | Computational identification of prion-like RNA-binding proteins that form liquid phase-separated condensatesabstractMOTIVATION: Eukaryotic cells contain different membrane-delimited compartments, which are crucial for the biochemical reactions necessary to sustain cell life. Recent studies showed that cells can also trigger the formation of membraneless organelles composed by phase-separated proteins to respond to various stimuli. These condensates provide new ways to control the reactions and phase-separation proteins (PSPs) are thus revolutionizing how cellular organization is conceived. The small number of experimentally validated proteins, and the difficulty in discovering them, remain bottlenecks in PSPs research. RESULTS: Here we present PSPer, the first in-silico screening tool for prion-like RNA-binding PSPs. We show that it can prioritize PSPs among proteins containing similar RNA-binding domains, intrinsically disordered regions and prions. PSPer is thus suitable to screen proteomes, identifying the most likely PSPs for further experimental investigation. Moreover, its predictions are fully interpretable in the sense that it assigns specific functional regions to the predicted proteins, providing valuable information for experimental investigation of targeted mutations on these regions. Finally, we show that it can estimate the ability of artificially designed proteins to form condensates (r=-0.87), thus providing an in-silico screening tool for protein design experiments. AVAILABILITY AND IMPLEMENTATION: PSPer is available at bio2byte.com/psp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriele Orlando, Daniele Raimondi, Francesco Tabaro, Francesco Codicè, Yves Moreau, Wim F. Vranken |
Bioinform. | 1 |
| 2018 | Ultra-fast global homology detection with Discrete Cosine Transform and Dynamic Time WarpingabstractMotivation: Evolutionary information is crucial for the annotation of proteins in bioinformatics. The amount of retrieved homologs often correlates with the quality of predicted protein annotations related to structure or function. With a growing amount of sequences available, fast and reliable methods for homology detection are essential, as they have a direct impact on predicted protein annotations. Results: We developed a discriminative, alignment-free algorithm for homology detection with quasi-linear complexity, enabling theoretically much faster homology searches. To reach this goal, we convert the protein sequence into numeric biophysical representations. These are shrunk to a fixed length using a novel vector quantization method which uses a Discrete Cosine Transform compression. We then compute, for each compressed representation, similarity scores between proteins with the Dynamic Time Warping algorithm and we feed them into a Random Forest. The WARP performances are comparable with state of the art methods. Availability and implementation: The method is available at http://ibsquare.be/warp. Supplementary information: Supplementary data are available at Bioinformatics online. Daniele Raimondi, Gabriele Orlando, Yves Moreau, Wim F. Vranken |
Bioinform. | 2 |
| 2017 | SVM-dependent pairwise HMM: an application to protein pairwise alignmentsabstractMOTIVATION: Methods able to provide reliable protein alignments are crucial for many bioinformatics applications. In the last years many different algorithms have been developed and various kinds of information, from sequence conservation to secondary structure, have been used to improve the alignment performances. This is especially relevant for proteins with highly divergent sequences. However, recent works suggest that different features may have different importance in diverse protein classes and it would be an advantage to have more customizable approaches, capable to deal with different alignment definitions. RESULTS: Here we present Rigapollo, a highly flexible pairwise alignment method based on a pairwise HMM-SVM that can use any type of information to build alignments. Rigapollo lets the user decide the optimal features to align their protein class of interest. It outperforms current state of the art methods on two well-known benchmark datasets when aligning highly divergent sequences. AVAILABILITY AND IMPLEMENTATION: A Python implementation of the algorithm is available at http://ibsquare.be/rigapollo. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriele Orlando, Daniele Raimondi, Taushif Khan, Tom Lenaerts, Wim F. Vranken |
Bioinform. | 1 |
| 2015 | Clustering-based model of cysteine co-evolution improves disulfide bond connectivity prediction and reduces homologous sequence requirementsabstractMOTIVATION: Cysteine residues have particular structural and functional relevance in proteins because of their ability to form covalent disulfide bonds. Bioinformatics tools that can accurately predict cysteine bonding states are already available, whereas it remains challenging to infer the disulfide connectivity pattern of unknown protein sequences. Improving accuracy in this area is highly relevant for the structural and functional annotation of proteins. RESULTS: We predict the intra-chain disulfide bond connectivity patterns starting from known cysteine bonding states with an evolutionary-based unsupervised approach called Sephiroth that relies on high-quality alignments obtained with HHblits and is based on a coarse-grained cluster-based modelization of tandem cysteine mutations within a protein family. We compared our method with state-of-the-art unsupervised predictors and achieve a performance improvement of 25-27% while requiring an order of magnitude less of aligned homologous sequences (∼10(3) instead of ∼10(4)). AVAILABILITY AND IMPLEMENTATION: The software described in this article and the datasets used are available at http://ibsquare.be/sephiroth. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online. Daniele Raimondi, Gabriele Orlando, Wim F. Vranken |
Bioinform. | 2 |