EDBT 2026 Demo / reviewers in the wild / expert
Wim F. Vranken
dblp:82/2295
· DBLP profile ↗
18ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-7470-4324ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 8 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Combining evolution and protein language models for an interpretable cancer driver mutation prediction with D2DeepabstractThe mutations driving cancer are being increasingly exposed through tumor-specific genomic data. However, differentiating between cancer-causing driver mutations and random passenger mutations remains challenging. State-of-the-art homology-based predictors contain built-in biases and are often ill-suited to the intricacies of cancer biology. Protein language models have successfully addressed various biological problems but have not yet been tested on the challenging task of cancer driver mutation prediction at a large scale. Additionally, they often fail to offer result interpretation, hindering their effective use in clinical settings. The AI-based D2Deep method we introduce here addresses these challenges by combining two powerful elements: (i) a nonspecialized protein language model that captures the makeup of all protein sequences and (ii) protein-specific evolutionary information that encompasses functional requirements for a particular protein. D2Deep relies exclusively on sequence information, outperforms state-of-the-art predictors, and captures intricate epistatic changes throughout the protein caused by mutations. These epistatic changes correlate with known mutations in the clinical setting and can be used for the interpretation of results. The model is trained on a balanced, somatic training set and so effectively mitigates biases related to hotspot mutations compared to state-of-the-art techniques. The versatility of D2Deep is illustrated by its performance on non-cancer mutation prediction, where most variants still lack known consequences. D2Deep predictions and confidence scores are available via https://tumorscope.be/d2deep to help with clinical interpretation and mutation prioritization. Konstantina Tzavella, Adrián Díaz, Catharina Olsen, Wim F. Vranken |
Briefings Bioinform. | 4 |
| 2025 | In silico identification of archaeal DNA-binding proteinsabstractMOTIVATION: The rapid advancement of next-generation sequencing technologies has generated an immense volume of genetic data. However, these data are unevenly distributed, with well-studied organisms being disproportionately represented, while other organisms, such as from archaea, remain significantly underexplored. The study of archaea is particularly challenging due to the extreme environments they inhabit and the difficulties associated with culturing them in the laboratory. Despite these challenges, archaea likely represent a crucial evolutionary link between eukaryotic and prokaryotic organisms, and their investigation could shed light on the early stages of life on Earth. Yet, a significant portion of archaeal proteins are annotated with limited or inaccurate information. Among the various classes of archaeal proteins, DNA-binding proteins are of particular importance. While they represent a large portion of every known proteome, their identification in archaea is complicated by the substantial evolutionary divergence between archaeal and the other better studied organisms. RESULTS: To address the challenges of identifying DNA-binding proteins in archaea, we developed Xenusia, a neural network-based tool capable of screening entire archaeal proteomes to identify DNA-binding proteins. Xenusia has proven effective across diverse datasets, including metagenomics data, successfully identifying novel DNA-binding proteins, with experimental validation of its predictions. AVAILABILITY AND IMPLEMENTATION: Xenusia is available as a PyPI package, with source code accessible at https://github.com/grogdrinker/xenusia, and as a Google Colab web server application at xenusia.ipynb. Linus Donvil, Joëlle A. J. Housmans, Eveline Peeters, Wim F. Vranken, Gabriele Orlando |
Bioinform. | 4 |
| 2024 | Large-scale structure-informed multiple sequence alignment of proteins with SIMSApiperabstractSUMMARY: SIMSApiper is a Nextflow pipeline that creates reliable, structure-informed MSAs of thousands of protein sequences faster than standard structure-based alignment methods. Structural information can be provided by the user or collected by the pipeline from online resources. Parallelization with sequence identity-based subsets can be activated to significantly speed up the alignment process. Finally, the number of gaps in the final alignment can be reduced by leveraging the position of conserved secondary structure elements. AVAILABILITY AND IMPLEMENTATION: The pipeline is implemented using Nextflow, Python3, and Bash. It is publicly available on github.com/Bio2Byte/simsapiper. Charlotte Crauwels, Sophie-Luise Heidig, Adrián Díaz, Wim F. Vranken |
Bioinform. | 4 |
| 2024 | bio2Byte Tools deployment as a Python package and Galaxy tool to predict protein biophysical propertiesabstractSUMMARY: We introduce a unified Python package for the prediction of protein biophysical properties, streamlining previous tools developed by the Bio2Byte research group. This suite facilitates comprehensive assessments of protein characteristics, incorporating predictors for backbone and sidechain dynamics, local secondary structure propensities, early folding, long disorder, beta-sheet aggregation, and fused in sarcoma (FUS)-like phase separation. Our package significantly eases the integration and execution of these tools, enhancing accessibility for both computational and experimental researchers. AVAILABILITY AND IMPLEMENTATION: The suite is available on the Python Package Index (PyPI): https://pypi.org/project/b2bTools/ and Bioconda: https://bioconda.github.io/recipes/b2btools/README.html for Linux and macOS systems, with Docker images hosted on Biocontainers: https://quay.io/repository/biocontainers/b2btools?tab=tags&tag=latest and Docker Hub: https://hub.docker.com/u/bio2byte. Online deployments are available on Galaxy Europe: https://usegalaxy.eu/root?tool_id=b2btools_single_sequence and our online server: https://bio2byte.be/b2btools/. The source code can be found at https://bitbucket.org/bio2byte/b2btools_releases. Jose Gavaldá-García, Adrián Díaz, Wim F. Vranken |
Bioinform. | 3 |
| 2023 | The ACPYPE web server for small-molecule MD topology generationabstractMOTIVATION: The generation of parameter files for molecular dynamics (MD) simulations of small molecules that are suitable for force fields commonly applied to proteins and nucleic acids is often challenging. The ACPYPE software and website aid the generation of such parameter files. RESULTS: ACPYPE uses OpenBabel and ANTECHAMBER to generate MD input files in Gromacs, AMBER, CHARMM, and CNS formats. It can now take a SMILES string as input, in addition to the original PDB or mol2 coordinate files, with GAFF2 support and GLYCAM force field conversion added. It can be installed locally via Anaconda, PyPI, and Docker distributions, while the web server at https://bio2byte.be/acpype/ was updated with an API, and provides visualization of results for uploaded molecules as well as a pre-generated set of 3738 drug molecules. AVAILABILITY AND IMPLEMENTATION: The web application is freely available at https://www.bio2byte.be/acpype/ and the open-source code can be found at https://github.com/alanwilter/acpype. Luciano Porto Kagami, Alan Wilter, Adrián Díaz, Wim F. Vranken |
Bioinform. | 4 |
| 2023 | Deciphering the RRM-RNA recognition code: A computational analysisabstractRNA recognition motifs (RRM) are the most prevalent class of RNA binding domains in eucaryotes. Their RNA binding preferences have been investigated for almost two decades, and even though some RRM domains are now very well described, their RNA recognition code has remained elusive. An increasing number of experimental structures of RRM-RNA complexes has become available in recent years. Here, we perform an in-depth computational analysis to derive an RNA recognition code for canonical RRMs. We present and validate a computational scoring method to estimate the binding between an RRM and a single stranded RNA, based on structural data from a carefully curated multiple sequence alignment, which can predict RRM binding RNA sequence motifs based on the RRM protein sequence. Given the importance and prevalence of RRMs in humans and other species, this tool could help design RNA binding motifs with uses in medical or synthetic biology applications, leading towards the de novo design of RRMs with specific RNA recognition. Joel Roca-Martínez, Hrishikesh Dhondge, Michael Sattler, Wim F. Vranken |
PLoS Comput. Biol. | 4 |
| 2021 | Computational resources for identifying and describing proteins driving liquid-liquid phase separationabstractOne of the most intriguing fields emerging in current molecular biology is the study of membraneless organelles formed via liquid-liquid phase separation (LLPS). These organelles perform crucial functions in cell regulation and signalling, and recent years have also brought about the understanding of the molecular mechanism of their formation. The LLPS field is continuously developing and optimizing dedicated in vitro and in vivo methods to identify and characterize these non-stoichiometric molecular condensates and the proteins able to drive or contribute to LLPS. Building on these observations, several computational tools and resources have emerged in parallel to serve as platforms for the collection, annotation and prediction of membraneless organelle-linked proteins. In this survey, we showcase recent advancements in LLPS bioinformatics, focusing on (i) available databases and ontologies that are necessary to describe the studied phenomena and the experimental results in an unambiguous way and (ii) prediction methods to assess the potential LLPS involvement of proteins. Through hands-on application of these resources on example proteins and representative datasets, we give a practical guide to show how they can be used in conjunction to provide in silico information on LLPS. Rita Pancsa, Wim F. Vranken, Bálint Mészáros |
Briefings Bioinform. | 2 |
| 2021 | MutaFrame - an interpretative visualization framework for deleteriousness prediction of missense variants in the human exomeabstractMOTIVATION: High-throughput experiments are generating ever increasing amounts of various -omics data, so shedding new light on the link between human disorders, their genetic causes and the related impact on protein behavior and structure. While numerous bioinformatics tools now exist that predict which variants in the human exome cause diseases, few tools predict the reasons why they might do so. Yet, understanding the impact of variants at the molecular level is a prerequisite for the rational development of targeted drugs or personalized therapies. RESULTS: We present the updated MutaFrame webserver, which aims to meet this need. It offers two deleteriousness prediction softwares, DEOGEN2 and SNPMuSiC, and is designed for bioinformaticians and medical researchers who want to gain insights into the origins of monogenic diseases. It contains information at two levels for each human protein: its amino acid sequence and its three-dimensional structure; we used the experimental structures whenever available, and modeled structures otherwise. MutaFrame also includes higher-level information, such as protein essentiality and protein-protein interactions. It has a user-friendly interface for the interpretation of results and a convenient visualization system for protein structures, in which the variant positions introduced by the user and other structural information are shown. In this way, MutaFrame aids our understanding of the pathogenic processes caused by single-site mutations and their molecular and contextual interpretation. AVAILABILITY AND IMPLEMENTATION: Mutaframe webserver at http://mutaframe.com/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. François Ancien, Fabrizio Pucci, Wim F. Vranken, Marianne Rooman |
Bioinform. | 3 |
| 2020 | Accurate prediction of protein beta-aggregation with generalized statistical potentialsabstractMOTIVATION: Protein beta-aggregation is an important but poorly understood phenomena involved in diseases as well as in beneficial physiological processes. However, while this task has been investigated for over 50 years, very little is known about its mechanisms of action. Moreover, the identification of regions involved in aggregation is still an open problem and the state-of-the-art methods are often inadequate in real case applications. RESULTS: In this article we present AgMata, an unsupervised tool for the identification of such regions from amino acidic sequence based on a generalized definition of statistical potentials that includes biophysical information. The tool outperforms the state-of-the-art methods on two different benchmarks. As case-study, we applied our tool to human ataxin-3, a protein involved in Machado-Joseph disease. Interestingly, AgMata identifies aggregation-prone residues that share the very same structural environment. Additionally, it successfully predicts the outcome of in vitro mutagenesis experiments, identifying point mutations that lead to an alteration of the aggregation propensity of the wild-type ataxin-3. AVAILABILITY AND IMPLEMENTATION: A python implementation of the tool is available at https://bitbucket.org/bio2byte/agmata. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriele Orlando, Alexandra Silva 0004, Sandra Macedo-Ribeiro, Daniele Raimondi, Wim F. Vranken |
Bioinform. | 5 |
| 2019 | Computational identification of prion-like RNA-binding proteins that form liquid phase-separated condensatesabstractMOTIVATION: Eukaryotic cells contain different membrane-delimited compartments, which are crucial for the biochemical reactions necessary to sustain cell life. Recent studies showed that cells can also trigger the formation of membraneless organelles composed by phase-separated proteins to respond to various stimuli. These condensates provide new ways to control the reactions and phase-separation proteins (PSPs) are thus revolutionizing how cellular organization is conceived. The small number of experimentally validated proteins, and the difficulty in discovering them, remain bottlenecks in PSPs research. RESULTS: Here we present PSPer, the first in-silico screening tool for prion-like RNA-binding PSPs. We show that it can prioritize PSPs among proteins containing similar RNA-binding domains, intrinsically disordered regions and prions. PSPer is thus suitable to screen proteomes, identifying the most likely PSPs for further experimental investigation. Moreover, its predictions are fully interpretable in the sense that it assigns specific functional regions to the predicted proteins, providing valuable information for experimental investigation of targeted mutations on these regions. Finally, we show that it can estimate the ability of artificially designed proteins to form condensates (r=-0.87), thus providing an in-silico screening tool for protein design experiments. AVAILABILITY AND IMPLEMENTATION: PSPer is available at bio2byte.com/psp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriele Orlando, Daniele Raimondi, Francesco Tabaro, Francesco Codicè, Yves Moreau, Wim F. Vranken |
Bioinform. | 6 |
| 2018 | RINspector: a Cytoscape app for centrality analyses and DynaMine flexibility predictionabstractMOTIVATION: Protein function is directly related to amino acid residue composition and the dynamics of these residues. Centrality analyses based on residue interaction networks permit to identify key residues in a protein that are important for its fold or function. Such central residues and their environment constitute suitable targets for mutagenesis experiments. Predicted flexibility and changes in flexibility upon mutation provide valuable additional information for the design of such experiments. RESULTS: We combined centrality analyses with DynaMine flexibility predictions in a Cytoscape app called RINspector. The app performs centrality analyses and directly visualizes the results on a graph of predicted residue flexibility. In addition, the effect of mutations on local flexibility can be calculated. AVAILABILITY AND IMPLEMENTATION: The app is publicly available in the Cytoscape app store. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Guillaume Brysbaert, Kevin Lorgouilloux, Wim F. Vranken, Marc F. Lensink |
Bioinform. | 3 |
| 2018 | Ultra-fast global homology detection with Discrete Cosine Transform and Dynamic Time WarpingabstractMotivation: Evolutionary information is crucial for the annotation of proteins in bioinformatics. The amount of retrieved homologs often correlates with the quality of predicted protein annotations related to structure or function. With a growing amount of sequences available, fast and reliable methods for homology detection are essential, as they have a direct impact on predicted protein annotations. Results: We developed a discriminative, alignment-free algorithm for homology detection with quasi-linear complexity, enabling theoretically much faster homology searches. To reach this goal, we convert the protein sequence into numeric biophysical representations. These are shrunk to a fixed length using a novel vector quantization method which uses a Discrete Cosine Transform compression. We then compute, for each compressed representation, similarity scores between proteins with the Dynamic Time Warping algorithm and we feed them into a Random Forest. The WARP performances are comparable with state of the art methods. Availability and implementation: The method is available at http://ibsquare.be/warp. Supplementary information: Supplementary data are available at Bioinformatics online. Daniele Raimondi, Gabriele Orlando, Yves Moreau, Wim F. Vranken |
Bioinform. | 4 |
| 2017 | Seeing the trees through the forest: sequence-based homo- and heteromeric protein-protein interaction sites prediction using random forestabstractMOTIVATION: Genome sequencing is producing an ever-increasing amount of associated protein sequences. Few of these sequences have experimentally validated annotations, however, and computational predictions are becoming increasingly successful in producing such annotations. One key challenge remains the prediction of the amino acids in a given protein sequence that are involved in protein-protein interactions. Such predictions are typically based on machine learning methods that take advantage of the properties and sequence positions of amino acids that are known to be involved in interaction. In this paper, we evaluate the importance of various features using Random Forest (RF), and include as a novel feature backbone flexibility predicted from sequences to further optimise protein interface prediction. RESULTS: We observe that there is no single sequence feature that enables pinpointing interacting sites in our Random Forest models. However, combining different properties does increase the performance of interface prediction. Our homomeric-trained RF interface predictor is able to distinguish interface from non-interface residues with an area under the ROC curve of 0.72 in a homomeric test-set. The heteromeric-trained RF interface predictor performs better than existing predictors on a independent heteromeric test-set. We trained a more general predictor on the combined homomeric and heteromeric dataset, and show that in addition to predicting homomeric interfaces, it is also able to pinpoint interface residues in heterodimers. This suggests that our random forest model and the features included capture common properties of both homodimer and heterodimer interfaces. AVAILABILITY AND IMPLEMENTATION: The predictors and test datasets used in our analyses are freely available ( http://www.ibi.vu.nl/downloads/RF_PPI/ ). CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qingzhen Hou, Paul F. G. De Geest, Wim F. Vranken, Jaap Heringa, K. Anton Feenstra |
Bioinform. | 3 |
| 2017 | SVM-dependent pairwise HMM: an application to protein pairwise alignmentsabstractMOTIVATION: Methods able to provide reliable protein alignments are crucial for many bioinformatics applications. In the last years many different algorithms have been developed and various kinds of information, from sequence conservation to secondary structure, have been used to improve the alignment performances. This is especially relevant for proteins with highly divergent sequences. However, recent works suggest that different features may have different importance in diverse protein classes and it would be an advantage to have more customizable approaches, capable to deal with different alignment definitions. RESULTS: Here we present Rigapollo, a highly flexible pairwise alignment method based on a pairwise HMM-SVM that can use any type of information to build alignments. Rigapollo lets the user decide the optimal features to align their protein class of interest. It outperforms current state of the art methods on two well-known benchmark datasets when aligning highly divergent sequences. AVAILABILITY AND IMPLEMENTATION: A Python implementation of the algorithm is available at http://ibsquare.be/rigapollo. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriele Orlando, Daniele Raimondi, Taushif Khan, Tom Lenaerts, Wim F. Vranken |
Bioinform. | 5 |
| 2016 | Multilevel biological characterization of exomic variants at the protein level significantly improves the identification of their deleterious effectsabstractMOTIVATION: There are now many predictors capable of identifying the likely phenotypic effects of single nucleotide variants (SNVs) or short in-frame Insertions or Deletions (INDELs) on the increasing amount of genome sequence data. Most of these predictors focus on SNVs and use a combination of features related to sequence conservation, biophysical, and/or structural properties to link the observed variant to either neutral or disease phenotype. Despite notable successes, the mapping between genetic variants and their phenotypic effects is riddled with levels of complexity that are not yet fully understood and that are often not taken into account in the predictions, despite their promise of significantly improving the prediction of deleterious mutants. RESULTS: We present DEOGEN, a novel variant effect predictor that can handle both missense SNVs and in-frame INDELs. By integrating information from different biological scales and mimicking the complex mixture of effects that lead from the variant to the phenotype, we obtain significant improvements in the variant-effect prediction results. Next to the typical variant-oriented features based on the evolutionary conservation of the mutated positions, we added a collection of protein-oriented features that are based on functional aspects of the gene affected. We cross-validated DEOGEN on 36 825 polymorphisms, 20 821 deleterious SNVs, and 1038 INDELs from SwissProt. The multilevel contextualization of each (variant, protein) pair in DEOGEN provides a 10% improvement of MCC with respect to current state-of-the-art tools. AVAILABILITY AND IMPLEMENTATION: The software and the data presented here is publicly available at http://ibsquare.be/deogen CONTACT: : [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniele Raimondi, Andrea M. Gazzo, Marianne Rooman, Tom Lenaerts, Wim F. Vranken |
Bioinform. | 5 |
| 2015 | Clustering-based model of cysteine co-evolution improves disulfide bond connectivity prediction and reduces homologous sequence requirementsabstractMOTIVATION: Cysteine residues have particular structural and functional relevance in proteins because of their ability to form covalent disulfide bonds. Bioinformatics tools that can accurately predict cysteine bonding states are already available, whereas it remains challenging to infer the disulfide connectivity pattern of unknown protein sequences. Improving accuracy in this area is highly relevant for the structural and functional annotation of proteins. RESULTS: We predict the intra-chain disulfide bond connectivity patterns starting from known cysteine bonding states with an evolutionary-based unsupervised approach called Sephiroth that relies on high-quality alignments obtained with HHblits and is based on a coarse-grained cluster-based modelization of tandem cysteine mutations within a protein family. We compared our method with state-of-the-art unsupervised predictors and achieve a performance improvement of 25-27% while requiring an order of magnitude less of aligned homologous sequences (∼10(3) instead of ∼10(4)). AVAILABILITY AND IMPLEMENTATION: The software described in this article and the datasets used are available at http://ibsquare.be/sephiroth. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online. Daniele Raimondi, Gabriele Orlando, Wim F. Vranken |
Bioinform. | 3 |
| 2012 | WeNMR: Structural Biology on the GridabstractThe WeNMR ( http://www.wenmr.eu ) project is a European Union funded international effort to streamline and automate analysis of Nuclear Magnetic Resonance (NMR) and Small Angle X-Ray scattering (SAXS) imaging data for atomic and near-atomic resolution molecular structures. Conventional calculation of structure requires the use of various software packages, considerable user expertise and ample computational resources. To facilitate the use of NMR spectroscopy and SAXS in life sciences the WeNMR consortium has established standard computational workflows and services through easy-to-use web interfaces, while still retaining sufficient flexibility to handle more specific requests. Thus far, a number of programs often used in structural biology have been made available through application portals. The implementation of these services, in particular the distribution of calculations to a Grid computing infrastructure, involves a novel mechanism for submission and handling of jobs that is independent of the type of job being run. With over 450 registered users (September 2012), WeNMR is currently the largest Virtual Organization (VO) in life sciences. With its large and worldwide user community, WeNMR has become the first Virtual Research Community officially recognized by the European Grid Infrastructure (EGI). Tsjerk A. Wassenaar, Marc van Dijk, Nuno Loureiro-Ferreira, Gijs van der Schot, Sjoerd Jacob de Vries, Christophe Schmitz, Johan van der Zwan, Rolf Boelens, Andrea Giachetti 0002, Lucio Ferella, Antonio Rosato, Ivano Bertini, Torsten Herrmann, Hendrik R. A. Jonker, Anurag Bagaria, Victor Jaravine, Peter Güntert, Harald Schwalbe, Wim F. Vranken, Jurgen F. Doreleijers, Gert Vriend, Geerten W. Vuister, Daniel Franke, Alexey Kikhney, Dmitri I. Svergun, Rasmus H. Fogh, John M. C. Ionides, Ernest D. Laue, Chris A. E. M. Spronk, Simonas Jurksa, Marco Verlato, Simone Badoer, Stefano Dal Pra, Mirco Mazzucato, Eric Frizziero, Alexandre M. J. J. Bonvin |
J. Grid Comput. | 19 |
| 2005 | A framework for scientific data modeling and automated software developmentabstractMOTIVATION: The lack of standards for storage and exchange of data is a serious hindrance for the large-scale data deposition, data mining and program interoperability that is becoming increasingly important in bioinformatics. The problem lies not only in defining and maintaining the standards, but also in convincing scientists and application programmers with a wide variety of backgrounds and interests to adhere to them. RESULTS: We present a UML-based programming framework for the modeling of data and the automated production of software to manipulate that data. Our approach allows one to make an abstract description of the structure of the data used in a particular scientific field and then use it to generate fully functional computer code for data access and input/output routines for data storage, together with accompanying documentation. This code can be generated simultaneously for different programming languages from a single model, together with, for example for format descriptions and I/O libraries XML and various relational databases. The framework is entirely general and could be applied in any subject area. We have used this approach to generate a data exchange standard for structural biology and analysis software for macromolecular NMR spectroscopy. AVAILABILITY: The framework is available under the GPL license, the data exchange standard with generated subroutine libraries under the LGPL license. Both may be found at http://www.ccpn.ac.uk; http://sourceforge.net/projects/ccpn CONTACT: [email protected]. Rasmus H. Fogh, Wayne Boucher, Wim F. Vranken, Anne Pajon, Tim J. Stevens, T. N. Bhat, John D. Westbrook, John M. C. Ionides, Ernest D. Laue |
Bioinform. | 3 |