Vera van Noort

dblp:28/3885 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-8436-6602ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 6 since 2021
YearPublicationVenuePosition
2025 AFragmenter: schema-free, tuneable protein domain segmentation for AlphaFold protein structures
abstract
SUMMARY: Protein domain segmentation is a crucial aspect of understanding protein functions and interactions, and it is vital for protein modelling exercises and evolutionary studies. Current segmentation methods often rely on predefined classification schemes, leading to inconsistencies and biases. AFragmenter provides a schema-free and tuneable approach to protein domain segmentation based on network analysis of AlphaFold-predicted structures. Utilizing Predicted Aligned Error values, AFragmenter constructs a fully connected network of protein residues and identifies distinct structural domains by using Leiden clustering. This method empowers users to adjust parameters including contrast threshold and resolution, providing control over the segmentation process. AVAILABILITY AND IMPLEMENTATION: AFragmenter is implemented in Python3 and freely available under an MIT license. It can be found as a Python library and command line tool at https://github.com/sverwimp/AFragmenter, pip, and Conda.
Stefaan Verwimp, Rob Lavigne, Cédric Lood, Vera van Noort
Bioinform.4
2025 SCARAP: scalable cross-species comparative genomics of prokaryotes
abstract
MOTIVATION: Much of prokaryotic comparative genomics currently relies on two critical computational tasks: pangenome inference and core genome inference. Pangenome inference involves clustering genes from a set of genomes into gene families, enabling genome-wide association studies and evolutionary history analysis. The core genome represents gene families present in nearly all genomes and is required to infer a high-quality phylogeny. For species-level datasets, fast pangenome inference tools have been developed. However, tools applicable to more diverse datasets are currently slow and scale poorly. RESULTS: Here, we introduce SCARAP, a program containing three modules for comparative genomics analyses: a fast and scalable pangenome inference module, a direct core genome inference module, and a module for subsampling representative genomes. When benchmarked against existing tools, the SCARAP pan module proved up to an order of magnitude faster with comparable accuracy. The core module was validated by comparing its result against a core genome extracted from a full pangenome. The sample module demonstrated the rapid sampling of genomes with decreasing novelty. Applied to a dataset of over 31 000 Lactobacillales genomes, SCARAP showcased its ability to derive a representative pangenome. Finally, we applied the novel concept of gene fixation frequency to this pangenome, showing that Lactobacillales genes that are prevalent but rarely fixate in species often encode bacteriophage functions. AVAILABILITY AND IMPLEMENTATION: The SCARAP toolkit is publicly available at https://github.com/swittouck/scarap.
Stijn Wittouck, Tom Eilers, Vera van Noort, Sarah Lebeer
Bioinform.3
2024 FLAMS: Find Lysine Acylations and other Modification Sites
abstract
SUMMARY: Today, hundreds of post-translational modification (PTM) sites are routinely identified at once, but the comparison of new experimental datasets to already existing ones is hampered by the current inability to search most PTM databases at the protein residue level. We present FLAMS (Find Lysine Acylations and other Modification Sites), a Python3-based command line and web-tool that enables researchers to compare their PTM sites to the contents of the CPLM, the largest dedicated protein lysine modification database, and dbPTM, the most comprehensive general PTM database, at the residue level. FLAMS can be integrated into PTM analysis pipelines, allowing researchers to quickly assess the novelty and conservation of PTM sites across species in newly generated datasets, aiding in the functional assessment of sites and the prioritization of sites for further experimental characterization. AVAILABILITY AND IMPLEMENTATION: FLAMS is implemented in Python3, and freely available under an MIT license. It can be found as a command line tool at https://github.com/hannelorelongin/FLAMS, pip and conda; and as a web service at https://www.biw.kuleuven.be/m2s/cmpg/research/CSB/tools/flams/.
Hannelore Longin, Nand Broeckaert, Maarten Langen, Roshan Hari, Anna Kramarska, Kasper Oikarinen, Hanne Hendrix, Rob Lavigne, Vera van Noort
Bioinform.9
2023 DARTpaths, an in silico platform to investigate molecular mechanisms of compounds
abstract
SUMMARY: Xpaths is a collection of algorithms that allow for the prediction of compound-induced molecular mechanisms of action by integrating phenotypic endpoints of different species; and proposes follow-up tests for model organisms to validate these pathway predictions. The Xpaths algorithms are applied to predict developmental and reproductive toxicity (DART) and implemented into an in silico platform, called DARTpaths. AVAILABILITY AND IMPLEMENTATION: All code is available on GitHub https://github.com/Xpaths/dartpaths-app under Apache license 2.0, detailed overview with demo is available at https://www.vivaltes.com/dartpaths/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Diksha Bhalla, Marvin N. Steijaert, Eefje S. Poppelaars, Marc Teunis, Monique van der Voet, Marie Corradi, Elisabeth Dévière, Luke Noothout, Wilco Tomassen, Martijn Rooseboom, Richard A. Currie, Cyrille Krul, Raymond H. H. Pieters, Vera van Noort, Marjolein Wildwater
Bioinform.14
2022 SASpector: analysis of missing genomic regions in draft genomes of prokaryotes
abstract
SUMMARY: Missing regions in short-read assemblies of prokaryote genomes are often attributed to biases in sequencing technologies and to repetitive elements, the former resulting in low sequencing coverage of certain loci and the latter to unresolved loops in the de novo assembly graph. We developed SASpector, a command-line tool that compares short-read assemblies (draft genomes) to their corresponding closed assemblies and extracts missing regions to analyze them at the sequence and functional level. SASpector allows to benchmark the need for resolved genomes, can be integrated into pipelines to control the quality of assemblies, and could be used for comparative investigations of missingness in assemblies for which both short-read and long-read data are available in the public databases. AVAILABILITY AND IMPLEMENTATION: SASpector is available at https://github.com/LoGT-KULeuven/SASpector. The tool is implemented in Python3 and available through pip and Docker (0mician/saspector). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Cédric Lood, Alejandro Correa Rojo, Deniz Sinar, Emma Verkinderen, Rob Lavigne, Vera van Noort
Bioinform.6
2021 Molecular dynamics shows complex interplay and long-range effects of post-translational modifications in yeast protein interactions
abstract
Post-translational modifications (PTMs) play a vital, yet often overlooked role in the living cells through modulation of protein properties, such as localization and affinity towards their interactors, thereby enabling quick adaptation to changing environmental conditions. We have previously benchmarked a computational framework for the prediction of PTMs' effects on the stability of protein-protein interactions, which has molecular dynamics simulations followed by free energy calculations at its core. In the present work, we apply this framework to publicly available data on Saccharomyces cerevisiae protein structures and PTM sites, identified in both normal and stress conditions. We predict proteome-wide effects of acetylations and phosphorylations on protein-protein interactions and find that acetylations more frequently have locally stabilizing roles in protein interactions, while the opposite is true for phosphorylations. However, the overall impact of PTMs on protein-protein interactions is more complex than a simple sum of local changes caused by the introduction of PTMs and adds to our understanding of PTM cross-talk. We further use the obtained data to calculate the conformational changes brought about by PTMs. Finally, conservation of the analyzed PTM residues in orthologues shows that some predictions for yeast proteins will be mirrored to other organisms, including human. This work, therefore, contributes to our overall understanding of the modulation of the cellular protein interaction networks in yeast and beyond.
Nikolina Sostaric, Vera van Noort
PLoS Comput. Biol.2
2017 Evolutionary conservation of Ebola virus proteins predicts important functions at residue level
abstract
MOTIVATION: The recent outbreak of Ebola virus disease (EVD) resulted in a large number of human deaths. Due to this devastation, the Ebola virus has attracted renewed interest as model for virus evolution. Recent literature on Ebola virus (EBOV) has contributed substantially to our understanding of the underlying genetics and its scope with reference to the 2014 outbreak. But no study yet, has focused on the conservation patterns of EBOV proteins. RESULTS: We analyzed the evolution of functional regions of EBOV and highlight the function of conserved residues in protein activities. We apply an array of computational tools to dissect the functions of EBOV proteins in detail: (i) protein sequence conservation, (ii) protein-protein interactome analysis, (iii) structural modeling and (iv) kinase prediction. Our results suggest the presence of novel post-translational modifications in EBOV proteins and their role in the modulation of protein functions and protein interactions. Moreover, on the basis of the presence of ATM recognition motifs in all EBOV proteins we postulate a role of DNA damage response pathways and ATM kinase in EVD. The ATM kinase is put forward, for further evaluation, as novel potential therapeutic target. AVAILABILITY AND IMPLEMENTATION: http://www.biw.kuleuven.be/CSB/EBOV-PTMs CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Ahmed Arslan, Vera van Noort
Bioinform.2
2017 yMap: an automated method to map yeast variants to protein modifications and functional regions
abstract
Summary: Recent advances in sequence technology result in large datasets of sequence variants. For the human genome, several tools are available to predict the impact of these variants on gene and protein functions. However, for model organisms such as yeast such tools are lacking, specifically to predict the effect of protein sequence altering variants on the protein level. We present a python framework that enables users to map in a fully automated fashion large set of variants to protein functional regions and post-translationally modified residues. Furthermore, we provide the user with the possibility to retrieve predicted functional information on modified residues from other resources for example that are predicted to play a role in protein-protein interactions. The results are complemented by statistical tests to highlight the significance of underlying functions and pathways affected by mutations. We show the application of this package on a yeast dataset derived from a recent evolutionary experiment on adaptation to ethanol. Availability and Implementation: The package is available from https://github.com/CSB-KUL/yMap and is implemented in Python. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Ahmed Arslan, Vera van Noort
Bioinform.2
2017 ARA-PEPs: a repository of putative sORF-encoded peptides in Arabidopsis thaliana
abstract
BACKGROUND: Many eukaryotic RNAs have been considered non-coding as they only contain short open reading frames (sORFs). However, there is increasing evidence for the translation of these sORFs into bioactive peptides with potent signaling, antimicrobial, developmental, antioxidant roles etc. Yet only a few peptides encoded by sORFs are annotated in the model organism Arabidopsis thaliana. RESULTS: To aid the functional annotation of these peptides, we have developed ARA-PEPs (available at http://www.biw.kuleuven.be/CSB/ARA-PEPs ), a repository of putative peptides encoded by sORFs in the A. thaliana genome starting from in-house Tiling arrays, RNA-seq data and other publicly available datasets. ARA-PEPs currently lists 13,748 sORF-encoded peptides with transcriptional evidence. In addition to existing data, we have identified 100 novel transcriptionally active regions (TARs) that might encode 341 novel stress-induced peptides (SIPs). To aid in identification of bioactivity, we add functional annotation and sequence conservation to predicted peptides. CONCLUSION: To our knowledge, this is the largest repository of plant peptides encoded by sORFs with transcript evidence, publicly available and this resource will help scientists to effortlessly navigate the list of experimentally studied peptides, the experimental and computational evidence supporting the activity of these peptides and gain new perspectives for peptide discovery.
Rashmi R. Hazarika, Barbara De Coninck, Lidia R. Yamamoto, Laura R. Martin, Bruno P. A. Cammue, Vera van Noort
BMC Bioinform.6
2016 CART - a chemical annotation retrieval toolkit
abstract
MOTIVATION: Data on bioactivities of drug-like chemicals are rapidly accumulating in public repositories, creating new opportunities for research in computational systems pharmacology. However, integrative analysis of these data sets is difficult due to prevailing ambiguity between chemical names and identifiers and a lack of cross-references between databases. RESULTS: To address this challenge, we have developed CART, a Chemical Annotation Retrieval Toolkit. As a key functionality, it matches an input list of chemical names into a comprehensive reference space to assign unambiguous chemical identifiers. In this unified space, bioactivity annotations can be easily retrieved from databases covering a wide variety of chemical effects on biological systems. Subsequently, CART can determine annotations enriched in the input set of chemicals and display these in tabular format and interactive network visualizations, thereby facilitating integrative analysis of chemical bioactivity data. AVAILABILITY AND IMPLEMENTATION: CART is available as a Galaxy web service (cart.embl.de). Source code and an easy-to-install command line tool can also be obtained from the web site. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Samy Deghou, Georg Zeller, Murat Iskar, Marja Driessen, Mercedes Castillo, Vera van Noort, Peer Bork
Bioinform.6
2014 Comments on "Human Dominant Disease Genes Are Enriched in Paralogs Originating from Whole Genome Duplication"
abstract
We previously showed that monogenic disease genes (MDs) are enriched in duplicates and hypothesized that functional redundancy among duplicates underlies this enrichment [1]. In their comment, Singh et al. refine this enrichment to genes resulting from whole genome duplications (WGDs) [2]; they, furthermore, “could not find any significant enrichment in duplicates in support of possible functional compensation for essential genes” [2] by using gene essentiality data from mouse (transferred to human through orthology). We appreciate the scientific argument, but we would like to point out that confounding factors and data biases can lead to seemingly opposing conclusions. For example, we carefully considered the duplication age of genes, which is a known confounder in such analyses [3], [4], as well as the use of gene subsets that have known biases such as the mouse essentiality data [3], which, in addition, have issues when conclusions are being transferred to human genes. First, when using the data of Singh et al. [2] and stratifying small-scale duplicates (SSDs) into old and young groups according to the duplication age relative to WGD, we found that MDs are enriched in old SSDs; limiting this analysis to recessive MDs produced similar results (Figure 1A). In contrast, MDs are depleted in young SSDs (Figure 1B), which is consistent with our hypothesis and with our findings that coexpression decreases with increased duplication age. Thus, when the duplication is old, the ability of the functional copy to compensate for the mutation-carrying malfunctioning copy could be easily disrupted because of random fluctuation in gene expression in a subpopulation; consequently, the gene is associated with a disease, but it will not be purged from the whole population. Therefore, functional compensation can promote the spreading of disease genes in duplicates. However, in young duplicates, the fluctuation in gene expression among duplicates may not be that huge; thus, deleterious mutations could be tolerated, and the corresponding genes are unlikely to associate with any diseases. Figure 1 Enrichment of MDs in old SSDs and distinct characteristics of the old SSDs as compared with the young ones. Second, mouse essentiality data are biased [5], e.g., towards developmental genes; i.e., they do not correspond to the full spectrum of MDs. Dividing the tested mouse genes into subgroups, the proportion of essential genes in young SSDs is significantly lower than that of singletons (Figure 1C), consistent with functional redundancy among duplicates; however, the opposite is found in old SSDs (Figure 1D). The latter has led to the somewhat counterintuitive conclusion that “duplicates are as essential as singletons” [6], which has been argued against by several follow-up studies [3]–[5]. These results, again, highlight the importance of taking duplication age into consideration. As previous studies suggested, it is not trivial to correct the biases [3]–[5], and hence, conclusions from this data regarding duplications have to be taken with caution. Furthermore, the essentiality status of mouse genes cannot be reliably transferred to human and vice versa. For example, using data from OGEE [7], an online gene essentiality database, 2,322 mouse essential genes have one-to-one orthologs in human; only 476 out of the 2,322 human genes (approximately 20%) were essential according to a genome-wide small interfering RNA (siRNA) experiment [8]. Finally, only less than 30% of the MDs we collected [1] were used in the analyses by Singh et al.; the intersection with the essentiality dataset is even smaller (approximately 18.6% of the MDs used in [1]) because, so far, only less than one-third (approximately 6,400) of mouse genes has been tested for essentiality [9]. Thus, extrapolating any observations on these data to the whole genome would be difficult; for example, some functional signals might only become statistically significant in larger datasets. Elucidating the molecular basis of human genetic disorders is one of the most important tasks in medical biology. With the relevant data, such as those from genome-wide association studies (GWAS), accumulated at an astonishing speed, integrative and comparative analyses through bioinformatics are much needed. In this regard, Singh et al. did provide an important contribution by refining the enrichment of dominant MDs in duplicates to those derived from WGD. However, we don't believe that they nullified our functional compensation hypothesis with the analyses performed, but they certainly encouraged further studies on more complete datasets, hopefully to be available in the near future.
Wei-Hua Chen, Xing-Ming Zhao, Vera van Noort, Peer Bork
PLoS Comput. Biol.3
2013 Human Monogenic Disease Genes Have Frequently Functionally Redundant Paralogs
abstract
Mendelian disorders are often caused by mutations in genes that are not lethal but induce functional distortions leading to diseases. Here we study the extent of gene duplicates that might compensate genes causing monogenic diseases. We provide evidence for pervasive functional redundancy of human monogenic disease genes (MDs) by duplicates by manifesting 1) genes involved in human genetic disorders are enriched in duplicates and 2) duplicated disease genes tend to have higher functional similarities with their closest paralogs in contrast to duplicated non-disease genes of similar age. We propose that functional compensation by duplication of genes masks the phenotypic effects of deleterious mutations and reduces the probability of purging the defective genes from the human population; this functional compensation could be further enhanced by higher purification selection between disease genes and their duplicates as well as their orthologous counterpart compared to non-disease genes. However, due to the intrinsic expression stochasticity among individuals, the deleterious mutations could still be present as genetic diseases in some subpopulations where the duplicate copies are expressed at low abundances. Consequently the defective genes are linked to genetic disorders while they continue propagating within the population. Our results provide insight into the molecular basis underlying the spreading of duplicated disease genes.
Wei-Hua Chen, Xing-Ming Zhao, Vera van Noort, Peer Bork
PLoS Comput. Biol.3
2011 Prediction of Drug Combinations by Integrating Molecular and Pharmacological Data
abstract
Combinatorial therapy is a promising strategy for combating complex disorders due to improved efficacy and reduced side effects. However, screening new drug combinations exhaustively is impractical considering all possible combinations between drugs. Here, we present a novel computational approach to predict drug combinations by integrating molecular and pharmacological data. Specifically, drugs are represented by a set of their properties, such as their targets or indications. By integrating several of these features, we show that feature patterns enriched in approved drug combinations are not only predictive for new drug combinations but also provide insights into mechanisms underlying combinatorial therapy. Further analysis confirmed that among our top ranked predictions of effective combinations, 69% are supported by literature, while the others represent novel potential drug combinations. We believe that our proposed approach can help to limit the search space of drug combinations and provide a new way to effectively utilize existing drugs for new purposes.
Xing-Ming Zhao, Murat Iskar, Georg Zeller, Michael Kuhn 0004, Vera van Noort, Peer Bork
PLoS Comput. Biol.5
2010 Drug-Induced Regulation of Target Expression
abstract
Drug perturbations of human cells lead to complex responses upon target binding. One of the known mechanisms is a (positive or negative) feedback loop that adjusts the expression level of the respective target protein. To quantify this mechanism systems-wide in an unbiased way, drug-induced differential expression of drug target mRNA was examined in three cell lines using the Connectivity Map. To overcome various biases in this valuable resource, we have developed a computational normalization and scoring procedure that is applicable to gene expression recording upon heterogeneous drug treatments. In 1290 drug-target relations, corresponding to 466 drugs acting on 167 drug targets studied, 8% of the targets are subject to regulation at the mRNA level. We confirmed systematically that in particular G-protein coupled receptors, when serving as known targets, are regulated upon drug treatment. We further newly identified drug-induced differential regulation of Lanosterol 14-alpha demethylase, Endoplasmin, DNA topoisomerase 2-alpha and Calmodulin 1. The feedback regulation in these and other targets is likely to be relevant for the success or failure of the molecular intervention.
Murat Iskar, Monica Campillos, Michael Kuhn 0004, Lars Juhl Jensen, Vera van Noort, Peer Bork
PLoS Comput. Biol.5
2007 Assessment of phylogenomic and orthology approaches for phylogenetic inference
abstract
MOTIVATION: Phylogenomics integrates the vast amount of phylogenetic information contained in complete genome sequences, and is rapidly becoming the standard for reliably inferring species phylogenies. There are, however, fundamental differences between the ways in which phylogenomic approaches like gene content, superalignment, superdistance and supertree integrate the phylogenetic information from separate orthologous groups. Furthermore, they all depend on the method by which the orthologous groups are initially determined. Here, we systematically compare these four phylogenomic approaches, in parallel with three approaches for large-scale orthology determination: pairwise orthology, cluster orthology and tree-based orthology. RESULTS: Including various phylogenetic methods, we apply a total of 54 fully automated phylogenomic procedures to the fungi, the eukaryotic clade with the largest number of sequenced genomes, for which we retrieved a golden standard phylogeny from the literature. Phylogenomic trees based on gene content show, relative to the other methods, a bias in the tree topology that parallels convergence in lifestyle among the species compared, indicating convergence in gene content. CONCLUSIONS: Complete genomes are no guarantee for good or even consistent phylogenies. However, the large amounts of data in genomes enable us to carefully select the data most suitable for phylogenomic inference. In terms of performance, the superalignment approach, combined with restrictive orthology, is the most successful in recovering a fungal phylogeny that agrees with current taxonomic views, and allows us to obtain a high-resolution phylogeny. We provide solid support for what has grown to be a common practice in phylogenomics during its advance in recent years. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bas E. Dutilh, Vera van Noort, René T. J. M. van der Heijden, T. Boekhout, Berend Snel, Martijn A. Huynen
Bioinform.2
2007 Orthology prediction at scalable resolution by phylogenetic tree analysis
abstract
BACKGROUND: Orthology is one of the cornerstones of gene function prediction. Dividing the phylogenetic relations between genes into either orthologs or paralogs is however an oversimplification. Already in two-species gene-phylogenies, the complicated, non-transitive nature of phylogenetic relations results in inparalogs and outparalogs. For situations with more than two species we lack semantics to specifically describe the phylogenetic relations, let alone to exploit them. Published procedures to extract orthologous groups from phylogenetic trees do not allow identification of orthology at various levels of resolution, nor do they document the relations between the orthologous groups. RESULTS: We introduce "levels of orthology" to describe the multi-level nature of gene relations. This is implemented in a program LOFT (Levels of Orthology From Trees) that assigns hierarchical orthology numbers to genes based on a phylogenetic tree. To decide upon speciation and gene duplication events in a tree LOFT can be instructed either to perform classical species-tree reconciliation or to use the species overlap between partitions in the tree. The hierarchical orthology numbers assigned by LOFT effectively summarize the phylogenetic relations between genes. The resulting high-resolution orthologous groups are depicted in colour, facilitating visual inspection of (large) trees. A benchmark for orthology prediction, that takes into account the varying levels of orthology between genes, shows that the phylogeny-based high-resolution orthology assignments made by LOFT are reliable. CONCLUSION: The "levels of orthology" concept offers high resolution, reliable orthology, while preserving the relations between orthologous groups. A Windows as well as a preliminary Java version of LOFT is available from the LOFT website http://www.cmbi.ru.nl/LOFT.
René T. J. M. van der Heijden, Berend Snel, Vera van Noort, Martijn A. Huynen
BMC Bioinform.3