Vladimir Gligorijevic

dblp:116/2862 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-5165-0973ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
11 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
5 papers
Generative modeling · 68% Optimization for machine learning · 19% Language models and text generation · 10%
Databases, data mining, and information retrieval
1 paper
Data mining · 75% Recommender systems · 25%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
protein design
1.632024
Protein Discovery with Discrete Walk-Jump Sampling · ICLR 2024
AbDiffuser: full-atom generation of in-vitro functioning antibodies · NeurIPS 2023
Implicitly Guided Design with PropEn: Match your Data to Follow the Gradient · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
1.522025
Unified all-atom molecule generation with neural fields · NeurIPS 2025
AbDiffuser: full-atom generation of in-vitro functioning antibodies · NeurIPS 2023
Machine learning › Optimization for machine learning
bilevel optimization
0.912025
Generalists vs. Specialists: Evaluating LLMs on Highly-Constrained Biophysical Sequence Optimization Tasks · ICML 2025
Natural language and speech › Language models and text generation › large language model
LLM-based optimization
0.912025
Generalists vs. Specialists: Evaluating LLMs on Highly-Constrained Biophysical Sequence Optimization Tasks · ICML 2025
Machine learning › Generative modeling › diffusion model
score-based generative model
0.912025
Unified all-atom molecule generation with neural fields · NeurIPS 2025
Bioinformatics and computational biology › protein sequence analysis
antibody sequence analysis
0.912025
deepNGS navigator: exploring antibody NGS datasets using deep contrastive learning · Bioinform. 2025
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecule generation
0.912025
Unified all-atom molecule generation with neural fields · NeurIPS 2025
Bioinformatics and computational biology › drug discovery › drug design
structure-based drug design
0.912025
Unified all-atom molecule generation with neural fields · NeurIPS 2025
Bioinformatics and computational biology
protein function prediction
0.822021
NetQuilt: deep multispecies network-based protein function prediction using homology-informed network similarity · Bioinform. 2021
deepNF: deep network fusion for protein function prediction · Bioinform. 2018
Machine learning › Generative modeling › generative model
discrete generative model
0.812024
Protein Discovery with Discrete Walk-Jump Sampling · ICLR 2024
Machine learning › Generative modeling
energy-based model
0.812024
Protein Discovery with Discrete Walk-Jump Sampling · ICLR 2024
Machine learning › Generative modeling › diffusion model › controllable generation
guided generation
0.812024
Implicitly Guided Design with PropEn: Match your Data to Follow the Gradient · NeurIPS 2024
Machine learning › Generative modeling › diffusion model › geometric diffusion model
equivariant diffusion model
0.712023
AbDiffuser: full-atom generation of in-vitro functioning antibodies · NeurIPS 2023
Machine learning › Generative modeling › protein design
protein structure generation
0.712023
AbDiffuser: full-atom generation of in-vitro functioning antibodies · NeurIPS 2023
Bioinformatics and computational biology › protein design
antibody design
0.712023
AbDiffuser: full-atom generation of in-vitro functioning antibodies · NeurIPS 2023
Bioinformatics and computational biology
immunoinformatics
0.712023
TCRconv: predicting recognition between T cell receptors and epitopes using contextualized motifs · Bioinform. 2023
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model
0.712023
TCRconv: predicting recognition between T cell receptors and epitopes using contextualized motifs · Bioinform. 2023
Data mining › structured data mining › graph mining
community detection
0.412019
Non-Negative Matrix Factorizations for Multiplex Network Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Data mining › clustering
graph clustering
0.412019
Non-Negative Matrix Factorizations for Multiplex Network Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Recommender systems › collaborative filtering
matrix factorization
0.412019
Non-Negative Matrix Factorizations for Multiplex Network Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Data mining › dimensionality reduction
nonnegative matrix factorization
0.412019
Non-Negative Matrix Factorizations for Multiplex Network Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network integration
0.322018
Integration of molecular network data reconstructs Gene Ontology · Bioinform. 2014
deepNF: deep network fusion for protein function prediction · Bioinform. 2018
Computer vision › 3D vision › implicit neural representation
neural field
0.312025
Unified all-atom molecule generation with neural fields · NeurIPS 2025
Bioinformatics and computational biology › protein analysis › protein-protein interaction › protein-protein interaction network analysis
multiple network alignment
0.212016
Fuse: multiple network alignment via data fusion · Bioinform. 2016
Bioinformatics and computational biology › protein analysis › protein-protein interaction › protein-protein interaction network analysis
protein interaction network alignment
0.212016
Fuse: multiple network alignment via data fusion · Bioinform. 2016
Bioinformatics and computational biology
protein structure prediction
0.212023
AbDiffuser: full-atom generation of in-vitro functioning antibodies · NeurIPS 2023
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction network
0.112018
deepNF: deep network fusion for protein function prediction · Bioinform. 2018
Bioinformatics and computational biology
functional similarity
0.112016
Fuse: multiple network alignment via data fusion · Bioinform. 2016
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction network analysis
0.112016
Fuse: multiple network alignment via data fusion · Bioinform. 2016
Bioinformatics and computational biology › network bioinformatics › biological network analysis
molecular network analysis
0.112014
Integration of molecular network data reconstructs Gene Ontology · Bioinform. 2014

Methods — techniques the papers use, named apart from their topics

score-based generative model · 1.7preference learning · 1.7neural field · 1.7black-box optimization · 1.7bi-level optimization · 1.7denoising · 1.5contrastive divergence · 1.5Langevin MCMC · 1.5language model · 0.9contrastive learning · 0.9score-based model · 0.8encoder-decoder architecture · 0.8low-dimensional feature representation · 0.4collective factorization · 0.4
YearPublicationVenuePosition
2025 Generalists vs. Specialists: Evaluating LLMs on Highly-Constrained Biophysical Sequence Optimization Tasks
abstract
Although large language models (LLMs) have shown promise in biomolecule optimization problems, they incur heavy computational costs and struggle to satisfy precise constraints. On the other hand, specialized solvers like LaMBO-2 offer efficiency and fine-grained control but require more domain expertise. Comparing these approaches is challenging due to expensive laboratory validation and inadequate synthetic benchmarks. We address this by introducing Ehrlich functions, a synthetic test suite that captures the geometric structure of biophysical sequence optimization problems. With prompting alone, off-the-shelf LLMs struggle to optimize Ehrlich functions. In response, we propose LLOME (Language Model Optimization with Margin Expectation), a bilevel optimization routine for online black-box optimization. When combined with a novel preference learning loss, we find LLOME can not only learn to solve some Ehrlich functions, but can even perform as well as or better than LaMBO-2 on moderately difficult Ehrlich variants. However, LLMs also exhibit some likelihood-reward miscalibration and struggle without explicit rewards. Our results indicate LLMs can occasionally provide significant benefits, but specialized solvers are still competitive and incur less overhead.
Angelica Chen, Samuel Stanton, Frances Ding, Robert G. Alberstein, Andrew M. Watkins, Richard Bonneau, Vladimir Gligorijevic, Kyunghyun Cho, Nathan C. Frey
ICML7
2025 Unified all-atom molecule generation with neural fields
abstract
Generative models for structure-based drug design are often limited to a specific modality, restricting their broader applicability. To address this challenge, we introduce FuncBind, a framework based on computer vision to generate target-conditioned, all-atom molecules across atomic systems. FuncBind uses neural fields to represent molecules as continuous atomic densities and employs score-based generative models with modern architectures adapted from the computer vision literature. This modality-agnostic representation allows a single unified model to be trained on diverse atomic systems, from small to large molecules, and handle variable atom/residue counts, including non-canonical amino acids. FuncBind achieves competitive in silico performance in generating small molecules, macrocyclic peptides, and antibody complementarity-determining region loops, conditioned on target structures. FuncBind also generated in vitro novel antibody binders via de novo redesign of the complementarity-determining region H3 loop of two chosen co-crystal structures. As a final contribution, we introduce a new dataset and benchmark for structure-conditioned macrocyclic peptide generation.
Matthieu Kirchmeyer, Pedro O. Pinheiro, Emma Willett, Karolis Martinkus, Joseph Kleinhenz, Emily K. Makowski, Andrew M. Watkins, Vladimir Gligorijevic, Richard Bonneau, Saeed Saremi
NeurIPS8
2025 deepNGS navigator: exploring antibody NGS datasets using deep contrastive learning
abstract
MOTIVATION: High-throughput sequencing uncovers how B-cells adapt in response to antigens by generating B-cell-receptor (BCR) sequences at an unprecedented scale. As BCR datasets grow to millions of sequences, using efficient computational methods becomes crucial. One important aspect of antibody sequence analysis is detecting clonal families or clusters of related sequences, whether they come from immunization, synthetic-libraries or even ML-generated datasets. RESULTS: We introduce deepNGS Navigator, a computational tool that leverages language models and contrastive learning to transform antibody sequences into intuitive 2D representations. The resulting 2D maps offer a visualization of overall diversity of input datasets, which can be clustered based on the sequence distances and their densities across the map. Beyond grouping related sequences, the 2D maps also represent mutational patterns inferred from sequence embeddings, enabling trajectory analysis and clustering within the projected space. By overlaying properties such as charge, the map helps identify clusters of interest for further investigation while also flagging potentially noisy or non-specific sequences with higher risk. We demonstrate deepNGS Navigator's utilities on several datasets, including: (i) a synthetic-library from a yeast-display targeting HER2, (ii) a machine learning-generated dataset with a hierarchical structure, (iii) NGS sequences from a llama immunized against COVID RBD, (iv) human naive and memory B-cell sequences, and (v) an in silico dataset simulating B-cell clonal lineages. AVAILABILITY AND IMPLEMENTATION: The deepNGS Navigator source code is available at: github.com/prescient-design/deepngs-navigator and github.com/prescient-design/deepngs-navigator-panel-app.
Homa Mohammadi Peyhani, Edith Lee, Richard Bonneau, Vladimir Gligorijevic, Jae Hyeon Lee
Bioinform.4
2024 Protein Discovery with Discrete Walk-Jump Sampling
abstract
We resolve difficulties in training and sampling from a discrete generative model by learning a smoothed energy function, sampling from the smoothed data manifold with Langevin Markov chain Monte Carlo (MCMC), and projecting back to the true data manifold with one-step denoising. Our $\textit{Discrete Walk-Jump Sampling}$ formalism combines the contrastive divergence training of an energy-based model and improved sample quality of a score-based model, while simplifying training and sampling by requiring only a single noise level. We evaluate the robustness of our approach on generative modeling of antibody proteins and introduce the $\textit{distributional conformity score}$ to benchmark protein generative models. By optimizing and sampling from our models for the proposed distributional conformity score, 97-100\% of generated samples are successfully expressed and purified and 70\% of functional designs show equal or improved binding affinity compared to known functional antibodies on the first attempt in a single round of laboratory experiments. We also report the first demonstration of long-run fast-mixing MCMC chains where diverse antibody protein classes are visited in a single MCMC chain.
Nathan C. Frey, Daniel Berenberg, Karina Zadorozhny, Joseph Kleinhenz, Julien Lafrance-Vanasse, Isidro Hötzel, Yan Wu 0027, Stephen Ra, Richard Bonneau, Kyunghyun Cho, Andreas Loukas, Vladimir Gligorijevic, Saeed Saremi
ICLR12
2024 Implicitly Guided Design with PropEn: Match your Data to Follow the Gradient
abstract
Across scientific domains, generating new models or optimizing existing ones while meeting specific criteria is crucial. Traditional machine learning frameworks for guided design use a generative model and a surrogate model (discriminator), requiring large datasets. However, real-world scientific applications often have limited data and complex landscapes, making data-hungry models inefficient or impractical. We propose a new framework, PropEn, inspired by ``matching'', which enables implicit guidance without training a discriminator. By matching each sample with a similar one that has a better property value, we create a larger training dataset that inherently indicates the direction of improvement. Matching, combined with an encoder-decoder architecture, forms a domain-agnostic generative framework for property enhancement. We show that training with a matched dataset approximates the gradient of the property of interest while remaining within the data distribution, allowing efficient design optimization. Extensive evaluations in toy problems and scientific applications, such as therapeutic protein design and airfoil optimization, demonstrate PropEn's advantages over common baselines. Notably, the protein design results are validated with wet lab experiments, confirming the competitiveness and effectiveness of our approach. Our code is available at https://github.com/prescient-design/propen.
Natasa Tagasovska, Vladimir Gligorijevic, Kyunghyun Cho, Andreas Loukas
NeurIPS2
2023 AbDiffuser: full-atom generation of in-vitro functioning antibodies
abstract
We introduce AbDiffuser, an equivariant and physics-informed diffusion model for the joint generation of antibody 3D structures and sequences. AbDiffuser is built on top of a new representation of protein structure, relies on a novel architecture for aligned proteins, and utilizes strong diffusion priors to improve the denoising process. Our approach improves protein diffusion by taking advantage of domain knowledge and physics-based constraints; handles sequence-length changes; and reduces memory complexity by an order of magnitude, enabling backbone and side chain generation. We validate AbDiffuser in silico and in vitro. Numerical experiments showcase the ability of AbDiffuser to generate antibodies that closely track the sequence and structural properties of a reference set. Laboratory experiments confirm that all 16 HER2 antibodies discovered were expressed at high levels and that 57.1% of the selected designs were tight binders.
Karolis Martinkus, Jan Ludwiczak, Wei-Ching Liang, Julien Lafrance-Vanasse, Isidro Hötzel, Arvind Rajpal, Yan Wu 0027, Kyunghyun Cho, Richard Bonneau, Vladimir Gligorijevic, Andreas Loukas
NeurIPS10
2023 TCRconv: predicting recognition between T cell receptors and epitopes using contextualized motifs
abstract
MOTIVATION: T cells use T cell receptors (TCRs) to recognize small parts of antigens, called epitopes, presented by major histocompatibility complexes. Once an epitope is recognized, an immune response is initiated and T cell activation and proliferation by clonal expansion begin. Clonal populations of T cells with identical TCRs can remain in the body for years, thus forming immunological memory and potentially mappable immunological signatures, which could have implications in clinical applications including infectious diseases, autoimmunity and tumor immunology. RESULTS: We introduce TCRconv, a deep learning model for predicting recognition between TCRs and epitopes. TCRconv uses a deep protein language model and convolutions to extract contextualized motifs and provides state-of-the-art TCR-epitope prediction accuracy. Using TCR repertoires from COVID-19 patients, we demonstrate that TCRconv can provide insight into T cell dynamics and phenotypes during the disease. AVAILABILITY AND IMPLEMENTATION: TCRconv is available at https://github.com/emmijokinen/tcrconv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Emmi Jokinen, Alexandru Dumitrescu, Jani Huuhtanen, Vladimir Gligorijevic, Satu Mustjoki, Richard Bonneau, Markus Heinonen, Harri Lähdesmäki
Bioinform.4
2021 NetQuilt: deep multispecies network-based protein function prediction using homology-informed network similarity
abstract
MOTIVATION: Transferring knowledge between species is challenging: different species contain distinct proteomes and cellular architectures, which cause their proteins to carry out different functions via different interaction networks. Many approaches to protein functional annotation use sequence similarity to transfer knowledge between species. These approaches cannot produce accurate predictions for proteins without homologues of known function, as many functions require cellular context for meaningful prediction. To supply this context, network-based methods use protein-protein interaction (PPI) networks as a source of information for inferring protein function and have demonstrated promising results in function prediction. However, most of these methods are tied to a network for a single species, and many species lack biological networks. RESULTS: In this work, we integrate sequence and network information across multiple species by computing IsoRank similarity scores to create a meta-network profile of the proteins of multiple species. We use this integrated multispecies meta-network as input to train a maxout neural network with Gene Ontology terms as target labels. Our multispecies approach takes advantage of more training examples, and consequently leads to significant improvements in function prediction performance compared to two network-based methods, a deep learning sequence-based method and the BLAST annotation method used in the Critial Assessment of Functional Annotation. We are able to demonstrate that our approach performs well even in cases where a species has no network information available: when an organism's PPI network is left out we can use our multi-species method to make predictions for the left-out organism with good performance. AVAILABILITY AND IMPLEMENTATION: The code is freely available at https://github.com/nowittynamesleft/NetQuilt. The data, including sequences, PPI networks and GO annotations are available at https://string-db.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Meet Barot, Vladimir Gligorijevic, Kyunghyun Cho, Richard Bonneau
Bioinform.2
2019 Non-Negative Matrix Factorizations for Multiplex Network Analysis
abstract
Networks have been a general tool for representing, analyzing, and modeling relational data arising in several domains. One of the most important aspect of network analysis is community detection or network clustering. Until recently, the major focus have been on discovering community structure in single (i.e., monoplex) networks. However, with the advent of relational data with multiple modalities, multiplex networks, i.e., networks composed of multiple layers representing different aspects of relations, have emerged. Consequently, community detection in multiplex network, i.e., detecting clusters of nodes shared by all layers, has become a new challenge. In this paper, we propose Network Fusion for Composite Community Extraction (NF-CCE), a new class of algorithms, based on four different non-negative matrix factorization models, capable of extracting composite communities in multiplex networks. Each algorithm works in two steps: first, it finds a non-negative, low-dimensional feature representation of each network layer; then, it fuses the feature representation of layers into a common non-negative, low-dimensional feature representation via collective factorization. The composite clusters are extracted from the common feature representation. We demonstrate the superior performance of our algorithms over the state-of-the-art methods on various types of multiplex networks, including biological, social, economic, citation, phone communication, and brain multiplex networks.
Vladimir Gligorijevic, Yannis Panagakis, Stefanos Zafeiriou
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 deepNF: deep network fusion for protein function prediction
abstract
Motivation: The prevalence of high-throughput experimental methods has resulted in an abundance of large-scale molecular and functional interaction networks. The connectivity of these networks provides a rich source of information for inferring functional annotations for genes and proteins. An important challenge has been to develop methods for combining these heterogeneous networks to extract useful protein feature representations for function prediction. Most of the existing approaches for network integration use shallow models that encounter difficulty in capturing complex and highly non-linear network structures. Thus, we propose deepNF, a network fusion method based on Multimodal Deep Autoencoders to extract high-level features of proteins from multiple heterogeneous interaction networks. Results: We apply this method to combine STRING networks to construct a common low-dimensional representation containing high-level protein features. We use separate layers for different network types in the early stages of the multimodal autoencoder, later connecting all the layers into a single bottleneck layer from which we extract features to predict protein function. We compare the cross-validation and temporal holdout predictive performance of our method with state-of-the-art methods, including the recently proposed method Mashup. Our results show that our method outperforms previous methods for both human and yeast STRING networks. We also show substantial improvement in the performance of our method in predicting gene ontology terms of varying type and specificity. Availability and implementation: deepNF is freely available at: https://github.com/VGligorijevic/deepNF. Supplementary information: Supplementary data are available at Bioinformatics online.
Vladimir Gligorijevic, Meet Barot, Richard Bonneau
Bioinform.1
2016 Fusion and community detection in multi-layer graphs
abstract
Relational data arising in many domains can be represented by networks (or graphs) with nodes capturing entities and edges representing relationships between these entities. Community detection in networks has become one of the most important problems having a broad range of applications. Until recently, the vast majority of papers have focused on discovering community structures in a single network. However, with the emergence of multi-view network data in many real-world applications and consequently with the advent of multilayer graph representation, community detection in multi-layer graphs has become a new challenge. Multi-layer graphs provide complementary views of connectivity patterns of the same set of vertices. Fusion of the network layers is expected to achieve better clustering performance. In this paper, we propose two novel methods, coined as WSSNMTF (Weighted Simultaneous Symmetric Non-Negative Matrix Tri-Factorization) and NG-WSSNMTF (Natural Gradient WSSNMTF), for fusion and clustering of multi-layer graphs. Both methods are robust with respect to missing edges and noise. We compare the performance of the proposed methods with two baseline methods, as well as with three state-of-the-art methods on synthetic and three real-world datasets. The experimental results indicate superior performance of the proposed methods.
Vladimir Gligorijevic, Yannis Panagakis, Stefanos Zafeiriou
ICPR1
2016 Fuse: multiple network alignment via data fusion
abstract
MOTIVATION: Discovering patterns in networks of protein-protein interactions (PPIs) is a central problem in systems biology. Alignments between these networks aid functional understanding as they uncover important information, such as evolutionary conserved pathways, protein complexes and functional orthologs. However, the complexity of the multiple network alignment problem grows exponentially with the number of networks being aligned and designing a multiple network aligner that is both scalable and that produces biologically relevant alignments is a challenging task that has not been fully addressed. The objective of multiple network alignment is to create clusters of nodes that are evolutionarily and functionally conserved across all networks. Unfortunately, the alignment methods proposed thus far do not meet this objective as they are guided by pairwise scores that do not utilize the entire functional and evolutionary information across all networks. RESULTS: To overcome this weakness, we propose Fuse, a new multiple network alignment algorithm that works in two steps. First, it computes our novel protein functional similarity scores by fusing information from wiring patterns of all aligned PPI networks and sequence similarities between their proteins. This is in contrast with the previous tools that are all based on protein similarities in pairs of networks being aligned. Our comprehensive new protein similarity scores are computed by Non-negative Matrix Tri-Factorization (NMTF) method that predicts associations between proteins whose homology (from sequences) and functioning similarity (from wiring patterns) are supported by all networks. Using the five largest and most complete PPI networks from BioGRID, we show that NMTF predicts a large number protein pairs that are biologically consistent. Second, to identify clusters of aligned proteins over all networks, Fuse uses our novel maximum weight k-partite matching approximation algorithm. We compare Fuse with the state of the art multiple network aligners and show that (i) by using only sequence alignment scores, Fuse already outperforms other aligners and produces a larger number of biologically consistent clusters that cover all aligned PPI networks and (ii) using both sequence alignments and topological NMTF-predicted scores leads to the best multiple network alignments thus far. AVAILABILITY AND IMPLEMENTATION: Our dataset and software are freely available from the web site: http://bio-nets.doc.ic.ac.uk/Fuse/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vladimir Gligorijevic, Noël Malod-Dognin, Natasa Przulj
Bioinform.1
2014 Integration of molecular network data reconstructs Gene Ontology
abstract
MOTIVATION: Recently, a shift was made from using Gene Ontology (GO) to evaluate molecular network data to using these data to construct and evaluate GO. Dutkowski et al. provide the first evidence that a large part of GO can be reconstructed solely from topologies of molecular networks. Motivated by this work, we develop a novel data integration framework that integrates multiple types of molecular network data to reconstruct and update GO. We ask how much of GO can be recovered by integrating various molecular interaction data. RESULTS: We introduce a computational framework for integration of various biological networks using penalized non-negative matrix tri-factorization (PNMTF). It takes all network data in a matrix form and performs simultaneous clustering of genes and GO terms, inducing new relations between genes and GO terms (annotations) and between GO terms themselves. To improve the accuracy of our predicted relations, we extend the integration methodology to include additional topological information represented as the similarity in wiring around non-interacting genes. Surprisingly, by integrating topologies of bakers' yeasts protein-protein interaction, genetic interaction (GI) and co-expression networks, our method reports as related 96% of GO terms that are directly related in GO. The inclusion of the wiring similarity of non-interacting genes contributes 6% to this large GO term association capture. Furthermore, we use our method to infer new relationships between GO terms solely from the topologies of these networks and validate 44% of our predictions in the literature. In addition, our integration method reproduces 48% of cellular component, 41% of molecular function and 41% of biological process GO terms, outperforming the previous method in the former two domains of GO. Finally, we predict new GO annotations of yeast genes and validate our predictions through GIs profiling. AVAILABILITY AND IMPLEMENTATION: Supplementary Tables of new GO term associations and predicted gene annotations are available at http://bio-nets.doc.ic.ac.uk/GO-Reconstruction/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vladimir Gligorijevic, Vuk Janjic, Natasa Przulj
Bioinform.1