VLDB 2026 Research / reviewers in the wild / expert
Georgios A. Pavlopoulos
dblp:36/4029
· DBLP profile ↗
19ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-4577-8276ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoding extremophiles: insights from bioinformatics, machine learning, and data-driven approachesabstractLife thrives in Earth's most inhospitable environments, from boiling hydrothermal vents to hypersaline lakes and frozen polar deserts, thanks to the remarkable adaptations of extremophilic microorganisms. The study of these organisms has rapidly evolved from early cultivation-based discoveries to a data-rich discipline powered by advanced omics technologies. This review comprehensively outlines the current landscape and future directions in extremophile research, emphasizing the pivotal role of bioinformatics, machine learning (ML), and data-driven approaches. We begin by charting the evolution of methodologies, from innovative in situ cultivation techniques and robust biomolecule extraction protocols to modern multi-omics workflows (metagenomics, transcriptomics, proteomics, and metabolomics) that decode the genetic and functional basis of extremophiles. We then catalogue essential bioinformatics resources and specialized databases critical for annotating extremophile genomes and uncovering their unique adaptive strategies, including protein stabilization and syntrophic metabolic relationships. Finally, we explore the transformative potential of artificial intelligence (AI) and ML in overcoming fundamental challenges in the field. These include predicting the functions of uncharacterized "hypothetical" proteins, identifying novel extremozymes, modeling complex genotype-phenotype relationships, and guiding the targeted engineering of industrially relevant strains. By synthesizing insights across these domains, this review highlights how integrating computational biology and AI is poised to unlock the full biotechnological potential of extremophiles and redefine the boundaries of life itself. Maria N. Chasapi, Nicholas Kontis, Robert Lehmann 0002, Ruqaiya Tasneem, Niketan S. Patel, Sumeer Ahmad Khan, Xabier Martinez de Morentin, Iro N. Chasapi, Eleni Aplakidou, Alexandros Galaras, Lila Aldakheel, Minjing Su, Fotis A. Baltoumas, Kasthuri Venkateswaran, Vincenzo Lagani, David Gomez-Cabrero, Jesper Tegnér, Georgios A. Pavlopoulos, Alexandre Soares Rosado |
Briefings Bioinform. | 18 |
| 2026 | Accelerating inference in genomic and proteomic foundation models via speculative decodingabstractMOTIVATION: Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose inference is relatively slow, long-sequence generation quickly becomes costly. RESULTS: In this work we adapt speculative decoding to a representative GFM: the DNA model DNAGPT and two representative PFMs: ProGen2 and ProtGPT2. We implement a probabilistic variant of speculative decoding, in which a lightweight draft model proposes short token spans and a larger target model verifies or corrects them in parallel, while preserving the target model's sampling distribution. Across all three models we systematically study the effect of speculation window length, temperature, draft architecture and prompt length, and we benchmark tokens per second over multiple runs per configuration. Speculative decoding yields consistent speedups over standard key-value cached decoding, with maximum observed speedup reaching 100% increase, while average gains across models ranging between 20% and 40% (e.g. 1.2×-1.4×), without changing the underlying target model predictions. Our results show that speculative decoding is a practical and model-agnostic strategy for accelerating genomic and proteomic sequence generation without sacrificing prediction quality. AVAILABILITY AND IMPLEMENTATION: All code and results are freely available at https://github.com/Georgakopoulos-Soares-lab/BioSpecDec. Kimonas Provatas, Aris Karatzikos, Charalampos Koilakos, Michail Patsakis, Alexandros Tzanakakis, Akshatha Nayak, Georgios A. Pavlopoulos, Ioannis Mouratidis, Evangelos Ioannis Avgoulas, Ilias Georgakopoulos-Soares |
Bioinform. | 7 |
| 2025 | MAFin: motif detection in multiple alignment filesabstractMOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin. Michail Patsakis, Kimonas Provatas, Fotis A. Baltoumas, Nikol Chantzi, Ioannis Mouratidis, Georgios A. Pavlopoulos, Ilias Georgakopoulos-Soares |
Bioinform. | 6 |
| 2023 | Flame (v2.0): advanced integration and interpretation of functional enrichment results from multiple sourcesabstractFunctional enrichment is the process of identifying implicated functional terms from a given input list of genes or proteins. In this article, we present Flame (v2.0), a web tool which offers a combinatorial approach through merging and visualizing results from widely used functional enrichment applications while also allowing various flexible input options. In this version, Flame utilizes the aGOtool, g: Profiler, WebGestalt, and Enrichr pipelines and presents their outputs separately or in combination following a visual analytics approach. For intuitive representations and easier interpretation, it uses interactive plots such as parameterizable networks, heatmaps, barcharts, and scatter plots. Users can also: (i) handle multiple protein/gene lists and analyse union and intersection sets simultaneously through interactive UpSet plots, (ii) automatically extract genes and proteins from free text through text-mining and Named Entity Recognition (NER) techniques, (iii) upload single nucleotide polymorphisms (SNPs) and extract their relative genes, or (iv) analyse multiple lists of differentially expressed proteins/genes after selecting them interactively from a parameterizable volcano plot. Compared to the previous version of 197 supported organisms, Flame (v2.0) currently allows enrichment for 14 436 organisms. AVAILABILITY AND IMPLEMENTATION: Web Application: http://flame.pavlopouloslab.info. Code: https://github.com/PavlopoulosLab/Flame. Docker: https://hub.docker.com/r/pavlopouloslab/flame. Evangelos Karatzas, Fotis A. Baltoumas, Eleni Aplakidou, Panagiota I. Kontou, Panos Stathopoulos, Leonidas Stefanis, Pantelis G. Bagos, Georgios A. Pavlopoulos |
Bioinform. | 8 |
| 2022 | Extreme-Scale Many-against-Many Protein Similarity SearchabstractSimilarity search is one of the most fundamental computations that are regularly performed on ever-increasing protein datasets. Scalability is of paramount importance for uncovering novel phenomena that occur at very large scales. We unleash the power of over 20,000 GPUs on the Summit system to perform all-vs-all protein similarity search on one of the largest publicly available datasets with 405 million proteins, in less than 3.5 hours, cutting the time-to-solution for many use cases from weeks. The variability of protein sequence lengths, as well as the sparsity of the space of pairwise comparisons, make this a challenging problem in distributed memory. Due to the need to construct and maintain a data structure holding indices to all other sequences, this application has a huge memory footprint that makes it hard to scale the problem sizes. We overcome this memory limitation by innovative matrix-based blocking techniques, without introducing additional load imbalance. Oguz Selvitopi, Saliya Ekanayake, Giulia Guidi, Muaaz Gul Awan, Georgios A. Pavlopoulos, Ariful Azad, Nikos Kyrpides, Leonid Oliker, Katherine A. Yelick, Aydin Buluç |
SC | 5 |
| 2020 | Distributed many-to-many protein sequence alignment using sparse matricesabstractIdentifying similar protein sequences is a core step in many computational biology pipelines such as detection of homologous protein sequences, generation of similarity protein graphs for downstream analysis, functional annotation, and gene location. Performance and scalability of protein similarity search have proven to be a bottleneck in many bioinformatics pipelines due to increase in cheap and abundant sequencing data. This work presents a new distributed-memory software PASTIS. PASTIS relies on sparse matrix computations for efficient identification of possibly similar proteins. We use distributed sparse matrices for scalability and show that the sparse matrix infrastructure is a great fit for protein similarity search when coupled with a fully-distributed dictionary of sequences that allow remote sequence requests to be fulfilled. Our algorithm incorporates the unique bias in amino acid sequence substitution in search without altering basic sparse matrix model, and in turn, achieves ideal scaling up to millions of protein sequences. Oguz Selvitopi, Saliya Ekanayake, Giulia Guidi, Georgios A. Pavlopoulos, Ariful Azad, Aydin Buluç |
SC | 4 |
| 2016 | DrugQuest - a text mining workflow for drug association discoveryabstractBACKGROUND: Text mining and data integration methods are gaining ground in the field of health sciences due to the exponential growth of bio-medical literature and information stored in biological databases. While such methods mostly try to extract bioentity associations from PubMed, very few of them are dedicated in mining other types of repositories such as chemical databases. RESULTS: Herein, we apply a text mining approach on the DrugBank database in order to explore drug associations based on the DrugBank "Description", "Indication", "Pharmacodynamics" and "Mechanism of Action" text fields. We apply Name Entity Recognition (NER) techniques on these fields to identify chemicals, proteins, genes, pathways, diseases, and we utilize the TextQuest algorithm to find additional biologically significant words. Using a plethora of similarity and partitional clustering techniques, we group the DrugBank records based on their common terms and investigate possible scenarios why these records are clustered together. Different views such as clustered chemicals based on their textual information, tag clouds consisting of Significant Terms along with the terms that were used for clustering are delivered to the user through a user-friendly web interface. CONCLUSIONS: DrugQuest is a text mining tool for knowledge discovery: it is designed to cluster DrugBank records based on text attributes in order to find new associations between drugs. The service is freely available at http://bioinformatics.med.uoc.gr/drugquest . Nikolas Papanikolaou, Georgios A. Pavlopoulos, Theodosios Theodosiou, Ioannis S. Vizirianakis |
BMC Bioinform. | 2 |
| 2015 | BioTextQuest+: a knowledge integration platform for literature mining and concept discoveryabstractBioinformatics (2014); 30(22), 3249–3256 doi: 10.1093/bioinformatics/btu524 The above article contained an incorrect email address for the corresponding author, Ioannis Iliopoulos, the correct email address is [email protected] Nikolas Papanikolaou, Georgios A. Pavlopoulos, Evangelos Pafilis, Theodosios Theodosiou, Reinhard Schneider 0002, Venkata P. Satagopam, Christos A. Ouzounis, Aristides G. Eliopoulos, Vasilis J. Promponas |
Bioinform. | 2 |
| 2014 | BioTextQuest+: a knowledge integration platform for literature mining and concept discoveryabstractSUMMARY: The iterative process of finding relevant information in biomedical literature and performing bioinformatics analyses might result in an endless loop for an inexperienced user, considering the exponential growth of scientific corpora and the plethora of tools designed to mine PubMed(®) and related biological databases. Herein, we describe BioTextQuest(+), a web-based interactive knowledge exploration platform with significant advances to its predecessor (BioTextQuest), aiming to bridge processes such as bioentity recognition, functional annotation, document clustering and data integration towards literature mining and concept discovery. BioTextQuest(+) enables PubMed and OMIM querying, retrieval of abstracts related to a targeted request and optimal detection of genes, proteins, molecular functions, pathways and biological processes within the retrieved documents. The front-end interface facilitates the browsing of document clustering per subject, the analysis of term co-occurrence, the generation of tag clouds containing highly represented terms per cluster and at-a-glance popup windows with information about relevant genes and proteins. Moreover, to support experimental research, BioTextQuest(+) addresses integration of its primary functionality with biological repositories and software tools able to deliver further bioinformatics services. The Google-like interface extends beyond simple use by offering a range of advanced parameterization for expert users. We demonstrate the functionality of BioTextQuest(+) through several exemplary research scenarios including author disambiguation, functional term enrichment, knowledge acquisition and concept discovery linking major human diseases, such as obesity and ageing. AVAILABILITY: The service is accessible at http://bioinformatics.med.uoc.gr/biotextquest. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nikolas Papanikolaou, Georgios A. Pavlopoulos, Evangelos Pafilis, Theodosios Theodosiou, Reinhard Schneider 0002, Venkata P. Satagopam, Christos A. Ouzounis, Aristides G. Eliopoulos, Vasilis J. Promponas |
Bioinform. | 2 |
| 2013 | OnTheFly 2.0: A tool for automatic annotation of files and biological information extractionabstractRetrieving all of the necessary information from databases about bioentities mentioned in an article is not a trivial or an easy task. Following the daily literature about a specific biological topic and collecting all the necessary information about the bioentities mentioned in the literature manually is tedious and time consuming. OnTheFly 2.0 is a web application mainly designed for non-computer experts which aims to automate data collection and knowledge extraction from biological literature in a user friendly and efficient way. OnTheFly 2.0 is able to extract bioentities from individual articles such as text, Microsoft Word, Excel and PDF files. With a simple drag-and-drop motion, the text of a document is extensively parsed for bioentities such as protein/gene names and chemical compound names. Utilizing high quality data integration platforms, OnTheFly allows the generation of informative summaries, interaction networks and at-a-glance popup windows containing knowledge related to the bioentities found in documents. OnTheFly 2.0 provides a concise application to automate the extraction of bioentities hidden in various documents and is offered as a web based application. It can be found at: http://onthefly.embl.de, http://onthefly.med.uoc.gr or http://onthefly.hcmr.gr. Evangelos Pafilis, Georgios A. Pavlopoulos, Venkata P. Satagopam, Nikolas Papanikolaou, Heiko Horn, Christos Arvanitidis, Lars Juhl Jensen, Reinhard Schneider 0002 |
BIBE | 2 |
| 2012 | Visualizing high dimensional datasets using parallel coordinates: Application to gene prioritizationabstractIn this paper, we introduce a visualization tool for interactive and efficient exploration of high dimensional data using parallel coordinates. An algorithm is developed to find an optimal permutation of dimensions, which allows the data miner to immediately see the most important features or irregularities in the dataset. This is implemented as a genetic algorithm based on the travelling salesman problem using maximal correlation as fitness. Other features of the tool include selection operators to group the data such as selection by intersection or by angle, orthogonal and density plots complementing the parallel coordinates plot, manual arrangement of permutation order of the dimensions, possibility to show all plots necessary to see all dimensional relations and displaying a certain number of standard deviations for each dimension separately. The tool is applied to multiple gene prioritization cases in search of genes that are relevant to certain genetic disorders. The used datasets are obtained with the MerKator and Endeavour tools and include a Breast cancer, Cataract, Charcoth-Marie-Tooth and Cardiomyopathy dataset, as well as a dataset relating 29 diseases with 22206 genes. Our tool, manual and data can be downloaded from http://www.toomas.be/parcoord/. Thomas Boogaerts, Léon-Charles Tranchevent, Georgios A. Pavlopoulos, Jan Aerts, Joos Vandewalle |
BIBE | 3 |
| 2012 | ReLiance: a machine learning and literature-based prioritization of receptor - ligand pairingsabstractMOTIVATION: The prediction of receptor-ligand pairings is an important area of research as intercellular communications are mediated by the successful interaction of these key proteins. As the exhaustive assaying of receptor-ligand pairs is impractical, a computational approach to predict pairings is necessary. We propose a workflow to carry out this interaction prediction task, using a text mining approach in conjunction with a state of the art prediction method, as well as a widely accessible and comprehensive dataset. Among several modern classifiers, random forests have been found to be the best at this prediction task. The training of this classifier was carried out using an experimentally validated dataset of Database of Ligand-Receptor Partners (DLRP) receptor-ligand pairs. New examples, co-cited with the training receptors and ligands, are then classified using the trained classifier. After applying our method, we find that we are able to successfully predict receptor-ligand pairs within the GPCR family with a balanced accuracy of 0.96. Upon further inspection, we find several supported interactions that were not present in the Database of Interacting Proteins (DIPdatabase). We have measured the balanced accuracy of our method resulting in high quality predictions stored in the available database ReLiance. AVAILABILITY: http://homes.esat.kuleuven.be/~bioiuser/ReLianceDB/index.php CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ernesto Iacucci, Léon-Charles Tranchevent, Dusan Popovic, Georgios A. Pavlopoulos, Bart De Moor, Reinhard Schneider 0002, Yves Moreau |
Bioinform. | 4 |
| 2012 | A bioinformatics e-dating story: computational prediction and prioritization of receptor-ligand pairsabstractRegulation of cellular events is initiated, often, via extracellular signaling when a circulating protein ligand interacts with one or more membrane-bound protein receptors. Identification of receptor-ligand pairs is thus an important and difficult task to address as this form of interaction is transient and not well studied. In order to address this problem, we collect the most readily available data from repositories (expression, domain, pathway, sequence, and text-based), and apply a high through-put analysis to this problem. We have worked on the receptor-ligand pairing problem in three main studies. In our first study, using a LS-SVM classifier, we show that we are able to more aptly match members of the chemokine and tgfβ families than a previously published method [ 1 ]. Notably, we are able to achieve an increase in recall of 0.76 over the 0.44 for the matching of receptor-ligands in the tgfβ family. In our subsequent study, we benchmarked several machine learning techniques, and essayed several parameters, on the receptior-ligand interaction prediction task. We found that we could reach a balanced accuracy of 0.84. In our final work, we produce a publicly available database of our results with respect to a text-based in silico prediction workflow. The resulting database, contains several key findings, particularly predictions in the GPCR family with a balanced accuracy of 0.96. The receptor-ligand prediction task is an essential one, as the challenge of predicting such pairs is an important issue in wet-labs, biotech, and pharmaceutical companies. Through several studies, we have determined the most appropriate methodology to predict the receptor-ligand pairs and have made available high-quality predictions at our ReLianceDB website ( http://homes.esat.kuleuven.be/~bioiuser/ReLianceDB ), a tool to aid in performing effective and targeted research. Ernesto Iacucci, Léon-Charles Tranchevent, Dusan Popovic, Georgios A. Pavlopoulos, Bart De Moor, Reinhard Schneider 0002, Yves Moreau |
BMC Bioinform. | 4 |
| 2012 | Arena3D: visualizing time-driven phenotypic differences in biological systemsabstractBACKGROUND: Elucidating the genotype-phenotype connection is one of the big challenges of modern molecular biology. To fully understand this connection, it is necessary to consider the underlying networks and the time factor. In this context of data deluge and heterogeneous information, visualization plays an essential role in interpreting complex and dynamic topologies. Thus, software that is able to bring the network, phenotypic and temporal information together is needed. Arena3D has been previously introduced as a tool that facilitates link discovery between processes. It uses a layered display to separate different levels of information while emphasizing the connections between them. We present novel developments of the tool for the visualization and analysis of dynamic genotype-phenotype landscapes. RESULTS: Version 2.0 introduces novel features that allow handling time course data in a phenotypic context. Gene expression levels or other measures can be loaded and visualized at different time points and phenotypic comparison is facilitated through clustering and correlation display or highlighting of impacting changes through time. Similarity scoring allows the identification of global patterns in dynamic heterogeneous data. In this paper we demonstrate the utility of the tool on two distinct biological problems of different scales. First, we analyze a medium scale dataset that looks at perturbation effects of the pluripotency regulator Nanog in murine embryonic stem cells. Dynamic cluster analysis suggests alternative indirect links between Nanog and other proteins in the core stem cell network. Moreover, recurrent correlations from the epigenetic to the translational level are identified. Second, we investigate a large scale dataset consisting of genome-wide knockdown screens for human genes essential in the mitotic process. Here, a potential new role for the gene lsm14a in cytokinesis is suggested. We also show how phenotypic patterning allows for extensive comparison and identification of high impact knockdown targets. CONCLUSIONS: We present a new visualization approach for perturbation screens with multiple phenotypic outcomes. The novel functionality implemented in Arena3D enables effective understanding and comparison of temporal patterns within morphological layers, to help with the system-wide analysis of dynamic processes. Arena3D is available free of charge for academics as a downloadable standalone application from: http://arena3d.org/. Maria Secrier, Georgios A. Pavlopoulos, Jan Aerts, Reinhard Schneider 0002 |
BMC Bioinform. | 2 |
| 2010 | LAITOR - Literature Assistant for Identification of Terms co-Occurrences and RelationshipsabstractBACKGROUND: Biological knowledge is represented in scientific literature that often describes the function of genes/proteins (bioentities) in terms of their interactions (biointeractions). Such bioentities are often related to biological concepts of interest that are specific of a determined research field. Therefore, the study of the current literature about a selected topic deposited in public databases, facilitates the generation of novel hypotheses associating a set of bioentities to a common context. RESULTS: We created a text mining system (LAITOR: Literature Assistant for Identification of Terms co-Occurrences and Relationships) that analyses co-occurrences of bioentities, biointeractions, and other biological terms in MEDLINE abstracts. The method accounts for the position of the co-occurring terms within sentences or abstracts. The system detected abstracts mentioning protein-protein interactions in a standard test (BioCreative II IAS test data) with a precision of 0.82-0.89 and a recall of 0.48-0.70. We illustrate the application of LAITOR to the detection of plant response genes in a dataset of 1000 abstracts relevant to the topic. CONCLUSIONS: Text mining tools combining the extraction of interacting bioentities and biological concepts with network displays can be helpful in developing reasonable hypotheses in different scientific backgrounds. Adriano Barbosa-Silva, Theodoros G. Soldatos, Ivan L. F. Magalhães, Georgios A. Pavlopoulos, Jean-Fred Fontaine, Miguel A. Andrade-Navarro, Reinhard Schneider 0002, José Miguel Ortega |
BMC Bioinform. | 4 |
| 2009 | jClust: a clustering and visualization toolboxabstractUNLABELLED: jClust is a user-friendly application which provides access to a set of widely used clustering and clique finding algorithms. The toolbox allows a range of filtering procedures to be applied and is combined with an advanced implementation of the Medusa interactive visualization module. These implemented algorithms are k-Means, Affinity propagation, Bron-Kerbosch, MULIC, Restricted neighborhood search cluster algorithm, Markov clustering and Spectral clustering, while the supported filtering procedures are haircut, outside-inside, best neighbors and density control operations. The combination of a simple input file format, a set of clustering and filtering algorithms linked together with the visualization tool provides a powerful tool for data analysis and information extraction. AVAILABILITY: http://jclust.embl.de/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Georgios A. Pavlopoulos, Charalampos N. Moschopoulos, Sean D. Hooper, Reinhard Schneider 0002, Sophia Kossida |
Bioinform. | 1 |
| 2009 | OnTheFly: a tool for automated document-based text annotation, data linking and network generationabstractUNLABELLED: OnTheFly is a web-based application that applies biological named entity recognition to enrich Microsoft Office, PDF and plain text documents. The input files are converted into the HTML format and then sent to the Reflect tagging server, which highlights biological entity names like genes, proteins and chemicals, and attaches to them JavaScript code to invoke a summary pop-up window. The window provides an overview of relevant information about the entity, such as a protein description, the domain composition, a link to the 3D structure and links to other relevant online resources. OnTheFly is also able to extract the bioentities mentioned in a set of files and to produce a graphical representation of the networks of the known and predicted associations of these entities by retrieving the information from the STITCH database. AVAILABILITY: http://onthefly.embl.de, http://onthefly.embl.de/FAQ.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Georgios A. Pavlopoulos, Evangelos Pafilis, Michael Kuhn 0004, Sean D. Hooper, Reinhard Schneider 0002 |
Bioinform. | 1 |
| 2009 | GIBA: a clustering tool for detecting protein complexesabstractBACKGROUND: During the last years, high throughput experimental methods have been developed which generate large datasets of protein - protein interactions (PPIs). However, due to the experimental methodologies these datasets contain errors mainly in terms of false positive data sets and reducing therefore the quality of any derived information. Typically these datasets can be modeled as graphs, where vertices represent proteins and edges the pairwise PPIs, making it easy to apply automated clustering methods to detect protein complexes or other biological significant functional groupings. METHODS: In this paper, a clustering tool, called GIBA (named by the first characters of its developers' nicknames), is presented. GIBA implements a two step procedure to a given dataset of protein-protein interaction data. First, a clustering algorithm is applied to the interaction data, which is then followed by a filtering step to generate the final candidate list of predicted complexes. RESULTS: The efficiency of GIBA is demonstrated through the analysis of 6 different yeast protein interaction datasets in comparison to four other available algorithms. We compared the results of the different methods by applying five different performance measurement metrices. Moreover, the parameters of the methods that constitute the filter have been checked on how they affect the final results. CONCLUSION: GIBA is an effective and easy to use tool for the detection of protein complexes out of experimentally measured protein - protein interaction networks. The results show that GIBA has superior prediction accuracy than previously published methods. Charalampos N. Moschopoulos, Georgios A. Pavlopoulos, Reinhard Schneider 0002, Spiridon D. Likothanassis, Sophia Kossida |
BMC Bioinform. | 2 |
| 2008 | An enhanced Markov clustering method for detecting protein complexesabstractWith the recent high-throughput methods, large datasets of experimentally detected pairwise protein-protein interactions are generated. However, these data suffer from noise, reducing the quality of the information they bring (identification of protein complexes). This paper introduces a novel methodology for detecting protein complexes in a protein-protein interaction graph. Our method initially uses the Markov clustering algorithm and then filters the derived results in order to obtain the best set of clusters that represent protein complexes. The efficiency of our method is shown in experimental results derived from 7 different yeast protein interaction datasets. Moreover, comparisons with 4 other algorithms are performed proving that our method predicts known protein complexes, recorded in the MIPS database, more accurately. Charalampos N. Moschopoulos, Georgios A. Pavlopoulos, Spiridon D. Likothanassis, Sophia Kossida |
BIBE | 2 |