Pantelis G. Bagos

dblp:56/533 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0003-4935-2325ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2026 PYRAMA: an open-source tool for advanced meta-analysis of genome wide association studies
abstract
MOTIVATION: Genome-wide association study (GWAS) meta-analysis tools are essential for integrating summary statistics across multiple cohorts, thereby increasing statistical power and validating genetic associations. Widely cited tools, such as METAL, PLINK, and GWAMA, have facilitated numerous significant discoveries in the field of GWAS. Nevertheless, these tools offer a limited set of meta-analysis methods and typically require users to have prior experience with command-line tools to be executed. RESULTS: We present here PYRAMA, an open-source tool which is designed for meta-analysis of genome wide association studies. This work introduces an easy-to-use software package that includes several meta-analysis methods that are absent in similar software packages. PYRAMA is faster compared to other tools, supports robust methods for analysis and meta-analysis, fixed-effects, random-effects and Bayesian meta-analysis and it is currently the only tool that supports meta-analysis with imputation of summary statistics. It is available both as a standalone tool and as a freely available web server. AVAILABILITY AND IMPLEMENTATION: https://github.com/pbagos/PYRAMA, https://doi.org/10.5281/zenodo.17830449.
Georgios A. Manios, Sophia Nteli, Panagiota I. Kontou, Pantelis G. Bagos
Bioinform.4
2025 PRED-LD: efficient imputation of GWAS summary statistics
abstract
BACKGROUND: Genome-wide association studies have identified connections between genetic variations and diseases, but they only examine a small portion of single nucleotide polymorphisms. To enhance genetic findings, researchers suggest imputing genotypes for unmeasured SNPs to improve coverage and statistical power. When this is not possible, summary statistics imputation can be used as an alternative. The available summary statistics imputation tools rely on reference panels, such as the 1000 Genomes Project, to estimate linkage disequilibrium (LD) between variants for accurate imputation. Tools like FAPI and SSIMP use these reference panels in variant call format (VCF) for this purpose, though this process can be time-consuming. A more effective approach for processing reference panels in summary statistics imputation was proposed in RAISS. In this approach, the LD among the variants is precomputed from the reference panel, prior to imputation, thereby reducing computational time. RESULTS: We present PRED-LD, an imputation method for GWAS summary statistics that aims to enhance the resolution of genetic association analyses. The proposed method uses precomputed linkage disequilibrium statistics from HapMap, Pheno Scanner and TOP-LD to impute summary statistics, given beta coefficients and standard errors. The single-point approach that we describe provides a fast and accurate way to estimate associations for untyped single nucleotide polymorphisms that exhibit high linkage disequilibrium (LD). The proposed method is faster, provides accurate imputation compared to existing tools, and has been implemented in both a web service ( https://compgen.dib.uth.gr/PRED-LD/ ) and a command-line tool ( https://github.com/pbagos/PRED-LD ), making it a useful resource for the research community. CONCLUSIONS: PRED-LD offers an efficient and accurate method for GWAS summary statistics imputation, providing faster performance, direct result interpretation, and the ability to use multiple reference panels. Also, the online version of PRED-LD simplifies obtaining LD information and performing imputation tasks without downloading reference panels and will be continuously updated to support tools for meta-analysis and fine-mapping in GWAS.
Georgios A. Manios, Aikaterini Michailidi, Panagiota I. Kontou, Pantelis G. Bagos
BMC Bioinform.4
2025 metacp: a versatile software package for combining dependent or independent p-values
abstract
BACKGROUND: We present metacp an open-source software package which implements an abundance of statistical methods for the combination of both independent p-values, with methods such as Fisher's, Stouffer's and Edgington's, and dependent p-values, with methods such as Brown's method and the Cauchy Combination Test. RESULTS: The tool is available in Python and STATA, it is very fast, and it is easy to use, requiring only minimal input. It offers a useful resource for combining both independent and dependent p-values, responding to diverse analytical needs for practitioners performing meta-analyses and bioinformaticians developing tools for a variety of applications. Depending on the input data it can be used for gene-based testing, for analysis of multiple traits in GWAS, or for combining diverse multi-omics data such as those of a TWAS, a colocalization or an RNA-seq study. CONCLUSIONS: Compared to other similar packages (like poolr or metap), metacp implements the largest collection of statistical methods for this problem, offering users the flexibility to choose from a wide variety of approaches. Being available both as a standalone Python tool and as a STATA command, metacp is accessible to a broad and diverse audience, including practitioners conducting meta-analyses across various fields and bioinformaticians developing new tools where p-value combination is a crucial component.
Evgenia K. Nikolitsa, Panagiota I. Kontou, Pantelis G. Bagos
BMC Bioinform.3
2023 Flame (v2.0): advanced integration and interpretation of functional enrichment results from multiple sources
abstract
Functional enrichment is the process of identifying implicated functional terms from a given input list of genes or proteins. In this article, we present Flame (v2.0), a web tool which offers a combinatorial approach through merging and visualizing results from widely used functional enrichment applications while also allowing various flexible input options. In this version, Flame utilizes the aGOtool, g: Profiler, WebGestalt, and Enrichr pipelines and presents their outputs separately or in combination following a visual analytics approach. For intuitive representations and easier interpretation, it uses interactive plots such as parameterizable networks, heatmaps, barcharts, and scatter plots. Users can also: (i) handle multiple protein/gene lists and analyse union and intersection sets simultaneously through interactive UpSet plots, (ii) automatically extract genes and proteins from free text through text-mining and Named Entity Recognition (NER) techniques, (iii) upload single nucleotide polymorphisms (SNPs) and extract their relative genes, or (iv) analyse multiple lists of differentially expressed proteins/genes after selecting them interactively from a parameterizable volcano plot. Compared to the previous version of 197 supported organisms, Flame (v2.0) currently allows enrichment for 14 436 organisms. AVAILABILITY AND IMPLEMENTATION: Web Application: http://flame.pavlopouloslab.info. Code: https://github.com/PavlopoulosLab/Flame. Docker: https://hub.docker.com/r/pavlopouloslab/flame.
Evangelos Karatzas, Fotis A. Baltoumas, Eleni Aplakidou, Panagiota I. Kontou, Panos Stathopoulos, Leonidas Stefanis, Pantelis G. Bagos, Georgios A. Pavlopoulos
Bioinform.7
2019 Semi-supervised learning of Hidden Markov Models for biological sequence analysis
abstract
MOTIVATION: Hidden Markov Models (HMMs) are probabilistic models widely used in applications in computational sequence analysis. HMMs are basically unsupervised models. However, in the most important applications, they are trained in a supervised manner. Training examples accompanied by labels corresponding to different classes are given as input and the set of parameters that maximize the joint probability of sequences and labels is estimated. A main problem with this approach is that, in the majority of the cases, labels are hard to find and thus the amount of training data is limited. On the other hand, there are plenty of unclassified (unlabeled) sequences deposited in the public databases that could potentially contribute to the training procedure. This approach is called semi-supervised learning and could be very helpful in many applications. RESULTS: We propose here, a method for semi-supervised learning of HMMs that can incorporate labeled, unlabeled and partially labeled data in a straightforward manner. The algorithm is based on a variant of the Expectation-Maximization (EM) algorithm, where the missing labels of the unlabeled or partially labeled data are considered as the missing data. We apply the algorithm to several biological problems, namely, for the prediction of transmembrane protein topology for alpha-helical and beta-barrel membrane proteins and for the prediction of archaeal signal peptides. The results are very promising, since the algorithms presented here can significantly improve the prediction performance of even the top-scoring classifiers. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ioannis A. Tamposis, Konstantinos D. Tsirigos, Margarita C. Theodoropoulou, Panagiota I. Kontou, Pantelis G. Bagos
Bioinform.5
2019 JUCHMME: a Java Utility for Class Hidden Markov Models and Extensions for biological sequence analysis
abstract
SUMMARY: JUCHMME is an open-source software package designed to fit arbitrary custom Hidden Markov Models (HMMs) with a discrete alphabet of symbols. We incorporate a large collection of standard algorithms for HMMs as well as a number of extensions and evaluate the software on various biological problems. Importantly, the JUCHMME toolkit includes several additional features that allow for easy building and evaluation of custom HMMs, which could be a useful resource for the research community. AVAILABILITY AND IMPLEMENTATION: http://www.compgen.org/tools/juchmme, https://github.com/pbagos/juchmme. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ioannis A. Tamposis, Konstantinos D. Tsirigos, Margarita C. Theodoropoulou, Panagiota I. Kontou, Georgios N. Tsaousis, Dimitra Sarantopoulou, Zoi I. Litou, Pantelis G. Bagos
Bioinform.8
2019 Ten simple rules for carrying out and writing meta-analyses
abstract
In the context of evidence-based medicine, meta-analyses provide novel and useful information [1], as they are at the top of the pyramid of evidence and consolidate previous evidence published in multiple previous reports [2]. Meta-analysis is a powerful tool to cumulate and summarize the knowledge in a research field [3]. Because of the significant increase in the published scientific literature in recent years, there has also been an important growth in the number of meta-analyses for a large number of topics [4]. It has been found that meta-analyses are among the types of publications that usually receive a larger number of citations in the biomedical sciences [5,6]. The methods and standards for carrying out meta-analyses have evolved in recent years [7–9]. Although there are several published articles describing comprehensive guidelines for specific types of meta-analyses, there is still the need for an abridged article with general and updated recommendations for researchers interested in the development of meta-analyses. We present here ten simple rules for carrying out and writing meta-analyses. Rule 1: Specify the topic and type of the meta-analysis Considering that a systematic review [10] is fundamental for a meta-analysis, you can use the Population, Intervention, Comparison, Outcome (PICO) model to formulate the research question. It is important to verify that there are no published meta-analyses on the specific topic in order to avoid duplication of efforts [11]. In some cases, an updated meta-analysis in a topic is needed if additional data become available. It is possible to carry out meta-analyses for multiple types of studies, such as epidemiological variables for case-control, cohort, and randomized clinical trials. As observational studies have a larger possibility of having several biases, meta-analyses of these types of designs should take that into account. In addition, there is the possibility to carry out meta-analyses for genetic association studies, gene expression studies, genome-wide association studies (GWASs), or data from animal experiments. It is advisable to preregister the systematic review protocols at the International Prospective Register of Systematic Reviews (PROSPERO; https://www.crd.york.ac.uk/Prospero) database [12]. Keep in mind that an increasing number of journals require registration prior to publication.
Diego A. Forero, Sandra Lopez-Leon, Yeimy González-Giraldo, Pantelis G. Bagos
PLoS Comput. Biol.4
2018 ECCB 2018: The 17th European Conference on Computational Biology
abstract
This volume of Bioinformatics includes the proceedings papers of the 17th European Conference in Computational Biology (ECCB), an annual international Conference for research in computational biology and bioinformatics. The Conference is being held jointly with ISMB in the odd-numbered years and independently in the even-numbered years. This year, the 17th ECCB (ECCB 2018) will take place in Athens, at the Stavros Niarchos Foundation Cultural Center (SNFCC), from September 8 to 12, 2018. SNFCC, which was opened in June 2017, is a multifunctional arts, education and entertainment complex located on the edge of Faliro Bay, 4.5 km south of the center of Athens. Information on the ECCB 2018 can be found at eccb18.org and will be later archived at www.ebi.ac.uk/eccb/2018/. With more than 1,000 participants from academia and industry, ECCB is the leading European Conference on computational biology and bioinformatics and the second largest internationally, next to ISMB (Intelligent Systems in Molecular Biology). The work presented in ECCB is related to all domains of the field of computational biology and bioinformatics, ranging from molecular level to systems biology level. The proceedings papers present new computational methodologies and tools for addressing challenging problems in the field, including research in single cell high-throughput data, microbiome data, 3 D organization of the genome, and others. The Conference also includes Highlights presentations, which showcase important papers published over the last year. Since 2016, it also features an Applications track that presents computational biology work applied in industry, clinics, governmental organizations, and other fields beyond academia. ECCB meetings are held every year in a different European country. In odd-numbered years, ECCB is co-organized with the ISMB Conference, which is held in Europe. Previous meetings were organized in Prague, Czech Republic (Beerenwinkel and Bromberg, 2017, ISMB/ECCB 2017); The Hague, Netherlands (Heringa and Reinders, ECCB 2016); Dublin, Ireland (Moreau and Beerenwinkel, 2015, ISMB/ECCB 2015); Strasbourg, France (Devignes and Moreau, 2014, ECCB 2014); Berlin, Germany (Ben-Tal, 2013, ISMB/ECCB 2013); Basel, Switzerland (Schwede and Iber, 2012, ECCB 2012); Vienna, Austria (Gaasterland and Vingron, 2011, ISMB/ECCB 2011); Ghent, Belgium (Moreau and Heringa, 2010, ECCB 2010); Stockholm, Sweden (Gusfield and Tramontano, 2009, ISMB/ECCB 2009); Cagliari, Italy (Tramontano, 2008, ECCB 2008); Vienna, Austria (Lengauer et al., 2009, ISMB/ECCB 2007); Eilat, Israel (Wolfson and Safer, 2006, ECCB 2006); Madrid, Spain (Guigo et al., 2005, ECCB 2005); Glasgow, United Kingdom (Thornton et al., 2004, ISMB/ECCB 2004); Paris, France (Lenhof and Sagot, 2003, ECCB 2003); and Saarbrücken, Germany (Lengauer, 2002). ECCB 2018 continues the 16-year tradition and is held under the auspices of the Hellenic Society for Computational Biology and Bioinformatics (HSCBB; www.hscbb.gr), which promotes bioinformatics research and training in Greece since 2009. HSCBB organizes an annual conference in a different Greek city each time (e.g., Athens, Thessaloniki, Alexandroupolis, Patras, Lamia, Heraklion), aiming to provide exposure to new field developments to graduate students and researchers. The broad participation in the annual HSCBB conferences (more than 130 participants each year) show that the bioinformatics community in Greek universities and research centers has now matured significantly. HSCBB has also a strong international presence in the field of computational biology. It is an observer in the Greek ELIXIR consortium and is affiliated with the International Society for Computational Biology (ISCB) and Global Organization for Bioinformatics Learning, Education & Training (GOBLET). The organization of ECCB 2018, along with the European Student Council Symposium 2018 (see below), inspired the formation this year of the Greek chapter of the ISCB Student Council. ECCB 2018 received a record high number of 280 applications for proceedings talks from institutions from 48 countries. The submissions were organized in five themes, according to their topic: (1) Data (organization, integration, knowledge discovery, multi-scale modeling), (2) Genes (expression, function, editing, geno/phenotype), (3) Genome (sequence analysis, evolution, phylogeny, microbiome), (4) Proteins and Structural Biology (structure, function, alterations, drug design), and (5) Systems (molecular pathways, signaling, metabolomics). All proceedings submissions were subjected to a rigorous peer-review process (2-4 reviews per paper), organized by the members of the Programme Committees of the corresponding theme. The Programme Committees consisted of a Theme Chair (Michael Krauthammer, Roderic Guigo, Martin Vingron, Ivet Bahar, Alfonso Valencia, respectively) and two to four co-chairs (depending on the number of submissions assigned to this Theme). For the review, the EasyChair system (www.easychair.org) was used by 398 reviewers, recruited by the Programme Committees. Reviewers evaluated the impact and reproducibility of the presented research, as well as its suitability for the ECCB audience. Once the review was completed, the committee chairs and co-chairs selected 48 papers to be included in the ECCB 2018 proceedings (acceptance rate: 17%), with the number of papers accepted in each being proportional to the number of submissions initially assigned to this Theme. These papers required minor revisions and the authors had two weeks to modify them accordingly. All Proceedings Track papers and any supplementary files accompanying them are freely available in the electronic form of Oxford University Press journal Bioinformatics, as a special issue of September 2018. Together with the Proceedings track ECCB 2018 features 24 Highlights talks about scientific work already published in high impact science journals. They are coordinately presented and managed across the five themes with the proceeding presentations. We continued with newly established Application and ELIXIR tracks with 15 Applications talks, and 12 ELIXIR talks, all of which were selected after review. The aim of the Application track is to give a voice to those who apply computational biology in industry, clinics, governmental organizations, and other fields beyond academia. The aim of the ELIXIR Track is to showcase to the community the latest outputs and services from across the initiative. The presentations focus on developments relating to services and infrastructure within ELIXIR. ECCB 2018 also features a Poster Track, with 617 accepted posters, in which researchers present their latest findings in the corresponding thematic areas (Data, Genes, Genome, Proteins and Structural Biology, Systems). Seven distinguished keynote speakers will present their work: Prof. Bonnie Berger from the Massachusetts Institute of Technology (MIT), Prof. Christos Davatzikos from University of Pennsylvania, Prof. John Ioannidis from Stanford University, Prof. Manolis Kellis from MIT, Prof. Jose Onuchic from Rice University, Dr. Janet Thornton, Director Emeritus of the European Bioinformatics Institute (EMBL-EBI), and Dr. Eleftheria Zeggini from the German Research Center for Environmental Health (GmbH). Furthermore, ECCB 2018 is hosting for the first time three special invited talks before the lunch breaks and a joint keynote talk. The invited speakers will present the funding opportunities of the European Research Council (ERC) funding mechanisms (Dr. Maria Siomos) and a talk on the ISCB Student Council Internship Program (Farzana Rahman). The joint keynote will be about ELIXIR: A Common European Infrastructure for Bioinformatics Research presented by the ELIXIR director Dr. Niklas Blomberg. For the first time in ECCB 2018, ELIXIR organizes eight short workshops, namely ELIXIR TeSS Usability Study, Beacons, BioSchemas, Galaxy, OpenEbench Fnd Bio.tools, Research under the GDPR, Implementation of DMPs and Data Stewardship in practice, Secure access to your services using ELIXIR AAI. The 5th European Student Council Symposium (ESCS) is taking place before the main conference organized by the Student Council of the ISCB (chair Daniele Parisi from KU Leuven and co-chair Yvonne Saara Gladbach from Rostock University). ESCS highlights will be published in F1000Research via the ISCB Student Council channel. For the first time the ISCB Student Council will present also an invited talk as stated above. We are really thankful to the ISCB Student Council for their enthusiasm, hard work and genuine scientific interest, which has made the ESCS an inseparable part of ECCB. During the ECCB 2018, ten exhibitor booths of all sizes will be set up on the Exhibit Hall floor, next to the central Conference venue and close to the Poster session and lunch area, namely: ELIXIR, BioExcel, EMBL-EBI, ERC, ISCB and ISCB Student Council, Goblet, Oxford University Press, Cambridge University Press, Peer Community In. The booths will be grouped together in a way that provides an enhanced networking environment for our exhibitors and delegates. Exhibitors will showcase the latest trends in computational biology and bioinformatics technology, in scientific literature, as well as in modelling and simulation. We invite participants to visit the exhibition area and support these community-minded organizations, who deliver a strong message in supporting the Computational Biology and Bioinformatics scientific field. The weekend before the conference 8 workshops (including two from Special Interest Groups – SIGs) and 13 Tutorials will take place. The aim of workshops is to provide participants the opportunity to discuss different perspectives on the cutting edge of a selected research field, through presenting technical issues, exchanging research ideas and sharing practical experiences. The two SIGs and six workshops preceding the ECCB 2018 main conference were selected out of a total of 12 applications, running all but one for a single day: (SIG-1) BioExcel 2nd SIG Meeting: Advanced Simulations for Biomolecular Research organized by Rossen Apostolov (KTH Royal Institute of Technology in Stockholm, Sweden), Zoe Cournia (Biomedical Research Foundation of the Academy of Athens, Greece), Vera Matser, (EMBL-EBI, UK), Anastas Mishev (UKIM Saints Cyril and Methodius University of Skopje, Republic of Macedonia), Hrachya Astsatryan (Institute for Informatics and Automation Problems, Armenia), Adam Carter (EPCC, UK) (SIG-2): Intrinsically disordered proteins: advances and state-of-the-art of the field organized by Silvio Tosatto (University of Padua, Italy), Zsuzsanna Dosztanyi (Eötvös Loránd University, Hungary), Norman Davey (University College Dublin, Ireland), Damiano Piovesan (University of Padua, Italy) (W1) Computational Epigenomics a two-day workshop organized by Yassen Assenov (German Cancer Research Center, Germany), Guido Sanguinetti (University of Edinburgh, UK), Jordana Bell (King’s College London, UK), Jörn Walter (Saarland University, Germany), Christoph Bock (CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, Austria) and Verena Wolf (Saarland University, Germany) (W2) BioNetVisA 2018 workshop: From biological network reconstruction to data visualization and analysis in molecular biology and medicine organized by Inna Kuperstein (Institut Curie, France), Emmanuel Barillot (Institut Curie, France), Andrei Zinovyev (Institut Curie, France), Luis Cristobal Monraz Gomez (Institut Curie, France), Hioraki Kitano (RIKEN Center for Integrative Medical Sciences, Japan), Minoru Kanehisa (Institute for Chemical Research, Kyoto University, Japan), Samik Ghosh (Systems Biology Institute, Tokyo, Japan), Nicolas Le Novère (Babraham Institute, UK), Robin Haw (Ontario Institute for Cancer Research, Canada), Alfonso Valencia (Spanish National Bioinformatics Institute, Madrid, Spain), Lodewyk Wessels (Netherlands Cancer Institute, Amsterdam, Netherlands), Patrick Kemmeren (Princess Maxima Center for Pediatric Oncology, Utrecht, Netherlands) (W3) CPW2018: Computational Pathology Workshop—second edition organized by Yves Sucaet (Vrije Universiteit Brussel, Belgium), Jeroen Van der Laak (UMC Radboud, Netherlands), Zev Leifer (New York College of Podiatric Medicine, USA), Yukako Yagi (Memorial Sloan Kettering Cancer Center, USA), Raphaël Marée (Université de Liège, Belgium), David Ameisen (IRIF, CNRS and Université Paris Diderot, France), Paul Van Diest (UMC Utrecht, Netherlands), Jeffrey Fine (Magee-Womens Hospital of UPMC) (W4) Interactive visualizations to guide health diagnostics and personalized medicine organized by Dimitrios Tzovaras, Konstantinos Votis and Kostas Stamatopoulos (all at Information Technologies Institute of the Center for Research and Technology Hellas, Greece) (W5) Logical modelling of cellular networks organized by Anna Niarakis (Univ Evry, Université Paris-Saclay, France) and Denis Thieffry (Ecole Normale Supérieure, Paris, France). (W6) Recent Computational Advances in Metagenomics organized by Valentin Loux, Mahendra Mariadassou, Pierre Peterlongo and Sophie Schbath (all at INRA, France) The purpose of the ECCB tutorial program is to provide participants with lectures and hands-on training to the most important and emerging topics in bioinformatics and computational biology research. The tutorials offer skills ranging from early and basic steps of computational analysis in recently introduced topics to advanced computational skills in important established topics. A total of thirteen ECCB 2018 tutorials running for half or a full day will be held at ECCB 2018. These were selected out of sixteen applications: (T1) Automated Machine Learning for Bioinformatics and Computational Biology organized by Ioannis Tsamardinos (University of Crete, Greece; Gnosis Data Analysis, Greece), Kleio Maria Verrou (University of Crete, Greece) and Vincenzo Lagani (Ilia State University, Georgia; Gnosis Data Analysis, Greece) (T2) Computational Mass Spectrometry with OpenMS—From Algorithms to Integrated Workflows organized by Julianus Pfeuffer (Freie Universität Berlin, Germany) and Timo Sachsenberg (Universität Tübingen, Germany) (T3)—Write your own R script to jointly analyze RNA-, ATAC-seq, DNA methylation, SNPs and find drug targets in signal transduction networks using TRANSFAC® organized by Philip Stegmaier (geneXplain GmbH, Germany), Olga Kel-Margoulis (geneXplain GmbH, Germany), Alexander Kel (geneXplain GmbH, Germany; Institute of Systems Biology, Russia) (T4) Deep learning for predicting protein–DNA and –RNA binding organized by Yaron Orenstein (Ben-Gurion University, Israel) (T5) Delving into non-coding RNA with RNAcentral and Rfam. Organized by Anton Petrov and Ioanna Kalvari (both at EMBL-EBI, UK) (T6) DIANA Tools and Databases: In silico investigation of miRNA functions organized by Artemis Hatzigeorgiou, Spyros Tastsoglou, Dimitra Karagkouni, Nikos Perdikopanis, Giorgos Skoufos, Ioannis Kavakiotis (all at University of Thessaly, Greece) (T7) Exploring programmatic access to Protein sequence, function and structure with UniProt and PDBe organized by Andrew Nightingale and Mihaly Varadi (both at EMBL-EBI, UK) (T8) Fifth International Hands-on Tutorial on Logical Modelling: Exploring the dynamics of biological systems organized by Tomas Helikar (University of Nebraska, USA) and Juilee Thakar (University of Rochester Medical Center, USA) (T9) From an application to a fully integrated workflow—Comprehensive software engineering with SeqAn organized by René Rahn and Hannes Hauswedell (both at Freie Universität Berlin, Germany) (T10) Hands-on on Protein Function Prediction with Machine Learning and Interactive Analytics organized by Rabie Saidi and Tunca Dogan (EMBL-EBI, UK) (T11) Single-cell RNA-Seq Data Analysis organized by Panagiotis Papasaikas (Friedrich Miescher Institute for Biomedical Research (FMI), Basel, Switzerland; Swiss Institute of Bioinformatics, Switzerland) and Atul Sethi (Friedrich Miescher Institute for Biomedical Research (FMI), Basel, Switzerland; University Hospital Basel, University of Basel, Switzerland; Swiss Institute of Bioinformatics, Switzerland) (T12) Modern and scalable tools for efficient analysis of very large metagenomic datasets organized by Alexander Sczyrba (Bielefeld University, Germany), Christian Henke (Bielefeld University, Germany), Clovis Galiez (MPI for Biophysical Chemistry, Germany), Milot Mirdita (PI for Biophysical Chemistry, Germany) and Johannes Soeding (MPI for Biophysical Chemistry, Germany) (T13) User Experience Design for Computational Biologists organized by Nikiforos Karamanis and Xavier Watkins (both at EMBL-EBI, UK) Finally, following the tradition of ECCB, a number of travel fellowships, sponsored by ISCB, were given to young participants. We received 103 applications from scientists from 31 countries. After careful review and consideration, we awarded 14 travel fellowships, mainly to PhD students from nine countries that had an accepted proceedings paper. Volunteering is another important way for young scientists to visit and support the conference, as well as connect with other young scientists. The skills, talent and dedication of the volunteers are expected to largely contribute to the overall quality of the Conference. We made our best to accept all 91 applicants from 16 countries who were eligible to attend the Conference at a very low registration fee. Next to the volunteers, we would like to thank all the people who helped through their work to make this Conference a success. The Theme Chairs and Co-chairs as well all the reviewers who have been the heart of the Conference and helped selecting excellent papers, making thus this Conference a success. We are truly indebted to them. We are also grateful to the ECCB Steering Committee for their advice and continuous the of the organization of ECCB 2018. This conference would be at all the support from Anna Steering Committee and Chair ECCB our application and advice about to make a of was It was our to and inspired by for science with to in a will be in our and in our During the one year of the conference the support of Steering Committee Chair ECCB ECCB and Yves ECCB 2010, ECCB was Yves and for your continuous and We are also grateful to the ISCB Society for Computational for their and in ECCB 2018 at the international level and for This Conference would be the support of our and ELIXIR level and BioExcel was our level Cambridge University Press, Oxford University Press, Community and were ISCB, the Student Council of the ISCB, and who as exhibitors at ECCB 2018 also A thank to all for their and for to make ECCB 2018 an and Conference. A number of people to the organization of ECCB 2018. They beyond their of for making this conference really Maria a as was for the the the venue the and was for and Panagiotis was for the technical support of the registration Dr. organized the members of the DIANA also this Spyros Tastsoglou, Dimitra Karagkouni, Perdikopanis, Dimitrios Finally, the of a conference is its presentations, the conference for its participants. We thank all of for to ECCB 2018. We will the and in the city of Athens, the city and were of
Artemis G. Hatzigeorgiou, Pantelis G. Bagos, Panayiotis V. Benos, Christoforos Nikolaou, Yves Moreau, Ioannis Kavakiotis
Bioinform.2
2017 GWAR: robust analysis and meta-analysis of genome-wide association studies
abstract
MOTIVATION: In the context of genome-wide association studies (GWAS), there is a variety of statistical techniques in order to conduct the analysis, but, in most cases, the underlying genetic model is usually unknown. Under these circumstances, the classical Cochran-Armitage trend test (CATT) is suboptimal. Robust procedures that maximize the power and preserve the nominal type I error rate are preferable. Moreover, performing a meta-analysis using robust procedures is of great interest and has never been addressed in the past. The primary goal of this work is to implement several robust methods for analysis and meta-analysis in the statistical package Stata and subsequently to make the software available to the scientific community. RESULTS: The CATT under a recessive, additive and dominant model of inheritance as well as robust methods based on the Maximum Efficiency Robust Test statistic, the MAX statistic and the MIN2 were implemented in Stata. Concerning MAX and MIN2, we calculated their asymptotic null distributions relying on numerical integration resulting in a great gain in computational time without losing accuracy. All the aforementioned approaches were employed in a fixed or a random effects meta-analysis setting using summary data with weights equal to the reciprocal of the combined cases and controls. Overall, this is the first complete effort to implement procedures for analysis and meta-analysis in GWAS using Stata. AVAILABILITY AND IMPLEMENTATION: A Stata program and a web-server are freely available for academic users at http://www.compgen.org/tools/GWAR. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Niki L. Dimou, Konstantinos D. Tsirigos, Arne Elofsson, Pantelis G. Bagos
Bioinform.4
2016 PRED-TMBB2: improved topology prediction and detection of beta-barrel outer membrane proteins
abstract
MOTIVATION: The PRED-TMBB method is based on Hidden Markov Models and is capable of predicting the topology of beta-barrel outer membrane proteins and discriminate them from water-soluble ones. Here, we present an updated version of the method, PRED-TMBB2, with several newly developed features that improve its performance. The inclusion of a properly defined end state allows for better modeling of the beta-barrel domain, while different emission probabilities for the adjacent residues in strands are used to incorporate knowledge concerning the asymmetric amino acid distribution occurring there. Furthermore, the training was performed using newly developed algorithms in order to optimize the labels of the training sequences. Moreover, the method is retrained on a larger, non-redundant dataset which includes recently solved structures, and a newly developed decoding method was added to the already available options. Finally, the method now allows the incorporation of evolutionary information in the form of multiple sequence alignments. RESULTS: The results of a strict cross-validation procedure show that PRED-TMBB2 with homology information performs significantly better compared to other available prediction methods. It yields 76% in correct topology predictions and outperforms the best available predictor by 7%, with an overall SOV of 0.9. Regarding detection of beta-barrel proteins, PRED-TMBB2, using just the query sequence as input, achieves an MCC value of 0.92, outperforming even predictors designed for this task and are much slower. AVAILABILITY AND IMPLEMENTATION: The method, along with all datasets used, is freely available for academic users at http://www.compgen.org/tools/PRED-TMBB2 CONTACT: [email protected].
Konstantinos D. Tsirigos, Arne Elofsson, Pantelis G. Bagos
Bioinform.3
2010 Combined prediction of Tat and Sec signal peptides with hidden Markov models
abstract
MOTIVATION: Computational prediction of signal peptides is of great importance in computational biology. In addition to the general secretory pathway (Sec), Bacteria, Archaea and chloroplasts possess another major pathway that utilizes the Twin-Arginine translocase (Tat), which recognizes longer and less hydrophobic signal peptides carrying a distinctive pattern of two consecutive Arginines (RR) in the n-region. A major functional differentiation between the Sec and Tat export pathways lies in the fact that the former translocates secreted proteins unfolded through a protein-conducting channel, whereas the latter translocates completely folded proteins using an unknown mechanism. The purpose of this work is to develop a novel method for predicting and discriminating Sec from Tat signal peptides at better accuracy. RESULTS: We report the development of a novel method, PRED-TAT, which is capable of discriminating Sec from Tat signal peptides and predicting their cleavage sites. The method is based on Hidden Markov Models and possesses a modular architecture suitable for both Sec and Tat signal peptides. On an independent test set of experimentally verified Tat signal peptides, PRED-TAT clearly outperforms the previously proposed methods TatP and TATFIND, whereas, when evaluated as a Sec signal peptide predictor compares favorably to top-scoring predictors such as SignalP and Phobius. The method is freely available for academic users at http://www.compgen.org/tools/PRED-TAT/.
Pantelis G. Bagos, Elisanthi P. Nikolaou, Theodore D. Liakopoulos, Konstantinos D. Tsirigos
Bioinform.1
2010 ExTopoDB: a database of experimentally derived topological models of transmembrane proteins
abstract
UNLABELLED: ExTopoDB is a publicly accessible database of experimentally derived topological models of transmembrane proteins. It contains information collected from studies in the literature that report the use of biochemical methods for the determination of the topology of α-helical transmembrane proteins. Transmembrane protein topology is highly important in order to understand their function and ExTopoDB provides an up to date, complete and comprehensive dataset of experimentally determined topologies of α-helical transmembrane proteins. Topological information is combined with transmembrane topology prediction resulting in more reliable topological models. AVAILABILITY: http://bioinformatics.biol.uoa.gr/ExTopoDB.
Georgios N. Tsaousis, Konstantinos D. Tsirigos, Xanthi D. Andrianou, Theodore D. Liakopoulos, Pantelis G. Bagos, Stavros J. Hamodrakas
Bioinform.5
2008 gpDB: a database of GPCRs, G-proteins, effectors and their interactions
abstract
UNLABELLED: gpDB is a publicly accessible, relational database, containing information about G-proteins, G-protein coupled receptors (GPCRs) and effectors, as well as information concerning known interactions between these molecules. The sequences are classified according to a hierarchy of different classes, families and subfamilies based on literature search. The main innovation besides the classification of G-proteins, GPCRs and effectors is the relational model of the database, describing the known coupling specificity of GPCRs to their respective alpha subunits of G-proteins, and also the specific interaction between G-proteins and their effectors, a unique feature not available in any other database. AVAILABILITY: http://bioinformatics.biol.uoa.gr/gpDB CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Margarita C. Theodoropoulou, Pantelis G. Bagos, Ioannis C. Spyropoulos, Stavros J. Hamodrakas
Bioinform.2
2006 Algorithms for incorporating prior topological information in HMMs: application to transmembrane proteins
abstract
BACKGROUND: Hidden Markov Models (HMMs) have been extensively used in computational molecular biology, for modelling protein and nucleic acid sequences. In many applications, such as transmembrane protein topology prediction, the incorporation of limited amount of information regarding the topology, arising from biochemical experiments, has been proved a very useful strategy that increased remarkably the performance of even the top-scoring methods. However, no clear and formal explanation of the algorithms that retains the probabilistic interpretation of the models has been presented so far in the literature. RESULTS: We present here, a simple method that allows incorporation of prior topological information concerning the sequences at hand, while at the same time the HMMs retain their full probabilistic interpretation in terms of conditional probabilities. We present modifications to the standard Forward and Backward algorithms of HMMs and we also show explicitly, how reliable predictions may arise by these modifications, using all the algorithms currently available for decoding HMMs. A similar procedure may be used in the training procedure, aiming at optimizing the labels of the HMM's classes, especially in cases such as transmembrane proteins where the labels of the membrane-spanning segments are inherently misplaced. We present an application of this approach developing a method to predict the transmembrane regions of alpha-helical membrane proteins, trained on crystallographically solved data. We show that this method compares well against already established algorithms presented in the literature, and it is extremely useful in practical applications. CONCLUSION: The algorithms presented here, are easily implemented in any kind of a Hidden Markov Model, whereas the prediction method (HMM-TM) is freely available for academic users at http://bioinformatics.biol.uoa.gr/HMM-TM, offering the most advanced decoding options currently available.
Pantelis G. Bagos, Theodore D. Liakopoulos, Stavros J. Hamodrakas
BMC Bioinform.1
2005 Prediction of the coupling specificity of GPCRs to four families of G-proteins using hidden Markov models and artificial neural networks
abstract
MOTIVATION: G-protein coupled receptors are a major class of eukaryotic cell-surface receptors. A very important aspect of their function is the specific interaction (coupling) with members of four G-protein families. A single GPCR may interact with members of more than one G-protein families (promiscuous coupling). To date all published methods that predict the coupling specificity of GPCRs are restricted to three main coupling groups G(i/o), G(q/11) and G(s), not including G(12/13)-coupled or other promiscuous receptors. RESULTS: We present a method that combines hidden Markov models and a feed-forward artificial neural network to overcome these limitations, while producing the most accurate predictions currently available. Using an up-to-date curated dataset, our method yields a 94% correct classification rate in a 5-fold cross-validation test. The method predicts also promiscuous coupling preferences, including coupling to G(12/13), whereas unlike other methods avoids overpredictions (false positives) when non-GPCR sequences are encountered. AVAILABILITY: A webserver for academic users is available at http://bioinformatics.biol.uoa.gr/PRED-COUPLE2
Nikolaos G. Sgourakis, Pantelis G. Bagos, Stavros J. Hamodrakas
Bioinform.2
2005 Evaluation of methods for predicting the topology of beta-barrel outer membrane proteins and a consensus prediction method
abstract
BACKGROUND: Prediction of the transmembrane strands and topology of beta-barrel outer membrane proteins is of interest in current bioinformatics research. Several methods have been applied so far for this task, utilizing different algorithmic techniques and a number of freely available predictors exist. The methods can be grossly divided to those based on Hidden Markov Models (HMMs), on Neural Networks (NNs) and on Support Vector Machines (SVMs). In this work, we compare the different available methods for topology prediction of beta-barrel outer membrane proteins. We evaluate their performance on a non-redundant dataset of 20 beta-barrel outer membrane proteins of gram-negative bacteria, with structures known at atomic resolution. Also, we describe, for the first time, an effective way to combine the individual predictors, at will, to a single consensus prediction method. RESULTS: We assess the statistical significance of the performance of each prediction scheme and conclude that Hidden Markov Model based methods, HMM-B2TMR, ProfTMB and PRED-TMBB, are currently the best predictors, according to either the per-residue accuracy, the segments overlap measure (SOV) or the total number of proteins with correctly predicted topologies in the test set. Furthermore, we show that the available predictors perform better when only transmembrane beta-barrel domains are used for prediction, rather than the precursor full-length sequences, even though the HMM-based predictors are not influenced significantly. The consensus prediction method performs significantly better than each individual available predictor, since it increases the accuracy up to 4% regarding SOV and up to 15% in correctly predicted topologies. CONCLUSIONS: The consensus prediction method described in this work, optimizes the predicted topology with a dynamic programming algorithm and is implemented in a web-based application freely available to non-commercial users at http://bioinformatics.biol.uoa.gr/ConBBPRED.
Pantelis G. Bagos, Theodore Liakopoulos, Stavros J. Hamodrakas
BMC Bioinform.1
2005 N-terminal sequence-based prediction of subcellular location
abstract
Different compartments in a cell perform diverse tasks, thus knowledge of the localization of a protein would be highly indicative of its function. Many proteins have an Nterminal sequence of approximately 20–50 residues, which is responsible for their targeting to the appropriate location and is cleaved off, after the protein has been inserted into the organelle. Such information can be used in computational methods predicting protein localization. There are several methods for prediction of protein subcellular location available on the web, with TargetP (Emanuelsson et al 2000) being the most widely used. PredSL is a software tool, which aims to classify proteins to five subcellular locations: chloroplast, thylakoid, mitochondrion, secretory pathway and other. It combines neural networks, Markov chains, HMMs and scoring matrices in order to identify a targeting sequence at the N-terminal of a protein sequence, and determine its type (chloroplast-cTP, mitochondrial-mTP, secreted-SP, thylakoidallTP) and the precise location of the cleavage site. PredSL was tested on a set of 732 plant protein sequences and 637 non-plant sequences and the overall accuracy was 90.4% for the plant set and 93.3% for the non-plant set. Compared to the results obtained by TargetP when tested on the same datasets, PredSL's performance was better by 1.6% for the plant sequences and slightly worse (0.2%) for the non-plant sequences. For the prediction of thylakoid proteins PredSL was tested by cross-validation and achieved 87.3% accuracy compared to 87% by LumenP. PredSL is the only method for protein subcellular localization prediction that is available by the authors as a free, stand-alone tool, or through the URL: http://bioinformat ics.biol.uoa.gr/PredSL. from BioSysBio: Bioinformatics and Systems Biology Conference Edinburgh, UK, 14–15 July 2005
Evangelia Petsalaki, Pantelis G. Bagos, Zoi I. Litou, Stavros J. Hamodrakas
BMC Bioinform.2
2005 A method for the prediction of GPCRs coupling specificity to G-proteins using refined profile Hidden Markov Models
abstract
BACKGROUND: G- Protein coupled receptors (GPCRs) comprise the largest group of eukaryotic cell surface receptors with great pharmacological interest. A broad range of native ligands interact and activate GPCRs, leading to signal transduction within cells. Most of these responses are mediated through the interaction of GPCRs with heterotrimeric GTP-binding proteins (G-proteins). Due to the information explosion in biological sequence databases, the development of software algorithms that could predict properties of GPCRs is important. Experimental data reported in the literature suggest that heterotrimeric G-proteins interact with parts of the activated receptor at the transmembrane helix-intracellular loop interface. Utilizing this information and membrane topology information, we have developed an intensive exploratory approach to generate a refined library of statistical models (Hidden Markov Models) that predict the coupling preference of GPCRs to heterotrimeric G-proteins. The method predicts the coupling preferences of GPCRs to Gs, Gi/o and Gq/11, but not G12/13 subfamilies. RESULTS: Using a dataset of 282 GPCR sequences of known coupling preference to G-proteins and adopting a five-fold cross-validation procedure, the method yielded an 89.7% correct classification rate. In a validation set comprised of all receptor sequences that are species homologues to GPCRs with known coupling preferences, excluding the sequences used to train the models, our method yields a correct classification rate of 91.0%. Furthermore, promiscuous coupling properties were correctly predicted for 6 of the 24 GPCRs that are known to interact with more than one subfamily of G-proteins. CONCLUSION: Our method demonstrates high correct classification rate. Unlike previously published methods performing the same task, it does not require any transmembrane topology prediction in a preceding step. A web-server for the prediction of GPCRs coupling specificity to G-proteins available for non-commercial users is located at http://bioinformatics.biol.uoa.gr/PRED-COUPLE.
Nikolaos G. Sgourakis, Pantelis G. Bagos, Panagiotis K. Papasaikas, Stavros J. Hamodrakas
BMC Bioinform.2
2004 TMRPres2D: high quality visual representation of transmembrane protein models
abstract
The 'TransMembrane protein Re-Presentation in 2-Dimensions' (TMRPres2D) tool, automates the creation of uniform, two-dimensional, high analysis graphical images/models of alpha-helical or beta-barrel transmembrane proteins. Protein sequence data and structural information may be acquired from public protein knowledge bases, emanate from prediction algorithms, or even be defined by the user. Several important biological and physical sequence attributes can be embedded in the graphical representation.
Ioannis C. Spyropoulos, Theodore Liakopoulos, Pantelis G. Bagos, Stavros J. Hamodrakas
Bioinform.3
2004 A Hidden Markov Model method, capable of predicting and discriminating beta-barrel outer membrane proteins
abstract
BACKGROUND: Integral membrane proteins constitute about 20-30% of all proteins in the fully sequenced genomes. They come in two structural classes, the alpha-helical and the beta-barrel membrane proteins, demonstrating different physicochemical characteristics, structure and localization. While transmembrane segment prediction for the alpha-helical integral membrane proteins appears to be an easy task nowadays, the same is much more difficult for the beta-barrel membrane proteins. We developed a method, based on a Hidden Markov Model, capable of predicting the transmembrane beta-strands of the outer membrane proteins of gram-negative bacteria, and discriminating those from water-soluble proteins in large datasets. The model is trained in a discriminative manner, aiming at maximizing the probability of correct predictions rather than the likelihood of the sequences. RESULTS: The training has been performed on a non-redundant database of 14 outer membrane proteins with structures known at atomic resolution; it has been tested with a jacknife procedure, yielding a per residue accuracy of 84.2% and a correlation coefficient of 0.72, whereas for the self-consistency test the per residue accuracy was 88.1% and the correlation coefficient 0.824. The total number of correctly predicted topologies is 10 out of 14 in the self-consistency test, and 9 out of 14 in the jacknife. Furthermore, the model is capable of discriminating outer membrane from water-soluble proteins in large-scale applications, with a success rate of 88.8% and 89.2% for the correct classification of outer membrane and water-soluble proteins respectively, the highest rates obtained in the literature. That test has been performed independently on a set of known outer membrane proteins with low sequence identity with each other and also with the proteins of the training set. CONCLUSION: Based on the above, we developed a strategy, that enabled us to screen the entire proteome of E. coli for outer membrane proteins. The results were satisfactory, thus the method presented here appears to be suitable for screening entire proteomes for the discovery of novel outer membrane proteins. A web interface available for non-commercial users is located at: http://bioinformatics.biol.uoa.gr/PRED-TMBB, and it is the only freely available HMM-based predictor for beta-barrel outer membrane protein topology.
Pantelis G. Bagos, Theodore Liakopoulos, Ioannis C. Spyropoulos, Stavros J. Hamodrakas
BMC Bioinform.1
2004 A database for G proteins and their interaction with GPCRs
abstract
BACKGROUND: G protein-coupled receptors (GPCRs) transduce signals from extracellular space into the cell, through their interaction with G proteins, which act as switches forming hetero-trimers composed of different subunits (alpha,beta,gamma). The alpha subunit of the G protein is responsible for the recognition of a given GPCR. Whereas specialised resources for GPCRs, and other groups of receptors, are already available, currently, there is no publicly available database focusing on G Proteins and containing information about their coupling specificity with their respective receptors. DESCRIPTION: gpDB is a publicly accessible G proteins/GPCRs relational database. Including species homologs, the database contains detailed information for 418 G protein monomers (272 Galpha, 87 Gbeta and 59 Ggamma) and 2782 GPCRs sequences belonging to families with known coupling to G proteins. The GPCRs and the G proteins are classified according to a hierarchy of different classes, families and sub-families, based on extensive literature searchs. The main innovation besides the classification of both G proteins and GPCRs is the relational model of the database, describing the known coupling specificity of the GPCRs to their respective alpha subunit of G proteins, a unique feature not available in any other database. There is full sequence information with cross-references to publicly available databases, references to the literature concerning the coupling specificity and the dimerization of GPCRs and the user may submit advanced queries for text search. Furthermore, we provide a pattern search tool, an interface for running BLAST against the database and interconnectivity with PRED-TMR, PRED-GPCR and TMRPres2D. CONCLUSIONS: The database will be very useful, for both experimentalists and bioinformaticians, for the study of G protein/GPCR interactions and for future development of predictive algorithms. It is available for academics, via a web browser at the URL: http://bioinformatics.biol.uoa.gr/gpDB.
Antigoni L. Elefsinioti, Pantelis G. Bagos, Ioannis C. Spyropoulos, Stavros J. Hamodrakas
BMC Bioinform.2