Artemis G. Hatzigeorgiou

dblp:95/1579 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-1414-5668ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2025 Bottlenecks in advancing and applying multiomic data integration - common data resources as rate-limiting drivers - the high-impact use case of atherosclerotic cardiovascular disease
abstract
Despite striking successes in identifying novel biomarkers for improved patient stratification and predicting disease progression, numerous challenges remain in the effective integration and exploitation of multiomic data in biomedical applications beyond cancer, for which most bioinformatics strategies are developed and validated. That focus on cancer severely limits the effective development and advancement of algorithms in machine learning and artificial intelligence that do not suffer degraded out-of-domain performance. Generalizability and interpretability of models, however, are also required for robust insights that may translate into clinical practice. Work across different independent datasets is critical for establishing models robust towards unwanted variation in assays, protocols, and cohort populations. Disease-specific context like ethnicity, socioeconomic background, sex, lifestyle, disease phase, and tissue type also strongly affect molecular profiles. We here discuss atherosclerotic cardiovascular disease (ASCVD) as a high-impact non-cancer use case for the challenges remaining in the development and application of the latest bioinformatics approaches to multiomics data integration. ASCVD remains the leading cause of death globally. Disease aetiology, progression, and therapy outcome depend on a complex interplay of genetic, environmental, and lifestyle factors. Integrating these diverse data types effectively remains a challenge but holds transformative potential for personalized medicine. Discovery and access to data of sufficient diversity and extent form key bottlenecks. We here compile a first comprehensive overview of key data sets in ASCVD to complement the established cancer-focused resources as a foundation for future effective development and application of state-of-the-art bioinformatics tools for multiomic data integration.
Stephanie Bezzina Wettinger, Kanita Karaduzovic-Hadziabdic, Ritienne Attard, Rosienne Farrugia, Brooke N. Wolford, Marco Chierici, Giuseppe Jurman, Panagiotis Alexiou, José L. Peñalvo, Rafael S. Costa, José Basilio, Frantisek Sabovcik, Rui Vitorino, Johannes A. Schmid, Rajesh Shigdel, Baiba Vilne, Artemis G. Hatzigeorgiou, Miron Sopic, Yvan Devaux, Paolo Magni, Maria Tellez-Plaza, David P. Kreil, Aleksandra Gruca
Briefings Bioinform.17
2025 microT-CNN: an avant-garde deep convolutional neural network unravels functional miRNA targets beyond canonical sites
abstract
microRNAs (miRNAs) are central post-transcriptional gene expression regulators in healthy and diseased states. Despite decades of effort, deciphering miRNA targets remains challenging, leading to an incomplete miRNA interactome and partially elucidated miRNA functions. Here, we introduce microT-CNN, an avant-garde deep convolutional neural network model that moves the needle by integrating hundreds of tissue-matched (in-)direct experiments from 26 distinct cell types, corresponding to a unique training and evaluation set of >60 000 miRNA binding events and ~30 000 unique miRNA-gene target pairs. The multilayer sequence-based design enables the prediction of both host and virus-encoded miRNA interactions, providing for the first time up to 67% of direct genuine Epstein-Barr virus- and Kaposi's sarcoma-associated herpesvirus-derived miRNA-target pairs corresponding to one out of four binding events of virus-encoded miRNAs. microT-CNN fills the existing gap of the miRNA-target prediction by providing functional targets beyond the canonical sites, including 3' compensatory miRNA pairings, prompting 1.4-fold more validated miRNA binding events compared to other implementations and shedding light on previously unexplored facets of the miRNA interactome.
Elissavet Zacharopoulou, Maria D. Paraskevopoulou, Spyros Tastsoglou, Athanasios Alexiou, Anna Karavangeli, Vasileios Pierros, Stefanos Digenis, Galatea Mavromati, Artemis G. Hatzigeorgiou, Dimitra Karagkouni
Briefings Bioinform.9
2024 A Machine Learning algorithm for Predicting Translation Initiation Start Positions
abstract
Translation Initiation Start (TIS) sites are specific regions in genes, where the protein synthesis starts. In this study, the specific problem of RNA sequence classification is addressed through the deployment of convolutional neural networks (CNNs) employing features extracted from genomic sequences. Translation initiation is a crucial step in protein synthesis, indicating the sites where the translation mechanism takes place and such, accurate prediction of its start sites is a significant step towards deeper understanding and harnessing of its complex characteristics. To this aim, a state-of-the-art deep learning model has been developed which predicts the presence/absence of translation initiation start sites within RNA sequences. The proposed deep learning model leverages the power of CNNs to effectively capture relevant features in RNA sequences. Trained on a comprehensive dataset of manually annotated translation initiation sites, the model learns to discern key patterns and motifs associated with initiation starts. Through rigorous evaluation, a model comprising sequence, conservation and hexamer features exhibits a remarkably high accuracy rate of 93.75%. This research holds promise for various biological applications, including gene annotation, drug target identification, and understanding the mechanisms of protein synthesis. The ability to reliably predict translation initiation start sites using deep learning provides a valuable tool for computational biologists and experimental researchers alike.
Stefanos Digenis, Dimitris Grigoriadis, Marios Miliotis, Artemis G. Hatzigeorgiou
CIBCB4
2024 Leveraging Large Language Models for Information Extraction: Identifying microRNA - Gene Interactions in Biomedical Literature
abstract
The rapid growth of biomedical literature necessitates efficient Information Extraction systems able to identify relevant knowledge for various biological applications, such as understanding gene regulation by microRNA (miRNA). In this study, we employed a Large Language Model, specifically GPT-3.5 (version 0301), in conjunction with BERN2 for miRNA-gene interaction extraction from paper titles and abstracts. We optimized our approach using an initial dataset of about a thousand molecular biology papers and subsequently evaluated its performance on a manually curated dataset of 400 papers, achieving an accuracy of 82-85%. Driven by the promising results and the practical utility of our method, we applied the system to a large dataset of 39,000 papers. The extracted miRNA-gene interactions, combined with a Natural Language Processing approach, were included in the TarBase v9 database. Our findings demonstrate the potential of Large Language Models in biomedical Information Extraction tasks and highlight the limitations of the current gene and miRNA recognition systems, which hinder further improvements in accuracy.
Steve Stavropoulos, Elissavet Zacharopoulou, Spiros V. Georgakopoulos, Sotiris K. Tasoulis, Vassilis P. Plagianakos, Artemis G. Hatzigeorgiou
CIBCB6
2022 DeepTSS: multi-branch convolutional neural network for transcription start site identification from CAGE data
abstract
BACKGROUND: The widespread usage of Cap Analysis of Gene Expression (CAGE) has led to numerous breakthroughs in understanding the transcription mechanisms. Recent evidence in the literature, however, suggests that CAGE suffers from transcriptional and technical noise. Regardless of the sample quality, there is a significant number of CAGE peaks that are not associated with transcription initiation events. This type of signal is typically attributed to technical noise and more frequently to random five-prime capping or transcription bioproducts. Thus, the need for computational methods emerges, that can accurately increase the signal-to-noise ratio in CAGE data, resulting in error-free transcription start site (TSS) annotation and quantification of regulatory region usage. In this study, we present DeepTSS, a novel computational method for processing CAGE samples, that combines genomic signal processing (GSP), structural DNA features, evolutionary conservation evidence and raw DNA sequence with Deep Learning (DL) to provide single-nucleotide TSS predictions with unprecedented levels of performance. RESULTS: To evaluate DeepTSS, we utilized experimental data, protein-coding gene annotations and computationally-derived genome segmentations by chromatin states. DeepTSS was found to outperform existing algorithms on all benchmarks, achieving 98% precision and 96% sensitivity (accuracy 95.4%) on the protein-coding gene strategy, with 96.66% of its positive predictions overlapping active chromatin, 98.27% and 92.04% co-localized with at least one transcription factor and H3K4me3 peak. CONCLUSIONS: CAGE is a key protocol in deciphering the language of transcription, however, as every experimental protocol, it suffers from biological and technical noise that can severely affect downstream analyses. DeepTSS is a novel DL-based method for effectively removing noisy CAGE signal. In contrast to existing software, DeepTSS does not require feature selection since the embedded convolutional layers can readily identify patterns and only utilize the important ones for the classification task. This study highlights the key role that DL can play in Molecular Biology, by removing the inherent flaws of experimental protocols, that form the backbone of contemporary research. Here, we show how DeepTSS can unleash the full potential of an already popular and mature method such as CAGE, and push the boundaries of coding and non-coding gene expression regulator research even further.
Dimitris Grigoriadis, Nikos Perdikopanis, Georgios K. Georgakilas, Artemis G. Hatzigeorgiou
BMC Bioinform.4
2018 ECCB 2018: The 17th European Conference on Computational Biology
abstract
This volume of Bioinformatics includes the proceedings papers of the 17th European Conference in Computational Biology (ECCB), an annual international Conference for research in computational biology and bioinformatics. The Conference is being held jointly with ISMB in the odd-numbered years and independently in the even-numbered years. This year, the 17th ECCB (ECCB 2018) will take place in Athens, at the Stavros Niarchos Foundation Cultural Center (SNFCC), from September 8 to 12, 2018. SNFCC, which was opened in June 2017, is a multifunctional arts, education and entertainment complex located on the edge of Faliro Bay, 4.5 km south of the center of Athens. Information on the ECCB 2018 can be found at eccb18.org and will be later archived at www.ebi.ac.uk/eccb/2018/. With more than 1,000 participants from academia and industry, ECCB is the leading European Conference on computational biology and bioinformatics and the second largest internationally, next to ISMB (Intelligent Systems in Molecular Biology). The work presented in ECCB is related to all domains of the field of computational biology and bioinformatics, ranging from molecular level to systems biology level. The proceedings papers present new computational methodologies and tools for addressing challenging problems in the field, including research in single cell high-throughput data, microbiome data, 3 D organization of the genome, and others. The Conference also includes Highlights presentations, which showcase important papers published over the last year. Since 2016, it also features an Applications track that presents computational biology work applied in industry, clinics, governmental organizations, and other fields beyond academia. ECCB meetings are held every year in a different European country. In odd-numbered years, ECCB is co-organized with the ISMB Conference, which is held in Europe. Previous meetings were organized in Prague, Czech Republic (Beerenwinkel and Bromberg, 2017, ISMB/ECCB 2017); The Hague, Netherlands (Heringa and Reinders, ECCB 2016); Dublin, Ireland (Moreau and Beerenwinkel, 2015, ISMB/ECCB 2015); Strasbourg, France (Devignes and Moreau, 2014, ECCB 2014); Berlin, Germany (Ben-Tal, 2013, ISMB/ECCB 2013); Basel, Switzerland (Schwede and Iber, 2012, ECCB 2012); Vienna, Austria (Gaasterland and Vingron, 2011, ISMB/ECCB 2011); Ghent, Belgium (Moreau and Heringa, 2010, ECCB 2010); Stockholm, Sweden (Gusfield and Tramontano, 2009, ISMB/ECCB 2009); Cagliari, Italy (Tramontano, 2008, ECCB 2008); Vienna, Austria (Lengauer et al., 2009, ISMB/ECCB 2007); Eilat, Israel (Wolfson and Safer, 2006, ECCB 2006); Madrid, Spain (Guigo et al., 2005, ECCB 2005); Glasgow, United Kingdom (Thornton et al., 2004, ISMB/ECCB 2004); Paris, France (Lenhof and Sagot, 2003, ECCB 2003); and Saarbrücken, Germany (Lengauer, 2002). ECCB 2018 continues the 16-year tradition and is held under the auspices of the Hellenic Society for Computational Biology and Bioinformatics (HSCBB; www.hscbb.gr), which promotes bioinformatics research and training in Greece since 2009. HSCBB organizes an annual conference in a different Greek city each time (e.g., Athens, Thessaloniki, Alexandroupolis, Patras, Lamia, Heraklion), aiming to provide exposure to new field developments to graduate students and researchers. The broad participation in the annual HSCBB conferences (more than 130 participants each year) show that the bioinformatics community in Greek universities and research centers has now matured significantly. HSCBB has also a strong international presence in the field of computational biology. It is an observer in the Greek ELIXIR consortium and is affiliated with the International Society for Computational Biology (ISCB) and Global Organization for Bioinformatics Learning, Education & Training (GOBLET). The organization of ECCB 2018, along with the European Student Council Symposium 2018 (see below), inspired the formation this year of the Greek chapter of the ISCB Student Council. ECCB 2018 received a record high number of 280 applications for proceedings talks from institutions from 48 countries. The submissions were organized in five themes, according to their topic: (1) Data (organization, integration, knowledge discovery, multi-scale modeling), (2) Genes (expression, function, editing, geno/phenotype), (3) Genome (sequence analysis, evolution, phylogeny, microbiome), (4) Proteins and Structural Biology (structure, function, alterations, drug design), and (5) Systems (molecular pathways, signaling, metabolomics). All proceedings submissions were subjected to a rigorous peer-review process (2-4 reviews per paper), organized by the members of the Programme Committees of the corresponding theme. The Programme Committees consisted of a Theme Chair (Michael Krauthammer, Roderic Guigo, Martin Vingron, Ivet Bahar, Alfonso Valencia, respectively) and two to four co-chairs (depending on the number of submissions assigned to this Theme). For the review, the EasyChair system (www.easychair.org) was used by 398 reviewers, recruited by the Programme Committees. Reviewers evaluated the impact and reproducibility of the presented research, as well as its suitability for the ECCB audience. Once the review was completed, the committee chairs and co-chairs selected 48 papers to be included in the ECCB 2018 proceedings (acceptance rate: 17%), with the number of papers accepted in each being proportional to the number of submissions initially assigned to this Theme. These papers required minor revisions and the authors had two weeks to modify them accordingly. All Proceedings Track papers and any supplementary files accompanying them are freely available in the electronic form of Oxford University Press journal Bioinformatics, as a special issue of September 2018. Together with the Proceedings track ECCB 2018 features 24 Highlights talks about scientific work already published in high impact science journals. They are coordinately presented and managed across the five themes with the proceeding presentations. We continued with newly established Application and ELIXIR tracks with 15 Applications talks, and 12 ELIXIR talks, all of which were selected after review. The aim of the Application track is to give a voice to those who apply computational biology in industry, clinics, governmental organizations, and other fields beyond academia. The aim of the ELIXIR Track is to showcase to the community the latest outputs and services from across the initiative. The presentations focus on developments relating to services and infrastructure within ELIXIR. ECCB 2018 also features a Poster Track, with 617 accepted posters, in which researchers present their latest findings in the corresponding thematic areas (Data, Genes, Genome, Proteins and Structural Biology, Systems). Seven distinguished keynote speakers will present their work: Prof. Bonnie Berger from the Massachusetts Institute of Technology (MIT), Prof. Christos Davatzikos from University of Pennsylvania, Prof. John Ioannidis from Stanford University, Prof. Manolis Kellis from MIT, Prof. Jose Onuchic from Rice University, Dr. Janet Thornton, Director Emeritus of the European Bioinformatics Institute (EMBL-EBI), and Dr. Eleftheria Zeggini from the German Research Center for Environmental Health (GmbH). Furthermore, ECCB 2018 is hosting for the first time three special invited talks before the lunch breaks and a joint keynote talk. The invited speakers will present the funding opportunities of the European Research Council (ERC) funding mechanisms (Dr. Maria Siomos) and a talk on the ISCB Student Council Internship Program (Farzana Rahman). The joint keynote will be about ELIXIR: A Common European Infrastructure for Bioinformatics Research presented by the ELIXIR director Dr. Niklas Blomberg. For the first time in ECCB 2018, ELIXIR organizes eight short workshops, namely ELIXIR TeSS Usability Study, Beacons, BioSchemas, Galaxy, OpenEbench Fnd Bio.tools, Research under the GDPR, Implementation of DMPs and Data Stewardship in practice, Secure access to your services using ELIXIR AAI. The 5th European Student Council Symposium (ESCS) is taking place before the main conference organized by the Student Council of the ISCB (chair Daniele Parisi from KU Leuven and co-chair Yvonne Saara Gladbach from Rostock University). ESCS highlights will be published in F1000Research via the ISCB Student Council channel. For the first time the ISCB Student Council will present also an invited talk as stated above. We are really thankful to the ISCB Student Council for their enthusiasm, hard work and genuine scientific interest, which has made the ESCS an inseparable part of ECCB. During the ECCB 2018, ten exhibitor booths of all sizes will be set up on the Exhibit Hall floor, next to the central Conference venue and close to the Poster session and lunch area, namely: ELIXIR, BioExcel, EMBL-EBI, ERC, ISCB and ISCB Student Council, Goblet, Oxford University Press, Cambridge University Press, Peer Community In. The booths will be grouped together in a way that provides an enhanced networking environment for our exhibitors and delegates. Exhibitors will showcase the latest trends in computational biology and bioinformatics technology, in scientific literature, as well as in modelling and simulation. We invite participants to visit the exhibition area and support these community-minded organizations, who deliver a strong message in supporting the Computational Biology and Bioinformatics scientific field. The weekend before the conference 8 workshops (including two from Special Interest Groups – SIGs) and 13 Tutorials will take place. The aim of workshops is to provide participants the opportunity to discuss different perspectives on the cutting edge of a selected research field, through presenting technical issues, exchanging research ideas and sharing practical experiences. The two SIGs and six workshops preceding the ECCB 2018 main conference were selected out of a total of 12 applications, running all but one for a single day: (SIG-1) BioExcel 2nd SIG Meeting: Advanced Simulations for Biomolecular Research organized by Rossen Apostolov (KTH Royal Institute of Technology in Stockholm, Sweden), Zoe Cournia (Biomedical Research Foundation of the Academy of Athens, Greece), Vera Matser, (EMBL-EBI, UK), Anastas Mishev (UKIM Saints Cyril and Methodius University of Skopje, Republic of Macedonia), Hrachya Astsatryan (Institute for Informatics and Automation Problems, Armenia), Adam Carter (EPCC, UK) (SIG-2): Intrinsically disordered proteins: advances and state-of-the-art of the field organized by Silvio Tosatto (University of Padua, Italy), Zsuzsanna Dosztanyi (Eötvös Loránd University, Hungary), Norman Davey (University College Dublin, Ireland), Damiano Piovesan (University of Padua, Italy) (W1) Computational Epigenomics a two-day workshop organized by Yassen Assenov (German Cancer Research Center, Germany), Guido Sanguinetti (University of Edinburgh, UK), Jordana Bell (King’s College London, UK), Jörn Walter (Saarland University, Germany), Christoph Bock (CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, Austria) and Verena Wolf (Saarland University, Germany) (W2) BioNetVisA 2018 workshop: From biological network reconstruction to data visualization and analysis in molecular biology and medicine organized by Inna Kuperstein (Institut Curie, France), Emmanuel Barillot (Institut Curie, France), Andrei Zinovyev (Institut Curie, France), Luis Cristobal Monraz Gomez (Institut Curie, France), Hioraki Kitano (RIKEN Center for Integrative Medical Sciences, Japan), Minoru Kanehisa (Institute for Chemical Research, Kyoto University, Japan), Samik Ghosh (Systems Biology Institute, Tokyo, Japan), Nicolas Le Novère (Babraham Institute, UK), Robin Haw (Ontario Institute for Cancer Research, Canada), Alfonso Valencia (Spanish National Bioinformatics Institute, Madrid, Spain), Lodewyk Wessels (Netherlands Cancer Institute, Amsterdam, Netherlands), Patrick Kemmeren (Princess Maxima Center for Pediatric Oncology, Utrecht, Netherlands) (W3) CPW2018: Computational Pathology Workshop—second edition organized by Yves Sucaet (Vrije Universiteit Brussel, Belgium), Jeroen Van der Laak (UMC Radboud, Netherlands), Zev Leifer (New York College of Podiatric Medicine, USA), Yukako Yagi (Memorial Sloan Kettering Cancer Center, USA), Raphaël Marée (Université de Liège, Belgium), David Ameisen (IRIF, CNRS and Université Paris Diderot, France), Paul Van Diest (UMC Utrecht, Netherlands), Jeffrey Fine (Magee-Womens Hospital of UPMC) (W4) Interactive visualizations to guide health diagnostics and personalized medicine organized by Dimitrios Tzovaras, Konstantinos Votis and Kostas Stamatopoulos (all at Information Technologies Institute of the Center for Research and Technology Hellas, Greece) (W5) Logical modelling of cellular networks organized by Anna Niarakis (Univ Evry, Université Paris-Saclay, France) and Denis Thieffry (Ecole Normale Supérieure, Paris, France). (W6) Recent Computational Advances in Metagenomics organized by Valentin Loux, Mahendra Mariadassou, Pierre Peterlongo and Sophie Schbath (all at INRA, France) The purpose of the ECCB tutorial program is to provide participants with lectures and hands-on training to the most important and emerging topics in bioinformatics and computational biology research. The tutorials offer skills ranging from early and basic steps of computational analysis in recently introduced topics to advanced computational skills in important established topics. A total of thirteen ECCB 2018 tutorials running for half or a full day will be held at ECCB 2018. These were selected out of sixteen applications: (T1) Automated Machine Learning for Bioinformatics and Computational Biology organized by Ioannis Tsamardinos (University of Crete, Greece; Gnosis Data Analysis, Greece), Kleio Maria Verrou (University of Crete, Greece) and Vincenzo Lagani (Ilia State University, Georgia; Gnosis Data Analysis, Greece) (T2) Computational Mass Spectrometry with OpenMS—From Algorithms to Integrated Workflows organized by Julianus Pfeuffer (Freie Universität Berlin, Germany) and Timo Sachsenberg (Universität Tübingen, Germany) (T3)—Write your own R script to jointly analyze RNA-, ATAC-seq, DNA methylation, SNPs and find drug targets in signal transduction networks using TRANSFAC® organized by Philip Stegmaier (geneXplain GmbH, Germany), Olga Kel-Margoulis (geneXplain GmbH, Germany), Alexander Kel (geneXplain GmbH, Germany; Institute of Systems Biology, Russia) (T4) Deep learning for predicting protein–DNA and –RNA binding organized by Yaron Orenstein (Ben-Gurion University, Israel) (T5) Delving into non-coding RNA with RNAcentral and Rfam. Organized by Anton Petrov and Ioanna Kalvari (both at EMBL-EBI, UK) (T6) DIANA Tools and Databases: In silico investigation of miRNA functions organized by Artemis Hatzigeorgiou, Spyros Tastsoglou, Dimitra Karagkouni, Nikos Perdikopanis, Giorgos Skoufos, Ioannis Kavakiotis (all at University of Thessaly, Greece) (T7) Exploring programmatic access to Protein sequence, function and structure with UniProt and PDBe organized by Andrew Nightingale and Mihaly Varadi (both at EMBL-EBI, UK) (T8) Fifth International Hands-on Tutorial on Logical Modelling: Exploring the dynamics of biological systems organized by Tomas Helikar (University of Nebraska, USA) and Juilee Thakar (University of Rochester Medical Center, USA) (T9) From an application to a fully integrated workflow—Comprehensive software engineering with SeqAn organized by René Rahn and Hannes Hauswedell (both at Freie Universität Berlin, Germany) (T10) Hands-on on Protein Function Prediction with Machine Learning and Interactive Analytics organized by Rabie Saidi and Tunca Dogan (EMBL-EBI, UK) (T11) Single-cell RNA-Seq Data Analysis organized by Panagiotis Papasaikas (Friedrich Miescher Institute for Biomedical Research (FMI), Basel, Switzerland; Swiss Institute of Bioinformatics, Switzerland) and Atul Sethi (Friedrich Miescher Institute for Biomedical Research (FMI), Basel, Switzerland; University Hospital Basel, University of Basel, Switzerland; Swiss Institute of Bioinformatics, Switzerland) (T12) Modern and scalable tools for efficient analysis of very large metagenomic datasets organized by Alexander Sczyrba (Bielefeld University, Germany), Christian Henke (Bielefeld University, Germany), Clovis Galiez (MPI for Biophysical Chemistry, Germany), Milot Mirdita (PI for Biophysical Chemistry, Germany) and Johannes Soeding (MPI for Biophysical Chemistry, Germany) (T13) User Experience Design for Computational Biologists organized by Nikiforos Karamanis and Xavier Watkins (both at EMBL-EBI, UK) Finally, following the tradition of ECCB, a number of travel fellowships, sponsored by ISCB, were given to young participants. We received 103 applications from scientists from 31 countries. After careful review and consideration, we awarded 14 travel fellowships, mainly to PhD students from nine countries that had an accepted proceedings paper. Volunteering is another important way for young scientists to visit and support the conference, as well as connect with other young scientists. The skills, talent and dedication of the volunteers are expected to largely contribute to the overall quality of the Conference. We made our best to accept all 91 applicants from 16 countries who were eligible to attend the Conference at a very low registration fee. Next to the volunteers, we would like to thank all the people who helped through their work to make this Conference a success. The Theme Chairs and Co-chairs as well all the reviewers who have been the heart of the Conference and helped selecting excellent papers, making thus this Conference a success. We are truly indebted to them. We are also grateful to the ECCB Steering Committee for their advice and continuous the of the organization of ECCB 2018. This conference would be at all the support from Anna Steering Committee and Chair ECCB our application and advice about to make a of was It was our to and inspired by for science with to in a will be in our and in our During the one year of the conference the support of Steering Committee Chair ECCB ECCB and Yves ECCB 2010, ECCB was Yves and for your continuous and We are also grateful to the ISCB Society for Computational for their and in ECCB 2018 at the international level and for This Conference would be the support of our and ELIXIR level and BioExcel was our level Cambridge University Press, Oxford University Press, Community and were ISCB, the Student Council of the ISCB, and who as exhibitors at ECCB 2018 also A thank to all for their and for to make ECCB 2018 an and Conference. A number of people to the organization of ECCB 2018. They beyond their of for making this conference really Maria a as was for the the the venue the and was for and Panagiotis was for the technical support of the registration Dr. organized the members of the DIANA also this Spyros Tastsoglou, Dimitra Karagkouni, Perdikopanis, Dimitrios Finally, the of a conference is its presentations, the conference for its participants. We thank all of for to ECCB 2018. We will the and in the city of Athens, the city and were of
Artemis G. Hatzigeorgiou, Pantelis G. Bagos, Panayiotis V. Benos, Christoforos Nikolaou, Yves Moreau, Ioannis Kavakiotis
Bioinform.1
2015 MirPub v2: Towards Ranking and Refining miRNA Publication Search Results
Ilias Kanellos, Vasiliki Vlachokyriakou, Thanasis Vergoulis, Georgios K. Georgakilas, Yannis Vassiliou, Artemis G. Hatzigeorgiou, Theodore Dalamagas 0001
TPDL6
2015 TarMiner: automatic extraction of miRNA targets from literature
abstract
MicroRNAs (miRNAs) are small RNA molecules that target particular genes and prohibit their expression. Since many important diseases are related to the expression or non-expression of particular genes, knowing the miRNAs that affect these genes can help in finding possible treatments. In the last decade, a large amount of experimental studies trying to reveal the targets of several miRNAs has been published. A handful of curated databases that collect miRNA targets from the literature have been developed to make this information more easily available. However, due to the large number of existing published articles, maintaining these databases up-to-date is a tedious task that requires important resources. In this work we introduce TarMiner, a pipeline for automatic extraction of miRNA targets that can facilitate the curation process of databases that maintain miRNA validated targets.
Rodothea-Myrsini Tsoupidi, Ilias Kanellos, Thanasis Vergoulis, Ioannis S. Vlachos, Artemis G. Hatzigeorgiou, Theodore Dalamagas 0001
SSDBM5
2015 mirPub: a database for searching microRNA publications
abstract
SUMMARY: Identifying, amongst millions of publications available in MEDLINE, those that are relevant to specific microRNAs (miRNAs) of interest based on keyword search faces major obstacles. References to miRNA names in the literature often deviate from standard nomenclature for various reasons, since even the official nomenclature evolves. For instance, a single miRNA name may identify two completely different molecules or two different names may refer to the same molecule. mirPub is a database with a powerful and intuitive interface, which facilitates searching for miRNA literature, addressing the aforementioned issues. To provide effective search services, mirPub applies text mining techniques on MEDLINE, integrates data from several curated databases and exploits data from its user community following a crowdsourcing approach. Other key features include an interactive visualization service that illustrates intuitively the evolution of miRNA data, tag clouds summarizing the relevance of publications to particular diseases, cell types or tissues and access to TarBase 6.0 data to oversee genes related to miRNA publications. AVAILABILITY AND IMPLEMENTATION: mirPub is freely available at http://www.microrna.gr/mirpub/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Thanasis Vergoulis, Ilias Kanellos, Nikos Kostoulas, Georgios K. Georgakilas, Timos K. Sellis, Artemis G. Hatzigeorgiou, Theodore Dalamagas 0001
Bioinform.6
2014 MR-microT: a MapReduce-based MicroRNA target prediction method
abstract
MicroRNAs (miRNAs) are small RNA molecules that inhibit the expression of particular genes, a function that makes them useful towards the treatment of many diseases. Computational methods that predict which genes are targeted by particular miRNA molecules are known as target prediction methods. In this paper, we present a MapReduce-based system, termed MR-microT, for one of the most popular and accurate, but computational intensive, prediction methods. MR-microT offers the highly requested by life scientists feature of predicting the targets of ad-hoc miRNA molecules in near-real time through an intuitive Web interface.
Ilias Kanellos, Thanasis Vergoulis, Dimitris Sacharidis, Theodore Dalamagas 0001, Artemis G. Hatzigeorgiou, Stelios Sartzetakis, Timos K. Sellis
SSDBM5
2012 TARCLOUD: A Cloud-Based Platform to Support miRNA Target Prediction
Thanasis Vergoulis, Michail Alexakis, Theodore Dalamagas 0001, Manolis Maragkakis, Artemis G. Hatzigeorgiou, Timos K. Sellis
SSDBM5
2012 Functional microRNA targets in protein coding sequences
abstract
Abstract Motivation: Experimental evidence has accumulated showing that microRNA (miRNA) binding sites within protein coding sequences (CDSs) are functional in controlling gene expression. Results: Here we report a computational analysis of such miRNA target sites, based on features extracted from existing mammalian high-throughput immunoprecipitation and sequencing data. The analysis is performed independently for the CDS and the 3′-untranslated regions (3′-UTRs) and reveals different sets of features and models for the two regions. The two models are combined into a novel computational model for miRNA target genes, DIANA-microT-CDS, which achieves higher sensitivity compared with other popular programs and the model that uses only the 3′-UTR target sites. Further analysis indicates that genes with shorter 3′-UTRs are preferentially targeted in the CDS, suggesting that evolutionary selection might favor additional sites on the CDS in cases where there is restricted space on the 3′-UTR. Availability: The results of DIANA-microT-CDS are available at www.microrna.gr/microT-CDS Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Martin Reczko, Manolis Maragkakis, Panagiotis Alexiou, Ivo Grosse, Artemis G. Hatzigeorgiou
Bioinform.5
2009 Lost in translation: an assessment and perspective for computational microRNA target identification
abstract
UNLABELLED: MicroRNAs (miRNAs) are a class of short endogenously expressed RNA molecules that regulate gene expression by binding directly to the messenger RNA of protein coding genes. They have been found to confer a novel layer of genetic regulation in a wide range of biological processes. Computational miRNA target prediction remains one of the key means used to decipher the role of miRNAs in development and disease. Here we introduce the basic idea behind the experimental identification of miRNA targets and present some of the most widely used computational miRNA target identification programs. The review includes an assessment of the prediction quality of these programs and their combinations. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Panagiotis Alexiou, Manolis Maragkakis, Giorgos L. Papadopoulos, Martin Reczko, Artemis G. Hatzigeorgiou
Bioinform.5
2009 DIANA-mirPath: Integrating human and mouse microRNAs in pathways
abstract
SUMMARY: DIANA-mirPath is a web-based computational tool developed to identify molecular pathways potentially altered by the expression of single or multiple microRNAs. The software performs an enrichment analysis of multiple microRNA target genes comparing each set of microRNA targets to all known KEGG pathways. The combinatorial effect of co-expressed microRNAs in the modulation of a given pathway is taken into account by the simultaneous analysis of multiple microRNAs. The graphical output of the program provides an overview of the parts of the pathway modulated by microRNAs, facilitating the interpretation and presentation of the analysis results. AVAILABILITY: The software is available at http://microrna.gr/mirpath and is free for all users with no login or download requirement.
Giorgos L. Papadopoulos, Panagiotis Alexiou, Manolis Maragkakis, Martin Reczko, Artemis G. Hatzigeorgiou
Bioinform.5
2009 Accurate microRNA target prediction correlates with protein repression levels
abstract
BACKGROUND: MicroRNAs are small endogenously expressed non-coding RNA molecules that regulate target gene expression through translation repression or messenger RNA degradation. MicroRNA regulation is performed through pairing of the microRNA to sites in the messenger RNA of protein coding genes. Since experimental identification of miRNA target genes poses difficulties, computational microRNA target prediction is one of the key means in deciphering the role of microRNAs in development and disease. RESULTS: DIANA-microT 3.0 is an algorithm for microRNA target prediction which is based on several parameters calculated individually for each microRNA and combines conserved and non-conserved microRNA recognition elements into a final prediction score, which correlates with protein production fold change. Specifically, for each predicted interaction the program reports a signal to noise ratio and a precision score which can be used as an indication of the false positive rate of the prediction. CONCLUSION: Recently, several computational target prediction programs were benchmarked based on a set of microRNA target genes identified by the pSILAC method. In this assessment DIANA-microT 3.0 was found to achieve the highest precision among the most widely used microRNA target prediction programs reaching approximately 66%. The DIANA-microT 3.0 prediction results are available online in a user friendly web server at http://www.microrna.gr/microT.
Manolis Maragkakis, Panagiotis Alexiou, Giorgos L. Papadopoulos, Martin Reczko, Theodore Dalamagas 0001, Giorgos Giannopoulos, Georgios I. Goumas, Evangelos Koukis, Kornilios Kourtis, Victor A. Simossis, Praveen Sethupathy, Thanasis Vergoulis, Nectarios Koziris, Timos K. Sellis, Panayiotis Tsanakas, Artemis G. Hatzigeorgiou
BMC Bioinform.16
2007 Global Discriminative Learning for Higher-Accuracy Computational Gene Prediction
abstract
Most ab initio gene predictors use a probabilistic sequence model, typically a hidden Markov model, to combine separately trained models of genomic signals and content. By combining separate models of relevant genomic features, such gene predictors can exploit small training sets and incomplete annotations, and can be trained fairly efficiently. However, that type of piecewise training does not optimize prediction accuracy and has difficulty in accounting for statistical dependencies among different parts of the gene model. With genomic information being created at an ever-increasing rate, it is worth investigating alternative approaches in which many different types of genomic evidence, with complex statistical dependencies, can be integrated by discriminative learning to maximize annotation accuracy. Among discriminative learning methods, large-margin classifiers have become prominent because of the success of support vector machines (SVM) in many classification tasks. We describe CRAIG, a new program for ab initio gene prediction based on a conditional random field model with semi-Markov structure that is trained with an online large-margin algorithm related to multiclass SVMs. Our experiments on benchmark vertebrate datasets and on regions from the ENCODE project show significant improvements in prediction accuracy over published gene predictors that use intrinsic features only, particularly at the gene level and on genes with long introns.
Axel Bernal, Koby Crammer, Artemis G. Hatzigeorgiou, Fernando Pereira 0003
PLoS Comput. Biol.3
2002 Finding Signal Peptides in Human Protein Sequences Using Recurrent Neural Networks
Martin Reczko, Petko Fiziev, Eike Staub, Artemis G. Hatzigeorgiou
WABI4
2002 Translation initiation start prediction in human cDNAs with high accuracy
abstract
MOTIVATION: Correct identification of the Translation Initiation Start (TIS) in cDNA sequences is an important issue for genome annotation. The aim of this work is to improve upon current methods and provide a performance guaranteed prediction. METHODS: This is achieved by using two modules, one sensitive to the conserved motif and the other sensitive to the coding/non-coding potential around the start codon. Both modules are based on Artificial Neural Networks (ANNs). By applying the simplified method of the ribosome scanning model, the algorithm starts a linear search at the beginning of the coding ORF and stops once the combination of the two modules predicts a positive score. RESULTS: According to the results of the test group, 94% of the TIS were correctly predicted. A confident decision is obtained through the use of the Las Vegas algorithm idea. The incorporation of this algorithm leads to a highly accurate recognition of the TIS in human cDNAs for 60% of the cases. AVAILABILITY: The program is available upon request from the author.
Artemis G. Hatzigeorgiou
Bioinform.1
2001 DIANA-EST: a statistical analysis
abstract
MOTIVATION: Expressed Sequence Tags (ESTs) are next to cDNA sequences as the most direct way to locate in silico the genes of the genome and determine their structure. Currently ESTs make up more than 60% of all the database entries. The goal of this work is the development of a new program called DNA Intelligent Analysis for ESTs (DIANA-EST) based on a combination of Artificial Neural Networks (ANN) and statistics for the characterization of the coding regions within ESTs and the reconstruction of the encoded protein. RESULTS: 89.7% of the nucleotides from an independent test set with 127 ESTs were predicted correctly as to whether they are coding or non coding. AVAILABILITY: The program is available upon request from the author. CONTACT: Present address: Department of Genetics, University of Pennsylvania, School of Medicine, 475 Clinical Research Building, 415 Curie Boulevard, Philadelphia, PA 19104-6145, USA. [email protected].
Artemis G. Hatzigeorgiou, Petko Fiziev, Martin Reczko
Bioinform.1
1999 Feature recognition on expressed sequence tags of human DNA
abstract
Expressed sequence tags (EST) are small parts of DNA, which are used to clone new genes. One main characteristic of EST is that they contain more than 1% sequencing errors. What we need to know is which parts of the EST contain information about proteins, the so called coding regions. In this paper we describe an error-tolerant program for the prediction of such coding regions in EST. The program is based on a combination of statistical methods and several artificial neural networks (ANN). 89.7% of the nucleotides of a independent test set with 127 EST's are predicted correctly as to whether they are coding or noncoding. These results are independent of the existence of homologous gene or protein sequences and representative for the application to the largest part of all EST.
Artemis G. Hatzigeorgiou, Martin Reczko
IJCNN1
1996 Computational analysis of transcriptional regulatory elements: a field in flux
abstract
Sequence analytic methods have played a role in the understanding of transcriptional regulation for many years (for example, an alignment of E. coli promoter regions showing conserved upstream regions was reported by Pribnow et al. in 1975). In recent years there has been a tremendous increase in experimental work aimed at understanding the fundamental biochemistry of transcription initiation as well as the mechanisms that regulate gene expression at the level of transcription. There is currently great interest in developing new computational methods as well. This interest is driven partly by the new biological understanding (and will hopefully contribute to it, by the use of quantitative models), and partly by the need to efficiently analyse newly determined genomic sequences. Computational biologists interested in transcription have made good progress in the last few years. For example, an improved ability to describe the DNA binding specificity of proteins involved in transcription lies at the foundation of much of the mathematical modeling in the field. It has been generally recognized that consensus sequences are usually inadequate to describe DNA-binding specificity, and it is now most common to describe the binding sites of a particular protein as the set of sequences scoring above a particular threshold with a Positional Weight Matrix (PWM). Considerable theoretical work, and some experimental effort, have gone into the development of algorithms to find a PWM from known binding sites, understand to what extent the description of specificity by means of a PWM is valid, and record PWMs for particular proteins. Other important areas of progress include computer programs for the recognition of eukaryotic promoters that for the first time have error rates low enough so that the program is of practical interest, and a number of recently developed data collections that are either more complete or more consistent than what is available in the primary sequence databases. It may also be counted as progress that some early errors of the field have now been corrected, so that (1) it is now widely recognized that transcriptional regulation is exceedingly complex, and that algorithms must take into account alternative pathways and the synergism of multiple transcription factors, and (2) cross-validation techniques are now commonly employed in the benchmarking of new algorithms for functional prediction. We feel that one of the main needs in the field is simply for better communication. Experimentalists often do not take advantage of the best computational techniques; algorithm developers often base their methods on an overly simplified view of the biology; computer scientists do not use the best data collections; and mathematicians sometimes show an aversion to learning about powerful machine learning techniques. In order to promote communication and collaboration, and to assess the state of the art, we organized the first International Workshop on Computational Analysis of Eukaryotic Transcriptional Regulatory Elements, at the Deutsches Krebsforschungszentrum in Heidelberg, in January of 1996. We were very pleased to have a highly interdisciplinary meeting of about 70 people, with participation from sequence analysts, pure experimentalists, computer scientists, researchers working on the nucleosome positioning problem, microscopists using computers for image processing, structural biologists interested in the 3D structure of promoters and gene regulatory proteins, and experts from the neighbouring field of prokaryotic gene transcription. It was particularly encouraging that there were several presentations at the meeting by groups comprising both experimental and computational biologists, and that communication between the experimental and the computational side seemed to be excellent. The interdisciplinary nature of the meeting also helped to focus attention on the primary scientific goal that all participants share: to understand, by modelling and model-testing, the transcription initiation event and its use in the regulation of gene expression. It remains a controversial issue whether function can be determined from the DNA sequence, at the level of symbol manipulation, without reference to 3D structure or other more biological representations of the data. (Most sequence analysis developers tacitly assume that such is possible, although it may be unwise to do so.) What is clear is that the work of a person in any one discipline will be much more effective if he or she is willing to understand and make use of the results of related disciplines. Experience to date suggests that in the analysis of transcriptional regulatory elements, more than in other domains of
Philipp Bucher, James W. Fickett, Artemis G. Hatzigeorgiou
Comput. Appl. Biosci.3
1995 A parallel neural network simulator on the connection machine CM-5
abstract
We here present a parallel implementation of artificial neural networks on the connection machine CM-5 and compare it with other parallel implementations on SIMD and MIMD architectures. This parallel implementation was developed with the goal of efficiently training large neural networks with huge training pattern sets for applications in molecular biology, in particular the prediction of coding regions in DNA sequences. The implementation uses training pattern parallelism and makes use of the parallel I/O facilities of the CM-5 and its efficient reduction operations available within the control network to achieve a high scalability. The parallel simulator obtains a maximum speed of 149.25 MCUPS for training feedforward networks with backpropagation on a 512 processor CM-5 system without using the CM-5 vector facility. The implementation poses no restriction on the type of network topology and works with different batch training algorithms like BP. Quickprop and Rprop.
Martin Reczko, Artemis G. Hatzigeorgiou, Niels Mache, Andreas Zell, Sándor Suhai
Comput. Appl. Biosci.2