Ilias Georgakopoulos-Soares

dblp:192/8270 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-3641-1488ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Accelerating inference in genomic and proteomic foundation models via speculative decoding
abstract
MOTIVATION: Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose inference is relatively slow, long-sequence generation quickly becomes costly. RESULTS: In this work we adapt speculative decoding to a representative GFM: the DNA model DNAGPT and two representative PFMs: ProGen2 and ProtGPT2. We implement a probabilistic variant of speculative decoding, in which a lightweight draft model proposes short token spans and a larger target model verifies or corrects them in parallel, while preserving the target model's sampling distribution. Across all three models we systematically study the effect of speculation window length, temperature, draft architecture and prompt length, and we benchmark tokens per second over multiple runs per configuration. Speculative decoding yields consistent speedups over standard key-value cached decoding, with maximum observed speedup reaching 100% increase, while average gains across models ranging between 20% and 40% (e.g. 1.2×-1.4×), without changing the underlying target model predictions. Our results show that speculative decoding is a practical and model-agnostic strategy for accelerating genomic and proteomic sequence generation without sacrificing prediction quality. AVAILABILITY AND IMPLEMENTATION: All code and results are freely available at https://github.com/Georgakopoulos-Soares-lab/BioSpecDec.
Kimonas Provatas, Aris Karatzikos, Charalampos Koilakos, Michail Patsakis, Alexandros Tzanakakis, Akshatha Nayak, Georgios A. Pavlopoulos, Ioannis Mouratidis, Evangelos Ioannis Avgoulas, Ilias Georgakopoulos-Soares
Bioinform.10
2026 Zimin patterns in genomes
abstract
Zimin words are words that have the same prefix and suffix. They are unavoidable patterns, with all sufficiently large strings encompassing them. Here, we examine for the first time the presence of k-mers not containing any Zimin patterns, defined hereafter as Zimin avoidmers, in the human genome. We report that in the reference human genome all k-mers above 104 base-pairs contain Zimin words. We find that Zimin avoidmers are most enriched in coding and Human Satellite 1 regions in the human genome. Zimin avoidmers display a depletion of germline insertions and deletions relative to surrounding genomic areas. We also apply our methodology in the genomes of another eight model organisms from all three domains of life, finding large differences in their Zimin avoidmer frequencies and their genomic localization preferences. We observe that Zimin avoidmers exhibit the highest genomic density in prokaryotic organisms, with E. coli showing particularly high levels, while the lowest density is found in eukaryotic organisms, with D. rerio having the lowest. Among the studied genomes the longest k-mer length at which Zimin avoidmers are observed is that of S. cerevisiae at k-mer length of 115 base-pairs. We conclude that Zimin avoidmers display inhomogeneous distributions in organismal genomes, have intricate properties including lower insertion and deletion rates, and disappear faster than the theoretical expected k-mer length, across the organismal genomes studied.
Nikol Chantzi, Ioannis Mouratidis, Ilias Georgakopoulos-Soares
PLoS Comput. Biol.3
2025 TaxaGO: a novel, phylogenetically informed gene ontology enrichment analysis tool
abstract
The functional interpretation of genes and their protein products across diverse species remains a central challenge in genomics, particularly as datasets grow in scale and complexity. The Gene Ontology (GO) knowledgebase offers a detailed resource of accessing a gene's function. While GO enrichment analysis tools are widely used to uncover biological insights, they are designed for single-species analyses and are not able to integrate phylogenetic relationships into the enrichment analyses. To address this, we created TaxaGO, a high-performance, multi-taxonomic GO enrichment analysis tool that incorporates evolutionary distances with species-level enrichment results to unravel GO enrichment profiles at a taxonomic level. Implemented in Rust for speed and scalability, TaxaGO enables robust cross-species GO enrichment analyses by combining species-specific results through phylogeny-aware statistical frameworks. It supports FASTA and CSV inputs, provides curated background populations for 12 131 species across Archaea, Bacteria, and Eukaryota, and offers advanced features such as count propagation, common ancestor analysis, semantic similarity calculation, and interactive visualizations. When benchmarking against established tools, TaxaGO demonstrates a maximum of 70.33× faster performance and 3.79× reduced memory usage. With an intuitive command-line interface and a user-friendly graphical interface, TaxaGO provides a powerful and accessible platform for functional genomics, evolutionary biology, and systems-level studies across the tree of life.
Eleftherios Bochalis, Antonios Papageorgiou, George Lagoumintzis, Dionysios V. Chartoumpekis, Ilias Georgakopoulos-Soares
Briefings Bioinform.5
2025 KmerCrypt: private k-mer search with homomorphic encryption
abstract
Outsourcing the storage and analysis of genomic data to third-party servers is often necessary due to the scale of modern datasets, but it introduces significant privacy challenges that must be addressed to ensure secure handling. K-mer-based analyses offer broad applications across genomics research, clinical diagnostics, pathogen surveillance, and metagenomic classification, though implementation requires careful ethical and technical considerations, particularly when processing human genomic data in clinical settings. We present a novel protocol utilizing homomorphic encryption that enables a client to store a fully encrypted version of a genome on an untrusted server and perform private k-mer searches. The protocol ensures the server never gains access to the client's non-encrypted genome sequence, nor does it learn the content of any k-mer query. After a one-time client-side encryption of the genome, the server performs all computations on ciphertext, returning only encrypted results that can be decrypted solely by the data owner. This framework transforms an honest but curious cloud server into a secure storage and computation system, enabling practical and confidential querying of encrypted, client-owned genomic data. The system supports exact k-mer searches on genomic data, as well as position weight matrix searches. Finally, we provide KmerCrypt, a private k-mer search toolkit that implements this protocol, offering researchers an efficient and secure solution for querying encrypted genomic datasets without compromising privacy.
Kimonas Provatas, Ioannis Mouratidis, Ilias Georgakopoulos-Soares
Briefings Bioinform.3
2025 ZSeeker: an optimized algorithm for Z-DNA detection in genomic sequences
abstract
Z-deoxyribonucleic acid (Z-DNA) is an alternative left-handed DNA structure with a zigzag-shaped backbone that differs from the right-handed canonical B-DNA helix. Z-DNA has been implicated in various biological processes, including transcription, replication, and DNA repair, and can induce genetic instability. Repetitive sequences of alternating purines and pyrimidines have the potential to adopt Z-DNA structures. ZSeeker is a novel computational tool developed for the accurate detection of potential Z-DNA-forming sequences in genomes, addressing key limitations of prior methods, such as computational inefficiency, difficult interpretability and usability, and lack of experimentally generated data. By introducing a novel methodology informed and validated by experimental data, ZSeeker enables the refined detection of potential Z-DNA-forming sequences. Built both as a standalone Python package and as an accessible web interface, ZSeeker allows users to input genomic sequences, adjust detection parameters, and view potential Z-DNA sequence distributions and Z-scores via downloadable visualizations. Our web platform provides a no-code solution for Z-DNA identification, with a focus on accessibility, user-friendliness, speed, and customizability. By providing efficient, high-throughput analysis, and enhanced detection accuracy, ZSeeker has the potential to support significant advancements in understanding the roles of Z-DNA in normal cellular functions, genetic instability, and its implications in human diseases.
Guliang Wang, Ioannis Mouratidis, Kimonas Provatas, Nikol Chantzi, Michail Patsakis, Ilias Georgakopoulos-Soares, Karen M. Vasquez
Briefings Bioinform.6
2025 MAFin: motif detection in multiple alignment files
abstract
MOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin.
Michail Patsakis, Kimonas Provatas, Fotis A. Baltoumas, Nikol Chantzi, Ioannis Mouratidis, Georgios A. Pavlopoulos, Ilias Georgakopoulos-Soares
Bioinform.7
2025 MAFcounter: an efficient tool for counting the occurrences of k-mers in MAF files
abstract
MOTIVATION: With the rapid expansion of large-scale biological datasets, DNA and protein sequence alignments have become essential for comparative genomics and proteomics. These alignments facilitate the exploration of sequence similarity patterns, providing valuable insights into sequence conservation, evolutionary relationships and for functional analyses. Typically, sequence alignments are stored in formats such as the Multiple Alignment Format (MAF). Counting k-mer occurrences is a crucial task in many computational biology applications, but currently, there is no algorithm designed for k-mer counting in alignment files. RESULTS: We have developed MAFcounter, the first k-mer counter dedicated to alignment files. MAFcounter is multithreaded, fast, and memory efficient, enabling k-mer counting in DNA and protein sequence alignment files with a wide variety of features for k-mer analysis. AVAILABILITY: MAFcounter is released under GPL license as a suite of binary C++ applications and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFcounter .
Michail Patsakis, Kimonas Provatas, Aris Karatzikos, Charalampos Koilakos, Ioannis Mouratidis, Ilias Georgakopoulos-Soares
BMC Bioinform.6
2017 MPRAnator: a web-based tool for the design of massively parallel reporter assay experiments
abstract
MOTIVATION: With the rapid advances in DNA synthesis and sequencing technologies and the continuing decline in the associated costs, high-throughput experiments can be performed to investigate the regulatory role of thousands of oligonucleotide sequences simultaneously. Nevertheless, designing high-throughput reporter assay experiments such as massively parallel reporter assays (MPRAs) and similar methods remains challenging. RESULTS: We introduce MPRAnator, a set of tools that facilitate rapid design of MPRA experiments. With MPRA Motif design, a set of variables provides fine control of how motifs are placed into sequences, thereby allowing the investigation of the rules that govern transcription factor (TF) occupancy. MPRA single-nucleotide polymorphism design can be used to systematically examine the functional effects of single or combinations of single-nucleotide polymorphisms at regulatory sequences. Finally, the Transmutation tool allows for the design of negative controls by permitting scrambling, reversing, complementing or introducing multiple random mutations in the input sequences or motifs. AVAILABILITY AND IMPLEMENTATION: MPRAnator tool set is implemented in Python, Perl and Javascript and is freely available at www.genomegeek.com and www.sanger.ac.uk/science/tools/mpranator The source code is available on www.github.com/hemberg-lab/MPRAnator/ under the MIT license. The REST API allows programmatic access to MPRAnator using simple URLs. CONTACT: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online.
Ilias Georgakopoulos-Soares, Naman Jain, Jesse M. Gray, Martin Hemberg
Bioinform.1