Nikolaos Alachiotis 0001

dblp:14/7434 · DBLP profile ↗
← Back
37ranked-venue papers
15as first author
14since 2021 · last 2026
0000-0001-8162-3792ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 12 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Empirical Study on the Realism Gap of Benchmarks for Multi-Objective Neural Architecture Search
Sebastian Bunda, Jeroen Rook, Nikolaos Alachiotis 0001, Luuk J. Spreeuwers
PPSN (2)3
2025 Accelerating Multiple Sequence Alignment Via Maximal Exact Match Identification
abstract
Multiple sequence alignment plays a crucial role in identifying evolutionary patterns and functional motifs, finding application in various fields such as genomics, phylogenetics, virology, and drug discovery. A major challenge in this domain is the scalability of alignment algorithms, particularly when aligning a large number of sequences. We introduce a novel algorithmic approach that builds on the concept of seed-andextend algorithms that were initially proposed for pairwise sequence alignment, and adapt it for multiple sequences. We describe a multiple sequence alignment approach for sequences that share conserved regions in the same order, dubbed MEMSA (MEM Extracting Multiple Sequence Aligner), that incorporates several pre-processing steps to reduce the alignment search space. MEMSA demonstrates promising results with datasets that exhibit high homology; however, it faces difficulties with highly divergent genomic sequences. When applied to a dataset comprising 500 genomes of the MERS (Middle East respiratory syndrome) coronavirus, MEMSA reduces alignment times by up to$27 \times$, while also improving alignment quality. MEMSA is available for download at https://github.com/timweh/MEMSA.
Tim Wehning, Nikolaos Alachiotis 0001
BIBE2
2025 Accelerated Phylogenetics on the AMD Versal Adaptive SoC
abstract
Phylogenetics study the evolutionary history of organisms using an iterative procedure of creating and evaluating phylogenetic trees. This procedure is highly compute-intensive; constructing a large phylogenetic tree requires hundreds to thousands of CPU hours. Most phylogenetic analyses today rely either on Maximum Likelihood (ML) or Bayesian Inference (BI) methods for inferring phylogenetic trees; the Phylogenetic Likelihood Function (PLF) is employed in both ML and BI approaches as the tree-evaluation function, accounting for up to 95% of the overall analysis time. In this work, we explore the AMD Versal Adaptive SoC architecture for accelerating the PLF that heavily relies on matrix multiplication operations. We find that the tight integration of domain-specific processors (AI Engines) with Programmable Logic (PL) in the Versal architecture is highly suitable for the needs of the PLF: we map the core operation of the PLF (matrix multiplication) to the AI Engines and deploy special-purpose units in the PL for PLF-specific data caching and input/output management, as well as numerical scaling that is a prerequisite for yielding numerically stable solutions for large-scale phylogenetic studies. We conducted a thorough performance analysis to leverage the platform capabilities to guide matrix-multiplication acceleration that requires close PL-AIE cooperation. We observed between \(23.8\times\) and \(47.0\times\) higher computational power of the Versal SoC than one x86 CPU core (both AMD and Intel) using AVX2 intrinsics, and between \(3.7\times\) and \(5.9\times\) higher performance than eight cores. For the full system, we observe comparable performance ( \(\pm\ 30%\) ) with eight CPU cores due to device memory access and PCIe limitations, showing the potential of the Versal architecture in accelerating a distinct application from DSP/AI, its primary design focus.
Geert Roks, Mario Ruiz Noguera, Nikolaos Alachiotis 0001
ACM Trans. Reconfigurable Technol. Syst.3
2024 Lightweight Instrumentation for Accurate Performance Monitoring in RTOSes
abstract
Evaluating performance metrics in embedded systems poses challenges, particularly due to the limited set of tools available for monitoring performance counters. In addition, performance evaluation frameworks for Real-Time Operating Systems (RTOSes) often lack the sophistication and capabilities available in general-purpose operating systems like Linux, which benefit from utilities such as perf_event. To bridge this gap, this paper presents an accurate and low-overhead instrumentation utility tailored for RTOSes. Our approach utilizes performance monitoring counters to observe individual user applications within the RTOS environment. Importantly, it enables comprehensive application monitoring by strategically placing probes at points of inherent system interference, thereby minimizing additional overhead. A pre-calibration of these probes allows for fine-grained measurements within user applications. This results in the elimination of 100 % of the overheads for most counters in our test configuration, impacting the context switch by only three additional instructions per monitored counter.
Bruno Endres Forlin, Kuan-Hsun Chen, Nikolaos Alachiotis 0001, Luca Cassano, Marco Ottavi
DATE3
2024 Exploring the Versal AI Engines for Signal Processing in Radio Astronomy
abstract
Nowadays, heterogeneous architectures are widely used to overcome the ongoing demand for increased computing performance at the edge, such as the pre-processing of raw antenna data in radio telescope systems. The Versal Adaptive SoC is a novel heterogeneous architecture that includes Programmable Logic (PL), a Processing System, and Artificial Intelligence Engines (AIEs) interconnected by a programmable Network-on-Chip. In this work, we explore the AIEs to evaluate their capabilities for real-time signal processing in radio telescope systems. We focus on the implementation of a Polyphase Filter Bank (PFB), which is a representative signal processing operation that consists of Finite Impulse Response (FIR) filters and the Fast-Fourier Transform (FFT) algorithm. We analyzed the performance of the AIEs with regard to the requirements of LOFAR, the world’s largest low-frequency radio telescope. By means of the roofline model, we reveal that the AIEs theoretically meet the LOFAR computational requirements, but the FIR vendor-provided library implementation did not reach the required performance. Therefore, we explore several optimization strategies for the FIR implementation on the AIEs and analyze the communication options between the PL and the AIE array. Finally, we have developed an efficient PFB implementation that requires only 12 AIEs. A prototype on a VC1902 device achieves a throughput of 437 MSPS, which is more than sufficient for a single antenna polarization of the LOFAR system. This work is open source and publicly available at: https://git.astron.nl/rd/acap.
Victor Van Wijhe, Vincent Sprave, Daniele Passaretti, Nikolaos Alachiotis 0001, Gerrit Grutzeck, Thilo Pionteck, Steven van der Vlugt
FPL4
2024 Scalable CNN-based classification of selective sweeps using derived allele frequencies
abstract
MOTIVATION: Selective sweeps can successfully be distinguished from neutral genetic data using summary statistics and likelihood-based methods that analyze single nucleotide polymorphisms (SNPs). However, these methods are sensitive to confounding factors, such as severe population bottlenecks and old migration. By virtue of machine learning, and specifically convolutional neural networks (CNNs), new accurate classification models that are robust to confounding factors have been recently proposed. However, such methods are more computationally expensive than summary-statistic-based ones, yielding them impractical for processing large-scale genomic data. Moreover, SNP data are frequently preprocessed to improve classification accuracy, further exacerbating the long analysis times. RESULTS: To this end, we propose a 1D CNN-based model, dubbed FAST-NN, that does not require any preprocessing while using only derived allele frequencies instead of summary statistics or raw SNP data, thereby yielding a sample-size-invariant, scalable solution. We evaluated several data fusion approaches to account for the variance of the density of genetic diversity across genomic regions (a selective sweep signature), and performed an extensive neural architecture search based on a state-of-the-art reference network architecture (SweepNet). The resulting model, FAST-NN, outperforms the reference architecture by up to 12% inference accuracy over all challenging evolutionary scenarios with confounding factors that were evaluated. Moreover, FAST-NN is between 30× and 259× faster on a single CPU core, and between 2.0× and 6.2× faster on a GPU, when processing sample sizes between 128 and 1000 samples. Our work paves the way for the practical use of CNNs in large-scale selective sweep detection. AVAILABILITY AND IMPLEMENTATION: https://github.com/SjoerdvandenBelt/FAST-NN.
Sjoerd van den Belt, Nikolaos Alachiotis 0001
Bioinform.3
2023 Effective Data Preprocessing Techniques for CNN-based Selective Sweep Detection
abstract
Identifying positive selection has been cast as a classification task, with Convolutional Neural Networks (CNNs) already delivering higher accuracy than summary statistics and likelihood-based approaches. While several CNN-based methods rearrange the pixels of images representing raw genomic data as a preprocessing technique to enhance classification accuracy, the effectiveness of such pixel-rearrangement methods has not been thoroughly studied in the presence of confounding factors such as population bottlenecks and recombination hotspots. Here, we present a series of pixel-rearrangement algorithms to increase CNN classification accuracy for selective sweep detection, and evaluate the performance of four CNN models that are specifically designed for detecting selective sweeps. We find that data preprocessing based on pixel-rearrangement algorithms significantly improves the overall classification accuracy of a given CNN for diverse datasets simulating confounding factors. We observe up to 24.55% higher top-1 accuracy than using the preprocessing algorithms proposed by the authors of each CNN architecture. Furthermore, our results suggest a correlation between the stability of the rearrangement algorithms (over the different CNN architectures and confounding factors) and their performance. Based on these findings, we make suggestions for the most suitable preprocessing technique per CNN architecture used in this study. We provide the data rearrangement algorithms as a distinct module available for download at: https://github.com/Zhaohq96/Genetic-data-rearrangement.
Nikolaos Alachiotis 0001
BIBM2
2023 Exploring Genomic Sequence Alignment for Improving Side-Channel Analysis
Heitor Uchoa, Vipul Arora 0004, Dennis Vermoen, Marco Ottavi, Nikolaos Alachiotis 0001
ESORICS (3)5
2023 An unprotected RISC-V Soft-core processor on an SRAM FPGA: Is it as bad as it sounds?
abstract
Fast development, low cost, and reconfigurability are becoming critical factors for aerospace applications, making SRAM FPGAs attractive. However, SRAM FPGAs are prone to errors in the user and on the configuration bits. For their correct functioning, they must be capable of withstanding failures without sacrificing much performance. When adjusting a soft core for these applications, it is essential to know where redundancies are necessary, to avoid unnecessary overhead. We characterize the reliability of an unprotected RISC-V microcontroller using an accelerated neutron beam. Our investigation shows that, for our chosen benchmark and processor, the user data in the memory banks is the leading cause of the total number of errors in the application. By reversing the benchmark operations, we could root cause the origin of the observed errors and found that most of the data corruption detected during the runs stem from previously corrupt input data or from output data that were corrupted while transmitting.
Bruno Endres Forlin, Wouter van Huffelen, Carlo Cazzaniga, Paolo Rech, Nikolaos Alachiotis 0001, Marco Ottavi
ETS5
2023 Genome-wide scans for selective sweeps using convolutional neural networks
abstract
MOTIVATION: Recent methods for selective sweep detection cast the problem as a classification task and use summary statistics as features to capture region characteristics that are indicative of a selective sweep, thereby being sensitive to confounding factors. Furthermore, they are not designed to perform whole-genome scans or to estimate the extent of the genomic region that was affected by positive selection; both are required for identifying candidate genes and the time and strength of selection. RESULTS: We present ASDEC (https://github.com/pephco/ASDEC), a neural-network-based framework that can scan whole genomes for selective sweeps. ASDEC achieves similar classification performance to other convolutional neural network-based classifiers that rely on summary statistics, but it is trained 10× faster and classifies genomic regions 5× faster by inferring region characteristics from the raw sequence data directly. Deploying ASDEC for genomic scans achieved up to 15.2× higher sensitivity, 19.4× higher success rates, and 4× higher detection accuracy than state-of-the-art methods. We used ASDEC to scan human chromosome 1 of the Yoruba population (1000Genomes project), identifying nine known candidate genes.
Matthijs Leon Souilljee, Pavlos Pavlidis, Nikolaos Alachiotis 0001
Bioinform.4
2023 A Survey of Processing Systems for Phylogenetics and Population Genetics
abstract
The COVID-19 pandemic brought Bioinformatics into the spotlight, revealing that several existing methods, algorithms, and tools were not well prepared to handle large amounts of genomic data efficiently. This led to prohibitively long execution times and the need to reduce the extent of analyses to obtain results in a reasonable amount of time. In this survey, we review available high-performance computing and hardware-accelerated systems based on FPGA and GPU technology. Optimized and hardware-accelerated systems can conduct more thorough analyses considerably faster than pure software implementations, allowing to reach important conclusions in a timely manner to drive scientific discoveries. We discuss the reasons that are currently hindering high-performance solutions from being widely deployed in real-world biological analyses and describe a research direction that can pave the way to enable this.
Reinout Corts, Nikolaos Alachiotis 0001
ACM Trans. Reconfigurable Technol. Syst.2
2022 Scalable Phylogeny Reconstruction with Disaggregated Near-memory Processing
abstract
Disaggregated computer architectures eliminate resource fragmentation in next-generation datacenters by enabling virtual machines to employ resources such as CPUs, memory, and accelerators that are physically located on different servers. While this paves the way for highly compute- and/or memory-intensive applications to potentially deploy all CPUs and/or memory resources in a datacenter, it poses a major challenge to the efficient deployment of hardware accelerators: input/output data can reside on different servers than the ones hosting accelerator resources, thereby requiring time- and energy-consuming remote data transfers that diminish the gains of hardware acceleration. Targeting a disaggregated datacenter architecture similar to the IBM dReDBox disaggregated datacenter prototype, the present work explores the potential of deploying custom acceleration units adjacently to the disaggregated-memory controller on memory bricks (in dReDBox terminology), which is implemented on FPGA technology, to reduce data movement and improve performance and energy efficiency when reconstructing large phylogenies (evolutionary relationships among organisms). A fundamental computational kernel is the Phylogenetic Likelihood Function (PLF), which dominates the total execution time (up to 95%) of widely used maximum-likelihood methods. Numerous efforts to boost PLF performance over the years focused on accelerating computation; since the PLF is a data-intensive, memory-bound operation, performance remains limited by data movement, and memory disaggregation only exacerbates the problem. We describe two near-memory processing models, one that addresses the problem of workload distribution to memory bricks, which is particularly tailored toward larger genomes (e.g., plants and mammals), and one that reduces overall memory requirements through memory-side data interpolation transparently to the application, thereby allowing the phylogeny size to scale to a larger number of organisms without requiring additional memory.
Nikolaos Alachiotis 0001, Panagiotis Skrimponis, Emmanouil Pissadakis, Dionisios N. Pnevmatikatos
ACM Trans. Reconfigurable Technol. Syst.1
2021 Integrative hierarchical ensemble clustering for improved disease subtype discovery
abstract
Multi-omics clustering methods are used for the stratification of patients into sub-groups of similar molecular characteristics. In recent years, a wide range of methods has been developed for this purpose. However, due to the high diversity of cancer-related data, a single method may not perform sufficiently well in all cases. Here, we propose a comprehensive framework for multi-omics hierarchical ensemble clustering. We provide a flexible environment that allows to build hierarchical clustering ensembles suitable for the available data and research goals. Survival analyses for data from The Cancer Genome Atlas (TCGA) indicate that our proposed ensembles provide more robust, and thus more reliable results than the state-of-the-art. We have implemented our architecture within the R-package HC-fused, which is freely available on Github.
Bastian Pfeifer, Andrei Voicu-Spineanu, Michael G. Schimek, Nikolaos Alachiotis 0001
BIBM4
2021 Increasing Flexibility of FPGA-based CNN Accelerators with Dynamic Partial Reconfiguration
abstract
Convolutional Neural Networks (CNN) are widely used for image classification and have achieved significantly accurate performance in the last decade. However, they require computationally intensive operations for embedded applications. In recent years, FPGA-based CNN accelerators have been proposed to improve energy efficiency and throughput. While dynamic partial reconfiguration (DPR) is increasingly used in CNN accelerators, the performance of dynamically reconfigurable accelerators is usually lower than the performance of pure static FPGA designs. This work presents a dynamically reconfigurable CNN accelerator architecture that does not sacrifice throughput performance or classification accuracy. The proposed accelerator is composed of reconfigurable macroblocks and dynamically utilizes the device resources according to model parameters. Moreover, we devise a novel approach, to the best of our knowledge, to hide the computations of the pooling layers inside the convolutional layers, thereby further improving throughput. Using the proposed architecture and DPR, different CNN architectures can be realized on the same FPGA with optimized throughput and accuracy. The proposed architecture is evaluated by implementing two different LeNet CNN models trained by different datasets and classifying different classes. Experimental results show that the implemented design achieves higher throughput than current LeNet FPGA accelerators.
Hasan Irmak, Daniel Ziener, Nikolaos Alachiotis 0001
FPL3
2020 Exploring Modern FPGA Platforms for Faster Phylogeny Reconstruction with RAxML
abstract
The Phylogenetic Likelihood Function (PLF) is one of the cornerstone functions in most phylogenetic inference tools; its execution represents the majority of time required to complete an analysis. This work proposes the acceleration of this function using reconfigurable hardware accelerators, focusing on system-on-chips that integrate Field Programmable Gate Array (FPGA) resources as well as traditional High Performance Computing (HPC) systems that use FPGA-based accelerator cards. Taking into account the specific properties of each platform in order to exploit their processing capabilities, the proposed solutions provide significant performance gains. The measured acceleration of PLF function is up to 8x while the overall time to complete a phylogenetic analysis using the popular RAxML software can be reduced up to 3.2 times (with respect to a pure software implementation on a high-end server processor). Compared to other similar solutions proposed in literature, our systems perform up to 65% faster.
Pavlos Malakonakis, Andreas Brokalakis, Nikolaos Alachiotis 0001, Euripides Sotiriades, Apostolos Dollas
BIBE3
2020 qLD: High-performance Computation of Linkage Disequilibrium on CPU and GPU
abstract
Linkage disequilibrium (LD) is the non-random association between alleles at different loci. Assessing LD in thousands of genomes and/or millions of single-nucleotide poly-morphisms (SNPs) exhibits excessive time and memory requirements that can potentially hinder future large-scale genomic analyses. To this end, we introduce qLD (quickLD) (https//lgithub.com/StrayLamb2lqLD), a highly optimized open-source software that assesses LD based on Pearson's correlation coefficient. qLD exploits the fact that the computational kernel for calculating LD can be cast in terms of dense linear algebra operations. In addition, the software employs memory-aware techniques to lower memory requirements, and parallel GPU architectures to further shorten analysis times. qLD delivers up to 5x faster processing than the current state-of-the-art software implementation when run on the same CPU, and up to 29x when computation is offloaded to a GPU. Furthermore, the software is designed to quantity allele associations between arbitrarily distant loci in a time-and memory-efficient way, thereby facilitating the evaluation of long-range LD and the detection of co-evolved genes. We showcase qLD on the analysis of 22,554 complete SARS-CoV-2 genomes.
Charalampos Theodoris, Nikolaos Alachiotis 0001, Tze Meng Low, Pavlos Pavlidis
BIBE2
2020 Near-memory Acceleration for Scalable Phylogenetic Inference
abstract
Phylogenetics study the evolutionary history of a collection of organisms based on observed heritable molecular traits, finding practical application in a wide range of domains, from conservation biology and epidemiology, to forensics and drug development. A fundamental computational kernel to evaluate evolutionary histories, also referred to as phylogenies, is the Phylogenetic Likelihood Function (PLF), which dominates the total execution time (by up to 95%) of widely used maximum-likelihood phylogenetic methods. Numerous efforts to boost PLF performance over the years mostly focused on accelerating computation; since the PLF is a data-intensive, memory-bound operation, performance remains limited by data movement. In this work, we employ near-memory computation units (NMUs) within a FPGA-based computing environment with disaggregated memory to alleviate the data movement problem and improve performance and energy efficiency when inferring large-scale phylogenies. NMUs were deployed on a multi-FPGA emulation platform for the IBM dReDBox disaggregated datacenter prototype. We find that performance and power efficiency improves by an order of magnitude when NMUs compute on local data that reside on the same server tray. This is achieved through an efficient data-allocation scheme that minimizes inter-tray data transfers (remote-data movement) when computing the PLF. More specifically, we observe up to 22x better FLOPS performance and 13x higher power efficiency (FLOPS/Watt) over the more traditional, accelerator-as-a-coprocessor model, which requires explicit remote-data transfers between disaggregated memory modules and accelerator units.
Nikolaos Alachiotis 0001, Panagiotis Skrimponis, Emmanouil Pissadakis, Sundeep Rangan, Dionisios N. Pnevmatikatos
FPGA1
2020 RAiSD-X: A Fast and Accurate FPGA System for the Detection of Positive Selection in Thousands of Genomes
abstract
Detecting traces of positive selection in genomes carries theoretical significance and has practical applications from shedding light on the forces that drive adaptive evolution to the design of more effective drug treatments. The size of genomic datasets currently grows at an unprecedented pace, fueled by continuous advances in DNA sequencing technologies, leading to ever-increasing compute and memory requirements for meaningful genomic analyses. The majority of existing methods for positive selection detection either are not designed to handle whole genomes or scale poorly with the sample size; they inevitably resort to a runtime versus accuracy tradeoff, raising an alarming concern for the feasibility of future large-scale scans. To this end, we present RAiSD-X, a high-performance system that relies on a decoupled access-execute processing paradigm for efficient FPGA acceleration and couples a novel, to our knowledge, sliding-window algorithm for the recently introduced μ statistic with a mutation-driven hashing technique to rapidly detect patterns in the data. RAiSD-X achieves up to three orders of magnitude faster processing than widely used software implementations, and more importantly, it can exhaustively scan thousands of human chromosomes in minutes, yielding a scalable full-system solution for future studies of positive selection in species of flora and fauna.
Nikolaos Alachiotis 0001, Charalampos Vatsolakis, Grigorios Chrysos 0001, Dionisios N. Pnevmatikatos
ACM Trans. Reconfigurable Technol. Syst.1
2018 REMAP: Remote mEmory Manager for disAggregated Platforms
abstract
Disaggregated computing is a new approach that promises to alleviate the problem of fixed resource proportionality in datacenter deployments. Two critical factors that affect the overall performance of disaggregated platforms are remote memory access latency and throughput. Previous works primarily expose remote data processing at the applcation level that (a) require code annotations and/or the use of custom user-level libraries, and (b) may hinder the overall system protection and functionality. In this paper, we are taking a different approach: we propose the Remote mEmory Manager for dis-Aggregated Platforms (REMAP), a hardware architecture that enables the hotplug of remote memory resources to processing nodes, as normal paged memory at the OS-level, without requiring application-level code modifications. REMAP tightly couples processing nodes with remote memory controllers. Our architecture “expands” system memory on demand, by dynamically attaching remote memory modules to unused Local Physical Address (LPA) ranges, where the memory access requests are tunneled over high-speed, low-latency serial links. To evaluate REMAP in terms of performance, we implemented a prototype using two zcul02 FPGA boards. REMAP provides a remote cache-line access latency of less than 750 nsec, and up to 1.3× overall system throughput, compared to a baseline CPU-memory configuration.
Dimitris Theodoropoulos 0001, Andrea Reale, Dimitris Syrivelis, Maciej Bielski, Nikolaos Alachiotis 0001, Dionisios N. Pnevmatikatos
ASAP5
2018 dReDBox: Materializing a full-stack rack-scale system prototype of a next-generation disaggregated datacenter
abstract
Current datacenters are based on server machines, whose mainboard and hardware components form the baseline, monolithic building block that the rest of the system software, middleware and application stack are built upon. This leads to the following limitations: (a) resource proportionality of a multi-tray system is bounded by the basic building block (mainboard), (b) resource allocation to processes or virtual machines (VMs) is bounded by the available resources within the boundary of the mainboard, leading to spare resource fragmentation and inefficiencies, and (c) upgrades must be applied to each and every server even when only a specific component needs to be upgraded. The dRedBox project (Disaggregated Recursive Datacentre-in-a-Box) addresses the above limitations, and proposes the next generation, low-power, across form-factor datacenters, departing from the paradigm of the mainboard-as-a-unit and enabling the creation of function-block-as-a-unit. Hardware-level disaggregation and software-defined wiring of resources is supported by a full-fledged Type-1 hypervisor that can execute commodity virtual machines, which communicate over a low-latency and high-throughput software-defined optical network. To evaluate its novel approach, dRedBox will demonstrate application execution in the domains of network functions virtualization, infrastructure analytics, and real-time video surveillance.
Maciej Bielski, Ilias Syrigos, Kostas Katrinis, Dimitris Syrivelis, Andrea Reale, Dimitris Theodoropoulos 0001, Nikolaos Alachiotis 0001, Dionisios N. Pnevmatikatos, E. H. Pap, Georgios Zervas, Vaibhawa Mishra, Arsalan Saljoghei, Alvise Rigo, Jose Fernando Zazo, Sergio López-Buedo, Martí Torrents, Ferad Zyulkyarov, Michael Enrico, Óscar González de Dios
DATE7
2018 Accelerated Inference of Positive Selection on Whole Genomes
abstract
Positive selection is the tendency of beneficial traits to increase in prevalence in a population. Its detection carries theoretical significance and has practical applications, from shedding light on the forces that drive adaptive evolution to identifying drug-resistant mutations in pathogens. With next-generation sequencing producing a plethora of genomic data for population genetic analyses, the increased computational complexity of existing methods and/or inefficient memory management hinders the efficient analysis of large-scale datasets. To this end, we devise a system-level solution that couples a generic out-of-core algorithm for parsing genomic data with a decoupled access/execute accelerator architecture, thereby providing a method-independent infrastructure for the rapid and scalable inference of positive selection. We employ a novel detection mechanism that mostly relies on integer arithmetic operations, which fit well to FPGA fabric, while yielding qualitatively superior results than current state-of-the-art methods. We deploy a high-end system that pairs Hybrid Memory Cube with a mid-range FPGA, forming a high-throughput streaming accelerator that achieves 751x, 62x, and 20x faster analyses of simulated genomes than the widely used software tools SweepFinder2 (1 thread), OmegaPlus (40 threads), and SweeD (40 threads), respectively. Importantly, our solution can scan thousands of human genomes and millions of genetic polymorphisms (1000 Genomes dataset, 5,008 samples) in a matter of hours, requiring between 4 and 22 minutes per autosome, depending on the chromosomal length.
Nikolaos Alachiotis 0001, Charalampos Vatsolakis, Grigorios Chrysos 0001, Dionisios N. Pnevmatikatos
FPL1
2018 ReFiRe: Efficient Deployment of Remote Fine-Grained Reconfigurable Accelerators
abstract
The need for specialized hardware acceleration in today's computing platforms is well established, due to power and efficiency reasons. Broadening an accelerator's scope of application is highly desirable, but requires a finer-grained architecture with basic primitives, which inevitably exhibits increased communication and synchronization requirements. In disaggregated-computing environ-ments, where data transfers between remote nodes are realized via datacenter-wide packet exchanges, reducing communication and synchronization is a prerequisite for the effective employment of remote acceleration. To this end, we present ReFiRe (Remote Fine-grained Reconfigurable acceleration), a generic deployment framework with native support for partial reconfiguration that allows to considerably reduce communication needs between a processor and remote accelerators. This is achieved by shifting control flow, partial reconfiguration, and execution decisions to the remote side through arbitrarily long instructions that encapsulate complex sequences of operations and their re-spective synchronization requirements. ReFiRe outperforms an SDSoC-generated accelerator system that employs the same accelerator cores to boost performance of a genomics application that detects positive selection.
Emmanouil Pissadakis, Nikolaos Alachiotis 0001, Panagiotis Skrimponis, Dimitris Theodoropoulos 0001, Thanasis Korakis, Dionisios N. Pnevmatikatos
FPT2
2017 Multi-FPGA Evaluation Platform for Disaggregated Computing
abstract
We present a versatile FPGA-based evaluation platform for exploring alternative execution strategies on disaggregated environments for applications, considering different processing block types: compute cores, memory, and accelerators. Developers can interconnect different blocks types in order to create optimal configurations. A user-level software library allows quick mapping of applications on real hardware. We have implemented a fully working prototype using three ZC706 FPGA boards, and evaluated different software / hardware configurations of a matrix multiplication benchmark.
Dimitris Theodoropoulos 0001, Nikolaos Alachiotis 0001, Dionisios N. Pnevmatikatos
FCCM2
2017 Versatile deployment of FPGA accelerators in disaggregated data centers: A bioinformatics case study
abstract
Important design considerations for the cost-effective employment of hardware accelerators in next-generation data centers involve a) the type of candidate applications that a proposed solution can accelerate (generality), and b) the required development effort to successfully deploy the available accelerators for a given application (adoption overhead). To address the problem of generality, we present a versatile and dynamically reconfigurable hardware architecture that exhibits several accelerator slots and programmable interconnect to create application-specific accelerator datapaths. The proposed architecture fits in the model of disaggregated data centers, where compute, memory, and accelerators are broadly regarded as large pools of resources, and subsets of these resource pools are dynamically allocated on an as-needed basis to cooperatively boost performance of a broad range of applications. Initial results for a bioinformatics application that we employ as a case study and deals with the detection of positive selection in large-scale genomic datasets reveal a speedup of up to 6.4X when custom hardware accelerators are mapped to the proposed versatile accelerator architecture and compared with a parallel and highly optimized software implementation executed on a multi-core processor.
Nikolaos Alachiotis 0001, Dimitris Theodoropoulos 0001, Dionisios N. Pnevmatikatos
FPL1
2017 Deploying FPGAs to future-proof genome-wide analyses based on linkage disequilibrium
abstract
The ever-increasing genomic dataset sizes, fueled by continuous advances in DNA sequencing technologies, are expected to bring new scientific achievements in several fields of biology. The fact that the demand for higher sequencing throughput has long outpaced Moore's law, however, presents a challenge for the efficient analysis of future large-scale datasets, suggesting the urgent need for custom solutions to keep up with the current trend of increasing sample sizes. In this work, we focus on a widely employed, yet prohibitively compute- and memory-intensive, measure that is called linkage disequilibrium (LD), defined as the non-random association between alleles. Modern microprocessor architectures are not well equipped to deliver high performance for LD due to the lack of a vectorized population counter (counting set bits in registers). We present a modular and highly parallel reconfigurable architecture that, in combination with a generic memory layout transform, allows to rapidly conduct large-scale pairwise calculations on arbitrarily large one- and two-dimensional binary vectors, exhibiting increased bit-counting capacity. We map the proposed architecture to all four reconfigurable devices of a multi-FPGA platform, and deploy them synergistically for the evaluation of LD on genomic datasets with up to 1,000,000 sequences, achieving between 12.7X (4 FPGAs vs. 12 cores) and 134.9X (4 FPGAs vs. 1 core) faster execution than state-of-the-art reference software running on multi-core workstations. For real-world analyses that employ LD, such as scanning the 22nd human chromosome for traces of positive selection, the proposed system can lead to 6X faster processing, thus enabling more thorough genome-wide scans.
Dimitrios Bozikas, Nikolaos Alachiotis 0001, Pavlos Pavlidis, Euripides Sotiriades, Apostolos Dollas
FPL2
2016 High Performance Linkage Disequilibrium: FPGAs Hold the Key
abstract
DNA sequencing technologies allow the rapid sequencing of full genomes in a cost-effective way, leading to ever-growing genomic datasets that comprise thousands of genomes and millions of genetic variants. In population genomics and genome-wide association studies, widely used statistics such as linkage disequilibrium become computationally demanding when thousands of whole genomes are investigated. Long analysis times and excessive memory requirements usually prevent researchers from conducting exhaustive analyses, sacrificing the ability to detect distant genetic associations. In this work, we describe a generic algorithmic approach for organizing arbitrarily distant computations on full genomes, and to offload operations from the host processor to accelerators. We explore FPGAs as accelerators for linkage disequilibrium because the bulk of required operations are discrete, making them a good fit for reconfigurable fabric. We describe a versatile and trivially expandable architecture, and develop an automatic RTL generation software to search the design space. We find that, when thousands of genomes from complex species such as humans, are analyzed, current FPGAs can achieve up to 50X faster processing than state-of-the-art software running on multi-core workstations.
Nikolaos Alachiotis 0001, Gabriel Weisz
FPGA1
2015 Generating FPGA accelerators for chemical similarity assessment
abstract
Numerical measures of similarity/distance between objects represented by binary vectors are common in a wide range of disciplines. Searching in large-scale chemical databases requires billions of comparisons between molecules that are represented by binary fingerprints to capture the atomic structure. The performance bottleneck here is the enumeration of set bits in vectors (population count). Due to the discrete representation, similarity measures between binary fingerprints should fit well to FPGAs. We present an architecture to accelerate binary similarity assessment, evaluate various design points, and compare performance to highly optimized CPU and GPU implementations. We implement an RTL generation software, SimGenRTL, to generate accelerators of various sizes based on the proposed architecture. We find that accelerators with fewer and wider population counters allow better distribution of the hardware resources, outperforming significantly accelerators with more and narrower bit-enumeration components. SimGenRTL is available for download to allow rapid design space exploration of the computational core ahead of a full custom system implementation.
Nikolaos Alachiotis 0001
FPL1
2015 Enabling portable energy efficiency with memory accelerated library
abstract
Over the last decade, the looming power wall has spurred a flurry of interest in developing heterogeneous systems with hardware accelerators. The questions, then, are what and how accelerators should be designed, and what software support is required. Our accelerator design approach stems from the observation that many efficient and portable software implementations rely on high performance software libraries with well-established application programming interfaces (APIs). We propose the integration of hardware accelerators on 3D-stacked memory that explicitly targets the memory-bounded operations within high performance libraries. The fixed APIs with limited configurability simplify the design of the accelerators, while ensuring that the accelerators have wide applicability. With our software support that automatically converts library APIs to accelerator invocations, an additional advantage of our approach is that library-based legacy code automatically gains the benefit of memory-side accelerators without requiring a reimplementation. On average, the legacy code using our proposed MEmory Accelerated Library (MEALib) improves performance and energy efficiency for individual operations in Intel's Math Kernel Library (MKL) by 38x and 75x, respectively. For a real-world signal processing application that employs Intel MKL, MEALib attains more than 10x better energy efficiency.
Qi Guo 0001, Tze Meng Low, Nikolaos Alachiotis 0001, Berkin Akin, Lawrence T. Pileggi, James C. Hoe, Franz Franchetti
MICRO3
2013 libgapmis: extending short-read alignments
abstract
BACKGROUND: A wide variety of short-read alignment programmes have been published recently to tackle the problem of mapping millions of short reads to a reference genome, focusing on different aspects of the procedure such as time and memory efficiency, sensitivity, and accuracy. These tools allow for a small number of mismatches in the alignment; however, their ability to allow for gaps varies greatly, with many performing poorly or not allowing them at all. The seed-and-extend strategy is applied in most short-read alignment programmes. After aligning a substring of the reference sequence against the high-quality prefix of a short read--the seed--an important problem is to find the best possible alignment between a substring of the reference sequence succeeding and the remaining suffix of low quality of the read--extend. The fact that the reads are rather short and that the gap occurrence frequency observed in various studies is rather low suggest that aligning (parts of) those reads with a single gap is in fact desirable. RESULTS: In this article, we present libgapmis, a library for extending pairwise short-read alignments. Apart from the standard CPU version, it includes ultrafast SSE- and GPU-based implementations. libgapmis is based on an algorithm computing a modified version of the traditional dynamic-programming matrix for sequence alignment. Extensive experimental results demonstrate that the functions of the CPU version provided in this library accelerate the computations by a factor of 20 compared to other programmes. The analogous SSE- and GPU-based implementations accelerate the computations by a factor of 6 and 11, respectively, compared to the CPU version. The library also provides the user the flexibility to split the read into fragments, based on the observed gap occurrence frequency and the length of the read, thereby allowing for a variable, but bounded, number of gaps in the alignment. CONCLUSIONS: We present libgapmis, a library for extending pairwise short-read alignments. We show that libgapmis is better-suited and more efficient than existing algorithms for this task. The importance of our contribution is underlined by the fact that the provided functions may be seamlessly integrated into any short-read alignment pipeline. The open-source code of libgapmis is available at http://www.exelixis-lab.org/gapmis.
Nikolaos Alachiotis 0001, Simon A. Berger, Tomás Flouri, Solon P. Pissis, Alexandros Stamatakis
BMC Bioinform.1
2012 Exploiting Multi-grain Parallelism for Efficient Selective Sweep Detection
Nikolaos Alachiotis 0001, Pavlos Pavlidis, Alexandros Stamatakis
ICA3PP (1)1
2012 OmegaPlus: a scalable tool for rapid detection of selective sweeps in whole-genome datasets
abstract
UNLABELLED: Recent advances in sequencing technologies have led to the rapid accumulation of molecular sequence data. Analyzing whole-genome data (as obtained from next-generation sequencers) from intra-species samples allows to detect signatures of positive selection along the genome and therefore identify potentially advantageous genes in the course of the evolution of a population. We introduce OmegaPlus, an open-source tool for rapid detection of selective sweeps in whole-genome data based on linkage disequilibrium. The tool is up to two orders of magnitude faster than existing programs for this purpose and also exhibits up to two orders of magnitude smaller memory requirements. AVAILABILITY: OmegaPlus is available under GNU GPL at http://www.exelixis-lab.org/software.html.
Nikolaos Alachiotis 0001, Alexandros Stamatakis, Pavlos Pavlidis
Bioinform.1
2012 Coupling SIMD and SIMT architectures to boost performance of a phylogeny-aware alignment kernel
abstract
BACKGROUND: Aligning short DNA reads to a reference sequence alignment is a prerequisite for detecting their biological origin and analyzing them in a phylogenetic context. With the PaPaRa tool we introduced a dedicated dynamic programming algorithm for simultaneously aligning short reads to reference alignments and corresponding evolutionary reference trees. The algorithm aligns short reads to phylogenetic profiles that correspond to the branches of such a reference tree. The algorithm needs to perform an immense number of pairwise alignments. Therefore, we explore vector intrinsics and GPUs to accelerate the PaPaRa alignment kernel. RESULTS: We optimized and parallelized PaPaRa on CPUs and GPUs. Via SSE 4.1 SIMD (Single Instruction, Multiple Data) intrinsics for x86 SIMD architectures and multi-threading, we obtained a 9-fold acceleration on a single core as well as linear speedups with respect to the number of cores. The peak CPU performance amounts to 18.1 GCUPS (Giga Cell Updates per Second) using all four physical cores on an Intel i7 2600 CPU running at 3.4 GHz. The average CPU performance (averaged over all test runs) is 12.33 GCUPS. We also used OpenCL to execute PaPaRa on a GPU SIMT (Single Instruction, Multiple Threads) architecture. A NVIDIA GeForce 560 GPU delivered peak and average performance of 22.1 and 18.4 GCUPS respectively. Finally, we combined the SIMD and SIMT implementations into a hybrid CPU-GPU system that achieved an accumulated peak performance of 33.8 GCUPS. CONCLUSIONS: This accelerated version of PaPaRa (available at http://www.exelixis-lab.org/software.html) provides a significant performance improvement that allows for analyzing larger datasets in less time. We observe that state-of-the-art SIMD and SIMT architectures deliver comparable performance for this dynamic programming kernel when the "competing programmer approach" is deployed. Finally, we show that overall performance can be substantially increased by designing a hybrid CPU-GPU system with appropriate load distribution mechanisms.
Nikolaos Alachiotis 0001, Simon A. Berger, Alexandros Stamatakis
BMC Bioinform.1
2011 Accelerating Phylogeny-Aware Short DNA Read Alignment with FPGAs
abstract
Recent advances in molecular sequencing technology have given rise to novel algorithms for simultaneously aligning short sequence reads to reference sequence alignments and corresponding evolutionary reference trees. We present a complete hardware/software implementation for the acceleration of a program called PaPaRa, a newly introduced dynamic programming algorithm for this purpose. We verify the correctness of the proposed architecture on a real FPGA and introduce a straight-forward communication protocol(using gigabit ethernet) for seamless integration with the encapsulating steering software that is executed on a PC processor. The hardware description and the software implementation are freely available for download. When mapped to a Virtex 6 FPGA, our reconfigurable architecture can compute 133.4 billion cell updates per second for the novel, tree-based alignment kernel of PaPaRa. Compared to PaPaRa, running on a 3.2GHz Intel Core i5 CPU, we obtain speedups for the alignment kernel, that range between 170 and 471. For the entire application, that is, the alignment kernel and the trace-back step, we obtain speedups between 74 and 118.
Nikolaos Alachiotis 0001, Simon A. Berger, Alexandros Stamatakis
FCCM1
2011 FPGA Acceleration of the Phylogenetic Parsimony Kernel?
abstract
The phylogenetic parsimony function is a popular, discrete criterion for reconstructing evolutionary trees based on molecular sequence data. Parsimony strives to find the phylogenetic tree that explains the evolutionary history of organisms by the least number of mutations. Because parsimony is a discrete function, it should fit well to FPGAs. We present a versatile FPGA implementation of the parsimony function and compare its performance to a highly optimized SSE3- and AVX-vectorized software implementation. We find that, because of a particular constellation in our lab, the speedups that can be achieved by using an FPGA, are substantially less impressive, than usually reported in papers on FPGA acceleration of bioinformatics kernels. We conclude that, a competitive spirit between SW and HW application developers can contribute toward obtaining more objective performance comparisons.
Nikolaos Alachiotis 0001, Alexandros Stamatakis
FPL1
2010 Time and memory efficient likelihood-based tree searches on phylogenomic alignments with missing data
abstract
MOTIVATION: The current molecular data explosion poses new challenges for large-scale phylogenomic analyses that can comprise hundreds or even thousands of genes. A property that characterizes phylogenomic datasets is that they tend to be gappy, i.e. can contain taxa with (many and disparate) missing genes. In current phylogenomic analyses, this type of alignment gappyness that is induced by missing data frequently exceeds 90%. We present and implement a generally applicable mechanism that allows for reducing memory footprints of likelihood-based [maximum likelihood (ML) or Bayesian] phylogenomic analyses proportional to the amount of missing data in the alignment. We also introduce a set of algorithmic rules to efficiently conduct tree searches via subtree pruning and re-grafting moves using this mechanism. RESULTS: On a large phylogenomic DNA dataset with 2177 taxa, 68 genes and a gappyness of 90%, we achieve a memory footprint reduction from 9 GB down to 1 GB, a speedup for optimizing ML model parameters of 11, and accelerate the Subtree Pruning Regrafting tree search phase by factor 16. Thus, our approach can be deployed to improve efficiency for the two most important resources, CPU time and memory, by up to one order of magnitude. AVAILABILITY: Current open-source version of RAxML v7.2.6 available at http://wwwkramer.in.tum.de/exelixis/software.html.
Alexandros Stamatakis, Nikolaos Alachiotis 0001
Bioinform.2
2009 A reconfigurable architecture for the Phylogenetic Likelihood Function
abstract
As FPGA devices become larger, more coarse-grain modules coupled with large scale reconfigurable fabric become available, thus enabling new classes of applications to run efficiently, as compared to a general-purpose computer. This paper presents an architecture that benefits from the large number of DSP modules in Xilinx technology to implement massive floating point arithmetic. Our architecture computes the Phylogenetic Likelihood Function (PLF) which accounts for approximately 95% of total execution time in all state-of-the-art Maximum Likelihood (ML) based programs for reconstruction of evolutionary relationships. We validate and assess performance of our architecture against a highly optimized and parallelized software implementation of the PLF that is based on RAxML, which is considered to be one of the fastest and most accurate programs for phylogenetic inference. Both software and hardware implementations use double precision floating point arithmetic. The new architecture achieves speedups ranging from 1.6 up to 7.2 compared to a high-end 8-way dual-core general-purpose computer running the aforementioned highly optimized OpenMP-based multi-threaded version of the PLF.
Nikolaos Alachiotis 0001, Alexandros Stamatakis, Euripides Sotiriades, Apostolos Dollas
FPL1
2009 Exploring FPGAs for accelerating the phylogenetic likelihood function
abstract
Driven by novel biological wet lab techniques such as pyrosequencing there has been an unprecedented molecular data explosion. The growth of biological sequence data has significantly out-paced Moore's law. This development also poses new computational and architectural challenges for the field of phylogenetic inference, i.e., the reconstruction of evolutionary histories (trees) for a set of organisms which are represented by respective molecular sequences. Phylogenetic trees are currently increasingly reconstructed from multiple genes or even whole genomes. The introduced term "phylogenomics" reflects this development. Hence, there is an urgent need to deploy and develop new techniques and computational solutions to calculate the computationally intensive scoring functions for phylogenetic trees. In this paper, we propose a dedicated computer architecture to compute the phylogenetic maximum likelihood (ML) function. The ML criterion represents one of the most accurate statistical models for phylogenetic inference and accounts for 85% to 95% of total execution time in all state-of-the-art ML-based phylogenetic inference programs. We present the implementation of our architecture on an FPGA (field programmable gate array) and compare the performance to an efficient C implementation of the ML function on a high-end multi-core architecture with 16 cores. Our results are two-fold: (i) the initial exploratory implementation of the ML function for trees comprising 4 up to 512 sequences on an FPGA yields speedups of a factor 8.3 on average compared to execution on a single-core and is faster than the OpenMP-based parallel implementation on up to 16 cores in all but one case; and (ii) we are able to show that current FPGAs are capable to efficiently execute floating point intensive computational kernels.
Nikolaos Alachiotis 0001, Euripides Sotiriades, Apostolos Dollas, Alexandros Stamatakis
IPDPS1