EDBT 2026 Demo / reviewers in the wild / expert
Pavlos Malakonakis
dblp:42/9860
· DBLP profile ↗
11ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0002-7265-7381ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | REBECCA: Reconfigurable Heterogeneous Highly Parallel Processing Platform for Safe and Secure AI
Andreas Brokalakis, Iakovos Mavroidis, Konstantinos Georgopoulos, Pavlos Malakonakis, Konstantinos Harteros, Dimitris Andronikou, Yannis Galanomatis, Charalampos Savvakos, Grigorios Chrysos 0001, Sotiris Ioannidis, Ioannis Papaefstathiou |
DSD | 4 |
| 2023 | Early Results of Mapping Industrial Applications on Heterogeneous HPC Systems: The OPTIMA ProjectabstractThe OPTIMA project aims to port and optimize industrial applications and a set of open-source libraries into two novel FPGA-populated HPC systems. Target applications are from the domains of robotics simulation, underground analysis and computational fluid dynamics (CFD), where data processing is based on differential equations, matrix-matrix and matrix-vector operations. Moreover, the OPTIMA OPen Source (OOPS) library will support basic linear algebraic operations, sparse matrix-vector arithmetic, as well as computer-aided engineering (CAE) solvers. The OPTIMA target platforms are JUMAX, an HPC system that couples an AMD Epyc Server with Maxeler FPGA-based Dataflow Engines (DFEs), and server class machines with Alveo FPGA cards installed. Experimental results show that performance on robotic simulation can be enhanced up to 1.2x, and CFD calculations up to 4.7x. Finally, BLAS L1 routines are improved up to 7x, with a performance-per-Watt ratio boost of more than 40x compared to multi-threaded software routines from the Intel Math Kernel Library (MKL) suite when executed on an Intel Xeon server-class machine. Dimitris Theodoropoulos 0001, Giorgos Pekridis, Panagiotis Miliadis, Chloe Alverti, Panagiotis Mpakos, Dionisios N. Pnevmatikatos, Pavlos Malakonakis, Konstantinos Georgopoulos, Iakovos Mavroidis, Gino Perna, Marisa Zanotti, Giovanni Isotton, Max Engelen, Aggelos Ioannou, Ioannis Papaefstathiou, Albert Kahira, Andreas Herten |
CF | 7 |
| 2023 | Optimizing Industrial Applications for Heterogeneous HPC Systems: The OPTIMA Project Intermediate stageabstractOPTIMA is an SME-driven project (intermediate stage) that aims to port and optimize industrial applications and a set of open-source libraries into two novel FPGA-populated HPC systems. Target applications are from the domain of robotics simulation, underground analysis and computational fluid dy-namics (CFD), where data processing is based on differential equations, matrix-matrix and matrix-vector operations. Moreover, the OPTIMA OPen Source (OOPS) library will support basic linear algebraic operations, sparse matrix-vector arithmetic, as well as computer-aided engineering (CAE) solvers. The OPTIMA target platforms are JUMAX, an HPC system that couples an AMD Epyc Server with Maxeler FPGA-based Dataflow Engines (DFEs), and server-class machines with Alveo FPGA cards in-stalled. Experimental results on applications up to now, show that performance on robotic simulation can be enhanced up to 1.2x, CFD calculations up to 4.7x, and BLAS routines up to 7x compared to optimized software implementations from OpenBLAS. Dimitris Theodoropoulos 0001, Pavlos Malakonakis, Konstantinos Georgopoulos, Giovanni Isotton, Dionisios N. Pnevmatikatos, Ioannis Papaefstathiou, Gino Perna, Marisa Zanotti, Panagiotis Miliadis, Panagiotis Mpakos, Chloe Alverti, Aggelos Ioannou, Max Engelen, Albert Kahira, Iakovos Mavroidis |
DATE | 3 |
| 2021 | An FPGA-Based Data Pre-Processing Architecture to Accelerate De-Novo Genome AssemblyabstractGenome assembly is a field of bioinformatics which refers to the process of taking small fragments of genetic material and putting them back together in order to reconstruct the original DNA sequence from which the fragments originated. As the DNA genome assembly input datasets in most cases have a very large amount of data, it is important to develop custom architectures in order to speed up these processes and gain significant execution time reduction. In this paper we present the Reads Matching Filter (RMF), an input dataset prefiltering process, based on string matching and implemented on Field Programmable Gate Array (FPGA) technology, in order to reduce the genome assembly execution time. The outputs of the RMF running on the FPGA as well as the original input dataset are given as input to the Velvet genome assembler which produces the assembly of the input sequences. The Velvet genome assembler is based on the manipulation of de Bruijn graphs, and produces its output via the removal of errors and the simplication of repeated regions. The FPGA-based RMF pre-filtering process manages to speedup the entire genome assembly processing, including I/O, by up to 6 times, while maintaining the quality of the output sequence contigs (i.e. the series of overlapping DNA sequences). Georgios Galanos, Pavlos Malakonakis, Apostolos Dollas |
BIBE | 2 |
| 2021 | Novel Reconfigurable Hardware Systems for Tumor Growth PredictionabstractAn emerging trend in biomedical systems research is the development of models that take full advantage of the increasing available computational power to manage and analyze new biological data as well as to model complex biological processes. Such biomedical models require significant computational resources, since they process and analyze large amounts of data, such as medical image sequences. We present a family of advanced computational models for the prediction of the spatio-temporal evolution of glioma and their novel implementation in state-of-the-art FPGA devices. Glioma is a rapidly evolving type of brain cancer, well known for its aggressive and diffusive behavior. The developed system simulates the glioma tumor growth in the brain tissue, which consists of different anatomic structures, by utilizing MRI slices. The presented models have been proved highly accurate in predicting the growth of the tumor, whereas the developed innovative hardware system, when implemented on a low-end, low-cost FPGA, is up to 85% faster than a high-end server consisting of 20 physical cores (and 40 virtual ones) and more than 28× more energy-efficient than it; the energy efficiency grows up to 50× and the speedup up to 14× if the presented designs are implemented in a high-end FPGA. Moreover, the proposed reconfigurable system, when implemented in a large FPGA, is significantly faster than a high-end GPU (i.e., from 80% and up to 250% faster), for the majority of the models, while it is also significantly better (i.e., from 80% to over 1,600%) in terms of power efficiency, for all the implemented models. Konstantinos Malavazos, Maria Papadogiorgaki, Pavlos Malakonakis, Ioannis Papaefstathiou |
ACM Trans. Comput. Heal. | 3 |
| 2020 | Exploring Modern FPGA Platforms for Faster Phylogeny Reconstruction with RAxMLabstractThe Phylogenetic Likelihood Function (PLF) is one of the cornerstone functions in most phylogenetic inference tools; its execution represents the majority of time required to complete an analysis. This work proposes the acceleration of this function using reconfigurable hardware accelerators, focusing on system-on-chips that integrate Field Programmable Gate Array (FPGA) resources as well as traditional High Performance Computing (HPC) systems that use FPGA-based accelerator cards. Taking into account the specific properties of each platform in order to exploit their processing capabilities, the proposed solutions provide significant performance gains. The measured acceleration of PLF function is up to 8x while the overall time to complete a phylogenetic analysis using the popular RAxML software can be reduced up to 3.2 times (with respect to a pure software implementation on a high-end server processor). Compared to other similar solutions proposed in literature, our systems perform up to 65% faster. Pavlos Malakonakis, Andreas Brokalakis, Nikolaos Alachiotis 0001, Euripides Sotiriades, Apostolos Dollas |
BIBE | 1 |
| 2020 | A novel FPGA-based system for Tumor Growth PredictionabstractAn emerging trend in the biomedical community is to create models that take advantage of the increasing available computational power, in order to manage and analyze new biological data as well as to model complex biological processes. Such biomedical software applications require significant computational resources since they process and analyze large amounts of data, such as medical image sequences. This paper presents a novel FPGA-based system that implements a novel model for the prediction of the spatio-temporal evolution of glioma. Glioma is a rapidly evolving type of brain cancer, well known for its aggressive and diffusive behavior. The developed system simulates the glioma tumor growth in the brain tissue, which consists of different anatomic structures, by utilizing individual MRI slices. The presented innovative hardware system is more than 60% faster than a high-end server consisting of 20 physical cores (and 40 virtual ones) and more than 28x more energy efficient. Konstantinos Malavazos, Maria Papadogiorgaki, Pavlos Malakonakis, Ioannis Papaefstathiou |
DATE | 3 |
| 2020 | UNILOGIC: A Novel Architecture for Highly Parallel Reconfigurable SystemsabstractOne of the main characteristics of High-performance Computing (HPC) applications is that they become increasingly performance and power demanding, pushing HPC systems to their limits. Existing HPC systems have not yet reached exascale performance mainly due to power limitations. Extrapolating from today’s top HPC systems, about 100–200 MWatts would be required to sustain an exaflop-level of performance. A promising solution for tackling power limitations is the deployment of energy-efficient reconfigurable resources (in the form of Field-programmable Gate Arrays (FPGAs)) tightly integrated with conventional CPUs. However, current FPGA tools and programming environments are optimized for accelerating a single application or even task on a single FPGA device. In this work, we present UNILOGIC (Unified Logic), a novel HPC-tailored parallel architecture that efficiently incorporates FPGAs. UNILOGIC adopts the Partitioned Global Address Space (PGAS) model and extends it to include hardware accelerators, i.e., tasks implemented on the reconfigurable resources. The main advantages of UNILOGIC are that (i) the hardware accelerators can be accessed directly by any processor in the system, and (ii) the hardware accelerators can access any memory location in the system. In this way, the proposed architecture offers a unified environment where all the reconfigurable resources can be seamlessly used by any processor/operating system. The UNILOGIC architecture also provides hardware virtualization of the reconfigurable logic so that the hardware accelerators can be shared among multiple applications or tasks. The FPGA layer of the architecture is implemented by splitting its reconfigurable resources into (i) a static partition, which provides the PGAS-related communication infrastructure, and (ii) fixed-size and dynamically reconfigurable slots that can be programmed and accessed independently or combined together to support both fine and coarse grain reconfiguration. 1 Finally, the UNILOGIC architecture has been evaluated on a custom prototype that consists of two 1U chassis, each of which includes eight interconnected daughter boards, called Quad-FPGA Daughter Boards (QFDBs); each QFDB supports four tightly coupled Xilinx Zynq Ultrascale+ MPSoCs as well as 64 Gigabytes of DDR4 memory, and thus, the prototype features a total of 64 Zynq MPSoCs and 1 Terabyte of memory. We tuned and evaluated the UNILOGIC prototype using both low-level (baremetal) performance tests, as well as two popular real-world HPC applications, one compute-intensive and one data-intensive. Our evaluation shows that UNILOGIC offers impressive performance that ranges from being 2.5 to 400 times faster and 46 to 300 times more energy efficient compared to conventional parallel systems utilizing only high-end CPUs, while it also outperforms GPUs by a factor ranging from 3 to 6 times in terms of time to solution, and from 10 to 20 times in terms of energy to solution. Aggelos Ioannou, Konstantinos Georgopoulos, Pavlos Malakonakis, Dionisios N. Pnevmatikatos, Vassilis Papaefstathiou, Ioannis Papaefstathiou, Iakovos Mavroidis |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2011 | Exploitation of Parallel Search Space Evaluation with FPGAs in Combinatorial Problems: The Eternity II CaseabstractThe Eternity II puzzle is a combinatorial search problem which qualifies as a computational grand challenge. As no known closed form solution exists, its solution is based on exhaustive search, making it an excellent candidate for FPGA-based architectures, in which complex data structures and non-trivial recursion are implemented in hardware. This paper presents such an architecture, which was designed and fully implemented on a Virtex5 FPGA (XUP ML505 board). Despite the serial nature of the recursion, as parallelism can be applied with the initiation of multiple searches, the system shows a measured speedup of 2.6 vs. a high-end multi-core compute server. Pavlos Malakonakis, Apostolos Dollas |
FPL | 1 |
| 2010 | GE3: A single FPGA client-server architecture for Golomb Ruler derivationabstractOptimal Golomb Rulers (OGR) are a discrete mathematics problem for which there is no known closed form solution. This problem is so computationally intensive that it is considered a “grand challenge problem”. Since the early 1990's FPGA-based OGR engines have been designed, with excellent performance vs. general-purpose computing. This paper presents a new, single FPGA clientserver architecture for OGR derivation. The new client architecture supports parallel evaluation of multiple hypotheses (up to 16), each implemented as a shift operation, and one server which can support many clients. The new architecture has a measured speedup of 8 against an Intel Core 2 Duo processor for a single client supporting up to eight “shifts” and running on a Virtex 2P FPGA, and a post place-and-route simulation-derived speedup of 160 with four clients on a Virtex 5 FPGA, each supporting up to sixteen “shifts”. The new architecture has been fully implemented and runs on actual hardware, whereas simulations have been used to project performance on FPGA's which were not available for experimentation. Pavlos Malakonakis, Euripides Sotiriades, Apostolos Dollas |
FPT | 1 |
| 2010 | CarlOthello : An FPGA-Based Monte Carlo Othello playerabstractIn the FPT 2010 International Conference an Othello competition has been announced, based on the popular game and with requirements for implementation of full designs on standardized FPGA platforms. This paper presents in detail the CarlOthello architecture and design, which is heavily pipelined in order to increase the expansion rate of the overall system, reaching a peak of 4×108expansions per second. The Monte Carlo-based Othello player was fully designed, implemented in hardware, and tested. This design wins every game against the FPT 2010 reference software. Miltiadis Smerdis, Pavlos Malakonakis, Apostolos Dollas |
FPT | 2 |