EDBT 2026 Demo / reviewers in the wild / expert
Dimitris Theodoropoulos 0001
dblp:35/5033-1
· DBLP profile ↗
22ranked-venue papers
10as first author
4since 2021 · last 2025
0000-0002-0707-9415ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 9 first-author · 4 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Nyx: Virtualizing dataflow execution on shared FPGA platformsabstractAs FPGAs become more widespread for improving computing performance within cloud infrastructure, researchers aim to equip them with virtualization features to enable resource sharing in both temporal and spatial domains, thereby improving hardware utilization.Existing multi-tenant solutions focus on task-parallel models, where tasks are assigned to distinct regions to process separate sets of data.However, this model introduces waiting times between dependent and pipelined tasks, leading to longer response times for applications.The root cause is the lack of support for dataflow execution -a key potential of FPGAs and a crucial optimization for applications.Dataflow allows direct data streaming between operators, forming a task-pipelined model that reduces application latency by overlapping task operations within its workflow.This paper presents Nyx, the first system to enable dataflow execution in a task-based virtualized and shared FPGA environment.Nyx enables efficient resource sharing by dividing the FPGA into distinct reconfigurable regions.At its core, Nyx employs virtual FIFOs, independent channels that allow seamless communication between pipelined tasks.Its approach ensures smooth task operation even when the predecessor or successor tasks are not simultaneously scheduled in the FPGA, making them agnostic to their dependencies, communication channels or data locations.An FPGA hypervisor is designed to handle all data dependencies and efficiently dispatch pipelined tasks across regions at high throughput.Nyx outperforms existing state of the art virtualized task-parallel approaches by 1.26x -8.87x across a series of real-world benchmarks.Furthermore, it reduces response times by 2.8x -3.28x during low-demand periods, decreasing also deadline violations by up to 76.5%.Under highdemand conditions, Nyx delivers 2x -2.75x reduction, 34.5% fewer violations, and up to 1.9x reduced tail response time. Panagiotis Miliadis, Dimitris Theodoropoulos 0001, Nectarios Koziris, Dionisios N. Pnevmatikatos |
ISCA | 2 |
| 2024 | Architectural Support for Sharing, Isolating and Virtualizing FPGA ResourcesabstractFPGAs are increasingly popular in cloud environments for their ability to offer on-demand acceleration and improved compute efficiency. Providers would like to increase utilization, by multiplexing customers on a single device, similar to how processing cores and memory are shared. Nonetheless, multi-tenancy still faces major architectural limitations including: (a) inefficient sharing of memory interfaces across hardware tasks (HT) exacerbated by technological limitations and peculiarities, (b) insufficient solutions for performance and data isolation and high quality of service, and (c) absent or simplistic allocation strategies to effectively distribute external FPGA memory across HT. This article presents a full-stack solution for enabling multi-tenancy on FPGAs. Specifically, our work proposes an intra-fpga virtualization layer to share FPGA interfaces and its resources across tenants. To achieve efficient inter-connectivity between virtual FPGAs (vFGPAs) and external interfaces, we employ a compact network-on-chip architecture to optimize resource utilization. Dedicated memory management units implement the concept of virtual memory in FPGAs, providing mechanisms to isolate the address space and enable memory protection. We also introduce a memory segmentation scheme to effectively allocate FPGA address space and enhance isolation through hardware-software support, while preserving the efficacy of memory transactions. We assess our solution on an Alveo U250 Data Center FPGA Card, employing 10 real-world benchmarks from the Rodinia and Rosetta suites. Our framework preserves the performance of HT from a non-virtualized environment, while enhancing the device aggregate throughput through resource sharing; up to 3.96x in isolated and up to 2.31x in highly congested settings, where an external interface is shared across four vFPGAs. Finally, our work ensures high-quality of service, with HT achieving up to 0.95x of their native performance, even when resource sharing introduces interference from other accelerators. Panagiotis Miliadis, Dimitris Theodoropoulos 0001, Dionisios N. Pnevmatikatos, Nectarios Koziris |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | Early Results of Mapping Industrial Applications on Heterogeneous HPC Systems: The OPTIMA ProjectabstractThe OPTIMA project aims to port and optimize industrial applications and a set of open-source libraries into two novel FPGA-populated HPC systems. Target applications are from the domains of robotics simulation, underground analysis and computational fluid dynamics (CFD), where data processing is based on differential equations, matrix-matrix and matrix-vector operations. Moreover, the OPTIMA OPen Source (OOPS) library will support basic linear algebraic operations, sparse matrix-vector arithmetic, as well as computer-aided engineering (CAE) solvers. The OPTIMA target platforms are JUMAX, an HPC system that couples an AMD Epyc Server with Maxeler FPGA-based Dataflow Engines (DFEs), and server class machines with Alveo FPGA cards installed. Experimental results show that performance on robotic simulation can be enhanced up to 1.2x, and CFD calculations up to 4.7x. Finally, BLAS L1 routines are improved up to 7x, with a performance-per-Watt ratio boost of more than 40x compared to multi-threaded software routines from the Intel Math Kernel Library (MKL) suite when executed on an Intel Xeon server-class machine. Dimitris Theodoropoulos 0001, Giorgos Pekridis, Panagiotis Miliadis, Chloe Alverti, Panagiotis Mpakos, Dionisios N. Pnevmatikatos, Pavlos Malakonakis, Konstantinos Georgopoulos, Iakovos Mavroidis, Gino Perna, Marisa Zanotti, Giovanni Isotton, Max Engelen, Aggelos Ioannou, Ioannis Papaefstathiou, Albert Kahira, Andreas Herten |
CF | 1 |
| 2023 | Optimizing Industrial Applications for Heterogeneous HPC Systems: The OPTIMA Project Intermediate stageabstractOPTIMA is an SME-driven project (intermediate stage) that aims to port and optimize industrial applications and a set of open-source libraries into two novel FPGA-populated HPC systems. Target applications are from the domain of robotics simulation, underground analysis and computational fluid dy-namics (CFD), where data processing is based on differential equations, matrix-matrix and matrix-vector operations. Moreover, the OPTIMA OPen Source (OOPS) library will support basic linear algebraic operations, sparse matrix-vector arithmetic, as well as computer-aided engineering (CAE) solvers. The OPTIMA target platforms are JUMAX, an HPC system that couples an AMD Epyc Server with Maxeler FPGA-based Dataflow Engines (DFEs), and server-class machines with Alveo FPGA cards in-stalled. Experimental results on applications up to now, show that performance on robotic simulation can be enhanced up to 1.2x, CFD calculations up to 4.7x, and BLAS routines up to 7x compared to optimized software implementations from OpenBLAS. Dimitris Theodoropoulos 0001, Pavlos Malakonakis, Konstantinos Georgopoulos, Giovanni Isotton, Dionisios N. Pnevmatikatos, Ioannis Papaefstathiou, Gino Perna, Marisa Zanotti, Panagiotis Miliadis, Panagiotis Mpakos, Chloe Alverti, Aggelos Ioannou, Max Engelen, Albert Kahira, Iakovos Mavroidis |
DATE | 1 |
| 2019 | Evaluation of a Rack-Scale Disaggregated Memory Prototype for Cloud Data CentersabstractDisaggregated data centers propose a modular architecture where memory and compute resources are utilized in a finer granularity, aiming for a better optimized capacity and power consumption. In contrast to a classic compute node in a data center, memory is not required to be co-located with the processor, disaggregation can enhance hardware elasticity, improve virtual machine (VM) migration, and reduce the total cost of ownership (TCO), when compared to current data center solutions. Josue V. Quiroga, Martí Torrents, Nehir Sönmez, Dimitris Theodoropoulos 0001, Ferad Zyulkyarov, Mario Nemirovsky |
RSP | 4 |
| 2018 | REMAP: Remote mEmory Manager for disAggregated PlatformsabstractDisaggregated computing is a new approach that promises to alleviate the problem of fixed resource proportionality in datacenter deployments. Two critical factors that affect the overall performance of disaggregated platforms are remote memory access latency and throughput. Previous works primarily expose remote data processing at the applcation level that (a) require code annotations and/or the use of custom user-level libraries, and (b) may hinder the overall system protection and functionality. In this paper, we are taking a different approach: we propose the Remote mEmory Manager for dis-Aggregated Platforms (REMAP), a hardware architecture that enables the hotplug of remote memory resources to processing nodes, as normal paged memory at the OS-level, without requiring application-level code modifications. REMAP tightly couples processing nodes with remote memory controllers. Our architecture “expands” system memory on demand, by dynamically attaching remote memory modules to unused Local Physical Address (LPA) ranges, where the memory access requests are tunneled over high-speed, low-latency serial links. To evaluate REMAP in terms of performance, we implemented a prototype using two zcul02 FPGA boards. REMAP provides a remote cache-line access latency of less than 750 nsec, and up to 1.3× overall system throughput, compared to a baseline CPU-memory configuration. Dimitris Theodoropoulos 0001, Andrea Reale, Dimitris Syrivelis, Maciej Bielski, Nikolaos Alachiotis 0001, Dionisios N. Pnevmatikatos |
ASAP | 1 |
| 2018 | dReDBox: Materializing a full-stack rack-scale system prototype of a next-generation disaggregated datacenterabstractCurrent datacenters are based on server machines, whose mainboard and hardware components form the baseline, monolithic building block that the rest of the system software, middleware and application stack are built upon. This leads to the following limitations: (a) resource proportionality of a multi-tray system is bounded by the basic building block (mainboard), (b) resource allocation to processes or virtual machines (VMs) is bounded by the available resources within the boundary of the mainboard, leading to spare resource fragmentation and inefficiencies, and (c) upgrades must be applied to each and every server even when only a specific component needs to be upgraded. The dRedBox project (Disaggregated Recursive Datacentre-in-a-Box) addresses the above limitations, and proposes the next generation, low-power, across form-factor datacenters, departing from the paradigm of the mainboard-as-a-unit and enabling the creation of function-block-as-a-unit. Hardware-level disaggregation and software-defined wiring of resources is supported by a full-fledged Type-1 hypervisor that can execute commodity virtual machines, which communicate over a low-latency and high-throughput software-defined optical network. To evaluate its novel approach, dRedBox will demonstrate application execution in the domains of network functions virtualization, infrastructure analytics, and real-time video surveillance. Maciej Bielski, Ilias Syrigos, Kostas Katrinis, Dimitris Syrivelis, Andrea Reale, Dimitris Theodoropoulos 0001, Nikolaos Alachiotis 0001, Dionisios N. Pnevmatikatos, E. H. Pap, Georgios Zervas, Vaibhawa Mishra, Arsalan Saljoghei, Alvise Rigo, Jose Fernando Zazo, Sergio López-Buedo, Martí Torrents, Ferad Zyulkyarov, Michael Enrico, Óscar González de Dios |
DATE | 6 |
| 2018 | ReFiRe: Efficient Deployment of Remote Fine-Grained Reconfigurable AcceleratorsabstractThe need for specialized hardware acceleration in today's computing platforms is well established, due to power and efficiency reasons. Broadening an accelerator's scope of application is highly desirable, but requires a finer-grained architecture with basic primitives, which inevitably exhibits increased communication and synchronization requirements. In disaggregated-computing environ-ments, where data transfers between remote nodes are realized via datacenter-wide packet exchanges, reducing communication and synchronization is a prerequisite for the effective employment of remote acceleration. To this end, we present ReFiRe (Remote Fine-grained Reconfigurable acceleration), a generic deployment framework with native support for partial reconfiguration that allows to considerably reduce communication needs between a processor and remote accelerators. This is achieved by shifting control flow, partial reconfiguration, and execution decisions to the remote side through arbitrarily long instructions that encapsulate complex sequences of operations and their re-spective synchronization requirements. ReFiRe outperforms an SDSoC-generated accelerator system that employs the same accelerator cores to boost performance of a genomics application that detects positive selection. Emmanouil Pissadakis, Nikolaos Alachiotis 0001, Panagiotis Skrimponis, Dimitris Theodoropoulos 0001, Thanasis Korakis, Dionisios N. Pnevmatikatos |
FPT | 4 |
| 2017 | Multi-FPGA Evaluation Platform for Disaggregated ComputingabstractWe present a versatile FPGA-based evaluation platform for exploring alternative execution strategies on disaggregated environments for applications, considering different processing block types: compute cores, memory, and accelerators. Developers can interconnect different blocks types in order to create optimal configurations. A user-level software library allows quick mapping of applications on real hardware. We have implemented a fully working prototype using three ZC706 FPGA boards, and evaluated different software / hardware configurations of a matrix multiplication benchmark. Dimitris Theodoropoulos 0001, Nikolaos Alachiotis 0001, Dionisios N. Pnevmatikatos |
FCCM | 1 |
| 2017 | Versatile deployment of FPGA accelerators in disaggregated data centers: A bioinformatics case studyabstractImportant design considerations for the cost-effective employment of hardware accelerators in next-generation data centers involve a) the type of candidate applications that a proposed solution can accelerate (generality), and b) the required development effort to successfully deploy the available accelerators for a given application (adoption overhead). To address the problem of generality, we present a versatile and dynamically reconfigurable hardware architecture that exhibits several accelerator slots and programmable interconnect to create application-specific accelerator datapaths. The proposed architecture fits in the model of disaggregated data centers, where compute, memory, and accelerators are broadly regarded as large pools of resources, and subsets of these resource pools are dynamically allocated on an as-needed basis to cooperatively boost performance of a broad range of applications. Initial results for a bioinformatics application that we employ as a case study and deals with the detection of positive selection in large-scale genomic datasets reveal a speedup of up to 6.4X when custom hardware accelerators are mapped to the proposed versatile accelerator architecture and compared with a parallel and highly optimized software implementation executed on a multi-core processor. Nikolaos Alachiotis 0001, Dimitris Theodoropoulos 0001, Dionisios N. Pnevmatikatos |
FPL | 2 |
| 2016 | mCluster: A Software Framework for Portable Device-Based Volunteer ComputingabstractRecent market forecasts predict that the portable computing trend will vastly spread, as by 2020 there will bemore than 3 billion LTE device users worldwide. Motivated by this fact, many companies and research institutes have already launched research projects that utilize portable devices, voluntarily provided by users, to perform the required computations. Many such projects employ Berkeley's BOINC middleware, since it can support a large variety of stationary and mobile devices. However, currently available BOINC high-level APIs, either do not support portable devices or lack advanced processing capabilities (such as inter-node task dependencies) and/or easiness of use. To resolve these issues, we propose the mCluster software framework for application execution powered by the BOINC middleware on portable devices. mCluster adopts a task-based programming model that requires simple, pragma-based annotations of the application software, in order to dynamically resolve task dependencies. To evaluate our framework, we have have mapped a scientific application from the neuroscience domain on an small-scaled network of portable devices. mCluster significantly reduces the required programming effort and complexity to efficiently map BOINC-powered applications with task dependencies on portable devices compared to previous approaches. Dimitris Theodoropoulos 0001, Grigorios Chrysos 0001, Iosif Koidis, George Charitopoulos, Emmanouil Pissadakis, Antonis Varikos, Dionisios N. Pnevmatikatos, Georgios Smaragdos, Christos Strydis, Nikolaos A. Zervos |
CCGrid | 1 |
| 2016 | Rack-scale disaggregated cloud data centers: The dReDBox project vision
Kostas Katrinis, Dimitris Syrivelis, Dionisios N. Pnevmatikatos, Georgios Zervas, Dimitris Theodoropoulos 0001, Iordanis Koutsopoulos, K. Hasharoni, Daniel Raho, Christian Pinto, Felix Espina, Sergio López-Buedo, Qianqiao Chen, Mario Nemirovsky, Damian Roca, H. Klos, T. Berends |
DATE | 5 |
| 2016 | AXIOM: A Hardware-Software Platform for Cyber Physical SystemsabstractCyber-Physical Systems (CPSs) are widely necessary for many applications that require interactions with the humans and the physical environment. A CPS integrates a set of hardware-software components to distribute, execute and manage its operations. The AXIOM project (Agile, eXtensible, fast I/O Module) aims at developing a hardware-software platform for CPS such that i) it can use an easy parallel programming model and ii) it can easily scale-up the performance by adding multiple boards (e.g., 1 to 10 boards can run in parallel). AXIOM supports task-based programming model based on OmpSs and leverage a high-speed, inexpensive communication interface called AXIOM-Link. Another key aspect is that the board provides programmable logic (FPGA) to accelerate portions of an application. We are using smart video surveillance, and smart home living applications to drive our design. Somnath Mazumdar, Eduard Ayguadé, Nicola Bettin, Javier Bueno, Sara Ermini, Antonio Filgueras, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Francesco Montefoschi, David Oro, Dionisios N. Pnevmatikatos, Antonio Rizzo, Dimitris Theodoropoulos 0001, Roberto Giorgi |
DSD | 14 |
| 2015 | The AXIOM Software LayersabstractPeople and objects will soon share the same digital network for information exchange in a world named as the age of the cyber-physical systems. The general expectation is that people and systems will interact in real-time. This poses pressure onto systems design to support increasing demands on computational power, while keeping a low power envelop. Additionally, modular scaling and easy programmability are also important to ensure these systems to become widespread. The whole set of expectations impose scientific and technological challenges that need to be properly addressed. The AXIOM project (Agile, eXtensible, fast I/O Module) will research new hardware/software architectures for cyber-physical systems to meet such expectations. The technical approach aims at solving fundamental problems to enable easy programmability of heterogeneous multi-core multi-board systems. AXIOM proposes the use of the task-based OmpSs programming model, leveraging low-level communication interfaces provided by the hardware. Modular scalability will be possible thanks to a fast interconnect embedded into each module. To this aim, an innovative ARM and FPGA-based board will be designed, with enhanced capabilities for interfacing with the physical world. Its effectiveness will be demonstrated with key scenarios such as Smart Video-Surveillance and Smart Living/Home (domotics). Carlos Álvarez 0001, Eduard Ayguadé, Javier Bueno, Antonio Filgueras, Daniel Jiménez-González, Xavier Martorell, Nacho Navarro, Dimitris Theodoropoulos 0001, Dionisios N. Pnevmatikatos, Davide Catani, Claudio Scordino, Paolo Gai, Carlos Segura, Carles Fernández, David Oro, Javier Rodríguez Saeta, Pierluigi Passera, Alberto Pomella, Antonio Rizzo, Roberto Giorgi |
DSD | 8 |
| 2013 | Custom architecture for multicore audio beamforming systemsabstractThe audio Beamforming (BF) technique utilizes microphone arrays to extract acoustic sources recorded in a noisy environment. In this article, we propose a new approach for rapid development of multicore BF systems. Research on literature reveals that the majority of such experimental and commercial audio systems are based on desktop PCs, due to their high-level programming support and potential of rapid system development. However, these approaches introduce performance bottlenecks, excessive power consumption, and increased overall cost. Systems based on DSPs require very low power, but their performance is still limited. Custom hardware solutions alleviate the aforementioned drawbacks, however, designers primarily focus on performance optimization without providing a high-level interface for system control and test. In order to address the aforementioned problems, we propose a custom platform-independent architecture for reconfigurable audio BF systems. To evaluate our proposal, we implement our architecture as a heterogeneous multicore reconfigurable processor and map it onto FPGAs. Our approach combines the software flexibility of General-Purpose Processors (GPPs) with the computational power of multicore platforms. In order to evaluate our system we compare it against a BF software application implemented to a low-power Atom 330, a middle-ranged Core2 Duo, and a high-end Core i3. Experimental results suggest that our proposed solution can extract up to 16 audio sources in real time under a 16-microphone setup. In contrast, under the same setup, the Atom 330 cannot extract any audio sources in real time, while the Core2 Duo and the Core i3 can process in real time only up to 4 and 6 sources respectively. Furthermore, a Virtex4-based BF system consumes more than an order less energy compared to the aforementioned GPP-based approaches. Dimitris Theodoropoulos 0001, Georgi Kuzmanov, Georgi Gaydadjiev |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2011 | Reconfigurable acceleration and dynamic partial self-reconfiguration in general purpose computingabstractIn this paper, we describe a generic approach for integrating a dynamically reconfigurable device into a general purpose system interconnected with a high-speed link. The system can dynamically install and execute hardware instances of functions to accelerate parts of a given software code. The hardware descriptions of the functions (bitstreams) are inserted into the executable binary running on the system. Our compiler further inserts system-calls to the software code to control the reconfigurable device. Thereby, the general purpose host-processor of the system manages the hardware reconfiguration and execution through a Linux device driver. The device has direct access to the main memory (DMA) operating in the virtual address space; it further supports memory mapped IO for data and control, and is able to raise and handle interrupts for synchronization. The above system is implemented on a general purpose machine providing a HyperTransport bus to connect a Xilinx Virtex4-100 FPGA, an AMD Opteron-244, and 1 GB of DDR main memory. We evaluate our proposal using a secure audio processing application. We accelerate in hardware the Audio processing kernel as well as the subsequent AES encryption function via dynamic partial self-reconfiguration. The proposed system achieves a 12× speedup over a software for the application at hand. Ioannis Sourdis, Abhijit Nandy, Venkatasubramanian Viswanathan, Anthony Brandon, Dimitris Theodoropoulos 0001, Georgi Gaydadjiev |
FPT | 5 |
| 2011 | Multi-Core Platforms for Beamforming and Wave Field SynthesisabstractImmersive-Audio technologies are widely used to build experimental and commercial audio systems. However, most of them are based on standard PCs, which introduce performance limitations and excessive power consumption. To address these drawbacks, we explore the implementation prospectives of two Immersive-Audio technologies: the beamforming (BF) and the wave field synthesis (WFS). We target two popular multi-core platforms, namely graphic processor units (GPUs) and field programmable gate arrays (FPGAs). We identify the most computationally intensive parts of both applications and employ the CUDA environment to map them onto a Quadro FX1700, a GeForce 8600GT, a GTX275, and a GTX460 GPU. Furthermore, we design our custom multi-core hardware accelerators for both algorithms and map them onto Virtex6 FPGAs. Both GPU and FPGA implementations are compared against OpenMP-annotated software running on a Core2 Duo at 3.0 GHz. Experimental results suggest that middle-range GPUs process data equally well as the Core2 Duo for the BF, and approximately two times faster for the WFS. However, high-end GPU and FPGA solutions provide an order of magnitude better performance for BF, and approximately two orders of magnitude better performance for WFS than the Core2 Duo. Ultimately, single-chip GPU and FPGA implementations can provide more power-effective solutions, since they can drive more complex microphone and loudspeaker setups than PC-based approaches. Dimitris Theodoropoulos 0001, Georgi Kuzmanov, Georgi Gaydadjiev |
IEEE Trans. Multim. | 1 |
| 2010 | A 3d-audio reconfigurable processorabstractVarious multimedia communication systems based on 3D-Audio algorithms have been proposed by researchers from the acoustic data processing domain. However, all systems reported in the literature follow a PC-based approach that introduces processing bottlenecks and excessive power consumption. In order to alleviate these problems, we propose a reconfigurable 3D-Audio processor that can record and render sound sources concurrently. Audio recording and rendering are performed by two hardware accelerators exploiting the beamforming and the Wave Field Synthesis algorithms. The theoretical scalability of the proposed processor is explored with respect to systems consisting of different microphone and loudspeaker arrays configurations. A working FPGA prototype is compared against a software implementation on a Core2 Duo system. Results suggest that the proposed reconfigurable hardware solution can process data up to 2.4x faster than the software approach, while power consumption is approximately 7 Watts according to the Xilinx XPower report. Dimitris Theodoropoulos 0001, Georgi Kuzmanov, Georgi Gaydadjiev |
FPGA | 1 |
| 2010 | A novel HDL coding style to reduce power consumption for reconfigurable devicesabstractPower consumption has become the major factor that has to be considered while designing systems using reconfigurable devices, especially for battery-operated applications. Minimizing transitions is one of the ways to reduce power consumption. Overwriting a register with the same value occurs frequently in real digital systems. Such unneeded transitions increase the power consumption. To avoid this, a new HDL coding style to reduce power consumption for reconfigurable devices is proposed. The idea is to “force” the CAD tool to configure the CLB flip-flop as a T flip-flop with its T input held constantly at logic one and drive its clock through the lookup table(LUT). Based on an extensive evaluation using MCNC benchmark circuits on a real FPGA and a real CAD tool, our proposal reduces total power consumption by 13-90 % and runs 2-20 % faster with 0-45 % area overhead compared to conventional coding style solutions. As a parallel activity we proposed a new logic element (LE) that implements the proposed design style directly. Thomas Marconi, Dimitris Theodoropoulos 0001, Koen Bertels, Georgi Gaydadjiev |
FPT | 2 |
| 2010 | Minimalistic architecture for reconfigurable audio BeamformingabstractIn this paper, we propose a minimal programming model that is tailored to audio Beamforming applications. The model consists of nine instructions that provide high flexibility to customize multi-core reconfigurable beamformers. We describe all instructions and demonstrate their functionality through pseudocode examples. We apply the proposed programming paradigm to a multi-core reconfigurable Beamforming architecture. Our approach combines software programming flexibility with improved hardware performance. Experimental results suggest that our Virtex4FX60-based solution at 100 MHz, can extract in real-time up to 12 acoustic sources 2.6x faster than a 3.0 GHz Core2 Duo OpenMP-based implementation. Dimitris Theodoropoulos 0001, Georgi Kuzmanov, Georgi Gaydadjiev |
FPT | 1 |
| 2009 | Algorithms for the automatic extension of an instruction-setabstractIn this paper, two general algorithms for the automatic generation of instruction-set extensions are presented. The basic instruction set of a reconfigurable architecture is specialized with new application-specific instructions. The paper proposes two methods for the generation of convex multiple input multiple output instructions, under hardware resource constraints, based on a two-step clustering process. Initially, the application is partitioned in single-output instructions of variable size and then, selected clusters are combined in convex multiple output clusters following different policies. Our results on well-known kernels show that the extended instructions-set allows to execute applications more efficiently and needing fewer cycles. Our results show that a significant overall application speed-up is achieved even for large kernels (for ADPCM decoder the speed-up is up to x2.2 and for TWOFISH encoder the speedup is up to x5.5). Carlo Galuzzi, Dimitris Theodoropoulos 0001, Roel Meeuws, Koen Bertels |
DATE | 2 |
| 2009 | Reconfigurable accelerator for WFS-based 3D-audioabstractIn this paper, we propose a reconfigurable and scalable hardware accelerator for 3D-audio systems based on the Wave Field Synthesis technology. Previous related work reveals that WFS sound systems are based on using standard PCs. However, two major obstacles are the relative low number of real-time sound sources that can be processed and the high power consumption. The proposed accelerator alleviates these limitations by its performance and energy efficient design. We propose a scalable organization comprising multiple rendering units (RUs), each of them independently processing audio samples. The processing is done in an environment of continuously varying number of sources and speakers. We provide a comprehensive study on the design trade-offs with respect to this multiplicity of sources and speakers. A hardware prototype of our proposal was implemented on a Virtex4FX60 FPGA operating at 200 MHz. A single RU can achieve up to 7× WFS processing speedup compared to a software implementation running on a Pentium D at 3.4 GHz, while consuming, according to Xilinx XPower, approximately 3 W of power only. Dimitris Theodoropoulos 0001, Georgi Kuzmanov, Georgi Gaydadjiev |
IPDPS | 1 |