EDBT 2026 Demo / reviewers in the wild / expert
George Lentaris
dblp:14/326
· DBLP profile ↗
26ranked-venue papers
5as first author
11since 2021 · last 2027
0000-0003-1664-8648ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | CECAIServe: Facilitating ML inference serving across the cloud-edge-continuumabstractMachine learning (ML) inference-serving has become a core operational component of MLOps, particularly as ML services are increasingly deployed across the cloud–edge continuum. However, existing inference-serving tools typically provide limited support for heterogeneous hardware, weak energy observability, and cumbersome integration of device-specific accelerated inference frameworks and model-specific pre-/post-processing. In this work, we present CECAIServe, an open-source, unified, and vendor-neutral framework that automatically generates deployment-ready Accelerated Inference Serving Containers (AISCs) from high-level TensorFlow and PyTorch models. Through its modular design, CECAIServe abstracts device- and framework-specific complexity while supporting CPUs, GPUs, edge accelerators, and FPGA-based systems. The framework further incorporates MLOps-oriented functionality, including fine-grained latency instrumentation, integrated power monitoring, and an interface for pre-/post-processing. Our comprehensive evaluation on 12 models demonstrates that CECAIServe can automatically generate AISCs across 7 diverse devices in less than 10 min. Additionally, we demonstrate that CECAIServe facilitates effective benchmarking and design space exploration on HW-accelerated inference-serving on devices across the cloud–edge-continuum. Finally, we validate the extensibility of the framework by adapting it to support Large Language Model inference-serving. Aimilios Leftheriotis, Achilleas Tzenetopoulos, George Lentaris, Dimitrios Soudris, George Theodoridis |
Future Gener. Comput. Syst. | 3 |
| 2024 | Seamless HW-accelerated AI serving in heterogeneous MEC Systems with AI@EDGEabstractThe advancement towards B5G/6G relies on the synthesis of connect-compute platforms and their use in highly heterogeneous clusters featuring hardware accelerators. While these accelerators offer improved computational efficiency, sill, they make development, deployment, and orchestration of services more complex, with limited flexibility, and necessitate domain-specific knowledge. In AI@EDGE we are targeting seamless integration of such diverse platforms for executing AI-related tasks. This paper focuses on acceleration aspects and presents a MEC system that facilitates AI servicing over a cluster of FPGA, GPU, and CPU nodes. To this end, we develop our custom tools for generating multi-variant AI models, informative function descriptors, flexible MEC orchestrators, and runtime resource managers. The results show successful interoperability, with generic Python models getting deployed/migrated across distinct platforms for performance gains in the area of 10x. Achilleas Tzenetopoulos, George Lentaris, Aimilios Leftheriotis, Panos Chrysomeris, Javier Palomares, Estefanía Coronado, Raman Kazhamiakin, Dimitrios Soudris |
HPDC | 2 |
| 2024 | Towards AI Onboard EO Satellites: Assessment of Virtualization Techniques for Extreme Edge ComputingabstractExtreme-edge computers coupled with sophisticated remote sensors are becoming a pivotal component for Earth Observation (EO) satellites. Such components benefit from the employment of virtualization techniques to enable the dynamic deployment of novel on-board Artificial Intelligence (AI) and to concurrently support multiple users on the same platform. This paper presents a comparative analysis of state-of-the-art virtualization techniques (namely Unikernels, Virtual Machines and Containers) specifically within the realm of AI-driven EO. Focused on five key criteria - security, scalability, resource efficiency, performance, and multitenancy - the study synthesizes existing literature to elucidate the strengths and limitations of each virtualization technology, and augments this understanding through a hands-on evaluation. Unikernels are distinguished for their minimalistic design and high efficiency, virtual machines for their robust isolation and stability, and containers for their flexibility. The comparative framework aims to guide engineers in selecting the most suitable virtualization technology according to the needs of their use case. This analysis clarifies the current state of virtualization technologies and provides a nuanced understanding of their applicability in advancing the capabilities of AI-driven EO systems. Antonis Karteris, Evgenios Tsigkanos, Mathieu Bernou, Alexis Chatzistylianos, George Lentaris |
IGARSS | 5 |
| 2024 | Evaluation of Resource-Efficient Crater Detectors on Embedded SystemsabstractReal-time analysis of Martian craters is crucial for mission-critical operations, including safe landings and geological exploration. This work leverages the latest breakthroughs for on-the-edge crater detection aboard spacecraft. We rigorously benchmark several YOLO networks using a Mars craters dataset, analyzing their performance on embedded systems with a focus on optimization for low-power devices. We optimize this process for a new wave of cost-effective, commercial-off-the-shelf-based smaller satellites. Implementations on diverse platforms, including Google Coral Edge TPU, AMD Versal SoC VCK190, Nvidia Jetson Nano and Jetson AGX Orin, undergo a detailed trade-off analysis. Our findings identify optimal network-device pairings, enhancing the feasibility of crater detection on resource-constrained hardware and setting a new precedent for efficient and resilient extraterrestrial imaging. Code at: https://github.com/billpsomas/mars_crater_detection. Simon Vellas, Bill Psomas, Kalliopi Karadima, Dimitrios Danopoulos, Alexandros Paterakis, George Lentaris, Dimitrios Soudris, Konstantinos Karantzalos |
IGARSS | 6 |
| 2022 | Improving the performance of RISC-V softcores on FPGA by exploiting PVT variability and DVFSabstractImproving RISC-V processors becomes important in a plethora of applications, many of which rely exclusively on FPGA fabric to achieve custom HW/SW co-processing. Our approach is to improve softcore implementations by also accounting for the PVT peculiarities of each underlying FPGA. The proposed method bases on a custom DVFS technique to overcome PVT-induced guardbands, in-the-field. We evaluate the potential gains of an example RISC-V HDL core on Zynq MPSoC while varying multiple parameters, i.e., Voltage, Frequency, SW benchmarks, and RISC-V configurations. Our exploration indicates up to 75–149% throughput increase and/or 40% power decrease, vs STA, along with a need for careful tuning of RISC-V memory size. Endri Taka, George Lentaris, Dimitrios Soudris |
ISCAS | 2 |
| 2022 | Towards Employing FPGA and ASIP Acceleration to Enable Onboard AI/ML in Space ApplicationsabstractThe success of AI/ML in terrestrial applications and the commercialization of space are now paving the way for the advent of AI/ML in satellites. However, the limited processing power of classical onboard processors drives the community towards extending the use of FPGAs in space with both rad-hard and Commercial-Off-The-Shelf devices. The increased performance of FPGAs can be complemented with VPU or TPU ASIP coprocessors to further facilitate high-level AI development and inflight reconfiguration. Thus, selecting the most suitable devices and designing the most efficient avionics architecture becomes crucial for the success of novel space missions. The current work presents industrial trends, comparative studies with inhouse benchmarking, as well as architectural designs utilizing FPGAs and AI accelerators towards enabling AI/ML in future space missions. Vasileios Leon, George Lentaris, Dimitrios Soudris, Simon Vellas, Mathieu Bernou |
VLSI-SoC | 2 |
| 2022 | Combining Fault Tolerance Techniques and COTS SoC Accelerators for Payload Processing in SpaceabstractThe ever-increasing demand for computational power and I/O throughput in space applications is transforming the landscape of on-board computing. A variety of Commercial-Off-The-Shelf (COTS) accelerators emerges as an attractive solution for payload processing to outperform the traditional radiation-hardened devices. Towards increasing the reliability of such COTS accelerators, the current paper explores and evaluates fault-tolerance techniques for the Zynq FPGA and the Myriad VPU, which are two device families being integrated in industrial space avionics architectures/boards, such as Ubotica’s CogniSat, Xiphos’ Q7S, and Cobham Gaisler’s GR-VPX-XCKU060. On the FPGA side, we combine techniques such as memory scrubbing, partial reconfiguration, triple modular redundancy, and watch-dogs. On the VPU side, we detect and correct errors in the instruction and data memories, as well as we apply redundancy at processor level (SHAVE cores). When considering FPGA with VPU co-processing, we also develop a fault-tolerant interface between the two devices based on the CIF/LCD protocols and our custom CRC error-detecting code. Vasileios Leon, Elissaios-Alexios Papatheofanous, George Lentaris, Charalampos Bezaitis, Nikolaos Mastorakis, Georgios Bampilis, Dionysios I. Reisis, Dimitrios Soudris |
VLSI-SoC | 3 |
| 2021 | ParalOS: A Scheduling & Memory Management Framework for Heterogeneous VPUsabstractEmbedded systems are presented today with the challenge of a very rapidly evolving application diversity followed by increased programming and computational complexity. Customised heterogeneous System-on-Chip (SoC) processors emerge as an attractive HW solution in various application domains, however, they still require sophisticated SW development to provide efficient implementations at the expense of slower adaptation to algorithmic changes. In this context, the current paper proposes a framework for accelerating the SW development of computationally intensive applications on Vision Processing Units (VPUs), while still enabling the exploitation of their full HW potential via low-level kernel optimisations. Our framework is tailored for heterogeneous architectures and integrates a dynamic task scheduler, a novel scratchpad memory management scheme, I/O & inter-process communication techniques, as well as a visual profiler. We evaluate our work on the Intel Movidius Myriad VPUs using synthetic benchmarks and real-world applications, which vary from Convolutional Neural Networks (CNNs) to computer vision algorithms. In terms of execution time, our results range from a limited ~8% performance overhead vs optimised CNN programs to 4.2× performance gain in content-dependent applications. We achieve up to 33% decrease in scratchpad memory usage vs well-established memory allocators and up to 6× smaller inter-process communication time. Evangelos Petrongonas, Vasileios Leon, George Lentaris, Dimitrios Soudris |
DSD | 3 |
| 2021 | A PVT-Aware Voltage Scaling Method for Energy Efficient FPGAsabstractThe paper proposes a method for guard-band customization to improve the energy efficiency of commercial FPGA. We deploy custom delay-based sensors alongside any user design to indirectly monitor, in real-time, the functional integrity of the target under voltage scaling. We develop a reliable sensing mechanism and regulate the FPGA operation by holistically considering process, voltage, and temperature variations during run-time. Tests with Xilinx Zynq SoC FPGAs and real benchmarks show significant power savings, in the area of 16-27%, while preserving nominal timing performance for only 1.6% resource overhead. Konstantinos Maragos 0001, George Lentaris, Dimitrios Soudris |
ISCAS | 2 |
| 2021 | Improving Performance-Power-Programmability in Space Avionics with Edge Devices: VBN on Myriad2 SoCabstractThe advent of powerful edge devices and AI algorithms has already revolutionized many terrestrial applications; however, for both technical and historical reasons, the space industry is still striving to adopt these key enabling technologies in new mission concepts. In this context, the current work evaluates an heterogeneous multi-core system-on-chip processor for use on-board future spacecraft to support novel, computationally demanding digital signal processors and AI functionalities. Given the importance of low power consumption in satellites, we consider the Intel Movidius Myriad2 system-on-chip and focus on SW development and performance aspects. We design a methodology and framework to accommodate efficient partitioning, mapping, parallelization, code optimization, and tuning of complex algorithms. Furthermore, we propose an avionics architecture combining this commercial off-the-shelf chip with a field programmable gate array device to facilitate, among others, interfacing with traditional space instruments via SpaceWire transcoding. We prototype our architecture in the lab targeting vision-based navigation tasks. We implement a representative computer vision pipeline to track the 6D pose of ENVISAT using megapixel images during hypothetical spacecraft proximity operations. Overall, we achieve 2.6 to 4.9 FPS with only 0.8 to 1.1 W on Myriad2 , i.e., 10-fold acceleration versus modern rad-hard processors. Based on the results, we assess various benefits of utilizing Myriad2 instead of conventional field programmable gate arrays and CPUs. Vasileios Leon, George Lentaris, Evangelos Petrongonas, Dimitrios Soudris, Gianluca Furano, Antonis Tavoularis, David Moloney |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | Process Variability Analysis in Interconnect, Logic, and Arithmetic Blocks of 16-nm FinFET FPGAsabstractIn the current work, we study the process variability of logic, interconnect, and arithmetic/DSP resources in commercial 16-nm FPGAs. We create multiple, soft-macro sensors for each distinct resource under evaluation, and we deploy them across the FPGA fabric to measure intra-die variation, as well as across multiple FPGAs to measure inter-die variation. The derived results are used to create device-signature variability maps characterizing the distribution of variability across the die. Our study includes decoupling of variability to systematic and stochastic parts, exploration of variability under various voltage and temperature conditions and correlation analysis between the variability maps of the different resources. Furthermore, we scrutinize the impact of variability on the performance of actual test circuits and correlate the retrieved results with the sensor-based maps. Our experimental results on four Zynq XCZU7EV FPGAs showed significant intra- and inter-die variability, up to 7.8% and 8.9%, respectively, with a small increase under certain operating conditions. The correlation analysis demonstrated a strong correlation between the logic and arithmetic resources, whereas the interconnects showed a slightly weaker correlation in specific devices. Finally, a relatively moderate correlation was calculated between the variability maps and performance of test circuits due their dissimilar operating behavior versus our sensors. Endri Taka, Konstantinos Maragos 0001, George Lentaris, Dimitrios Soudris |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2020 | Fast Packet Classification using RISC-V and HyperSplit Acceleration on FPGAabstractPerformance demands in communications technology is driving research towards advanced network processors, which are able to handle huge rates of incoming packets via application-specific circuits, however, without sacrificing all of the conventional CPU flexibility. At the same time, the advent of RISC-V is disrupting the industry & academia by opening computer architecture to a broader research community. Combining the above, the current paper considers placing dedicated VHDL accelerators next to a RISC-V processor to accommodate network functions via customized HW/SW co-processing. We extend the ISA with a new instruction to perform search tree operations that accelerate Packet Classification tasks in routers. For rapid prototyping and design exploration, we implement the binary search of HyperSplit algorithm on an Xilinx Ultrascale xcku060 FPGA. Our design achieves up to 118× faster classification than RISC-V alone and sustains up to 25.4M packets/sec throughput. Arsinoe Pnevmatikou, George Lentaris, Dimitrios Soudris, Nikos Kokkalis |
ISCAS | 2 |
| 2020 | High-Performance Vision-Based Navigation on SoC FPGA for Spacecraft Proximity OperationsabstractFuture autonomous spacecraft rendezvous with uncooperative or unprepared objects will be enabled by vision-based navigation, which imposes great computational challenges. Targeting short duration missions in low Earth orbit, this paper develops high-performance avionics supporting custom computer vision algorithms of increased complexity for satellite pose tracking. At algorithmic level, we track 6D pose by rendering a depth image from an object mesh model and robustly matching edges detected in the depth and intensity images. At system level, we devise an architecture to exploit the structure of commercial system-on-chip FPGAs, i.e., Zynq7000, and the benefits of tightly coupling VHDL accelerators with CPU-based functions. At implementation level, we employ our custom HW/SW co-design methodology and an elaborate combination of digital circuit design techniques to optimize and map efficiently all functions to a compact embedded device. Providing significant performance per watt improvement, the resulting VBN system achieves a throughput of 10-14 FPS for 1 Mpixel images, with only 4.3 watts mean power and 1U size, while tracking ENVISAT in real-time with only 0.5% mean positional error. George Lentaris, Ioannis Stratakos, Ioannis Stamoulias, Dimitrios Soudris, Manolis I. A. Lourakis, Xenophon Zabulis |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | PVT-Aware Sensing and Voltage Scaling for Energy Efficient FPGAsabstractIn this work we introduce a method to improve the energy efficiency of the FPGA devices by reducing the pessimistic operation guardbands posed by the commercial EDA tools. The proposed method bases on a voltage scaling scheme that reliably decreases the supply voltage. We deploy a uniform network of delay-based sensors across the fabric of the FPGA to sense all process, voltage and temperature variation (PVT) effects. The delay of all the sensors is calibrated to match the worst critical path delay of the target application. In that respect, the monitoring of the sensor network enables the indirect assessment of the functional integrity of the target application. The distributed placement of the sensors provides the desired sensitivity with appropriate granularity across the fabric and allows us to consider the worst-case scenario. The sensor network is integrated during the development cycle as ready-to-use software IP with negligible resource overhead, for example, 1-2% of a Zynq XC7Z020 FPGA for 10 sensors. The sensitivity of the sensors to all PVT variations and the correlation with the application operation is verified through extensive testing by using multiple FPGAs and realistic benchmarks. The aforementioned approach facilitates a closed-loop voltage scaling scheme to regulate the supply voltage and reduce the power of the system. In our experiments on a set of 28nm Xilinx XC7Z020 SoC FPGAs and realistic digital signal processing (DSP) benchmarks, we demonstrate up to 27.2% decrease in power for 13% decrease in voltage, while retaining the nominal timing performance. Konstantinos Maragos 0001, George Lentaris, Dimitrios Soudris, Vasilis F. Pavlidis |
FPGA | 2 |
| 2019 | Analysis of Performance Variation in 16nm FinFET FPGA DevicesabstractProcess variability is a challenging fabrication issue impacting, mainly, the reliability and performance of chips. Variability is already present in current technology nodes and is expected to become even more significant in the future. In this work, we focus on the study of performance variation in 16nm FinFET FPGAs. We devise a comprehensive assessment methodology based on multiple programmable sensors with diverse resource and delay characteristics. Additionally, we consider various voltage and temperature conditions and decouple variability to systematic and stochastic. The experimental results on Zynq XCZU7EV show up to 7.3% intra-die variation increasing to 9.9% for certain operating conditions. Our approach demonstrates that logic and interconnect resources present different variability, slightly uncorrelated, which highlights the necessity and way towards more sophisticated mitigation methods/tools. Konstantinos Maragos 0001, Endri Taka, George Lentaris, Ioannis Stratakos, Dimitrios Soudris |
FPL | 3 |
| 2019 | In-the-Field Mitigation of Process Variability for Improved FPGA PerformanceabstractThe mitigation of process variability becomes paramount as chip fabrication advances deeper into the sub-micron regime. Conservative guard-bands result in considerable performance loss, while most low-level solutions impede dynamic customization at application level. This paper exploits the existing process variability of commercial off-the-shelf FPGAs to improve the operating frequency of a design, in-the-field, at anytime during the lifetime of a chip. We begin by measuring variability in prevalent FPGAs and assessing its impact on the performance of common DSP benchmarks. For the former, we develop a custom sensing network of Ring-Oscillators to generate detailed 2D maps per chip. For the latter, we perform intensive testing and statistical analysis to establish the relation between variability maps and benchmark frequencies. Accordingly, we propose a framework to automatically characterize the user's devices, place the design on the most efficient region, and scale its frequency based on user requirements and functional verification. Experimental results on 20 FPGAs of 28 nm Xilinx technology show up to 13 percent intra-die and 30 percent inter-die variability; with limited cost, our framework provides 10-14.7 percent average gain by exploiting such variability, or up to 56-138 percent by also customizing the guard-band. Konstantinos Maragos 0001, George Lentaris, Dimitrios Soudris |
IEEE Trans. Computers | 2 |
| 2019 | Single- and Multi-FPGA Acceleration of Dense Stereo Vision for Planetary RoversabstractIncreased mobile autonomy is a vital requisite for future planetary exploration rovers. Stereo vision is a key enabling technology in this regard, as it can passively reconstruct in three dimensions the surroundings of a rover and facilitate the selection of science targets and the planning of safe routes. Nonetheless, accurate dense stereo algorithms are computationally demanding. When executed on the low-performance, radiation-hardened CPUs typically installed on rovers, slow stereo processing severely limits the driving speed and hence the science that can be conducted in situ . Aiming to decrease execution time while increasing the accuracy of stereo vision embedded in future rovers, this article proposes HW/SW co-design and acceleration on resource-constrained, space-grade FPGAs. In a top-down approach, we develop a stereo algorithm based on the space sweep paradigm, design its parallel HW architecture, implement it with VHDL, and demonstrate feasible solutions even on small-sized devices with our multi-FPGA partitioning methodology. To meet all cost, accuracy, and speed requirements set by the European Space Agency for this system, we customize our HW/SW co-processor by design space exploration and testing on a Mars-like dataset. Implemented on Xilinx Virtex technology, or European NG-MEDIUM devices, the FPGA kernel processes a 1,120 × 1,120 stereo pair in 1.7s−3.1s, utilizing only 5.4−9.3 LUT6 and 200−312 RAMB18. The proposed system exhibits up to 32× speedup over desktop CPUs, or 2,810× over space-grade LEON3, and achieves a mean reconstruction error less than 2cm up to 4m depth. Excluding errors exceeding 2cm (which are less than 4% of the total), the mean error is under 8mm. George Lentaris, Konstantinos Maragos 0001, Dimitrios Soudris, Xenophon Zabulis, Manolis I. A. Lourakis |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2018 | A Framework Exploiting Process Variability to Improve Energy Efficiency in FPGA ApplicationsabstractAs technology node scales-down and process variability increases, the vendors impose even more conservative guard-bands to prevent potential malfunction of their microchips. However, this approach introduces considerable amounts of unexploited performance to individual chips, which can be harvested by developing novel customization tools. In the current work, we focus on the exploitation of process variability in modern FPGA chips to provide more energy efficient solutions. We propose a framework that i) generates variability maps characterizing the energy efficiency of commercial chips and ii) combines voltage and frequency scaling to limit the power dissipation of any given design for a given set of performance constraints. Experimental results on Zynq XC7Z020 28nm FPGAs show that the developed framework achieves up to 28.3% power reduction while maintaining the performance and functional integrity of realistic benchmarks. Moreover, by selecting the most efficient chip, we achieve up to 5.1% additional power savings. Konstantinos Maragos 0001, George Lentaris, Ioannis Stratakos, Dimitrios Soudris |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Carrier Phase Recovery of 64 GBd Optical 16-QAM Using Extensive Parallelization on an FPGAabstractCarrier phase recovery (CPR) for phase noise mitigation is a vital part of modern coherent optical systems. As higher order M-QAM modulation formats are being employed to increase network capacity, CPR complexity grows as well, and more powerful chips are needed to cope with the advanced DSP. In this paper we present an FPGA-based flexible architecture of the Nonlinear Least Squares (NLS) CPR algorithm, targeting present-day and future generation systems. We describe optimization approaches at the algorithmic and HW levels to facilitate HW efficiency, while we combine multiple parallelization techniques to achieve high-throughput processing. Considering various FPGA devices, we perform fine-grain exploration with respect to cost, throughput and accuracy to support up to 64 GBd 16-QAM links with less than 1 dB SNR penalty. Vatistas Kostalampros, Konstantinos Maragos 0001, George Lentaris, Dimitrios Soudris, Christos Spatharakis, Nikolaos Argyris 0002, Hercules Avramopoulos, Stefanos Dris, André Richter |
ISCAS | 3 |
| 2017 | Application performance improvement by exploiting process variability on FPGA devicesabstractProcess variability is known to be increasing with technology scaling in IC fabrication, thereby degrading the overall performance of the manufactured devices. The current paper focuses on the variability effect in FPGAs and the possibility to boost the performance of each device at run-time, after fabrication, based on the individual characteristics of this device. First, we develop a sensing infrastructure involving a wide network of customized ring oscillators to measure intra-chip and inter-chip variability in 28nm FPGAs, i.e., in eight Xilinx Zynq XC7Z020T-1CSG324 devices. Second, we develop a closed-loop framework based on dynamic reconfiguration of clock tiles, I/O data sniffing, HW/SW communication, and verification with test vectors, to dynamically increase the operating frequency in Zynq while preserving its correctness. Our results show intra-chip variability in the area of 5.2% to 7.7% and inter-chip variability up to 17%. Our framework improves the performance of example FIR designs by up to 90.3% compared to the SW tool reports and shows speed difference among devices by up to 12.4%. Konstantinos Maragos 0001, George Lentaris, Dimitrios Soudris, Kostas Siozios, Vasilis F. Pavlidis |
DATE | 2 |
| 2017 | FPGA acceleration of hyperspectral image processing for high-speed detection applicationsabstractRecent advances in photonics and imaging technology allow the development of cutting-edge, lightweight hyperspectral sensors, both push-broom/line-scanning and snapshot/frame. At the same time, emerging applications in robotics, food inspection, medicine and earth observation are posing critical challenges on real-time processing and computational efficiency, both in terms of accuracy and power consumption. In this direction, in the current paper, we accelerate hyperspectral processing kernels by utilizing FPGAs, i.e., Zynq-7000 SoC, to perform similarity-based matching of spectral signatures. We propose a custom HW architecture based on multi-level parallelization, modularity, and parametric VHDL coding, which allows for in-depth design space exploration and trade-off analysis. Depending on configuration, our implementation processes 22-107 Megapixels per second providing an acceleration of 40-355x vs Intel-i3 CPU and 360-104x vs the embedded ARM Cortex A9, whereas the overall detection quality ranges from 56% to 97% when evaluated with multiple objects and images of 285 spectral channels. Simon Vellas, George Lentaris, Konstantinos Maragos 0001, Dimitrios Soudris, Zacharias Kandylakis, Konstantinos Karantzalos |
ISCAS | 2 |
| 2016 | Reduced Complexity Superresolution for Low-Bitrate Video CompressionabstractEvolving video applications impose requirements for high image quality, low bitrate, and/or small computational cost. This paper combines state-of-the-art coding and superresolution (SR) techniques to improve video compression both in terms of coding efficiency and complexity. The proposed approach improves a generic decimation–quantization compression scheme by introducing low complexity single-image SR techniques for rescaling the data at the decoder side and by jointly exploring/optimizing the downsampling/upsampling processes. The enhanced scheme achieves improvement of the quality and system’s complexity compared with conventional codecs and can be easily modified to meet various diverse requirements, such as effectively supporting any off-the-shelf video codec, for instance H.264/Advanced Video Coding or High Efficiency Video Coding. Our approach builds on studying the generic scheme’s parameterization with common rescaling techniques to achieve 2.4-dB peak signal-to-noise ratio (PSNR) quality improvement at low-bitrates compared with the conventional codecs and proposes a novel SR algorithm to advance the critical bitrate at the level of 10 Mb/s. The evaluation of the SR algorithm includes the comparison of its performance to other image rescaling solutions of the literature. The results show quality improvement by 5-dB PSNR over straightforward interpolation techniques and computational time reduction by three orders of magnitude when compared with the highly involved methods of the field. Therefore, our algorithm proves to be most suitable for use in reduced complexity downsampled compression schemes. Georgios Georgis, George Lentaris, Dionysios I. Reisis |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | HW/SW Codesign and FPGA Acceleration of Visual Odometry Algorithms for Rover Navigation on MarsabstractFuture Mars exploration missions rely heavily on high-mobility autonomous rovers equipped with sophisticated scientific instruments and possessing advanced navigational capabilities. Increasing their navigation velocity and localization accuracy is essential for enabling these rovers to explore large areas on Mars. Contemporary Mars rovers move slowly, partially due to the long execution time of complex computer vision algorithms running on their slow space-grade CPUs. This paper exploits the advent of high-performance space-grade field-programmable gate arrays (FPGAs) to accelerate the navigation of future rovers. Specifically, it focuses on visual odometry (VO) and performs HW/SW codesign to achieve one order of magnitude faster execution and improved accuracy. Conforming to the specifications of the European Space Agency, we build a proof-of-concept system on an HW/SW platform with processing power resembling that to be available onboard future rovers. We develop a codesign methodology adapted to the rover's specifications, design parallel architectures, and customize several feature extraction, matching, and motion estimation algorithms. We implement and evaluate five distinct HW/SW pipelines on a Virtex6 FPGA and a 150 MIPS CPU. We provide a detailed analysis of their cost-time-accuracy tradeoffs and quantify the benefits of employing FPGAs for implementing VO. Our solution achieves a speedup factor of 16× over a CPU-only implementation, handling a stereo image pair in less than 1 s, with a 1.25% mean positional error after a 100 m traverse and an FPGA cost of 54 K LUTs and 1.46-MB RAM. George Lentaris, Ioannis Stamoulias, Dimitrios Soudris, Manolis I. A. Lourakis |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Hardware implementation of stereo correspondence algorithm for the ExoMars missionabstractComputer vision algorithms exhibit increased complexity introducing significant implementation problems in conventional computing systems, especially whenever real-time constraints are imposed. This paper describes the ESA compatible VHDL development of a stereo correspondence algorithm for rover navigation in the SPARTAN system. The design is implemented on a Xilinx Virtex-6 FPGA and the evaluation results validate the efficiency of the applied methodology by showing real-time performance with minimal hardware utilization. George Lentaris, Dionysios Diamantopoulos, Kostas Siozios, Dimitrios Soudris, Marcos Avilés |
FPL | 1 |
| 2010 | A Graphics Parallel Memory Organization Exploiting Request CorrelationsabstractReal-time graphics applications require memory organizations featuring parallel pixel access and low-cost implementation. This work bases on a nonlinear skew mapping scheme and exploits the correlation between consecutive requests for pixels to design an efficient parallel memory organization. The mapping achieves parallel access, of mn pixels in various shapes, to the memory organized with mn banks. The proposed design technique combines the mapping properties and the spatial correlations among pixel requests to eliminate conflicts by spending at most one extra cycle every mn consecutive parallel pixel accesses. Consequently, the technique ensures that any pixel pattern-among these commonly used in graphics-can be accessed in a single cycle from any image location. The address computations become straightforward as the numbers of the requested pixels and the banks-apart from equal-can be powers of 2. George Lentaris, Dionysios I. Reisis |
IEEE Trans. Computers | 1 |
| 2006 | An approach for efficient design of digital amplifiersabstractThis paper presents a control-theoretic approach to the design of digital to analog converters and digital amplifiers leading to improved performance in audio and multimedia applications. The design involves oversampling and noise compression blocks as a pulse width modulation class-D amplifier, introduces an output-filter estimator block and applies a different modulation scheme. The theoretical model results in a family of digital circuits verified by software simulation and validated by a FPGA implementation with best performance 147 dB signal to noise ratio Nikolaos Vlassopoulos, Dionysios I. Reisis, George Lentaris, George S. Tombras, Evangelos A. Prosalentis, N. Ritas, Konstantinos S. Tsakalis |
ISCAS | 3 |