EDBT 2026 Demo / reviewers in the wild / expert
Lester Kalms
dblp:183/8643
· DBLP profile ↗
8ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0003-0638-0510ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | EuFRATE: European FPGA Radiation-hardened Architecture for TelecommunicationsabstractThe EuFRATE project aims to research, develop and test radiation-hardening methods for telecommunication payloads deployed for Geostationary-Earth Orbit (GEO) using Commercial-Off- The-Shelf Field Programmable Gate Arrays (FPGAs). This project is conducted by Argotec Group (Italy) with the collaboration of two partners: Politecnico di Torino (Italy) and Technische Universität Dresden (Germany). The idea of the project focuses on high-performance telecommunication algorithms and the design and implementation strategies for connecting an FPGA device into a robust and efficient cluster of multi-FPGA systems. The radiation-hardening techniques currently under development are addressing both device and cluster levels, with redundant datapaths on multiple devices, comparing the results and isolating fatal errors. This paper introduces the current state of the project's hardware design description, the composition of the FPGA cluster node, the proposed cluster topology, and the radiation hardening techniques. Intermediate stage experimental results of the FPGA communication layer performance and fault detection techniques are presented. Finally, a wide summary of the project's impact on the scientific community is provided.1 Ludovica Bozzoli, Antonino Catanese, Emilio Fazzoletto, Eugenio Scarpa, Diana Göhringer, Sergio A. Pertuz 0001, Lester Kalms, Cornelia Wulf, Najdet Charaf, Luca Sterpone, Sarah Azimi, Daniele Rizzieri, Salvatore Gabriele La Greca, David Merodio Codinachs |
DATE | 7 |
| 2022 | High-Performance AKAZE Implementation Including Parametrizable and Generic HLS ModulesabstractThe amount of image data to be processed has increased tremendously over the last decades. One major computer vision task is the extraction of information to find patterns in and between images. One well-studied pattern recognition algorithm is AKAZE which builds a nonlinear scale space to detect features. While being more efficient compared to its predecessor KAZE, the computational demands of AKAZE are still high. Since many real-world computer vision applications require fast computations, sometimes under hard power and time constraints, FPGAs became a focus as a suitable target platform. This work presents a highly modularized and parameterizable implementation of the AKAZE feature detection algorithm integrated into HiFlipVX, which is a High-Level Synthesis library based on the OpenVX standard. The fine granular modularization and the generic design of the implemented functions allows them to be easily reused, increasing the workflow for other computer vision algorithms. The high degree of parameterization and extension of the library enables also a fast and extensive exploration of the design space. The proposed design achieved a high repeatability and frame rate of up to 480 frames per second for an image resolution of 1920×1080 compared to related work. Matthias Nickel, Lester Kalms, Tim Haering, Diana Göhringer |
ASAP | 2 |
| 2022 | A cross-platform OpenVX library for FPGA acceleratorsabstractFPGAs are an excellent platform to implement computer vision applications, since these applications tend to offer a high level of parallelism with many data-independent operations. However, the freedom in the solution design space of FPGAs represents a problem because each solution must be individually designed, verified, and tuned. The emergence of High Level Synthesis (HLS) helps solving this problem and has allowed the implementation of open programming standards as OpenVX for computer vision applications on FPGAs, such as the HiFlipVX library developed exclusively for Xilinx devices. Although with the HiFlipVX library, designers can develop solutions efficiently on Xilinx, they do not have an approach to port and run their code on FPGAs from other manufacturers. This work extends the HiFlipVX capabilities in two significant ways: supporting Intel FPGA devices and enabling execution on discrete FPGA accelerators. To provide both without affecting user-facing code, the new carried out implementation combines two HLS programming models: C++, using Intel’s system of tasks, and OpenCL, which provides the CPU interoperability. Comparing with pure OpenCL implementations, this work reduces kernel dispatch resources, saving up to 24% of ALUT resources for each kernel in a graph, and improves performance 2.6 × and energy consumption 1.6 × on average for a set of representative applications, compared with state-of-the-art frameworks. Maria Angelica Davila Guzman, Lester Kalms, Ruben Gran Tejero, María Villarroya-Gaudó, Darío Suárez Gracia, Diana Göhringer |
J. Syst. Archit. | 2 |
| 2021 | A Cross-Platform OpenVX Library for FPGA AcceleratorsabstractIn Computer Vision, open programming standards such as OpenVX have emerged to bring together portability and acceleration across devices. Unfortunately, achieving both goals on FPGAs remains a challenge because FPGAs still require to adapt the code with proprietary extensions. Exclusively for Xilinx devices, the HiFlipVX open source library partially solves this problem by offering a clean C++ OpenVX API that offers the performance of proprietary extensions without exposing its complexity to programmer. While HiFlipVX enables portability within Xilinx devices, portability between FPGA manufacturers remains an open challenge. This work extends the HiFlipVX's capabilities with a twofold goal: i) to support Intel FPGA devices with different memory configurations, and ii) to enable execution on FPGAs as discrete accelerators. To accomplish these goals, the proposed implementation combines two HLS programming models: C++, using Intel's system of tasks that enables to coalesce nodes and reduce control overhead, and OpenCL, which provides efficient compute kernel nodes. On Intel FPGAs, compared with pure OpenCL implementations, the proposed implementation reduces kernel dispatch resources, saving up to 24% of ALUT resources for each kernel in a graph, and improves performance. Gains are 2.6× on average for representative applications, such as Canny edge detector, or Census transform, compared with state-of-the-art frameworks. Maria Angelica Davila Guzman, Ruben Gran Tejero, María Villarroya-Gaudó, Darío Suárez Gracia, Lester Kalms, Diana Göhringer |
PDP | 5 |
| 2019 | Efficient Pattern Recognition Algorithm Including a Fast Retina Keypoint FPGA ImplementationabstractThe field of computer vision is continuously increasing and becoming more complex and power demanding. Using feature detection and description allows a fast object detection without needing big databases. FPGAs are predestined for different requirements, like real-time and power constraints, which are important in many application areas. This work proposes a new pattern recognition algorithm, based on an improved Accelerated KAZE (AKAZE) detector and Fast Retina Keypoint (FREAK) descriptor. Our software implementation increased the repeatability in comparison to the original algorithm using optimized configurations. The percentage of correct matching features between two images (repeatability) increased from 85.7% to 91.4%, while the computation time decreases from 70.3ms to 24.9ms. Furthermore, we present an efficient FPGA implementation of the FREAK descriptor. The accelerator processes 2048 features at 73.4 frames per second; achieving a repeatability of 90.9%, while being optimized for resource utilization and memory bandwidth consumption. Additionally, we show an efficient Integral Image implementation that processes four image pixels per clock cycle at a high frequency (204 MHz on xc7z020clg484-1) consuming minimum resources. Lester Kalms, Maximilian Hajduk, Diana Göhringer |
FPL | 1 |
| 2019 | Scalable clustering and mapping algorithm for application distribution on heterogeneous and irregular FPGA clusters
Lester Kalms, Diana Göhringer |
J. Parallel Distributed Comput. | 1 |
| 2017 | Exploration of OpenCL for FPGAs using SDAccel and comparison to GPUs and multicore CPUsabstractDue to energy efficiency, heterogeneous computing is gaining more and more attention. Since FPGA implementations are time consuming, high-level synthesis (HLS) is used to close the productivity gap. OpenCL has become accepted as a good programming model for HLS, due to its portability, good capability of design verification and rich instruction set. This work implements different optimization strategies using OpenCL for a heterogeneous system containing CPU, integrated GPU, GPU and FPGA. Energy efficiency and performance of the architectures are compared using a feature detection algorithm. It is shown how to maximize performance while hitting the maximum memory bandwidth and keeping the resource utilization low for the SDAccel tool from Xilinx. The evaluation shows the great streaming capability of OpenCL for FPGAs. The FPGA achieves a speed up of 62.8 and consumes 49 times less energy for the application in comparison to an optimized single threaded CPU implementation in full HD. Lester Kalms, Diana Göhringer |
FPL | 1 |
| 2016 | FPGA based hardware accelerator for KAZE feature extraction algorithmabstractProcessing and understanding of visual data has a significant importance in many applications such as robotics and vision aid devices. Extracting image features is one of the important tasks in computer vision. This paper focuses on KAZE features algorithm, due to its good performance. KAZE features is a multi-scale 2D feature detection and description algorithm. It describes 2D features in a non-linear scale space by means of non-linear diffusion filtering. In this paper, the algorithm was optimized for speed, memory usage and portability. The paper presents a hardware accelerator for the scale-space analysis part of the algorithm on FPGA. A high speed-up has been achieved by this accelerator by parallelizing several parts of the algorithm and reducing the memory bandwidth. Lester Kalms, Ahmed Elhossini, Ben H. H. Juurlink |
FPT | 1 |