Tomasz Kryjak

dblp:55/8121 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-6798-4444ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-author
YearPublicationVenuePosition
2026 SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
abstract
While neural network quantization effectively reduces the cost of matrix multiplications, aggressive quantization can expose non-matrix-multiply operations as significant performance and resource bottlenecks on embedded systems. Addressing such bottlenecks requires a comprehensive approach to tailoring the precision across operations in the inference computation. To this end, we introduce scaled-integer range analysis ( SIRA ), a static analysis technique employing interval arithmetic to determine the range, scale, and bias for tensors in quantized neural networks. We show how this information can be exploited to reduce the resource footprint of FPGA dataflow neural network accelerators via tailored bitwidth adaptation for accumulators and downstream operations, aggregation of scales and biases, and conversion of consecutive elementwise operations to thresholding operations. We integrate SIRA -driven optimizations into the open source FINN framework, then evaluate their effectiveness across a range of quantized neural network workloads and compare implementation alternatives for non-matrix-multiply operations. We demonstrate an average reduction of 17% for LUTs, 66% for DSPs, and 22% for accumulator bitwidths with SIRA optimizations, providing detailed benchmark analysis and analytical models to guide the implementation style for non-matrix layers. Finally, we open source SIRA to facilitate community exploration of its benefits across various applications and hardware platforms.
Yaman Umuroglu, Christoph Berganski, Felix Paul Jentzsch, Michal Danilowicz, Tomasz Kryjak, Charalampos Bezaitis, Magnus Själander, Ian Colbert, Thomas B. Preußer, Jakoba Petri-Koenig, Michaela Blott
ACM Trans. Reconfigurable Technol. Syst.5
2025 Hardware-Aware Feature Extraction Quantisation for Real-Time Visual Odometry on FPGA Platforms
abstract
Accurate position estimation is essential for modern navigation systems deployed in autonomous platforms, including ground vehicles, marine vessels, and aerial drones. In this context, Visual Simultaneous Localisation and Mapping (VSLAM) which includes Visual Odometry - relies heavily on the reliable extraction of salient feature points from the visual input data. In this work, we propose an embedded implementation of an unsupervised architecture capable of detecting and describing feature points. It is based on a quantised SuperPoint convolutional neural network. Our objective is to minimise the computational demands of the model while preserving high detection quality, thus facilitating efficient deployment on platforms with limited resources, such as mobile or embedded systems. We implemented the solution on an FPGA System-on-Chip (SoC) platform, specifically the AMD/Xilinx Zynq UltraScale+, where we evaluated the performance of Deep Learning Processing Units (DPUs) and we also used the Brevitas library and the FINN framework to perform model quantisation and hardware-aware optimisation. This allowed us to process $640 \times 480$ pixel images at up to 54 fps on an FPGA platform, outperforming state-of-the-art solutions in the field. We conducted experiments on the TUM dataset to demonstrate and discuss the impact of different quantisation techniques on the accuracy and performance of the model in a visual odometry task.
Mateusz Wasala, Mateusz Smolarczyk, Michal Danilowicz, Tomasz Kryjak
DSD4
2025 Live Demonstration: Continuous Processing of Event-Data with Graph Convolutional Neural Networks Implemented for SoC FPGA
abstract
For perception systems in mobile robotics - particularly in demanding, highly dynamic environments - event cameras (DVS - Dynamic Vision Sensors) are being increasingly utilised as an alternative to conventional vision sensors. The real-time processing of the registered spatio-temporal sparse point cloud must satisfy stringent requirements in terms of latency, throughput, and energy efficiency. In this demonstration, we present our method for implementing Graph Convolutional Neural Networks on a heterogeneous SoC FPGA platform, aimed at meeting these constraints. We address the challenges associated with integrating event-based sensors with reconfigurable hardware, as well as the intricate relationship between latency, performance, and hardware resource utilisation.
Piotr Wzorek, Kamil Jeziorek, Marcin Kowalczyk, Krzysztof Blachut, Tomasz Kryjak, Marek Gorgon
FPL5
2024 Event-Based Vision on FPGAs - a Survey
abstract
In recent years there has been a growing interest in event cameras, i.e. vision sensors that record changes in illumination independently for each pixel. This type of operation ensures that acquisition is possible in very adverse lighting conditions, both in low light and high dynamic range, and reduces average power consumption. In addition, the independent operation of each pixel results in low latency, which is desirable for robotic solutions. Nowadays, Field Programmable Gate Arrays (FPGAs), along with general-purpose processors (GPPs/CPUs) and programmable graphics processing units (GPUs), are popular architectures for implementing and accelerating computing tasks. In particular, their usefulness in the embedded vision domain has been repeatedly demonstrated over the past 30 years, where they have enabled fast data processing (even in real-time) and energy efficiency. Hence, the combination of event cameras and reconfigurable devices seems to be a good solution, especially in the context of energy-efficient real-time embedded systems. This paper gives an overview of the most important works, where FPGAs have been used in different contexts to process event data. It covers applications in the following areas: filtering, stereovision, optical flow, acceleration of AI-based algorithms (including spiking neural networks) for object classification, detection and tracking, and applications in robotics and inspection systems. Current trends and challenges for such systems are also discussed.
Tomasz Kryjak
DSD1
2024 PowerYOLO: Mixed Precision Model for Hardware Efficient Object Detection with Event Data
abstract
The performance of object detection systems in automotive solutions must be as high as possible, with minimal response time and, due to the often battery-powered operation, low energy consumption. When designing such solutions, we therefore face challenges typical for embedded vision systems: the problem of fitting algorithms of high memory and computational complexity into small low-power devices. In this paper we propose PowerYOLO - a mixed precision solution, which targets three essential elements of such application. First, we propose a system based on a Dynamic Vision Sensor (DVS), a novel sensor, that offers low power requirements and operates well in conditions with variable illumination. It is these features that may make event cameras a preferential choice over frame cameras in some applications. Second, to ensure high accuracy and low memory and computational complexity, we propose to use 4-bit width Powers-of- Two (PoT) quantisation for convolution weights of the YOLO detector, with all other parameters quantised linearly. Finally, we embrace from PoT scheme and replace multiplication with bit-shifting to increase the efficiency of hardware acceleration of such solution, with a special convolution-batch normalisation fusion scheme. The use of specific sensor with PoT quantisation and special batch normalisation fusion leads to a unique system with almost 8x reduction in memory complexity and vast computational simplifications, with relation to a standard approach. This efficient system achieves high accuracy of mAP 0.301 on the GENt DVS dataset, marking the new state-of-the-art for such compressed model.
Dominika Przewlocka-Rus, Tomasz Kryjak, Marek Gorgon
DSD2
2023 Power-of- Two Quantized YOLO Network for Pedestrian Detection with Dynamic Vision Sensor
abstract
Pedestrian detection algorithms have a wide range of applications: from video surveillance, to driver assistance systems and autonomous vehicles. The performance of these systems must be as high as possible, with minimal response time and, due to the often battery-powered operation (like in electric vehicles), low energy consumption. When designing such solutions, we therefore face challenges typical for embedded vision systems: the problem of fitting algorithms of high memory and computational complexity into small low-power devices. In this paper, we propose a system based on a Dynamic Vision Sensor (DVS), which has low power requirements and operates well in conditions with variable illumination. It is these features that may make event cameras a preferential choice over frame cameras in some applications. To ensure high accuracy, for pedestrian detection we use the YOLO (You Only Look Once) deep network on event data representation. Due to the high complexity of the applied algorithm, we propose a low precision architecture: the weights of the convolution layers are quantized logarithmically to 4 bits powers-of-two values (PoT quantization). Such compression reduces not only the memory complexity by almost 8x, but also the computational complexity by replacing most multiplication operations with bit-shifting, due to use of powers-of-two weights. At the same time, the proposed system achieves the accuracy on par with the floating point baseline, of mAP0.5 0.708.
Dominika Przewlocka-Rus, Tomasz Kryjak
DSD2
2022 Hardware architecture for high throughput event visual data filtering with matrix of IIR filters algorithm
abstract
Neuromorphic vision is a rapidly growing field with numerous applications in the perception systems of autonomous vehicles. Unfortunately, due to the sensors working principle, there is a significant amount of noise in the event stream. In this paper we present a novel algorithm based on an IIR filter matrix for filtering this type of noise and a hardware architecture that allows its acceleration using an SoC FPGA. Our method has a very good filtering efficiency for uncorrelated noise - over 99% of noisy events are removed. It has been tested for several event data sets with added random noise. We designed the hardware architecture in such a way as to reduce the utilisation of the FPGA's internal BRAM resources. This enabled a very low latency and a throughput of up to 385.8 MEPS million events per second. The proposed hardware architecture was verified in simulation and in hardware on the Xilinx Zynq Ultrascale+ MPSoC chip on the Mercury+ XU9 module with the Mercury+ ST1 base board.
Marcin Kowalczyk, Tomasz Kryjak
DSD2
2021 A Connected Component Labelling algorithm for a multi-pixel per clock cycle video stream
abstract
This work describes the hardware implementation of a connected component labelling (CCL) module in reprogammable logic. The main novelty of the design is the "full", i.e. without any simplifications, support of a 4 pixel per clock format (4 ppc) and real-time processing of a 4K/UltraHD video stream (3840 x 2160 pixels) at 60 frames per second. To achieve this, a special labelling method was designed and a functionality that stops the input data stream in order to process pixel groups which require writing more than one merger into the equivalence table. The proposed module was verified in simulation and in hardware on the Xilinx Zynq Ultrascale+ MPSoC chip on the ZCU104 evaluation board.
Marcin Kowalczyk, Tomasz Kryjak
DSD2
2021 A comparison of real-time 4K/UltraHD connected component labelling architectures
abstract
This work presents a comparison of hardware architectures realising a connected component labelling (CCL) algorithm in reprogrammable logic. The architectures are capable of processing a video stream with 4K/UltraHD resolution at 60 frames per second in real-time. The modules were verified in simulation and in hardware on the Xilinx Zynq Ultrascale+ MPSoC chip on the ZCU104 evaluation board.
Marcin Kowalczyk, Tomasz Kryjak
FPL2
2021 Quantised Siamese Tracker for 4K/UltraHD Video Stream - a demo
abstract
This demo presents a hardware architecture for an object tracker based on a quantised Siamese neural network. The system is designed to work with a 4K/UltraHD input video stream and to meet the real time and energy efficiency constraints. The network is designed using Brevitas and implemented with FINN tool from Xilinx. The Siamese tracker runs on the ZCU 104 board, with the Zynq UltraScale+ MPSoC device from Xilinx.
Dominika Przewlocka-Rus, Tomasz Kryjak
FPL2
2021 Hardware-software implementation of a DNN for 3D object detection using FINN - a demo
abstract
In this demo, we present a hardware-software system for 3D object detection in LiDAR point clouds based on a deep neural network. The PointPillars architecture was used in this research, as it is a reasonable compromise between detection accuracy and computational complexity. The Brevitas tool was used for network quantisation and the FINN tool for hardware-software implementation in the reprogrammable Zynq UltraScale+ MPSoC device.
Joanna Stanisz, Konrad Lis, Tomasz Kryjak, Marek Gorgon
FPL3
2016 A compact deep convolutional neural network architecture for video based age and gender estimation
abstract
In this paper research on a compact deep convolutional neural network (DCNN) architecture for age and gender estimation from facial images has been presented.The proposed solution was tested on the FERET and the Adience Benchmark databases.In the first case a 98.6% accuracy for gender and 86.4% for age estimation was obtained.For the Adience database, which contains images recorded in unconstrained conditions and is much more demanding, a 62.0% for gender and 42.0% for age accuracy was obtained.When compared to the reference results on a much larger network, the performance should be considered as satisfactory.The research shows that a compact DCNN with small input images can provide quite good classification results.
Bartlomiej Hebda, Tomasz Kryjak
FedCSIS2
2015 Shape and colour recognition of dishes for the purpose of customer service process automation in a self-service canteen
abstract
In the article a vision system for shape and colour recognition of dishes (plates, bowls, mugs), which can be used to automate the process of customer service in a self-service canteen is described.In consists of three basic components: object segmentation using so-called background model subtraction, shape recognition using geometric invariant moments and SVM classifier, as well as colour recognition using a Gaussian model.In addition, recognition in case of close or abut objects using a distance transform like approach is presented.The solution was evaluated on a dedicated test stand with controlled LED lightning.A 98% accuracy was obtained on over 100 test images, which indicates that the solution could be used in business practise.
Tomasz Kryjak, Damian Krol
FedCSIS1
2013 Real-time Implementation of the ViBe Foreground Object Segmentation Algorithm
Tomasz Kryjak, Marek Gorgon
FedCSIS1
2009 Pipeline implementation of the 128-bit block cipher CLEFIA in FPGA
abstract
The article presents a pipeline implementation of the block cipher CLEFIA. The article examines three known methods of implementing a single encryption round and proposes a new fourth method. The article proposes the implementation of a key scheduler, which is highly compatible with pipeline encryption. The article contains a detailed analysis of the data processing path for the 128-bit key version of the algorithm and verifies its operation on two FPGA cards in practice. On the basis of one of these cards, the article proposes a prototype of an effective supercomputer-compatible hardware accelerator (High Performance Computing Application).
Tomasz Kryjak, Marek Gorgon
FPL1