Felix Paul Jentzsch

dblp:276/2439 · also Felix Jentzsch · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0003-4987-5708ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
abstract
While neural network quantization effectively reduces the cost of matrix multiplications, aggressive quantization can expose non-matrix-multiply operations as significant performance and resource bottlenecks on embedded systems. Addressing such bottlenecks requires a comprehensive approach to tailoring the precision across operations in the inference computation. To this end, we introduce scaled-integer range analysis ( SIRA ), a static analysis technique employing interval arithmetic to determine the range, scale, and bias for tensors in quantized neural networks. We show how this information can be exploited to reduce the resource footprint of FPGA dataflow neural network accelerators via tailored bitwidth adaptation for accumulators and downstream operations, aggregation of scales and biases, and conversion of consecutive elementwise operations to thresholding operations. We integrate SIRA -driven optimizations into the open source FINN framework, then evaluate their effectiveness across a range of quantized neural network workloads and compare implementation alternatives for non-matrix-multiply operations. We demonstrate an average reduction of 17% for LUTs, 66% for DSPs, and 22% for accumulator bitwidths with SIRA optimizations, providing detailed benchmark analysis and analytical models to guide the implementation style for non-matrix layers. Finally, we open source SIRA to facilitate community exploration of its benefits across various applications and hardware platforms.
Yaman Umuroglu, Christoph Berganski, Felix Paul Jentzsch, Michal Danilowicz, Tomasz Kryjak, Charalampos Bezaitis, Magnus Själander, Ian Colbert, Thomas B. Preußer, Jakoba Petri-Koenig, Michaela Blott
ACM Trans. Reconfigurable Technol. Syst.3
2025 Analytical Buffer Sizing for Neural Network Inference Applications on FPGAs
abstract
A crucial and challenging step of implementing neural network inference on FPGAs is the sizing of intermediate buffers between the pipeline stages. Under-sizing buffers can lead to throughput loss and deadlocks, over-sizing wastes valuable device resources. Typically, buffer sizing is performed either with exhaustive search or solvers, accompanied by simulations. With modern neural networks rapidly growing in size, these approaches become infeasibly time consuming. In this work, we show how given the specialized domain of machine learning on FPGAs, it is feasible to dramatically speedup buffer sizing by avoiding the use of simulations and solvers. This is done by replacing them with analytical modeling and heuristics respectively. We incorporate our approach into the FINN compiler for neural network inference on FPGAs and show over three orders of magnitude speed up with buffer sizes closely matching exhaustive search results for a wide range of example models such as mobilenet-v1 and resnet50.
Lukas Stasytis, Felix Paul Jentzsch, Zsolt István
FPL2
2023 Hardware-Aware AutoML for Exploration of Custom FPGA Accelerators for RadioML
abstract
This PhD project aims to tackle the laborious design process of FPGA-based deep neural network (DNN) inference solutions by combining automatic machine learning (AutoML) algorithms with a compiler for custom-tailored dataflow accelerator architectures. Manual experiments on a first use case from the RadioML domain suggests great potential for joint optimization of DNN and accelerator, but emphasize the need for empirical quality-of-result (QoR) estimation models, which are the current focus of the project. Future work will entail design space modeling and evaluation of state-of-the-art exploration techniques.
Felix Paul Jentzsch
FPL1
2020 DeepWind: An Accurate Wind Turbine Condition Monitoring Framework via Deep Learning on Embedded Platforms
abstract
Condition monitoring of critical components in wind turbines is of utmost importance to enable low cost and preventive maintenance and to minimize downtime. In this paper we propose DEEPWIND, an end-to-end condition monitoring and fault detection framework that can be implemented on resource-constrained embedded platforms. DEEPWIND exploits multi-channel convolutional neural networks to automatically extract features from sensor data without any need of human feature engineering, and utilizes the features to classify faults occurring in rotor blades of wind turbines. Experiments on a real-world dataset provided by Weidmüller Monitoring Systems GmbH reveal an average F1-score of 0.94 and thus underline the suitability of the approach for fault detection in wind turbines.
Hassan Ghasemzadeh Mohammadi, Rahil Arshad, Sneha Rautmare, Suraj Manjunatha, Maurice Kuschel, Felix Paul Jentzsch, Marco Platzner, Alexander Boschmann, Dirk Schollbach
ETFA6