EDBT 2026 Demo / reviewers in the wild / expert
Jonas Ney
dblp:295/9171
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0002-1075-4589ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | eFAirWrite: Bringing energy efficient text entry to next generation smart devicesabstractText entry tasks for emerging wireless Augmented Reality (AR) and Virtual Reality (VR) devices can be realized in many ways, one of the most promising methods is based on an Inertial Measurement Unit (IMU) sensor, which resembles a human writing style and is called air-writing. An existing air-writing Deep Neural Network (DNN) based algorithm called FAirWrite achieves state-of-the-art accuracy. However, this algorithm is optimized only for accuracy without considering the implementation constraints. State-of-the-art implementation executes the algorithm in a cloud, which is associated with large communication latency; privacy issues with respect to data transfer; and a need for reliable internet connection. A solution that tackles all three challenges is executing the algorithm locally at the edge in the closest proximity to the sensor. However, inference at the edge is challenging due to limited memory and computing resources, which can impact accuracy and increase latency. Additionally, battery-powered edge devices must adhere to strict power and energy consumption limits. All these constraints collectively restrict the model size and computational complexity that can be deployed near the sensor. In this work, we explore various optimizations required to enable state-of-the-art FAirWrite algorithm for a real-world deployment scenario, i.e. executing on an edge device. We perform a multi-layer design-space exploration considering multiple levels of design hierarchy spanning from optimizations applied on algorithmic level down to hardware level, considering various deployment run-times and edge platforms including embedded micro-controller, embedded Central Processing Unit (CPU), embedded Graphics Processing Unit (GPU), Neural Processing Unit (NPU), and Field-Programmable Gate Array (FPGA). The complexity reduction optimizations result in a smaller eFAirWrite model without any accuracy degradation. We propose and implement a custom hardware architecture of the algorithm on an FPGA by utilizing custom data-types and memory hierarchy. We demonstrate that the FPGA implementation achieves , , , and higher energy efficiency as compared to NPU, embedded GPU, embedded CPU, and micro-controller, respectively, which makes it a suitable edge platform for portable and real-time air-writing text entry. Muhammad Mohsin Ghaffar, Ahmad Abdullah, Junaid Younas, Vladimir Rybalkin, Jonas Ney, Paul Lukowicz, Norbert Wehn |
Expert Syst. Appl. | 5 |
| 2024 | Achieving High Throughput with a Trainable Neural-Network-Based Equalizer for Communications on FPGAabstractThe ever-increasing data rates of modern communication systems lead to severe distortions of the communication signal, imposing great challenges to state-of-the-art signal processing algorithms. In this context, neural network (NN)-based equalizers are a promising concept since they can compensate for impairments introduced by the channel. However, due to the large computational complexity, efficient hardware implementation of NNs is challenging. Especially the backpropagation algorithm, required to adapt the NN's parameters to varying channel conditions, is highly complex, limiting the throughput on resource-constrained devices like field programmable gate arrays (FPGAs). In this work, we present an FPGA architecture of an NN-based equalizer that exploits batch-level parallelism of the convolutional layer to enable a custom mapping scheme of two multiplication to a single digital signal processor (DSP). Our implementation achieves a throughput of up to 20 GBd, which enables the equalization of high-data-rate nonlinear optical fiber channels while providing adaptation capabilities by retraining the NN using backpropagation. As a result, our FPGA implementation outperforms an embedded graphics processing unit (GPU) in terms of throughput by two orders of magnitude. Further, we achieve a higher energy efficiency and throughput as state-of-the-art NN training FPGA implementations. Thus, this work fills the gap of high-throughput NN-based equalization while enabling adaptability by NN training on the edge FPGA. Jonas Ney, Norbert Wehn |
DSD | 1 |
| 2024 | FPGA Onboard Processing of Tiny-ML Models Using Radar-Sensing for Oil Spill MonitoringabstractDeep learning models have been widely used recently for environmental monitoring. Unless designed carefully, they are well known for their large computational complexity and long computation time, which limit their direct utilization as onsite embedded solutions. Therefore, to assure their suitability for onboard processing on drones during tactical responses, we propose in this paper energy-efficient and highly accurate tiny machine learning (TinyML) models in the field of radar remote sensing for oil spill monitoring. More precisely, the proposed signal processing models are based on artificial neural networks and use radar signals to detect oil slicks on top of the seawater and estimate their thicknesses using drones in calm ocean conditions. Despite their extremely small size (7-37 Bytes), their accuracy in detection and parameter estimation exceeds 94%. Moreover, the proposed models are characterized by low-power (32-173 mW) and low-latency (0.35-0.59 µs) performance when implemented on the FPGA computing platform. Bilal Hammoud, Jonas Ney, Charbel Bou-Mosleh, Norbert Wehn |
IGARSS | 2 |
| 2022 | FPGA-based Trainable Autoencoder for Communication SystemsabstractIn communication systems, autoencoder refers to a system that replaces parts of the traditional transmitter and receiver of the baseband processing chain with artificial neural networks (ANNs). This allows to jointly train the system for an underlying channel model by reconstructing the input symbols at the output. Since the actual behavior of a real communication channel cannot be perfectly reproduced by an abstract model, it is necessary for the autoencoder to adapt to the changing conditions at runtime. Thus, online fine-tuning, in the form of ANN-retraining is of great importance. A platform able to satisfy the low-latency and low-power requirements of embedded communication systems are Field-programmable gate arrays (FPGAs). In this paper, we present an online-trainable low-power FPGA architecture for the receiver of an autoencoder-based communication chain. The architecture is embedded into an exploration framework that automatically determines the optimal degree of parallelism to minimize latency or power consumption. Our solutions achieve 2000×higher throughput than a high-performance GPU, draw 5×less power than an embedded CPU and are 5800×more energy efficient compared to an embedded GPU, for a batch size of one. To the best of our knowledge, this is the first FPGA-based autoencoder implementation for communication systems. Jonas Ney, Sebastian Dörner, Matthias Herrmann, Mohammad Hassani Sadi, Jannis Clausius, Stephan ten Brink, Norbert Wehn |
FPGA | 1 |
| 2022 | Blind and Channel-agnostic Equalization Using Adversarial NetworksabstractDue to the rapid development of autonomous driving, the Internet of Things and streaming services, modern communication systems have to cope with varying channel conditions and a steadily rising number of users and devices. This, and the still rising bandwidth demands, can only be met by intelligent network automation, which requires highly flexible and blind transceiver algorithms. To tackle those challenges, we propose a novel adaptive equalization scheme, which exploits the prosperous advances in deep learning by training an equalizer with an adversarial network. The learning is only based on the statistics of the transmit signal, so it is blind regarding the actual transmit symbols and agnostic to the channel model. The proposed approach is independent of the equalizer topology and enables the application of powerful neural network based equalizers. In this work, we prove this concept in simulations of different―both linear and nonlinear―transmission channels and demonstrate the capability of the proposed blind learning scheme to approach the performance of non-blind equalizers. Furthermore, we provide a theoretical perspective and highlight the challenges of the approach. Vincent Lauinger, Manuel Dossinger, Jonas Ney, Norbert Wehn, Laurent Schmalen |
GLOBECOM | 3 |
| 2022 | When Massive GPU Parallelism Ain't Enough: A Novel Hardware Architecture of 2D-LSTM Neural NetworkabstractMultidimensional Long Short-Term Memory (MD-LSTM) neural network is an extension of one-dimensional LSTM for data with more than one dimension. MD-LSTM achieves state-of-the-art results in various applications, including handwritten text recognition, medical imaging, and many more. However, its implementation suffers from the inherently sequential execution that tremendously slows down both training and inference compared to other neural networks. The main goal of the current research is to provide acceleration for inference of MD-LSTM. We advocate that Field-Programmable Gate Array (FPGA) is an alternative platform for deep learning that can offer a solution when the massive parallelism of GPUs does not provide the necessary performance required by the application. In this article, we present the first hardware architecture for MD-LSTM. We conduct a systematic exploration to analyze a tradeoff between precision and accuracy. We use a challenging dataset for semantic segmentation, namely historical document image binarization from the DIBCO 2017 contest and a well-known MNIST dataset for handwritten digit recognition. Based on our new architecture, we implement FPGA-based accelerators that outperform Nvidia Geforce RTX 2080 Ti with respect to throughput by up to 9.9 and Nvidia Jetson AGX Xavier with respect to energy efficiency by up to 48 . Our accelerators achieve higher throughput, energy efficiency, and resource efficiency than FPGA-based implementations of convolutional neural networks (CNNs) for semantic segmentation tasks. For the handwritten digit recognition task, our FPGA implementations provide higher accuracy and can be considered as a solution when accuracy is a priority. Furthermore, they outperform earlier FPGA implementations of one-dimensional LSTMs with respect to throughput, energy efficiency, and resource efficiency. Vladimir Rybalkin, Jonas Ney, Menbere Tekleyohannes, Norbert Wehn |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2021 | HALF: Holistic Auto Machine Learning for FPGAsabstractDeep Neural Networks (DNNs) are capable of solving complex problems in domains related to embedded systems, such as image and natural language processing. To efficiently implement DNNs on a specific FPGA platform for a given cost criterion, e.g., energy efficiency, an enormous amount of design parameters must be considered from the topology down to the final hardware implementation. Interdependencies between the different design layers must be taken into account and explored efficiently, making it hardly possible to find optimized solutions manually. An automatic, holistic design approach can improve the quality of DNN implementations on FPGA significantly. To this end, we present a cross-layer design space exploration methodology. It comprises optimizations starting from a hardware-aware topology search for DNNs down to the final optimized implementation for a given FPGA platform. The methodology is implemented in our Holistic Auto machine Learning for FPGAs (HALF) framework, which combines an evolutionary search algorithm, various optimization steps, and a library of parametrizable hardware DNN modules. HALF automates both the exploration process and the implementation of optimized solutions on a target FPGA platform for various applications. We demonstrate the performance of HALF on a medical use case for arrhythmia detection for three different design goals, i.e., low-energy, low-power, and high-throughput. Our FPGA implementation outperforms a TensorRT optimized model on an Nvidia Jetson platform in both throughput and energy consumption. Jonas Ney, Dominik Marek Loroch, Vladimir Rybalkin, Nico Weber, Jens Krüger 0004, Norbert Wehn |
FPL | 1 |