EDBT 2026 Demo / reviewers in the wild / expert
Matthias Wess
dblp:189/5615
· DBLP profile ↗
9ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-1877-4114ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lanes Are Not Enough: Enhancing Trajectory Prediction in Intralogistics Through Detailed Environmental ContextabstractTrajectory prediction is an essential component of the perception stack in autonomous mobile robots (AMRs). AMRs operate in complex environments where their movements are influenced by various environment elements, such as racks and storage locations. Therefore, accurate and efficient trajectory prediction for intralogistics requires detailed environment modeling that goes beyond the lane-based context mainly used in road traffic methods. We propose the addition of a new environment context encoder module that can be seamlessly integrated into state-of-the-art autonomous driving systems. Our approach, tailored to the specific challenges of intralogistics, achieves highly accurate predictions using compact and efficient baseline networks. Alexander Prutsch, Matthias Wess, Horst Possegger |
IROS | 2 |
| 2025 | Efficient and interpretable raw audio classification with diagonal state space modelsabstractAbstract State Space Models have achieved good performance on long sequence modeling tasks such as raw audio classification. Their definition in continuous time allows for discretization and operation of the network at different sampling rates. However, this property has not yet been utilized to decrease the computational demand on a per-layer basis. We propose a family of hardware-friendly S-Edge models with a layer-wise downsampling approach to adjust the temporal resolution between individual layers. Applying existing methods from linear control theory allows us to analyze state/memory dynamics and provides an understanding of how and where to downsample. Evaluated on the Google Speech Command dataset, our autoregressive/causal S-Edge models range from 8–141k parameters at 90–95% test accuracy in comparison to a causal S5 model with 208k parameters at 95.8% test accuracy. Using our C++17 header-only implementation on an ARM Cortex-M4F the largest model requires 103 sec. inference time with 95.19% test accuracy, and the smallest model with 88.01% test accuracy, requires 0.29 sec. Our solutions cover a design space that spans 17x in model size, 358x in inference latency, and 7.18 percentage points in accuracy. Matthias Bittner, Daniel Schnöll, Matthias Wess, Axel Jantsch |
Mach. Learn. | 3 |
| 2024 | Special Session: Estimation and Optimization of DNNs for Embedded PlatformsabstractSeveral state of the art estimation and optimization techniques for CNNs and LLMs on embedded devices are summarized. For LLMs an Activation-aware Weight Quantization and on-the-fly dequantization techniques is presented. For CNNs various pruning algorithms and an integrated optimization and implementation flow is discussed. To estimate inference latency of CNNs on specific hardware platforms, three different techniques are reviewed: A mixed analytic-stochastic model, an analytic model based on step-wise linear functions, and a method that uses a detailed architecture description of the hardware. Axel Jantsch, Song Han 0003, Lin Meng 0001, Oliver Bringmann 0001, Haotian Tang, Shang Yang, Matthias Wess, Martin Lechner |
CODES+ISSS | 8 |
| 2023 | Fast, Quantization Aware DNN Training for Efficient HW ImplementationabstractQuantization of Deep Neural Networks is a central technique to reduce the computation load in embedded devices. Even in quantized Deep Neural Networks (DNNs), the scaler/rescaler following a convolution or dense layer often requires a high bit width multiplication and a shift. Previous work has proposed to remove the multiplier by restricting the quantization method. We propose a Quantisation Aware Training (QAT) approach, which explicitly models the rescaler during training, eliminating the limitations of quantization functions and achieving a 30–35% improvement in training time and a significant reduction in memory requirements compared to the state-of-the-art. GitHub: https://github.com/embedded-machine-learning/FastQATforPOTRescaler Daniel Schnöll, Matthias Wess, Matthias Bittner, Maximilian Götzinger, Axel Jantsch |
DSD | 2 |
| 2023 | Energy Profiling of DNN AcceleratorsabstractThis paper introduces a novel methodology for assessing the energy efficiency of neural network accelerators at both layer and network granularity. The approach involves extracting per-layer timing reports from recorded power profiles. The power and energy consumption of three prominent neural network accelerators, namely the Intel Neural Compute Stick 2, the Coral Edge TPU, and the NXP i.MX8M Plus is evaluated for three different Deep Neural Networks (DNNs) using this method. The study investigates the relationship between decreasing sampling frequencies and the average error, as well as the detailed energy consumption of individual DNN layers and layer types. The findings reveal that latency outperforms the number of operations per layer as a predictor for both overall and dynamic energy, with errors of 10 % and 100 % respectively. The main conclusions are: a sampling frequency of 200 kHz is necessary to achieve an average error of 5 %; the number of operations is an inadequate predictor of energy consumption; and specific hardware settings significantly influence power and energy consumption, emphasizing the need for their consideration in estimation. Matthias Wess, Dominik Dallinger, Daniel Schnöll, Matthias Bittner, Maximilian Götzinger, Axel Jantsch |
DSD | 1 |
| 2022 | The SELENE Deep Learning Acceleration Framework for Safety-related ApplicationsabstractThe goal of the H2020 SELENE project is the development of a flexible computing platform for autonomous applications that includes built-in hardware support for safety. The SELENE computing platform is an open-source RISC-V heterogeneous multicore system-on-chip (SoC) that includes 6 NOEL-V RISC-V cores and artificial intelligence accelerators. In this paper, we describe the approach followed in the SELENE project to accelerate neural network inference processes. Our intermediate results show that both the FPGA and ASIC accel-erators provide real-time inference performance for the analyzed network models at a reasonable implementation cost. Laura Medina, Salva Carrion, Pablo Andreu, Tomás Picornell, José Flich, Carles Hernández 0001, Michael Sandoval, Markel Sainz, Charles-Alexis Lefebvre, Martin Rönnbäck, Martin Matschnig, Matthias Wess, Herbert Taucher |
DATE | 12 |
| 2018 | Weighted Quantization-Regularization in DNNs for Weight Memory Minimization Toward HW ImplementationabstractDeployment of deep neural networks on hardware platforms is often constrained by limited on-chip memory and computational power. The proposed weight quantization offers the possibility of optimizing weight memory alongside transforming the weights to hardware friendly data types. We apply dynamic fixed point (DFP) and power-of-two (Po2) quantization in conjunction with layer-wise precision scaling to minimize the weight memory. To alleviate accuracy degradation due to precision scaling, we employ quantization-aware fine-tuning. For fine-tuning, quantization-regularization (QR) and weighted QR are introduced to force the trained quantization by adding the distance of the weights to the desired quantization levels as a regularization term to the loss-function. While DFP quantization performs better when allowing different bit-widths for each layer, Po2 quantization in combination with retraining allows higher compression rates for equal bit-width quantization. The techniques are verified on an all-convolutional network. With accuracy degradation of 0.10% points, for DFP with layer-wise precision scaling we achieve compression ratios of 7.34 for CIFAR-10, 4.7 for CIFAR-100, and 9.33 for SVHN dataset. Matthias Wess, Sai Manoj Pudukotai Dinakarrao, Axel Jantsch |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | Neural network based ECG anomaly detection on FPGA and trade-off analysisabstractThis paper presents FPGA-based ECG arrhythmia detection using an Artificial Neural Network (ANN). The objective is to implement a neural network based machine learning algorithm on FPGA to detect anomalies in ECG signals, with a better performance and accuracy, compared to statistical methods. An implementation with Principal Component Analysis (PCA) for feature reduction and a multi-layer perceptron (MLP) for classification, proved superior to other algorithms. For implementation on FPGA, the effects of several parameters and simplification on performance, accuracy and power consumption were studied. Piecewise linear approximation for activation functions and fixed point implementation were effective methods to reduce the amount of needed resources. The resulting neural network with twelve inputs and six neurons in the hidden layer, achieved, in spite of the simplifications, the same overall accuracy as simulations with floating point number representation. An accuracy of 99.82% was achieved on average for the MIT-BIH database. Matthias Wess, Sai Manoj Pudukotai Dinakarrao, Axel Jantsch |
ISCAS | 1 |
| 2016 | ICT emulation platform setup demonstration of smart grid component prototype examplesabstractThe shift towards massively distributed energy generation demands more decentralized flexibility to meet strict power quality constraints of the electric grid. A cyber-physical system such as a smart grid can provide increased flexibility by utilizing additional information and communication technologies to better monitor the medium and low voltage distribution networks and to actively control grid-connected resources, ranging from loads to distributed generation, to electric mobility but at the cost of increased complexity. Essential future functionalities such as dynamic management of line use, fault detection and fast service restoration are only possible with appropriate sensors and actuators in place. These missing sensors and actuators on the distribution level are being developed today. This paper presents a standards based, low cost, open source, ICT emulation platform setup to test necessary networking concepts of these smart grid component prototypes already in various stages of development. Preliminary development results of the first example applications chosen: Customer Energy Management System, Smart Breaker, and Smart Meter, are shown in this work in progress paper. Marcus Meisel, Stefan Wilker, Matthias Wess, Alexander Wendt, Thilo Sauter, Georg Kienesberger |
ETFA | 3 |