EDBT 2026 Demo / reviewers in the wild / expert
Elias Trommer
dblp:307/3116
· DBLP profile ↗
5ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-5671-3291ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fast Retraining of Approximate CNNs for High AccuracyabstractOne technique for approximating neural networks (NNs) when deploying to resource-constrained systems is the use of approximate multiplications. Giving up full mathematical accuracy opens new opportunities for more efficient hardware implementations. Modeling the effects of inaccurate hardware already in the training stage improves performance but significantly slows down the training due to expensive type conversions and memory access operations. We propose a method to speed up the simulation of inaccurate hardware by using a composition of floating-point functions. Both an analytical and a data-driven method for finding these functions are provided. We further provide a study and implementation of per-channel quantization, a scheme that enhances the granularity of converting NN parameters to integers. This helps boost the application’s accuracy. In our evaluation, our floating-point models achieve up to a$4 \times $speed-up over the commonly used lookup table implementation, while providing a high-fidelity simulation of the target function. Extending quantization with per-channel granularity yields a median accuracy improvement of 0.87 p.p. for ResNet8/CIFAR10 with 4-bit weight quantization in combination with hardware using approximate multipliers (AMs). Our extended software toolkit for the study of AMs in PyTorch is publicly available and provides a variety of building blocks for applying inaccurate product functions to NNs. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Smaller Together: Groupwise Encoding of Sparse Neural NetworksabstractWith the drive towards ever more intelligent devices, neural networks are deployed on smaller and smaller systems. For these embedded microcontrollers, memory consumption becomes a significant challenge. We propose multiple encoding schemes that convert the decrease in parameter counts, achieved through unstructured pruning, into tangible memory savings. We first discuss a sparse encoding scheme for arbitrary sparse matrices that is based on encoding offsets from a predicted even spacing of elements in a row. The compression rate of this scheme is improved further by identifying groups of elements which can be encoded with even lower overhead. Both methods are combined into a hybrid scheme which encodes arbitrary sparse matrices with low overhead, while allowing for parallel access to multiple elements in a row at once—an important feature for using the scheme on the latest generation of microcontrollers with parallel SIMD capabilities. Our scheme compresses sparse models to below the size of their dense counterparts for sparsities as low as 30% and reduces model size by 32.4% and 26.4% at less than one percentage point of accuracy loss for two convolutional neural network tasks in our evaluation. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | High-Throughput Approximate Multiplication Models in PyTorchabstractApproximate multipliers can reduce the resource consumption of neural network accelerators. To study their effects on an application, they need to be simulated during network training. We develop simulation models for a common class of approximate multipliers. Our models speed up execution by replacing time-consuming type conversions and memory accesses with fast floating-point arithmetic. Across six different neural network architectures, these models increase throughput by 2.7× over the commonly used array lookup while recreating behavioral simulation with high fidelity. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
DDECS | 1 |
| 2022 | Combining Gradients and Probabilities for Heterogeneous Approximation of Neural NetworksabstractThis work explores the search for heterogeneous approximate multiplier configurations for neural networks that produce high accuracy and low energy consumption. We discuss the validity of additive Gaussian noise added to accurate neural network computations as a surrogate model for behavioral simulation of approximate multipliers. The continuous and differentiable properties of the solution space spanned by the additive Gaussian noise model are used as a heuristic that generates meaningful estimates of layer robustness without the need for combinatorial optimization techniques. Instead, the amount of noise injected into the accurate computations is learned during network training using backpropagation. A probabilistic model of the multiplier error is presented to bridge the gap between the domains; the model estimates the standard deviation of the approximate multiplier error, connecting solutions in the additive Gaussian noise space to actual hardware instances. Our experiments show that the combination of heterogeneous approximation and neural network retraining reduces the energy consumption for multiplications by 70% to 79% for different ResNet variants on the CIFAR-10 dataset with a Top-1 accuracy loss below one percentage point. For the more complex Tiny ImageNet task, our VGG16 model achieves a 53 % reduction in energy consumption with a drop in Top-5 accuracy of 0.5 percentage points. We further demonstrate that our error model can predict the parameters of an approximate multiplier in the context of the commonly used additive Gaussian noise (AGN) model with high accuracy. Our software implementation is available under https://github.com/etrommer/agn-approx. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
ICCAD | 1 |
| 2021 | dCSR: A Memory-Efficient Sparse Matrix Representation for Parallel Neural Network InferenceabstractReducing the memory footprint of neural networks is a crucial prerequisite for deploying them in small and low-cost embedded devices. Network parameters can often be reduced significantly through pruning. We discuss how to best represent the indexing overhead of sparse networks for the coming generation of Single Instruction, Multiple Data (SIMD)-capable microcontrollers. From this, we develop Delta-Compressed Storage Row (dCSR), a storage format for sparse matrices that allows for both low overhead storage and fast inference on embedded systems with wide SIMD units. We demonstrate our method on an ARM Cortex-M55 MCU prototype with M-Profile Vector Extension (MVE). A comparison of memory consumption and throughput shows that our method achieves competitive compression ratios and increases throughput over dense methods by up to$2.9\times$for sparse matrix-vector multiplication (SpMV)-based kernels and$1.06\times$for sparse matrix-matrix multiplication (SpMM). This is accomplished through handling the generation of index information directly in the SIMD unit, leading to an increase in effective memory bandwidth. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
ICCAD | 1 |