VLDB 2026 Research / reviewers in the wild / expert
Bernd Waschneck
dblp:195/4539
· DBLP profile ↗
13ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0003-0294-8594ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fast Retraining of Approximate CNNs for High AccuracyabstractOne technique for approximating neural networks (NNs) when deploying to resource-constrained systems is the use of approximate multiplications. Giving up full mathematical accuracy opens new opportunities for more efficient hardware implementations. Modeling the effects of inaccurate hardware already in the training stage improves performance but significantly slows down the training due to expensive type conversions and memory access operations. We propose a method to speed up the simulation of inaccurate hardware by using a composition of floating-point functions. Both an analytical and a data-driven method for finding these functions are provided. We further provide a study and implementation of per-channel quantization, a scheme that enhances the granularity of converting NN parameters to integers. This helps boost the application’s accuracy. In our evaluation, our floating-point models achieve up to a$4 \times $speed-up over the commonly used lookup table implementation, while providing a high-fidelity simulation of the target function. Extending quantization with per-channel granularity yields a median accuracy improvement of 0.87 p.p. for ResNet8/CIFAR10 with 4-bit weight quantization in combination with hardware using approximate multipliers (AMs). Our extended software toolkit for the study of AMs in PyTorch is publicly available and provides a variety of building blocks for applying inaccurate product functions to NNs. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | Automating application-driven customization of ASIPs: A survey
Eslam Hussein, Bernd Waschneck, Christian Mayr 0001 |
J. Syst. Archit. | 2 |
| 2024 | Smaller Together: Groupwise Encoding of Sparse Neural NetworksabstractWith the drive towards ever more intelligent devices, neural networks are deployed on smaller and smaller systems. For these embedded microcontrollers, memory consumption becomes a significant challenge. We propose multiple encoding schemes that convert the decrease in parameter counts, achieved through unstructured pruning, into tangible memory savings. We first discuss a sparse encoding scheme for arbitrary sparse matrices that is based on encoding offsets from a predicted even spacing of elements in a row. The compression rate of this scheme is improved further by identifying groups of elements which can be encoded with even lower overhead. Both methods are combined into a hybrid scheme which encodes arbitrary sparse matrices with low overhead, while allowing for parallel access to multiple elements in a row at once—an important feature for using the scheme on the latest generation of microcontrollers with parallel SIMD capabilities. Our scheme compresses sparse models to below the size of their dense counterparts for sparsities as low as 30% and reduces model size by 32.4% and 26.4% at less than one percentage point of accuracy loss for two convolutional neural network tasks in our evaluation. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | High-Throughput Approximate Multiplication Models in PyTorchabstractApproximate multipliers can reduce the resource consumption of neural network accelerators. To study their effects on an application, they need to be simulated during network training. We develop simulation models for a common class of approximate multipliers. Our models speed up execution by replacing time-consuming type conversions and memory accesses with fast floating-point arithmetic. Across six different neural network architectures, these models increase throughput by 2.7× over the commonly used array lookup while recreating behavioral simulation with high fidelity. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
DDECS | 2 |
| 2022 | Industry-track: Towards Agile Design of Neural Processing UnitabstractMore and more specialized processors, known as Neural Processing Units (NPUs), have been or are being built for deep neural network inference. Design and optimization of this kind of processor are inseparable from the deep learning ecosystem and corresponding underlying software. This HW/SW co-design requirement poses challenges for designers. Therefore, in this work, we experiment with an agile development method to shorten the development cycles of NPUs. We utilize Chisel for hardware design and develop a custom Chisel backend for generating cycle-accurate simulators with C++/Python APIs. On top of the simulator, we built a Python software stack for software development, performance evaluation, and simulation-based verification. The proposed method is purely software and does not involve real hardware, thus allowing the integration of software agile development methods into digital designs. In the experiments, we show how it helps us identify inherent hardware limitations and how it shortens our development cycles. Binyi Wu, Wolfgang Furtner, Bernd Waschneck, Christian Mayr 0001 |
CODES+ISSS | 3 |
| 2022 | Neural Architecture Search for Low-Precision Neural Networks
Binyi Wu, Bernd Waschneck, Christian Mayr 0001 |
ICANN (4) | 2 |
| 2022 | Combining Gradients and Probabilities for Heterogeneous Approximation of Neural NetworksabstractThis work explores the search for heterogeneous approximate multiplier configurations for neural networks that produce high accuracy and low energy consumption. We discuss the validity of additive Gaussian noise added to accurate neural network computations as a surrogate model for behavioral simulation of approximate multipliers. The continuous and differentiable properties of the solution space spanned by the additive Gaussian noise model are used as a heuristic that generates meaningful estimates of layer robustness without the need for combinatorial optimization techniques. Instead, the amount of noise injected into the accurate computations is learned during network training using backpropagation. A probabilistic model of the multiplier error is presented to bridge the gap between the domains; the model estimates the standard deviation of the approximate multiplier error, connecting solutions in the additive Gaussian noise space to actual hardware instances. Our experiments show that the combination of heterogeneous approximation and neural network retraining reduces the energy consumption for multiplications by 70% to 79% for different ResNet variants on the CIFAR-10 dataset with a Top-1 accuracy loss below one percentage point. For the more complex Tiny ImageNet task, our VGG16 model achieves a 53 % reduction in energy consumption with a drop in Top-5 accuracy of 0.5 percentage points. We further demonstrate that our error model can predict the parameters of an approximate multiplier in the context of the commonly used additive Gaussian noise (AGN) model with high accuracy. Our software implementation is available under https://github.com/etrommer/agn-approx. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
ICCAD | 2 |
| 2022 | Prototyping of Low-Cost Configurable Sparse Neural Processing Unit with Buffer and Mixed-Precision Reshapeable MAC ArrayabstractMore recently, it has become possible to run deep learning algorithms on edge devices such as microcontrollers due to continuous improvements in neural network optimization algorithms such as quantization and neural architecture search. Nonetheless, most of the embedded hardware available today still falls short of the requirements of running deep neural networks. As a result, specialized processors have emerged to improve the inference efficiency of deep learning algorithms. However, most are not for edge applications that require efficient and low-cost hardware. Therefore, we design and prototype a low-cost configurable sparse Neural Processing Unit (NPU). The NPU has a built-in buffer and a reshapable mixed-precision multiply-accumulator (MAC) array. The computing and memory resources of the NPU are parameterized, and different NPUs can be derived. Besides, users can also conFigure the NPU at runtime to fully utilize the resources. In our experiments, the 200MHz NPU with only 32 MACs is more than 32 times faster than the 400MHzSTM32H7 when inferring MobileNet-Vl. Besides, the yielded NPUs can achieve roofline or even beyond roofline performance. The buffer and reshapeable MAC array push the NPU’s attainable performance to the roofline, while the feature of supporting sparsity allows the NPU to obtain performance beyond the roofline. Binyi Wu, Wolfgang Furtner, Bernd Waschneck, Christian Mayr 0001 |
ICPADS | 3 |
| 2022 | Hardware-Efficient Ultrasonic Entrance Counting: Comparing Different Machine Learning ApproachesabstractIn this work, the classification of walking direction based on ultrasonic signals has been examined for entrance counting. Feed-forward and recurrent neural network architectures as well as simpler machine learning techniques have been investigated and compared with classical signal processing techniques.Using only a single ultrasonic receiver, the focus was set on the development of a hardware-efficient system concept. Different ultrasonic measurement methods in time and frequency domain have been compared with the perspective of a holistic energy optimization. The analysis of the system’s hardware efficiency was completed by an estimation of algorithmic latency, energy and storage consumption based on the arithmetic of the classification algorithms. All algorithms showed an estimated energy consumption of less than 10 μJ for a single inference on a state-of-the-art implementation of an ARM® Cortex® M4F micro-controller, which was found to be negligible compared to the energy of the measurement principle. Compared to other sensor types and multi-sensor systems, a state-of-the-art test accuracy of 99.72% could be achieved for differentiating between the two entrance directions of a present person and the absence of a person. Tim Langer, Bernd Waschneck, Johannes Partzsch, Florian Kelber, Christian Mayr 0001 |
ICPR | 2 |
| 2022 | Convolutional Neural Networks Quantization with Double-Stage Squeeze-and-ThresholdabstractIt has been proven that, compared to using 32-bit floating-point numbers in the training phase, Deep Convolutional Neural Networks (DCNNs) can operate with low-precision during inference, thereby saving memory footprint and power consumption. However, neural network quantization is always accompanied by accuracy degradation. Here, we propose a quantization method called double-stage Squeeze-and-Threshold (double-stage ST) to close the accuracy gap with full-precision models. While accurate colors in pictures can be pleasing to the viewer, they are not necessary for distinguishing objects. The era of black and white television proves this idea. As long as the limited colors are filled reasonably for different objects, the objects can be well identified and distinguished. Our method utilizes the attention mechanism to adjust the activations and learn the thresholds to distinguish objects (features). We then divide the numerically rich activations into intervals (a limited variety of numerical values) by the learned thresholds. The proposed method supports both binarization and multi-bit quantization. Our method achieves state-of-the-art results. In binarization, ReActNet [Z. Liu, Z. Shen, S. Li, K. Helwegen, D. Huang and K. Cheng, arXiv:abs/2106.11309 ] trained with our method outperforms the previous state-of-the-art result by 0.2 percentage points. Whereas in multi-bit quantization, the top-1 accuracy of the 3-bit ResNet-18 [K. He, X. Zhang, S. Ren and J. Sun, Deep residual learning for image recognition, 2016 IEEE Conf. Computer Vision and Pattern Recognition, CVPR 2016, 27–30 June 2016, Las Vegas, NV, USA (IEEE Computer Society, 2016), pp. 770–778] model exceeds the top-1 accuracy of its full-precision baseline model by 0.4 percentage points. The double-stage ST activation quantization method is easy to apply by inserting it before the convolution. Besides, the double-stage ST is detachable after training and introducing no computational cost in inference. Binyi Wu, Bernd Waschneck, Christian Mayr 0001 |
Int. J. Neural Syst. | 2 |
| 2021 | Squeeze-and-Threshold Based Quantization for Low-Precision Neural Networks
Binyi Wu, Bernd Waschneck, Christian Mayr 0001 |
EANN | 2 |
| 2021 | dCSR: A Memory-Efficient Sparse Matrix Representation for Parallel Neural Network InferenceabstractReducing the memory footprint of neural networks is a crucial prerequisite for deploying them in small and low-cost embedded devices. Network parameters can often be reduced significantly through pruning. We discuss how to best represent the indexing overhead of sparse networks for the coming generation of Single Instruction, Multiple Data (SIMD)-capable microcontrollers. From this, we develop Delta-Compressed Storage Row (dCSR), a storage format for sparse matrices that allows for both low overhead storage and fast inference on embedded systems with wide SIMD units. We demonstrate our method on an ARM Cortex-M55 MCU prototype with M-Profile Vector Extension (MVE). A comparison of memory consumption and throughput shows that our method achieves competitive compression ratios and increases throughput over dense methods by up to$2.9\times$for sparse matrix-vector multiplication (SpMV)-based kernels and$1.06\times$for sparse matrix-matrix multiplication (SpMM). This is accomplished through handling the generation of index information directly in the SIMD unit, leading to an increase in effective memory bandwidth. Elias Trommer, Bernd Waschneck, Akash Kumar 0001 |
ICCAD | 2 |
| 2020 | Small-Footprint Keyword Spotting on Raw Audio Data with Sinc-ConvolutionsabstractKeyword Spotting (KWS) enables speech-based user interaction on smart devices. Always-on and battery-powered application scenarios for smart devices put constraints on hardware resources and power consumption, while also demanding high accuracy as well as real-time capability. Previous architectures first extracted acoustic features and then applied a neural network to classify keyword probabilities, optimizing towards memory footprint and execution time.Compared to previous publications, we took additional steps to reduce power and memory consumption without reducing classification accuracy. Power-consuming audio preprocessing and data transfer steps are eliminated by directly classifying from raw audio. For this, our end-to-end architecture extracts spectral features using parametrized Sinc-convolutions. Its memory footprint is further reduced by grouping depthwise separable convolutions. Our network achieves the competitive accuracy of 96.4% on Google's Speech Commands test set with only 62k parameters. Simon Mittermaier, Ludwig Kurzinger, Bernd Waschneck, Gerhard Rigoll |
ICASSP | 3 |