Jens Mehnert

dblp:09/5738 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-0079-0036ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks
abstract
Vision Transformers (ViTs) have emerged as the state-of-the-art models in various Computer Vision (CV) tasks, but their high computational and resource demands pose significant challenges. While Mixture-of-Experts (MoE) can make these models more efficient, they often require costly retraining or even training from scratch. Recent developments aim to reduce these computational costs by leveraging pretrained networks. These have been shown to produce sparse activation patterns in the Multi-Layer Perceptrons (MLPs) of the encoder blocks, allowing for conditional activation of only relevant subnetworks for each sample.Building on this idea, we propose a new method to construct MoE variants from pretrained models. Our approach extracts expert subnetworks from the model’s MLP layers post-training in two phases. First, we cluster output activations to identify distinct activation patterns. In the second phase, we use these clusters to extract the corresponding subnetworks responsible for producing them. On ImageNet-1k recognition tasks, we demonstrate that these extracted experts can perform surprisingly well out of the box and require only minimal fine-tuning to regain 98% of the original performance, all while reducing MACs and model size, by up to 36% and 32% respectively.
Uranik Berisha, Jens Mehnert, Alexandru Condurache
CVPR2
2025 PROM: Prioritize Reduction of Multiplications Over Lower Bit-Widths for Efficient CNNs
abstract
Convolutional neural networks (CNNs) are crucial for computer vision tasks on resource-constrained devices. Quantization effectively compresses these models, reducing storage size and energy cost. However, in modern depthwise-separable architectures, the computational cost is distributed unevenly across its components, with pointwise operations being the most expensive. By applying a general quantization scheme to this imbalanced cost distribution, existing quantization approaches fail to fully exploit potential efficiency gains. To this end, we introduce PROM, a straightforward approach for quantizing modern depthwise-separable convolutional networks by selectively using two distinct bit-widths. Specifically, pointwise convolutions are quantized to ternary weights, while the remaining modules use 8-bit weights, which is achieved through a simple quantization-aware training procedure. Additionally, by quantizing activations to 8-bit, our method transforms pointwise convolutions with ternary weights into int8 additions, which enjoy broad support across hardware platforms and effectively eliminates the need for expensive multiplications. Applying PROM to MobileNetV2 reduces the model’s energy cost by more than an order of magnitude (23.9×) and its storage size by 2.7× compared to the float16 baseline while retaining similar classification performance on ImageNet. Our method advances the Pareto frontier for energy consumption vs. top-1 accuracy for quantized convolutional models on ImageNet. PROM addresses the challenges of quantizing depthwise-separable convolutional networks to both ternary and 8-bit weights, offering a simple way to reduce energy cost and storage size.
Lukas Meiner, Jens Mehnert, Alexandru Condurache
ECAI2
2025 Variance-Based Pruning for Accelerating and Compressing Trained Networks
abstract
Increasingly expensive training of ever larger models such as Vision Transfomers motivate reusing the vast library of already trained state-of-the-art networks. However, their latency, high computational costs and memory demands pose significant challenges for deployment, especially on resource-constrained hardware. While structured pruning methods can reduce these factors, they often require costly retraining, sometimes for up to hundreds of epochs, or even training from scratch to recover the lost accuracy resulting from the structural modifications. Maintaining the provided performance of trained models after structured pruning and thereby avoiding extensive retraining remains a challenge. To solve this, we introduce Variance-Based Pruning, a simple and structured one-shot pruning technique for efficiently compressing networks, with minimal finetuning. Our approach first gathers activation statistics, which are used to select neurons for pruning. Simultaneously the mean activations are integrated back into the model to preserve a high degree of performance. On ImageNet-1k recognition tasks, we demonstrate that directly after pruning DeiT-Base retains over 70% of its original performance and requires only 10 epochs of fine-tuning to regain 99% of the original accuracy while simultaneously reducing MACs by 35% and model size by 36%, thus speeding up the model by 1.44x. The code is available at: https://github.com/boschresearch/variance-based-pruning
Uranik Berisha, Jens Mehnert, Alexandru Condurache
ICCV2
2024 CNN Mixture-of-Depths
Rinor Cakaj, Jens Mehnert, Bin Yang 0009
ACCV (7)2
2024 Squeeze-and-Remember Block
abstract
Convolutional Neural Networks (CNNs) are important for many machine learning tasks. They are built with different types of layers: convolutional layers that detect features, dropout layers that help to avoid over-reliance on any single neuron, and residual layers that allow the reuse of features. However, CNNs lack a dynamic feature retention mechanism similar to the human brain's memory, limiting their ability to use learned information in new contexts. To bridge this gap, we introduce the “Squeeze-and-Remember” (SR) block, a novel architectural unit that gives CNNs dynamic memory-like functionalities. The SR block selectively memorizes important features during training, and then adaptively re-applies these features during inference. This improves the network's ability to make contextually informed predictions. Empirical results on ImageNet and Cityscapes datasets demonstrate the SR block's efficacy: integration into ResNet50 improved top-1 validation accuracy on ImageNet by 0.52% over dropout2d alone, and its application in DeepLab v3 increased mean Intersection over Union in Cityscapes by 0.20%. These improvements are achieved with minimal computational overhead. This show the SR block's potential to enhance the capabilities of CNNs in image processing tasks.
Rinor Cakaj, Jens Mehnert, Bin Yang 0009
ICMLA2
2024 Spectral Wavelet Dropout: Regularization in the Wavelet Domain
abstract
Regularization techniques help prevent overfitting and therefore improve the ability of convolutional neural net-works (CNNs) to generalize. One reason for overfitting is the complex co-adaptations among different parts of the network, which make the CNN dependent on their joint response rather than encouraging each part to learn a useful feature representation independently. Frequency domain manipulation is a powerful strategy for modifying data that has temporal and spatial coherence by utilizing frequency decomposition. This work intro-duces Spectral Wavelet Dropout (SWD), a novel regularization method that includes two variants: ID-SWD and 2D-SWD. These variants improve CNN generalization by randomly dropping detailed frequency bands in the discrete wavelet decomposition of feature maps. Our approach distinguishes itself from the pre-existing Spectral “Fourier” Dropout (2D-SFD), which eliminates coefficients in the Fourier domain. Notably, SWD requires only a single hyperparameter, unlike the two required by SFD. We also extend the literature by implementing a one-dimensional version of Spectral “Fourier” Dropout (lD-SFD), setting the stage for a comprehensive comparison. Our evaluation shows that both ID and 2D SWD variants have competitive performance on CIFAR-IO/IOO benchmarks relative to both ID-SFD and 2D-SFD. Specifically, ID-SWD has a significantly lower computational complexity compared to ID/2D-SFD. In the Pascal VOC Object Detection benchmark, SWD variants surpass ID-SFD and 2D-SFD in performance and demonstrate lower computational complexity during training.
Rinor Cakaj, Jens Mehnert, Bin Yang 0009
ICMLA2
2023 Weight Compander: A Simple Weight Reparameterization for Regularization
abstract
Regularization is a set of techniques that are used to improve the generalization ability of deep neural networks. In this paper, we introduce weight compander (WC), a novel effective method to improve generalization by reparameterizing each weight in deep neural networks using a nonlinear function. It is a general, intuitive, cheap and easy to implement method, which can be combined with various other regularization techniques. Large weights in deep neural networks are a sign of a more complex network that is overfitted to the training data. Moreover, regularized networks tend to have a greater range of weights around zero with fewer weights centered at zero. We introduce a weight reparameterization function which is applied to each weight and implicitly reduces overfitting by restricting the magnitude of the weights while forcing them away from zero at the same time. This leads to a more democratic decision-making in the network. Firstly, individual weights cannot have too much influence in the prediction process due to the restriction of their magnitude. Secondly, more weights are used in the prediction process, since they are forced away from zero during the training. This promotes the extraction of more features from the input data and increases the level of weight redundancy, which makes the network less sensitive to statistical differences between training and test data. From an optimizational point of view, the second effect of WC can be seen as a reactivation of “dead” (near zero) weights to participate in the training. This increases the probability to find an ensemble of weights which performs better in the given task. We extend our method to learn the hyperparameters of the introduced weight reparameterization function. This avoids hyperparameter search and gives the network the opportunity to align the weight reparameterization with the training progress. We show experimentally that using weight compander in addition to standard regularization methods improves the performance of neural networks. Furthermore, we empirically analyze the weight distribution with and without weight compander after training to confirm the companding effects of our method on the weights.
Rinor Cakaj, Jens Mehnert, Bin Yang 0009
IJCNN2
2023 Spectral Batch Normalization: Normalization in the Frequency Domain
abstract
Regularization is a set of techniques that are used to improve the generalization ability of deep neural networks. In this paper, we introduce spectral batch normalization (SBN), a novel effective method to improve generalization by normalizing feature maps in the frequency (spectral) domain. The activations of residual networks without batch normalization (BN) tend to explode exponentially in the depth of the network at initialization. This leads to extremely large feature map norms even though the parameters are relatively small. These explosive dynamics can be very detrimental to learning. BN makes weight decay regularization on the scaling factors$\gamma,\beta$approximately equivalent to an additive penalty on the norm of the feature maps, which prevents extremely large feature map norms to a certain degree. It was previously shown that preventing explosive growth at the final layer at initialization and during training in ResNets can recover a large part of Batch Normalization's generalization boost. However, we show experimentally that, despite the approximate additive penalty of BN, feature maps in deep neural networks (DNNs) tend to explode at the beginning of the training and that feature maps of DNNs contain large values during the whole training. This phenomenon also occurs in a weakened form in non-residual networks. Intuitively, it is not preferred to have large values in feature maps since they have too much influence on the prediction in contrast to other parts of the feature map. SBN addresses large feature maps by normalizing them in the frequency domain. In our experiments, we empirically show that SBN prevents exploding feature maps at initialization and large feature map values during the training. Moreover, the normalization of feature maps in the frequency domain leads to more uniform distributed frequency components. This discourages the DNNs to rely on single frequency components of feature maps. These, together with other effects (e.g. noise injection, scaling and shifting of the feature map) of SBN, have a regularizing effect on the training of residual and non-residual networks. We show experimentally that using SBN in addition to standard regularization methods improves the performance of DNNs by a relevant margin, e.g. ResNet50 on CIFAR-100 by 2.31%, on ImageNet by 0.71% (from 76.80% to 77.51%) and VGG19 on CIFAR-100 by 0.66%.
Rinor Cakaj, Jens Mehnert, Bin Yang 0009
IJCNN2
2022 Interspace Pruning: Using Adaptive Filter Representations to Improve Training of Sparse CNNs
abstract
Unstructured pruning is well suited to reduce the memory footprint of convolutional neural networks (CNNs), both at training and inference time. CNNs contain parameters arranged in K x K filters. Standard unstructured pruning (SP) reduces the memory footprint of CNNs by setting filter elements to zero, thereby specifying a fixed subspace that constrains the filter. Especially if pruning is applied before or during training, this induces a strong bias. To overcome this, we introduce interspace pruning (IP), a general tool to improve existing pruning methods. It uses filters represented in a dynamic interspace by linear combinations of an underlying adaptive filter basis (FB). For IP, FB coefficients are set to zero while un-pruned coefficients and FBs are trained jointly. In this work, we provide mathematical evidence for IP's superior performance and demonstrate that IP outperforms SP on all tested state-of-the-art unstructured pruning methods. Especially in challenging situations, like pruning for ImageNet or pruning to high sparsity, IP greatly exceeds SP with equal runtime and parameter costs. Finally, we show that advances of IP are due to improved trainability and superior generalization ability.
Paul Wimmer, Jens Mehnert, Alexandru Condurache
CVPR2
2021 COPS: Controlled Pruning Before Training Starts
abstract
State-of-the-art deep neural network (DNN) pruning techniques, applied one-shot before training starts, evaluate sparse architectures with the help of a single criterion-called pruning score. Pruning weights based on a solitary score works well for some architectures and pruning rates but may also fail for other ones. As a common baseline for pruning scores, we introduce the notion of a generalized synaptic score (GSS). In this work we do not concentrate on a single pruning criterion, but provide a framework for combining arbitrary GSSs to create more powerful pruning strategies. These COmbined Pruning Scores (COPS) are obtained by solving a constrained optimization problem. Optimizing for more than one score prevents the sparse network to overly specialize on an individual task, thus COntrols Pruning before training Starts. The combinatorial optimization problem given by COPS is relaxed on a linear program (LP). This LP is solved analytically and determines a solution for COPS. Furthermore, an algorithm to compute it for two scores numerically is proposed and evaluated. Solving COPS in such a way has lower complexity than the best general LP solver. In our experiments we compared pruning with COPS against state-of-the-art methods for different network architectures and image classification tasks and obtained improved results.
Paul Wimmer, Jens Mehnert, Alexandru Condurache
IJCNN2
2020 FreezeNet: Full Performance by Reduced Storage Costs
Paul Wimmer, Jens Mehnert, Alexandru Condurache
ACCV (6)2
2011 Evaluation of different approaches for road course estimation using imaging radar
abstract
This work presents three imaging radar sensor approaches to estimate road courses, needed by intelligent vehicle systems such as active cruise control or collision avoidance. Two of the approaches use gridmap data. A gridmap integrates each measurement in a chronological order. The third approach analyzes moving objects ahead of the ego vehicle. One approach has been published previously, the other two are new. A range estimation is necessary on country roads and completes each approach. All approaches are evaluated using a huge dataset of country roads. The driven trajectory is taken as ground truth for the evaluation. The advantages and disadvantages are determined for each approach. The results show the new approach based on gridmap data performs up to 78% better than the known one. The other new approach using moving objects as input information yields estimations which are about three times more accurate than the ones from the known approach.
Frederik Sarholz, Jens Mehnert, Jens Klappstein, Jürgen Dickmann, Bernd Radig
IROS2
2011 Analysis of the rate of convergence of least squares neural network regression estimates in case of measurement errors
Michael Kohler, Jens Mehnert
Neural Networks2
2008 Automatic recognition of German news focusing on future-directed beliefs and intentions
Judith Eckle-Kohler, Michael Kohler, Jens Mehnert
Comput. Speech Lang.3