VLDB 2026 Research / reviewers in the wild / expert
Walter Stechele
dblp:37/3374
· DBLP profile ↗
68ranked-venue papers
1as first author
16since 2021 · last 2025
0000-0002-7455-8483ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 53 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 15 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LA-MTL: Latency-Aware Automated Multi-Task LearningabstractMulti-Task Learning (MTL) aims to unify a variety of tasks into a single network for improved training and inference efficiency. This is particularly attractive for real-time applications that require simultaneous execution of multiple workloads in resource-constrained embedded environments. However, most MTL approaches focus on enhancing parameters efficiency and overall tasks metrics, often lacking explicit inference latency awareness in the optimization loop. The design space exploration should not compromise on the parameters efficiency or task accuracy objectives in order to meet latency requirements. To address this, we propose LA-MTL, an automated layer-level MTL policy search that incorporates a novel analytical latency factor (ALF). By accounting for local and global latencies during the MTL policy search, we derive solutions that balance task metrics, parameters efficiency and latency constraints. LA-MTL search on ResNet34 yields solutions with up to 50% lower latency on the Jetson AGX Orin while maintaining competitive metrics in semantic segmentation and depth estimation tasks with a +/-2 p.p., on the CityScapes dataset. Additionally, we achieve a superior parameters efficiency, surpassing the state-of-theart MTL parameters reduction by over 20 p.p. Experiments on benchmark datasets (CityScapes, NYUv2) demonstrate the effectiveness of our approach across various backbones including ResNet34, MobileNetV2, and MobileOne in its expanded form. Code is available at https://github.com/shamvbs/LA-MTL.1 Shambhavi Balamuthu Sampath, Sami Sawani, Moritz Thoma, Lukas Frickenstein, Pierpaolo Morì, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Ulf Schlichtmann, Claudio Passerone, Walter Stechele |
DAC | 11 |
| 2025 | HiFi-SAGE: High Fidelity GraphSAGE-Based Latency Estimators for DNN OptimizationabstractAs deep neural networks (DNNs) are increasingly deployed on resource-constrained edge devices, optimizing and compressing them for real-time performance becomes crucial. Traditional hardware-aware DNN search methods often rely on inaccurate proxy metrics, expensive latency lookup tables, or slow hardware-in-the-Iloop (HIL) evaluations. To address this, quasi-generalized latency estimators, typically meta-learning-based, were proposed to replace HIL evaluations and accelerate the search. These come with a one-time data collection and training cost and can adapt to new hardware with few measurements. However, they still have some drawbacks: (1) They increase complexity by trying to generalize across a range of diverse hardware types; (2) They depend on handcrafted hardware descriptors, which may fail to capture hardware characteristics; (3) They often perform poorly on new, unseen hardware that significantly differs from their initial training set. To overcome these challenges, this paper turns to the more straightforward platform-specific estimators that do not require hardware descriptors and can be easily trained on any hardware. We introduce HiFi-SAGE, a high fidelity GraphSAGE-based platform-specific latency estimator. When trained from scratch on only 100 latency measurements, our novel dual-head estimator design surpasses the state-of-the-art (SoTA) on the 10% error bound metric by up to 17.4 p.p. while achieving an impressive fidelity score of 99% on the diverse LatBench dataset. We demonstrate that applying HiFi-SAGE to a genetic algorithm-based DNN compression search, achieved a Pareto front comparable to real HIL feedback with a mean absolute percentage error (MAPE) of 2.54%, 2.48%, and 4.16%, for InceptionV3, DenseNet169, and ResNet50 respectively. Compared to existing platform-specific works, the lower number of latency measurements and higher fidelity scores positions HiFi-SAGE as an attractive alternative to replace expensive HIL setups. Code is available at: https://github.com/shamvbs/HiFi-SAGE * Shambhavi Balamuthu Sampath, Leon Hecht, Moritz Thoma, Lukas Frickenstein, Pierpaolo Morì, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
DATE | 9 |
| 2025 | HotShot: A Loss-Guided Data Augmentation and Curriculum Learning Technique for the Task of Semantic SegmentationabstractSemantic segmentation is an important computer vision task that requires costly pixel-level annotations to train deep neural networks (DNNs) for. Especially for applications like autonomous driving, precise pixel-level understanding of scenes is a decisive factor between success and failure of the application. It follows that every labeled sample of an existing dataset is highly valuable and should be optimally used during training to maximize its value. This is achieved using (1) augmentation of the same labeled sample to help the model learn it in different ways, and (2) curriculum learning to introduce training samples to the model in an strategic order to ease the learning process. In this work, we present HotShot, a loss-guided cropping technique that assesses the DNN's prediction capability during the training to derive probability scores of potential cropping regions. This effectively combines augmentation and curriculum learning in one technique, where a single sample is cropped (augmentation) in regions selected based on the DNN's loss throughout the training (curriculum learning). For UperNet using a ConvNeXt-tiny backbone and DeepLabV3+ architecture using a ResNet-50 backbone, applying HotShot provides a +0.41 p.p. and +0.43 p.p. mIoU improvement over randomly cropping regions on the CityScapes and BDD100K datasets respectively. More interestingly, the analysis shows HotShot primarily boosts the classes that are most challenging for the model. For example, the rider and motorcycle classes on the BDD100K dataset improve by 163% and 129% using DeepLabV3+ with a ResNet-50 backbone. HotShot achieves improved mIoU in almost all cases and normalizes imbalances in learning challenging classes in datasets. Lukas Frickenstein, Moritz Thoma, Pierpaolo Morì, Shambhavi Balamuthu Sampath, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Christian Unger, Claudio Passerone, Walter Stechele |
IV | 10 |
| 2024 | MATAR: Multi-Quantization-Aware Training for Accurate and Fast Hardware RetargetingabstractQuantization of deep neural networks (DNNs) reduces their memory footprint and simplifies their hardware arithmetic logic, enabling efficient inference on edge devices. Different hardware targets can support different forms of quantization, e.g. full 8-bit, or 8/4/2-bit mixed-precision combinations, or fully-flexible bit-serial solutions. This makes standard quantization-aware training (QAT) of a DNN for different targets challenging, as there needs to be careful consideration of the supported quantization-levels of each target at training time. In this paper, we propose a generalized QAT solution that results in a DNN which can be retargeted to different hardware, without any retraining or prior knowledge of the hardware's supported quantization policy. First, we present the novel training scheme which makes the model aware of multiple quantization strategies. Then we demonstrate the retargeting capabilities of the resulting DNN by using a genetic algorithm to search for layer-wise, mixed-precision solutions that maximize performance and/or accuracy on the hardware target, without the need of fine-tuning. By making the DNN agnostic of the final hardware target, our method allows DNNs to be distributed to many users on different hardware platforms, without the need for sharing the training loop or dataset of the DNN developers, nor detailing the hardware capabilities ahead of time by the end-users of the efficient quantized solution. Models trained with our approach can generalize on multiple quantization policies with minimal accuracy degradation compared to target-specific quantization counterparts. Pierpaolo Morì, Moritz Thoma, Lukas Frickenstein, Shambhavi Balamuthu Sampath, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
DATE | 8 |
| 2024 | Wino Vidi Vici: Conquering Numerical Instability of 8-bit Winograd Convolution for Accurate Inference Acceleration on EdgeabstractWinograd-based convolution can reduce the total number of operations needed for convolutional neural network (CNN) inference on edge devices. Most edge hardware accelerators use low-precision, 8-bit integer arithmetic units to improve energy efficiency and latency. This makes CNN quantization a critical step before deploying the model on such an edge device. To extract the benefits of fast Winograd-based convolution and efficient integer quantization, the two approaches must be combined. Research has shown that the transform required to execute convolutions in the Winograd domain results in numerical instability and severe accuracy degradation when combined with quantization, making the two techniques incompatible on edge hardware. This paper proposes a novel training scheme to achieve efficient Winograd-accelerated, quantized CNNs. 8-bit quantization is applied to all the intermediate results of the Winograd convolution without sacrificing task-related accuracy. This is achieved by introducing clipping factors in the intermediate quantization stages as well as using the complex numerical system to improve the transform. We achieve 2.8× and 2.1× reduction in MAC operations on ResNet-20-CIFAR-10 and ResNet-18-ImageNet, respectively, with no accuracy degradation. Pierpaolo Morì, Lukas Frickenstein, Shambhavi Balamuthu Sampath, Moritz Thoma, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Christian Unger, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
WACV | 9 |
| 2023 | WinoTrain: Winograd-Aware Training for Accurate Full 8-bit Convolution AccelerationabstractEfficient inference is critical in realizing a low-power, real-time implementation of convolutional neural networks (CNNs) on compute and memory-constrained embedded platforms. Using quantization techniques and fast convolutional algorithms like Winograd, CNN inference can achieve benefits in latency and in energy consumption. Performing Winograd convolution involves (1) transforming the weights and activations to the Winograd domain, (2) performing element-wise multiplication on the transformed tensors, and (3) transforming the results back to the conventional spatial domain. Combining Winograd with quantization of all its steps results in severe accuracy degradation due to numerical instability. In this paper we propose a simple quantization-aware training technique, which quantizes all three steps of the Winograd convolution, while using a minimal number of scaling factors. Additionally, we propose an FPGA accelerator employing tiling and unrolling methods to highlight the performance benefits of using the full 8-bit quantized Winograd algorithm. We achieve 2× reduction in inference time compared to standard convolution on ResNet-18 for the ImageNet dataset, while improving the Top-1 accuracy by 55.7 p.p. compared to a standard post-training quantized Winograd variant of the network. Pierpaolo Morì, Shambhavi Balamuthu Sampath, Lukas Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Alexander Frickenstein, Walter Stechele, Claudio Passerone |
DAC | 7 |
| 2022 | Accelerating and pruning CNNs for semantic segmentation on FPGAabstractSemantic segmentation is one of the popular tasks in computer vision, providing pixel-wise annotations for scene understanding. However, segmentation-based convolutional neural networks require tremendous computational power. In this work, a fully-pipelined hardware accelerator with support for dilated convolution is introduced, which cuts down the redundant zero multiplications. Furthermore, we propose a genetic algorithm based automated channel pruning technique to jointly optimize computational complexity and model accuracy. Finally, hardware heuristics and an accurate model of the custom accelerator design enable a hardware-aware pruning framework. We achieve 2.44X lower latency with minimal degradation in semantic prediction quality (−1.98 pp lower mean intersection over union) compared to the baseline DeepLabV3+ model, evaluated on an Arria-10 FPGA. The binary files of the FPGA design, baseline and pruned models can be found in github.com/pierpaolomori/SemanticSegmentationFPGA Pierpaolo Morì, Manoj Rohit Vemparala, Nael Fasfous, Saptarshi Mitra, Sreetama Sarkar, Alexander Frickenstein, Lukas Frickenstein, Domenik Helms, Naveen Shankar Nagaraja, Walter Stechele, Claudio Passerone |
DAC | 10 |
| 2022 | Mind the Scaling Factors: Resilience Analysis of Quantized Adversarially Robust CNNsabstractAs more deep learning algorithms enter safety-critical application domains, the importance of analyzing their resilience against hardware faults cannot be overstated. Most existing works focus on bit-flips in memory, fewer focus on compute errors, and almost none study the effect of hardware faults on adversarially trained convolutional neural networks (CNNs). In this work, we show that adversarially trained CNNs are more susceptible to failure due to hardware errors when compared to vanilla-trained models. We identify large differences in the quantization scaling factors of the CNNs which are resilient to hardware faults and those which are not. As adversarially trained CNNs learn robustness against input attack perturbations, their internal weight and activation distributions open a backdoor for injecting large magnitude hardware faults. We propose a simple weight decay remedy for adversarially trained models to maintain adversarial robustness and hardware resilience in the same CNN. We improve the fault resilience of an adversarially trained ResNet56 by 25% for large-scale bit-flip benchmarks on activation data while gaining slightly improved accuracy and adversarial robustness. Nael Fasfous, Lukas Frickenstein, Michael Neumeier, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Maurizio Martina, Walter Stechele |
DATE | 8 |
| 2022 | AnaCoNGA: Analytical HW-CNN Co-Design Using Nested Genetic AlgorithmsabstractWe present AnaCoNGA, an analytical co-design methodology, which enables two genetic algorithms to evaluate the fitness of design decisions on layer-wise quantization of a neural network and hardware (HW) resource allocation. We embed a hardware architecture search (HAS) algorithm into a quantization strategy search (QSS) algorithm to evaluate the hardware design Pareto-front of each considered quantization strategy. We harness the speed and flexibility of analytical HW-modeling to enable parallel HW-CNN co-design. With this approach, the QSS is focused on seeking high-accuracy quantization strategies which are guaranteed to have efficient hardware designs at the end of the search. Through AnaCoNGA, we improve the accuracy by 2.88 p.p. with respect to a uniform 2-bit ResNet20 on CIFAR-10, and achieve a 35% and 37% improvement in latency and DRAM accesses, while reducing LUT and BRAM resources by 9% and 59% respectively, when compared to a standard edge variant of the accelerator. The nested genetic algorithm formulation also reduces the search time by 51% compared to an equivalent, sequential co-design formulation. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Julian Höfer, Anmol Singh, Naveen Shankar Nagaraja, Hans-Jörg Vögel, Nguyen Anh Vu Doan, Maurizio Martina, Jürgen Becker 0001, Walter Stechele |
DATE | 13 |
| 2022 | Region of interest based non-dominated sorting genetic algorithm-II: an invite and conquer approachabstractEvolutionary multi-objective optimization plays a vital role in solving many complex real-world optimization problems. Numerous approaches have been proposed over the years, and popular methods such as NSGA and its variants incorporate non-dominated sorting selection into evolutionary genetic algorithms to extract competing Pareto-optimal solutions from all over the objective space. However, in applications where the decision-maker is interested in a region of interest, a global optimization wastes effort to find irrelevant solutions outside of the preferred region. In this work, we propose an approach named ROI-NSGA-II to limit the optimization effort to a region of interest defined by the boundaries provided by the decision-maker. The ROI-NSGA-II invites the classical NSGA-II algorithm into the desired region using a modified dominance relation and conquers solutions within this region using a modified crowding distance based selection. The effectiveness of our approach is demonstrated on a set of benchmark problems with up to ten objectives and a real-world application, and the results are compared to a state-of-the-art R-NSGA-II. Manu Manuel, Benjamin Hien, Simon Conrady, Arne Kreddig, Nguyen Anh Vu Doan, Walter Stechele |
GECCO | 6 |
| 2021 | Hardware-Aware Mixed-Precision Neural Networks using In-Train Quantization
Manoj Rohit Vemparala, Nael Fasfous, Lukas Frickenstein, Alexander Frickenstein, Anmol Singh, Driton Salihu, Christian Unger, Naveen Shankar Nagaraja, Walter Stechele |
BMVC | 9 |
| 2021 | A Framework for Hardware-Accelerated Design Space Exploration for Approximate Computing on FPGAabstractThe demands for both increased performance and low power consumption on computing devices are outpacing technological improvements. Approximate computing is a design paradigm to leverage inherent error resilience of applications and trades in quality to reduce resource usage. Numerous approaches for approximation on FPGAs have been proposed in recent years and combining different methods can increase the resulting benefits in complex systems. Interactions between system components and error propagation necessitate a global design space exploration for the optimization of the approximation parameters. The loss of quality can be assessed by employing application-specific reference error metrics like PSNR or CIELAB ΔE which are well understood by designers and take error propagation implicitly into account. However, using a reference error metric can be very time-consuming, slowing down the design space exploration. To overcome this problem, we propose a framework for fast design space exploration of approximated FPGA designs in which the quality estimation is offloaded to an FPGA-based accelerator while the rest of the design space exploration is handled by a workstation PC. We evaluate the proposed framework on an image processing pipeline which is used to adapt image colors to be displayed correctly on a monitor. Our experiments show that using the accelerator yields similar results for the design in terms of the achieved quality-power trade-off compared to a software-only setup and can speed up the exploration by a factor of over 200x. Arne Kreddig, Simon Conrady, Manu Manuel, Walter Stechele |
DSD | 4 |
| 2021 | Sensor Fusion Neural Networks for Gesture Recognition on Low-power Edge Devices
Gábor Balázs 0001, Mateusz Chmurski, Walter Stechele, Mariusz Zubert |
ICAART (2) | 3 |
| 2021 | Binary-LoRAX: Low-Latency Runtime Adaptable XNOR Classifier for Semi-Autonomous Grasping with Prosthetic HandsabstractIntelligent, semi-autonomous prostheses take ad-vantage of combining autonomous functions and traditional myoelectric control. With the help of visual and environment sensors, intelligent prostheses achieve a level of autonomy which relieves the user from generating elaborate electromyographic (EMG) signals for grasp type and trajectory. To achieve the desired functionality, the semi-autonomous prosthesis must efficiently process the incoming environmental data at a high rate, with low power and high accuracy. In this paper, we propose Binary-LoRAX, a low-latency runtime adaptable classifier for the semi-autonomous grasping task of prosthetic hands. We offload the classification task to an efficient binary neural network accelerator which performs high-throughput XNOR operations on digital signal processing (DSP) blocks. To tailor the classifier’s performance to the current application scenario, we propose a frequency scaling approach which dynamically switches between two modes of operation, high-performance and power-saving. At high-performance, classifications are performed with a low latency of 0.45ms, high-throughput of 4999 FPS and power consumption of ∼ 2.15 W. This enables functions such as object localization and batch classification. Switching to power-saving mode, a latency of 80 ms is maintained, with up to 19% improved classifier battery-life. Our prototypes achieve a high accuracy of up to 99.82% on a 25 class problem from the YCB graspable object dataset. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Mohamed Badawy, Felix Hundhausen, Julian Höfer, Naveen Shankar Nagaraja, Christian Unger, Hans-Jörg Vögel, Jürgen Becker 0001, Tamim Asfour, Walter Stechele |
ICRA | 12 |
| 2021 | Investigating Binary Neural Networks for Traffic Sign Detection and RecognitionabstractTraffic sign detection is crucial for enabling autonomous vehicles to navigate in real-world streets, which must be carried out with high accuracy and in real-time. CNNs have become one of the standard approaches for traffic sign detection research in recent years. The use of CNNs has allowed the development of traffic sign detectors that are capable of achieving prediction accuracies similar to those of human drivers. However, most CNNs do not run in real-time due to the high number of computational operations involved during the inference phase. This hinders the deployment of CNNs in autonomous vehicles despite their high prediction accuracy. In this paper, we explore BNNs to tackle this problem. BNNs binarize the full-precision weights and activations of a CNN, drastically reducing the complexity of the computational operations required for inference, while at the same time maintaining the architectural parameters, as well as spatial dimensions of the input image. This reduces the memory required to run the model and enables faster inference time. We carry out in-depth studies on applying BNNs for traffic sign detection using real-world datasets. We observe an improvement of 11.63 × for normalized compute complexity, while suffering only 3.93 pp in detection accuracy on GTSDB dataset. Ee Heng Chen, Manoj Rohit Vemparala, Nael Fasfous, Alexander Frickenstein, Ahmed Mzid, Naveen Shankar Nagaraja, Jöran Zeisler, Walter Stechele |
IV | 8 |
| 2021 | HW-FlowQ: A Multi-Abstraction Level HW-CNN Co-design Quantization MethodologyabstractModel compression through quantization is commonly applied to convolutional neural networks (CNNs) deployed on compute and memory-constrained embedded platforms. Different layers of the CNN can have varying degrees of numerical precision for both weights and activations, resulting in a large search space. Together with the hardware (HW) design space, the challenge of finding the globally optimal HW-CNN combination for a given application becomes daunting. To this end, we propose HW-FlowQ, a systematic approach that enables the co-design of the target hardware platform and the compressed CNN model through quantization. The search space is viewed at three levels of abstraction, allowing for an iterative approach for narrowing down the solution space before reaching a high-fidelity CNN hardware modeling tool, capable of capturing the effects of mixed-precision quantization strategies on different hardware architectures (processing unit counts, memory levels, cost models, dataflows) and two types of computation engines (bit-parallel vectorized, bit-serial). To combine both worlds, a multi-objective non-dominated sorting genetic algorithm (NSGA-II) is leveraged to establish a Pareto-optimal set of quantization strategies for the target HW-metrics at each abstraction level. HW-FlowQ detects optima in a discrete search space and maximizes the task-related accuracy of the underlying CNN while minimizing hardware-related costs. The Pareto-front approach keeps the design space open to a range of non-dominated solutions before refining the design to a more detailed level of abstraction. With equivalent prediction accuracy, we improve the energy and latency by 20% and 45% respectively for ResNet56 compared to existing mixed-precision search methods. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Nguyen Anh Vu Doan, Christian Unger, Naveen Shankar Nagaraja, Maurizio Martina, Walter Stechele |
ACM Trans. Embed. Comput. Syst. | 10 |
| 2020 | ALF: Autoencoder-based Low-rank Filter-sharing for Efficient Convolutional Neural NetworksabstractClosing the gap between the hardware requirements of state-of-the-art convolutional neural networks and the limited resources constraining embedded applications is the next big challenge in deep learning research. The computational complexity and memory footprint of such neural networks are typically daunting for deployment in resource constrained environments. Model compression techniques, such as pruning, are emphasized among other optimization methods for solving this problem. Most existing techniques require domain expertise or result in irregular sparse representations, which increase the burden of deploying deep learning applications on embedded hardware accelerators. In this paper, we propose the autoencoder-based low-rank filter-sharing technique (ALF). When applied to various networks, ALF is compared to state-of-the-art pruning methods, demonstrating its efficient compression capabilities on theoretical metrics as well as on an accurate, deterministic hardware-model. In our experiments, ALF showed a reduction of 70% in network parameters, 61% in operations and 41% in execution time, with minimal loss in accuracy. Alexander Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Laura Hauenschild, Naveen Shankar Nagaraja, Christian Unger, Walter Stechele |
DAC | 7 |
| 2020 | OrthrusPE: Runtime Reconfigurable Processing Elements for Binary Neural NetworksabstractRecent advancements in Binary Neural Networks (BNNs) have yielded promising results, bringing them a step closer to their full-precision counterparts in terms of prediction accuracy. These advancements were brought about by additional arithmetic and binary operations, in the form of scale and shift operations (fixed-point) and convolutions with multiple weight and activation bases (binary). In this paper, we propose OrthrusPE, a runtime reconfigurable processing element (PE) which is capable of executing all the operations required by modern BNNs while improving resource utilization and power efficiency. More precisely, we exploit DSP48 blocks on off-the-shelf FPGAs to compute binary Hadamard products (for binary convolutions) and fixed-point arithmetic (for scaling, shifting, batch norm, and non-binary layers), thereby utilizing the same hardware resource for two distinct, critical modes of operation. Our experiments show that common PE implementations increase dynamic power consumption by 67%, while requiring 39% more lookup tables, when compared to an OrthrusPE implementation. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele |
DATE | 4 |
| 2020 | Neural Architecture Search for Automotive Grid Fusion Networks Under Embedded Hardware ConstraintsabstractThe goal of automotive sensor fusion is to generate an environmental model with low-cost, mass-produced sensors-typically camera and radar. The central fusion controller needs sophisticated post-processing algorithms to achieve this goal despite extensive pre-processing on the sensor side. Recent advances show the capabilities of deep convolutional auto-encoders for sensor fusion but are still limited to high-performance hardware. Limited memory and processing capabilities of embedded devices lead to the multi-objective optimization problem addressed in this work. The proposed evolutionary neural architecture search is configured to find novel sensor fusion network architectures optimized for embedded hardware. The discovered encoding and decoding cells improve fusion performance compared to that of corresponding cells from literature. These lightweight network architectures allow the deployment on constrained embedded devices, thus leverage environmental perception of level 2/2+ automated cars. Gábor Balázs 0001, Walter Stechele |
ICMLA | 2 |
| 2020 | Binary DAD-Net: Binarized Driveable Area Detection Network for Autonomous DrivingabstractDriveable area detection is a key component for various applications in the field of autonomous driving (AD), such as ground-plane detection, obstacle detection and maneuver planning. Additionally, bulky and over-parameterized networks can be easily forgone and replaced with smaller networks for faster inference on embedded systems. The driveable area detection, posed as a two class segmentation task, can be efficiently modeled with slim binary networks. This paper proposes a novel binarized driveable area detection network (binary DAD-Net), which uses only binary weights and activations in the encoder, the bottleneck, and the decoder part. The latent space of the bottleneck is efficiently increased (×32→×16 downsampling) through binary dilated convolutions, learning more complex features. Along with automatically generated training data, the binary DAD-Net outperforms state-of-the-art semantic segmentation networks on public datasets. In comparison to a full-precision model, our approach has a ×14.3 reduced compute complexity on an FPGA and it requires only 0.9MB memory resources. Therefore, commodity SIMD-based AD-hardware is capable of accelerating the binary DAD-Net. Alexander Frickenstein, Manoj Rohit Vemparala, Jakob Mayr, Naveen Shankar Nagaraja, Christian Unger, Federico Tombari, Walter Stechele |
ICRA | 7 |
| 2019 | Mixed Frame-/Event-Driven Fast Pedestrian DetectionabstractPedestrian detection has attracted enormous research attention in the field of Intelligent Transportation System (ITS) due to that pedestrians are the most vulnerable traffic participants. So far, almost all pedestrian detection solutions are based on the conventional frame-based camera. However, they cannot perform very well in scenarios with bad light condition and high-speed motion. In this work, a Dynamic and Active Pixel Sensor (DAVIS), whose two channels concurrently output conventional gray-scale frames and asynchronous low-latency temporal contrast events of light intensity, was first used to detect pedestrians in a traffic monitoring scenario. Data from two camera channels were fed into Convolutional Neural Networks (CNNs) including three YOLOv3 models and three YOLO-tiny models to gather bounding boxes of pedestrians with respective confidence map. Furthermore, a confidence map fusion method combining the CNN-based detection results from both DAVIS channels was proposed to obtain higher accuracy. The experiments were conducted on a custom dataset collected on TUM campus. Benefiting from the high speed, low latency and wide dynamic range of the event channel, our method achieved higher frame rate and lower latency than those only using a conventional camera. Additionally, it reached higher average precision by using the fusion approach. Zhuangyi Jiang, Kai Huang 0001, Walter Stechele, Guang Chen 0001, Zhenshan Bing, Alois C. Knoll |
ICRA | 4 |
| 2018 | Efficient hardware acceleration of CNNs using logarithmic data representation with arbitrary log-baseabstractEfficient acceleration of Deep Neural Networks is a manifold task. In order to save memory requirements and reduce energy consumption we propose the use of dedicated accelerators with novel arithmetic processing elements which use bit shifts instead of multipliers. While a regular power-of-2 quantization scheme allows for multiplierless computation of multiply-accumulate-operations, it suffers from high accuracy losses in neural networks. Therefore, we evaluate the use of powers-of-arbitrary-log-bases and confirmed their suitability for quantization of pre-trained neural networks. The presented method works without retraining of the neural network and therefore is suitable for applications in which no labeled training data is available. In order to verify our proposed method, we implement the log-based processing elements into a neural network accelerator on an FPGA. The hardware efficiency is evaluated in terms of FPGA utilization and energy requirements in comparison to regular 8-bit-fixed-point multiplier based acceleration. Using this approach hardware resources are minimized and power consumption is reduced by 22.3%. Sebastian Vogel, Mengyu Liang, Andre Guntoro, Walter Stechele, Gerd Ascheid |
ICCAD | 4 |
| 2017 | Hardware-accelerated CCD readout smear correction for Fast Solar PolarimeterabstractShutterless frame store charge-coupled devices (CCDs) are commonly used in ground-based solar observations, but the characteristical readout smear error of such devices hinders an application of frame store CCDs to autonomous missions. The combination of polarimetric modulation and image accumulation disables a correction of this error via software-based post-facto processing if in addition microvibrations occur during flight. This paper presents the first FPGA-based architecture for online smear correction of images from frame store CCDs, which allows for the usage of a certain frame store CCD camera on a balloon-borne solar observatory. First, we explore fast convolution-based algorithms with respect to their properties for an implementation. Afterwards, a hardware architecture is derived and implemented. Our results show that 400 frames of megapixel size can be corrected per second, maintaining an acceptable power consumption of less than 12 Watt. Finally, we discuss the circuit and show the degrees of freedom for further designs. Stefan Tabel, Korbinian Weikl, Walter Stechele |
ASAP | 3 |
| 2016 | PrefaceabstractAnother year and another step forward in reconfigurable computing as it joins the computing technology mainstream. For the past 26 years FPL has reflected this progress in its technical program, its keynote addresses, and the programs of its workshops and tutorials. This year FPL is hosted by the École polytechnique fédérale de Lausanne (EPFL) on the shores of the largest Alpine lake, Lake Geneva. Paolo Ienne, Walid A. Najjar, Jason Helge Anderson, Philip Brisk, Walter Stechele |
FPL | 5 |
| 2016 | Tackling long duration transients in sequential logicabstractSingle Event Transients (SETs) in combinational logic remain to be an important topic in the reliability domain. SETs were traditionally relatively short in comparison to the clock period. The majority of the countermeasures utilizes this property. However, advances in technology scaling will reverse the ratio. Investigations show that SETs may last up to multiple clock cycles in the future. So called Long Duration Transients (LDTs) corrupt almost all available countermeasures. This work presents a new methodology to tackle LDTs. Dual Modular Redundancy (DMR) is used to detect any corruption of the application logic. A new micro-rollback scheme is introduced, that expands DMR-Architectures with correction capabilities. The concept is also capable of handling single event upsets and timing violations. The scheme utilizes a newly designed History Cell. The History Cell introduces an area overhead of 74% and a power overhead of 106% compared to a standard cell D-flip-flop. Area and power overheads for expanding an existing DMR-Architecture with correction capabilities are approximately 24% and 28%, respectively. Erol Koser, Walter Stechele |
IOLTS | 2 |
| 2016 | Integrated Soft Error Resilience and Self-TestabstractMany VLSI-SoC include both, protection against soft errors and Built-In-Self-Test (BIST). This work investigates on the combination of both domains. The proposed approach offers additional functionality for BIST, i.e. self-test capabilities for the test logic itself and fault localization. A snapshot mode is offered as well. It enables the taking of snapshots of the current system state concurrently to task execution. The new approach was implemented and verified in various test circuits. The resource overhead is approx. 15 % in registers and 37 % in LUTs when synthesized for a FPGA, as compared to simple superposition of the initial approaches. Erol Koser, Sebastian Krosche, Walter Stechele |
VLSI-SoC | 3 |
| 2015 | A soft-core processor array for relational operatorsabstractDespite the performance and power efficiency gains achieved by FPGAs for text analytics queries, analysis shows a low utilization of the custom hardware operator modules. Furthermore the long synthesis times limit the accelerator's use in enterprise systems to static queries. To overcome these limitations we propose the use of an overlay architecture to share area resources among multiple operators and reduce compilation times. In this paper we present a novel soft-core architecture tailored to efficiently perform relational operations of text analytics queries on multiple virtual streams. It combines the ability to perform efficient streaming based operations while adding the flexibility of an instruction programmable core. It is used as a processing element in an array of cores to execute large query graphs and has access to shared co-processors to perform string-and context-based operations. We evaluate the core architecture in terms of area and performance compared to the custom hardware modules, and show how a minimum number of cores can be calculated to avoid stalling the document processing. Raphael Polig, Heiner Giefers, Walter Stechele |
ASAP | 3 |
| 2015 | Matching Detection and Correction Schemes for Soft Error Handling in Sequential LogicabstractThis paper addresses two common problems of soft error handling schemes for sequential logic. The first issue is a race condition between the correction of soft errors and their propagation to following stages. The second issue concerns erroneous write accesses to external memory. Both problems are thoroughly examined. The key idea in order to resolve the problems is to distinguish between different types of transient faults (soft errors, timing violations) and explicitly utilize their individual characteristics. Based on that principle a low-cost shadow cell is proposed. Transient faults are detected by shadowing the interfaces of a flip-flop. Timing violations and early soft errors are corrected by architectural replay. Late soft errors are instantly corrected within the shadow cell. The introduced area overhead is 57% and the power overhead 32% compared to a standard cell D-flip-flop. Erol Koser, Felix Miller, Walter Stechele |
DSD | 3 |
| 2015 | Design of fine-grained sequential approximate circuits using probability-aware fault emulationabstractApproximate Computing has recently drawn interest due to its promise to substantially decrease the power consumption of integrated circuits. By tolerating a certain imprecision at a circuit output, the circuit can be operated at a more resource-saving state. For instance, parts of the circuit could be switched off or driven at sub-threshold voltage. Clearly, not all applications are suitable for this approach. Especially applications from the signal and image processing domain are applicable, due to their intrinsic tolerance to imprecision. But even for these circuits, one has to be very careful where to approximate a circuit and to what extent, in order not to fall below a minimum required QoS. In this paper we are presenting an approach to generate approximate circuits from existing deterministic implementations. The flow reaches from application-driven QoS definition down to approximated RTL. We are employing FPGA-based fault emulation of the circuit in order to find out how faults, i.e. imprecisions in the circuit, affect the overall circuit behavior. Most existing approaches only consider combinational circuits. Our proposed methodology is able to approximate complete sequential circuits. Due to the FPGA-based emulation, our approach is very fast and accurate. And furthermore, it allows us to fine-granular tune the resulting precision to the required QoS, in order to get the most out of the approximation. David May 0003, Walter Stechele |
ISLPED | 2 |
| 2015 | Self-adaptive corner detection on MPSoC through resource-aware programming
Johny Paul, Benjamin Oechslein, Christoph Erhardt, Jens Schedel, Manfred Kröhnert, Daniel Lohmann, Walter Stechele, Tamim Asfour, Wolfgang Schröder-Preikschat |
J. Syst. Archit. | 7 |
| 2015 | Resource-awareness on heterogeneous MPSoCs for image processing
Johny Paul, Walter Stechele, Benjamin Oechslein, Christoph Erhardt, Jens Schedel, Daniel Lohmann, Wolfgang Schröder-Preikschat, Manfred Kröhnert, Tamim Asfour, Éricles Sousa, Vahid Lari, Frank Hannig, Jürgen Teich, Artjom Grudnitsky, Lars Bauer, Jörg Henkel |
J. Syst. Archit. | 2 |
| 2014 | Towards low-cost fault detection strategy of FPGA configuration memory in real-time systemsabstractAs a result of the recent advancements in technology, FPGAs are more often used for automotive applications. They must therefore meet industrial requirements like a fast and very low cost fault detection strategy for their configuration memory. Cyclic memory tests are the state of the art approach for this task. They do, however, violate fault detection times, especially for the latest FPGA devices. The approach presented in this paper splits the configuration memory in several parts and prioritizes their test execution depending on the application data flow. Using two conclusive examples this adaptive strategy is compared to state of the art memory tests on a Xilinx FPGA. They show that our approach is a useful means to efficiently meet requirements on automotive fault detection times. Michael Frischke, Andreas J. Rohatschek, Walter Stechele |
IOLTS | 3 |
| 2014 | Improving the significance of probabilistic circuit fault emulationsabstractRecent trends go into the direction to tolerate a certain number of faults in integrated circuits, at locations where the error remains unnoticed by the user, trying to further extend Moore's law. In order to enable such an approach, a comprehensive analysis of the probabilistic behavior of the circuit is required. We focus on FPGA-based probability-aware circuit fault emulation in order to do this analysis. Faults are injected at run-time, based on error probabilities, into the circuit. The influence on the circuit outputs is observed by comparing the fault-free with the faulty circuit. The probability-awareness allows us to model, for instance, different voltage-domains or channel widths, within the same circuit. Previous work is usually restricted to a non-probabilistic treatment of circuit faults, due to its complexity in terms of resource requirements and the difficulty to generate statistical significant results. In this work we will present two fundamental strategies, in order to tackle this latter issue. First, we will present methods to separate the data and control path of the circuit. Errors can only be tolerated in the data path of a circuit. Hence, it is necessary to reliably identify the control path inside the netlist, in order to make it robust against faults. Consequently, the control path hast to be excluded from the emulations. Secondly, we will present algorithms to score the stability of the measured output error probability. We will present ways to determine how many clock cycles have to be emulated in order to get a stable simulation result. By applying those methods we are able to generate trustworthy information about the probabilistic behavior of a circuit. David May 0003, Walter Stechele |
IOLTS | 2 |
| 2014 | Automatic denoising parameter estimation using gradient histogramsabstractState-of-the-art denoising methods provide denoising results that can be considered close to optimal. The denoising methods usually have one or more parameters regulating denoising strength that can be adapted for a specific image. To obtain the optimal denoising result, the correct parameter setting is crucial. In this paper, we therefore propose a method that can automatically estimate the optimal parameter of a denoising algorithm. Our approach compares the gradient histogram of a denoised image to an estimated reference gradient histogram. The reference gradient histogram is estimated based on down- and upsampling of the noisy image, thus our method works without a reference and is image-adaptive. We evaluate our propsed down-/upsampling-based gradient histogram method (DUG) based on a subjective test with 20 participants. In the test data, we included images from both the Kodak data set and the more realistic ARRI data set and we used the state-of-the-art denoising method BM3D. Based on the test results we can show that the parameter estimated by our method is very close to the human perception. Despite being very fast and simple to implement, our method shows a lower error than all other suitable no-reference metrics we found. Tamara Seybold, Florian Kuhn, Julian Habigt, Mark Hartenstein, Walter Stechele |
VCIP | 5 |
| 2013 | Dynamic Noise Estimation Approach for X-Ray Detectors on FPGAsabstractThe proposed algorithm enables real-time noise estimation directly on the raw data stream of an X-ray detection system. This allows for better detector monitoring and increases the measurement accuracy, especially in long term measurements and under changing environment conditions. The noise value is used as an event extraction threshold and is therefore the most critical parameter for the performance in X-ray data processing. Inaccurately estimated noise values can result into significantly poorer energy resolutions. The new algorithm is designed for low memory bandwidth, precise estimation results and a low processing complexity. This allows real-time data processing even for large and fast detector systems with an output rate of hundreds of MPixels/s. Compared to the calculation method used up to now the memory bandwidth can be reduced down to a factor of a hundred with the new algorithm. The new developed real-time Dynamic Noise Estimation (DNE) algorithm uses an iterative approach for the estimation. For the calculation is only an increment/decrement operation and an integer comparison required. This simple approach allows a resource efficient hardware implementation on FPGAs. Simulation results confirm the excellent estimation accuracy and stability with real-time performance. Measurements with Fano-limited energy resolution, the theoretically best achievable energy resolution with semiconductor detectors, confirm this outstanding performance on real detector systems. The reliability and stability of the system was demonstrated during a long term irradiation campaign for Mercury Imaging X-ray Spectrometer (MIXS) [5], [3]. MIXS is an instrument onboard ESA's corner stone mission BepiColombo [1] and will be launched 2015 into an orbit around Mercury. Florian Aschauer, Walter Stechele, Johannes Treis |
DSD | 2 |
| 2013 | FPGA Based Real-Time Data Processing DAQ System for the Mercury Imaging X-Ray SpectrometerabstractThis paper presents a new DAQ system with real-time data processing capability for X-ray detector systems. The data processing algorithms are working directly on the serialised input data stream from the detector system and run in real-time on a FPGA processing system. The DAQ system includes new algorithms for dynamic adaptation of important processing parameters during a measurement. These dynamic algorithms enable better measurement precisions, autonomous system operation and avoid detector recalibration cycles. The developed and presented data processing algorithms are optimised for low memory bandwidth and hardware resource utilisation. High data reduction and live monitoring of important system parameters are achieved with the real-time processing. The concept of the new DAQ system was tested and evaluated with the Mercury Imaging X-ray Spectrometer (MIXS). MIXS is an instrument on board of the ESA corner stone mission BepiColombo which will be launched 2015 into an orbit around Mercury. The 64x64 detector matrix used in the MIXS detector system for X-ray detection delivers 6000 frames per second and provides an optimal evaluation platform for the new DAQ system. Florian Aschauer, Walter Stechele, Johannes Treis |
DSD | 2 |
| 2013 | A Design Space Exploration Framework For Automotive Embedded Systems And Their Power ManagementabstractThe E/E (electric/electronic) architecture of a modern vehicle is a complex distributed system, where up to 80 electronic control units (ECUs), interconnected by several communication buses, need to collaborate with each other in order to implement the various comfort and safety features. The presented E/E design space exploration framework supports engineers during the development process of new E/E architectures by providing a graphical modeling and simulation environment with an interactive visualization of simulation-based results. Furthermore, it contains an advanced business logic which administrates the modeling, storing, retrieving and cloning of evaluation sessions consisting of complex experiments. Due to a high-level modeling approach, future architectures can be evaluated in respect to power consumption and performance values already in an early stage of the design process and design alternatives can be easily compared with each other. Gregor Walla, Zaur Molotnikov, Hans-Ulrich Michel, Walter Stechele, Andreas Barthels, Andreas Herkersdorf |
ECMS | 4 |
| 2013 | Weighted partitioning of sequential processing chains for dynamically reconfigurable FPGASabstractTemporal runtime-reconfiguration of FPGAs allows for a resource-efficient sequential execution of signal processing modules. Approaches for partitioning processing chains into modules have been derived in various previous works. We will present a metric for weighted partitioning of pre-defined processing element sequences. The proposed method yields a set of reconfigurable partitions, which are balanced in terms of resources, while jointly have a minimal data throughput. Using this metric, we will formulate a partitioning algorithm with linear complexity and will compare our approach to the state of the art. Michael Feilen, Andreas Iliopoulos, Michael Vonbun, Walter Stechele |
FPL | 4 |
| 2013 | A resource-efficient probabilistic fault simulatorabstractThe reduction of CMOS structures into the nanometer regime, as well as the high demand for low-power applications, animating to further reduce the supply voltages towards the threshold, results in an increased susceptibility of integrated circuits to soft errors. Hence, circuit reliability has become a major concern in today's VLSI design process. A new approach to further support these trends is to relax the reliability requirements of a circuit, while ensuring that the functionality of the circuit remains unaffected, or effects remain unnoticed by the user. To realize such an approach it is necessary to determine the probability of an error at the output of a circuit, given an error probability distribution at the circuits' elements. Purely software-based simulation approaches are unsuitable due to the large simulation times. Hardware-accelerated approaches exist, but lack the ability to inject errors based on probabilities, are slow or have a large area overhead. In this paper we propose a novel approach for FPGA-based, probabilistic, circuit fault simulation. The proposed system is a mainly hardware-based, which makes the simulation fast, but also keeps the hardware overhead on the FPGA low by exploiting FPGA specific features. David May 0003, Walter Stechele |
FPL | 2 |
| 2013 | Towards an Evaluation of Denoising Algorithms with Respect to Realistic Camera NoiseabstractThe development and tuning of denoising algorithms is usually based on readily processed test images that are artificially degraded with additive white Gaussian noise (AWGN). While AWGN allows us to easily generate test data in a repeatable manner, it does not reflect the noise characteristics in a real digital camera. Realistic camera noise is signal-dependent and spatially correlated due to the demosaicking step required to obtain full-color images. Hence, the noise characteristic is fundamentally different from AWGN. Using such unrealistic data to test, optimize and compare denoising algorithms may lead to incorrect parameter tuning or sub optimal choices in research on denoising algorithms. In this paper, we therefore propose an approach to evaluate denoising algorithms with respect to realistic camera noise: we describe a new camera noise model that includes the full processing chain of a single sensor camera. We determine the visual quality of noisy and denoised test sequences using a subjective test with 18 participants. We show that the noise characteristics have a significant effect on visual quality. Quality metrics, which are required to compare denoising results, are applied, and we evaluate the performance of 10 full-reference metrics and one no-reference metric with our realistic test data. We conclude that a more realistic noise model should be used in future research to improve the quality estimation of digital images and videos and to improve the research on denoising algorithms. Tamara Seybold, Christian Keimel, Marion Knopp, Walter Stechele |
ISM | 4 |
| 2012 | Invasive Computing for robotic visionabstractMost robotic vision algorithms are computationally intensive and operate on millions of pixels of real-time video sequences. But they offer a high degree of parallelism that can be exploited through parallel computing techniques like Invasive Computing. But the conventional way of multi-processing alone (with static resource allocation) is not sufficient enough to handle a scenario like robotic maneuver, where processing elements have to be shared between various applications and the computing requirements of such applications may not be known entirely at compile-time. Such static mapping schemes leads to inefficient utilization of resources. At the same time it is difficult to dynamically control and distribute resources among different applications running on a single chip, achieving high resource utilization under high-performance constraints. Invasive Computing obtains more importance under such circumstances, where it offers resource awareness to the application programs so that they can adapt themselves to the changing conditions, at run-time. In this paper we demonstrate the resource aware and self-organizing behavior of invasive applications using three widely used applications from the area of robotic vision - Optical Flow, Object Recognition and Disparity Map Computation. The applications can dynamically acquire and release hardware resources, considering the level of parallelism available in the algorithm and time-varying load. Johny Paul, Walter Stechele, Manfred Kröhnert, Tamim Asfour, Rüdiger Dillmann |
ASP-DAC | 2 |
| 2012 | A low-overhead monitoring ring interconnect for MPSoC parameter optimizationabstractMPSoCs need to integrate self-x properties in order to get rid of the worst-case design style which is no longer affordable in large SoCs. Integrating self-x properties in SoCs is possible through a monitoring interconnect which carries monitor information to evaluators that decide on actions that will tune the SoC operation mode. We have designed a customized interconnect for SoC monitoring/actuation. We have implemented it in VHDL and tested it in FPGA. The prototype proved that this customized interconnect provides good results regarding latency and area overheads and is a key component in enabling self-optimization in our FPGA MPSoC prototype. Abdelmajid Bouajila, Abdallah Lakhtel, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf |
DDECS | 4 |
| 2012 | Efficient DVB-T2 decoding accelerator design by time-multiplexing FPGA resourcesabstractDemodulation and decoding of second generation terrestrial digital video broadcasting (DVB-T2) signals on general purpose processor platforms is challenging in terms of complexity and in terms of power. FPGA-based runtime acceleration for DVB-T2 allows for unwrapping the iterative structures of modern channel decoding schemes by using parallel hardware designs. Additionally, due to the sequential nature of the DVB-T2 receiver chain we can use partial reconfiguration to switch between different decoding modules. We will show in a theoretical analysis that this time-multiplexing approach can be used to realize resource-efficient DVB-T2 receiver chains at a much lower resource and power consumption as compared to solely processor-based solutions. Michael Feilen, Matthias Ihmig, Christian Schwarzbauer, Walter Stechele |
FPL | 4 |
| 2011 | An architecture and an FPGA prototype of a reliable processor pipeline towards multiple soft- and timing errorsabstractThis paper presents a reliable processor pipeline architecture resilient to multiple soft- and timing errors. It also presents a probabilistic quantification of its performance overheads. This reliable processor pipeline architecture has been implemented in the Leon3 VHDL open source processor. An FPGA prototype running under random fault injection has also been developed. This reliable processor pipeline has low performance overheads (relative CPI of 1.06 at an error injection rate of 3 %) and is therefore much better than techniques based on flushing. Abdelmajid Bouajila, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf |
DDECS | 3 |
| 2011 | A reasoning approach to enable abductive semantic explanation upon collected observations for forensic visual surveillanceabstractThis paper proposes an approach to enable automatic generation of probable semantic hypotheses for a given set of collected observations for forensic visual surveillance. As video analytic power exploited in visual surveillance is getting matured, the more automatically generated intermediate semantic metadata became available. In the sense of forensic reuse of such data, the majority of approaches have been focused on specific semantic query based scene analysis. However, in reality, there are often cases in which it is more natural to reason about the most probable semantic explanation of a scene given a collection of specific semantic evidences. In general, this type of diagnostic reasoning is known as abduction. To enable such a semantic reasoning, in this paper, we propose a layered reasoning pipeline that combines abductive logic programming together with backward and forward chaining based deductive logic programming. To rate derived hypotheses, we apply subjective logic. We present a conceptual case study in a distributed camera based scenario. The case study shows the potential and feasibility of the proposed approach for forensic analysis of visual surveillance data. Seunghan Han, Andreas Hutter, Walter Stechele |
ICME | 3 |
| 2010 | Subjective Logic Based Hybrid Approach to Conditional Evidence Fusion for Forensic Visual SurveillanceabstractIn forensic analysis of visual surveillance data, conditional knowledge representation and inference under uncertainty play an important role for deriving new contextual cues by fusing relevant evidential patterns. To address this aspect, both rule-based (aka. extensional) and state based (aka. intensional) approaches have been adopted for situation or visual event analysis. The former provides flexible expressive power and computational efficiency but typically allows only one directional inference. The latter is computationally expensive but allows bidirectional interpretation of conditionals by treating antecedent and consequent of conditionals as mutually relevant states. In visual surveillance, considering the varying semantics and potentially ambiguous causality in conditionals, it would be useful to combine the expressive power of rule-based system with the ability of bidirectional interpretation. In this paper, we propose a hybrid approach that, while relying mainly on a rule-based architecture, also provides an intensional way of on-demand conditional modeling using conditional operators in subjective logic. We first show how conditionals can be assessed via explicit representation of ignorance in subjective logic. We then describe the proposed hybrid conditional handling framework. Finally we present an experimental case study from a typical airport scene taken from visual surveillance data. Seunghan Han, Bonjung Koo, Andreas Hutter, Vinay D. Shet, Walter Stechele |
AVSS | 5 |
| 2010 | A rapid prototyping system for error-resilient multi-processor systems-on-chipabstractStatic and dynamic variations, which have negative impact on the reliability of microelectronic systems, increase with smaller CMOS technology. Thus, further downscaling is only profitable if the costs in terms of area, energy and delay for reliability keep within limits. Therefore, the traditional worst case design methodology will become infeasible. Future architectures have to be error resilient, i.e., the hardware architecture has to tolerate autonomously transient errors. In this paper, we present an FPGA based rapid prototyping system for multi-processor systems-on-chip composed of autonomous hardware units for error-resilient processing and interconnect. This platform allows the fast architectural exploration of various error protection techniques under different failure rates on the microarchitectural level while keeping track of the system behavior. We demonstrate its applicability on a concrete wireless communication system. Matthias May 0001, Norbert Wehn, Abdelmajid Bouajila, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf, Daniel Ziener, Jürgen Teich |
DATE | 5 |
| 2010 | Architectural Vulnerability Factor Estimation with Backwards AnalysisabstractSingle-Event-Upsets in synchronous register-based designs are a severe problem for safety-critical applications. Exact and detailed error rate estimations are needed to determine a system's level of reliability. Available methods for estimation consider only special effects, use special reliability models or are computationally intensive. We present an innovative method that is able to calculate the architectural vulnerability factor (AVF)of any RT-level circuit description by applying time-reversed stimulus values. This method, which we call Backwards Analysis, considers all major masking effects (logic masking, information lifetime, timing derating, transitive masking) in a single algorithm and delivers results in several levels of detail from average AVF through sensitivity waveforms. The results show the critical parts and states of a design, which could be used for reliability assessment and selective hardening of the circuit to reach a target failure rate. Robert Hartl, Andreas J. Rohatschek, Walter Stechele, Andreas Herkersdorf |
DSD | 3 |
| 2009 | Optimizing the SUSAN corner detection algorithm for a high speed FPGA implementationabstractIn many embedded systems for video surveillance distinctive features are used for the detection of objects. In this contribution a real-time FPGA implementation of a feature detector, namely the SUSAN algorithm is described. As the original SUSAN algorithm performs poorly on non-synthetic images a significant quality improvement of this algorithm is presented. The hardware accelerator outperforms a comparable software version running on an Intel Core2Duo E8400 core at 3.00 GHz and delivers almost the same execution time compared to an implementation of the Harris corner detector running on an Nvidia GeForce 8800 GTX GPU. Christopher Claus, Robert Huitl, Joachim Rausch, Walter Stechele |
FPL | 4 |
| 2009 | Wire Topology Optimization for Low Power CMOSabstractAn increasing fraction of dynamic power consumption can be attributed to switched interconnect capacitances. Non-uniform wire spacing depending on activity had shown promising power reductions for on-chip buses. In this paper, a new and fast routing optimization methodology based on non-uniform spacing is proposed for entire circuits. No area investment is required, since whitespace remaining after detailed routing is exploited. The proposed methodology has been implemented and tapped into an industry-proven design flow. Wire power reductions of up to 9.55% for modern multiprocessor benchmarks with tight area constraints are demonstrated, twice as much as approaches that do not take switching activities into account. Timing is not adversely affected, and the yield limit is slightly improved. Paul Zuber, Othman Bahlous, Thomas Ilnseher, Michael Ritter, Walter Stechele |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2008 | Hardware/software architecture of an algorithm for vision-based real-time vehicle detection in dark environmentsabstractHardware/software partitioning of algorithms is gaining more and more importance in order to benefit from the advantages of both worlds. Pure software implementations are easy to change but the processing time is rather high. By contrast pure hardware implementations usually result in faster processing due to inherent parallelism but they do not offer the necessary flexibility for quick changes and adaptions. In this paper the hardware/software co-design of a self-developed algorithm to detect cars by their taillights as well as its implementation on an embedded system (FPGA) is presented. Instead of utilizing expensive sensors such as RADAR which also can be used to detect obstacles in dark environments, the detection method presented here is based solely on grayscale images taken by a low-cost on-board camera which was mounted on a moving vehicle. Only computationally intense parts - namely pixel or sliding window operations - are implemented in hardware to achieve the necessary real-time requirements. The remainder of the algorithm - the so called higher level application code - is running on standard embedded CPU cores. With this architecture it is possible to process the incoming video-stream (25 frames/s) and detect cars in real-time on an embedded system. Nicolas Alt, Christopher Claus, Walter Stechele |
DATE | 3 |
| 2008 | Design Flows, Communication Based Design and Architectures in Automotive Electronic SystemsabstractSummary form only given. The complete presentation was not made available for publication as part of the conference proceedings. A steadily increasing number of microprocessors and electronic components with the heavy demand of computation performance in automotive electronic systems affect substantially the design of networked ECUs in today as well as future cars. Novel approaches, based on heterogeneous hardware (Coarse- fine Grained reconfigurable Hardware, Microprocessors) could be a solution to handle the computation intensive tasks, e.g. for driver-assistance systems. The challenge here is to find an optimal trade-off between power consumption, cost, performance and flexibility which leads to the question which technology and which distribution (automotive function centralisation - decentralisation trade-offs!) will be targeted in future car electronics. Introducing novel architecture topologies and corresponding tool flows with standardised specification and verification are here severe challenges. A first approach to meet these challenges is the AUTOSAR development partnership, which aims at a standardisation of automotive software architecture. The purpose of this tutorial is to evaluate and discuss new concepts for communication based design of automotive electronic and car network systems, as well as to discuss and envisage future system design in automotive electronics. Both aspects, hardware / software design and tool-integration will be discussed. The main emphasis in this session is design-flow, tool-development, applications and system design. The tutorial is addressed to hardware and system engineers as well as to researchers. A set of presentations intended to set the stage for the discussion, will be followed by a panel where selected world-wide specialists in the field of automotive electronics will discuss the demands and interests of industry on novel technologies and systems and research activities for future automotive systems. Jürgen Becker 0001, Michael Hübner 0001, Robert Esser, Andreas Herkersdorf, Walter Stechele, Vera Lauer |
DATE | 5 |
| 2008 | Fine grain reconfigurable architecturesabstractIn this booth on fine grain reconfigurable architectures, several research groups demonstrate their joint work on operating concepts for managing dynamic and partial reconfiguration, visualization of bitstreams and routing, presenting an application applying dynamic reconfiguration for video engines as well as work on minimization of reconfiguration data. Unique is that all the above four projects present their work using the same reconfigurable FPGA-based fabric called Erlangen slot machine that has also been built within one project just the purpose of experimenting with dynamic fine grain reconfiguration as an interdisciplinary platform. Josef Angermeier, Mateusz Majer, Jürgen Teich, Lars Braun, Tobias Schwalb, Philipp Graf, Michael Hübner 0001, Jürgen Becker 0001, Enno Lübbers, Marco Platzner, Christopher Claus, Walter Stechele, Andreas Herkersdorf, Markus Rullmann, Renate Merker |
FPL | 12 |
| 2008 | A comparison of embedded reconfigurable video-processing architecturesabstractUsing field programmable gate arrays (FPGAs) as accelerators for image or video processing operations and algorithms has gained increasing attention over the last few years. One reason for that is FPGAs are able to exploit both temporal and spatial parallelism. In this paper two platforms for FPGA-based real-time image and video processing are presented and compared against each other. With both of these platforms it is possible to update the physical resources during run-time by exploiting the dynamic partial reconfiguration capabilities of Xilinx Virtex FPGAs. The analysis of both platforms with respect to their benefits and draw-backs has led to the concept of an optimal FPGA-based dynamically and partially reconfigurable platform for real-time video and image processing. Christopher Claus, Walter Stechele, Matthias Kovatsch, Josef Angermeier, Jürgen Teich |
FPL | 2 |
| 2008 | A multi-platform controller allowing for maximum Dynamic Partial Reconfiguration throughputabstractDynamic and Partial Reconfiguration (DPR) is a special feature offered by Xilinx Field Programmable Gate Arrays (FPGAs), giving the designer the ability to reconfigure a certain portion of the FPGA during run-time without influencing the other parts. This feature allows the hardware to be adaptable to any potential situation. For some applications, such as video-based driver assistance [1], the time needed to exchange a certain portion of the device might be critical. This paper addresses problems, limitations and results of on-chip reconfiguration that enable the user to decide whether DPR is suitable for a certain design prior to its implementation. A method is therefore introduced to calculate the expected reconfiguration throughput and latency. In addition, an IP core is presented that enables fast on-chip DPR close to the maximum achievable speed. Compared to an alternative state-of-the art realization, an increase in speed by a factor of 58 can be obtained. Christopher Claus, Walter Stechele, Lars Braun, Michael Hübner 0001, Jürgen Becker 0001 |
FPL | 3 |
| 2007 | Using partial-run-time reconfigurable hardware to accelerate video processing in driver assistance systemabstractIn this paper we show a reconfigurable hardware architecture for the acceleration of video-based driver assistance applications in future automotive systems. The concept is based on a separation of pixel-level operations and high level application code. Pixel-level operations are accelerated by coprocessors, whereas high level application code is implemented fully programmable on standard PowerPC CPU cores to allow flexibility for new algorithms. In addition, the application code is able to dynamically reconfigure the coprocessors available on the system, allowing for a much larger set of hardware accelerated functionality than would normally fit onto a device. This process makes use of the partial dynamic reconfiguration capabilities of Xilinx Virtex FPGAs Christopher Claus, Johannes Zeppenfeld, Florian Helmut Müller, Walter Stechele |
DATE | 4 |
| 2007 | A new framework to accelerate Virtex-II Pro dynamic partial self-reconfigurationabstractThe Xilinx Virtex family of FPGAs provides the ability to perform partial run-time reconfiguration, also known as dynamic partial reconfiguration (DPR). Taking this concept one step further, partial dynamic self-reconfiguration becomes possible through the internal configuration access port (ICAP). In this paper a framework for lowering reconfiguration times using the combitgen tool to reduce the overhead found within bitstreams, along with a completely new, very simple and area efficient ICAP controller that is connected directly to the processor local bus (PLB) and is equipped with direct memory access (DMA) capabilities is presented. Using this PLB Master ICAP controller, it is possible to reach the maximum practical throughput that can be achieved with the ICAP interface of Virtex-II Pro devices. Compared to an alternative realization using the OPBHWICAP provided by Xilinx (a slave attachment on the on-chip peripheral bus), it is possible to achieve improvements concerning reconfiguration times by a factor of 20. Christopher Claus, Florian Helmut Müller, Johannes Zeppenfeld, Walter Stechele |
IPDPS | 4 |
| 2006 | AutoVision: flexible processor architecture for video-assisted drivingabstractSummary form only given. Future automotive security systems will benefit from visual scene analysis based on a fusion of video, infrared, and radar images. Today we have already functions like lane departure warning and automatic cruise control (ACC) for pretty well defined driving environments, such as highways and primary roads. Recent research activities concentrate on more complex environments, such as city traffic with a wide variety of traffic participants moving in an unpredictable manner, e.g. bikes, pedestrians, children, and even animals, and under changing weather and lighting conditions. The ITRS semiconductor roadmap for microelectronics forecasts a continued doubling of transistor capacity per chip every 2 to 2.5 years enabling billion transistor ASIC designs in the near future. Multi processor system on chip (MPSoC) solutions with 8, 16 or even more standard RISC CPU cores, mega-bytes of fast (ns access latencies) on-chip SRAM memories, giga-byte per second interconnect buses or NoC (network on chip) meshes, high-speed serial I/Os and, last but not least, million gate equivalent dedicated hardware accelerator functions in eFPGA (embedded field programmable gate array) logic are becoming reality on a single silicon substrate. Examples of current research projects shall illustrate our perception on how this tremendous increase in functionality and computational performance per chip area may impact automotive control unit (ACU) architectures for driver assistance applications. The AutoVision processor is a dynamically reconfigurable MPSoC prototype where video-specific pixel processing engines are on-the-fly loaded or exchanged without interrupting regular system operations. For the time being, pixel processing engines cover functions such as object edge detection or luminance segmentation, and are implemented as dedicated hardware accelerators to ensure real-time frame processing capabilities of the AutoVision processor. Dynamic replacement of processing engines ensures an automatic and area efficient adaptation to various driving conditions. Segmented objects are, in a subsequent step, characterized by means of standard MPEG-7 descriptors and entered as search criteria into traffic scene analysis databases. Goal is to obtain a clean distinction between passenger cars, trucks, and big rectangular traffic signs, and to identify pedestrians or bikers in complex traffic situations. The AutoVision processor project is supported by the German Research Foundation (DFG) in the special emphasis research programme "reconfigurable computing" Andreas Herkersdorf, Walter Stechele |
DATE | 2 |
| 2006 | Multithreaded virtual-memory-enabled reconfigurable hardware acceleratorsabstractAlthough naturally belonging to the user process, hardware parts of codesigned reconfigurable applications execute outside of the operating system (OS) process: they have neither unified memory abstraction with software nor system services provided by the OS. This imposes limitations on hardware and software interfacing, narrows available programming paradigms, and affects application portability. Advanced programming concepts, such as multithreading, usually demand additional activities on the programmer side, to perform memory transfers and enforce memory consistency. In this paper, we introduce a system layer (an OS extension relying on a system hardware extension) that provides: (1) unified virtual memory, (2) platform-agnostic interfacing, and (3) multithreaded execution, for hardware accelerators running within the same OS process with user software. The system layer releases software programmer and hardware designer from interfacing burdens and, still, achieves significant speedups over software with only limited overheads. Virtual-memory-enabled hardware accelerators benefit from all abstractions and services already available to software. To prove our concept in practice and demonstrate the ease of programming, we execute image processing and cryptography applications on reconfigurable systems-on-chip running GNU/Linux that supports virtual memory for multithreaded hardware accelerators Miljan Vuletic, Paolo Ienne, Christopher Claus, Walter Stechele |
FPT | 4 |
| 2006 | Organic Computing at the System on Chip LevelabstractThe evolution of CMOS technologies leads to integrated circuits with ever smaller device sizes, lower supply voltage, higher clock frequency and more process variability. Intermittent faults effecting logic and timing are becoming a major challenge for future integrated circuit designs. This paper presents an organic computing inspired SoC architecture which applies self-organization and self-calibration concepts to build reliable SoCs with lower overheads and a broader fault coverage than classical fault-tolerance techniques. We demonstrate the feasibility of this approach by example on the processing pipeline of a public-domain RISC CPU core Abdelmajid Bouajila, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf, Andreas Bernauer, Oliver Bringmann 0001, Wolfgang Rosenstiel |
VLSI-SoC | 3 |
| 2005 | A Coprocessor for Accelerating Visual Information ProcessingabstractVisual information processing will play an increasingly important role in future electronics systems. In many applications, e.g. video surveillance cameras, data throughput of microprocessors is not sufficient and power consumption is too high. Instruction profiling on a typical test algorithm has shown that pixel address calculations are the dominant operations to be optimized. Therefore AddressLib, a structured scheme for pixel addressing was developed, that can be accelerated by AddressEngine, a coprocessor for visual information processing. In this paper, the architectural design of AddressEngine is described, which in the first step supports a subset of the AddressLib. Dataflow and memory organization are optimized during architectural design. AddressEngine was implemented in an FPGA and was tested with the MPEG-7 global motion estimation algorithm. Results on processing speed and circuit complexity are given and compared to a pure software implementation. The next step will be the support for the full AddressLib, including segment addressing. An outlook on further investigations on dynamic reconfiguration capabilities is given. Walter Stechele, L. Alvado Cárcel, Stephan Herrmann 0002, J. Lidón Simón |
DATE | 1 |
| 2005 | Reduction of CMOS Power Consumption and Signal Integrity Issues by Routing OptimizationabstractThis paper suggests a methodology to decrease the power of a static CMOS standard cell design at layout level by focusing on switched capacitance. The term switched is the key: if a capacitance is not switched often, it may be high. If it is frequently switched, it should be minimized in order to reduce power consumption. This can be done by an algorithm based on forces that automatically optimizes the position and length of every single wire segment in a routed design. The forces are proportional to the toggle activities derived from a gate level simulation. The novelty is that this allows us to iteratively find a new topology for the wire segments. Our algorithm takes as input an already given, grid routed layout. Paul Zuber, Armin Windschiegl, Raúl Medina Beltrán de Otálora, Walter Stechele, Andreas Herkersdorf |
DATE | 4 |
| 2002 | MPEG-7 Binary Format for XML DatabstractSummary form only given. For the MPEG-7 standard, a binary format for the encoding of XML data has been developed that meets a set of requirements that was derived from a wide range of targeted applications. The resulting key features of the binary format are: high data compression (up to 98% for the document structure), provision of streaming, dynamic update of the document structure, random order of transmission of XML elements as well as fast random access of data entities in the compressed stream. To provide these functionalities, a novel, schema-aware approach was taken that exploits the knowledge of standardized MPEG-7 schema. The XML schema definition is used to assign codes to the individual children of an XML element. These codes are signalled in binary format to select nodes in the XML description tree. The binary format bit stream is organized as a sequence of access units. Each access unit can be decoded independently and contains information about a fragment of the description (fragment payload) and where to place the fragment in the current tree (context path). Compared to the standard text compressor ZIP, or the XML-optimized tool XMill, the MPEG-7 binary format achieves a 2-5 times better compression of the document structure and provides additional functionalities. These increase the flexibility and make it useful in broadcast applications and scenarios with limited bandwidth. Ulrich Niedermeier, Jörg Heuer, Andreas Hutter, Walter Stechele |
DCC | 4 |
| 2002 | An MPEG-7 tool for compression and streaming of XML dataabstractIn the course of work on the MPEG-7 standard, a binary format with special features for the encoding of XML data was required. These required key features are a high data compression ratio, provision for streaming, dynamic update of the document structure and fast random access of data entities in the compressed stream. To support these features, we propose a novel, schema-aware approach which exploits the knowledge of the standardized MPEG-7 syntax definition of the encoded XML document on the encoder and decoder side. The technique is part of the MPEG-7 standard. This paper gives an overview of the coding algorithm, including a comparison to standard (XML) compression tools. Ulrich Niedermeier, Jörg Heuer, Andreas Hutter, Walter Stechele, André Kaup |
ICME (1) | 4 |
| 2002 | Novel modeling techniques for RTL power estimationabstractIn this work, we propose efficient macromodeling techniques for RTL power estimation, based only on word and bit level switching information of the module inputs. We present practicable combinations of these two properties for the construction of power macro-models. It is demonstrated, that our developed models reduce the estimation error compared to the Hamming-distance model at least by 64%. The total average errors (compared to PowerMill) achieved over a wide range of test modules and input stimuli are less than 4.6%. This is comparable to complex models, which however, have to make use of several more signal properties. Michael Eiermann, Walter Stechele |
ISLPED | 2 |
| 2000 | A coprocessor architecture implementing the MPEG-4 visual core profile for mobile multimedia applicationsabstractThis paper describes a new VLSI coprocessor architecture implementing the MPEG-4 visual core profile for advanced mobile multimedia terminals and applications. An analysis of the algorithmic requirements and of the constraints in the mobile multimedia applications field leads to an architectural approach based on a RISC core processor and programmable yet specialized coprocessors. We focus on one coprocessor called MacroBlockEngine which supports the tools for the texture processing part in the coding algorithm. The architecture of the MacroBlockEngine is presented in detail and experimental results from the simulations and from the hardware synthesis are given. Andreas Hutter, Georg Giebel, Walter Stechele |
ISCAS | 3 |
| 1999 | A video segmentation algorithm for hierarchical object representations and its implementationabstractThis paper describes a segmentation algorithm for generating hierarchical object representations of images and image sequences. Starting from an object model, we describe the structure of the corresponding segmentation algorithm including all analysis methods applied. Besides the well-known color and motion analysis, we also show how to utilize shape information. Furthermore, we discuss the tradeoff between reducing the computational complexity and the quality of the segmentation results. Last, we present the implementation concept for our analysis model, which uses a special toolbox model. The toolbox provides a set of addressing schemes that are needed by low-level video processing tools. The low-level tools are functions that apply a single operation to all pixels in one frame. Using these addressing functions makes it easy to implement new video processing tools, which, when combined, form new analysis methods. The toolbox exists in C-code and is partially transferred into VHDL. Stephan Herrmann 0002, Hubert Mooshofer, Harald Dietrich, Walter Stechele |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 1997 | A flexible VLSI architecture for variable block size segment matching with luminance correctionabstractThis paper describes a flexible 25.6 Giga operations per second exhaustive search segment matching VLSI architecture to support evolving motion estimation algorithms as well as block matching algorithms of established video coding standards. The architecture is based on a 16/spl times/16 processor element (PE) array and a 12 kbyte on-chip search area RAM and allows concurrent calculation of motion vectors for 32/spl times/32, 16/spl times/16, 8/spl times/8 and 4/spl times/4 blocks and partial quadtrees (called segments)for a +/-32 pel search range with 100% PE utilization. This architecture supports object based algorithms by excluding pixels outside of video objects from the segment matching process as well as advanced algorithms like variable blocksize segment matching with luminance correction. A preprocessing unit is included to support halfpel interpolation and pixel decimation. The VLSI has been designed using VHDL synthesis and a 0.5 /spl mu/m CMOS technology. The chip will have a clock rate of 100 MHz (min.) allowing realtime variable blocksize segment matching of 4CIF video (704/spl times/576 pel) at 15 fps or luminance corrected variable blocksize segment matching at above CIF (352/spl times/288), 15 fps resolution. Peter M. Kuhn, Andreas Weisgerber, Robert Poppenwimmer, Walter Stechele |
ASAP | 4 |