VLDB 2026 Research / reviewers in the wild / expert
Holger Blume
dblp:71/5087
· DBLP profile ↗
49ranked-venue papers
7as first author
17since 2021 · last 2025
0000-0002-0640-6875ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Noise Reduction in Hearing-Aid Processors: Traditional Methods vs. Neural NetworksabstractMany deep neural networks (DNNs) have been applied lately in the field of speech enhancement. One particular subfield, where DNNs have shifted the boundaries of what is considered possible, is noise reduction, where the degrading effects of sounds interfering with speech are minimized. This is especially relevant for hearing impaired listeners, as their ability to understand speech in noisy circumstances is reduced. In contrast to traditional methods, which are known to improve speech quality, DNNs promise to also improve speech intelligibility. Due to the high computational complexity, DNNs have not yet been deployed on a hearing aid processor, constrained by frequencies up to 50 MHz and memory up to 2 MB. In this work we deploy a convolutional neural network (CNN) trained for noise reduction to a hearing-aid system-on-chip (SoC) developed at our institute. Real time capability is achieved by thorough optimization of the C -Code, leading to a speed up by a factor of 88 for the inference relevant layers when compared to a naïve C-Code implementation. The CNN approach is compared to an implementation of a traditional noise reduction method regarding their speech enhancement performance on white and complex noise and their computational cost. While both methods improve the speech quality measured with Perceptual Evaluation of Speech Quality (PESQ), only the CNN achieves a Short-Time Objective Intelligibility (STOI) improvement of 0.077 for complex noise. On the other hand, the CNN has a higher processor utilization of 60.1% compared to 23.5% for the traditional approach. Nonetheless, both methods are real time capable and consume only 3.3 mW for the CNN and 1.78 mW for the traditional approach, respectively. Simon C. Klein, Lando Rossol, Finn Venema, Sven Schönewald, Jens Karrenbauer, Holger Blume |
ASAP | 6 |
| 2025 | RRNS Arith Lib - An Open-Source Redundant Residue Number System Arithmetic VHDL LibraryabstractA Residue Number System (RNS) represents integers through a set of residues obtained by integer division using a predefined set of pairwise coprime moduli. RNS is well-known for enabling efficient carry-free arithmetic and representing large numbers with multiple shorter numerical values. By incorporating additional moduli, the system can be extended into an over-determined Redundant Residue Number System (RRNS), introducing redundancy that facilitates the detection and correction of bit errors. Tim Oberschulte, Enno Sievers, Holger Blume |
FPGA | 3 |
| 2025 | Modified Parabolic Synthesis for Hardware-Oriented Approximation of Unary FunctionsabstractThe approximation of unary functions such as sine and arctangent is an essential part of many digital signal processing applications, often requiring significant computational power and resources. Traditionally, trigonometric functions are computed using software approximations or dedicated hardware accelerators like CORDIC. This paper presents a novel methodology to efficiently compute unary functions in VLSI designs, based on parabolic synthesis by E. Hertz et al. In contrast to traditional parabolic synthesis, all parabolic sub-functions are combined using adders instead of costly multipliers. A corresponding hardware architecture is developed, demonstrating its effectiveness. Taking the sine function as a reference, approximation accuracy is characterized over a wide range of configurations and compared with alternative approaches. Finally, area and critical path is evaluated using the Skywater 130 nm technology node, demonstrating its strength against traditional parabolic synthesis and CORDIC. The results show significant improvements in both critical path and area across the full range of configurations, achieving more than 2× reduction in area while simultaneously doubling frequency compared to traditional parabolic synthesis having similar approximation accuracy. Viktor Schneider, Sven Schönewald, Holger Blume |
ISCAS | 3 |
| 2025 | ZuSE-KI-Mobil: AI Chip Design Platform for Automotive and Industrial Applications
Shaown Mojumder, Simon Friedrich, Emil Matús, Matthias Lüders, Martin Friedrich, Oliver Renke, Holger Blume, Markus Kock, Gregor Schewior, Darius Grantz, Jens Benndorf, Julian Höfer, Patrick Schmidt 0003, Jürgen Becker 0001, Nael Fasfous, Pierpaolo Morì, Hans-Jörg Vögel, Samira Ahmadifarsani, Leonidas Kontopoulos, Ulf Schlichtmann, Yun-Jin Li, Gerhard P. Fettweis |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | Enhancing a Hearing Aid Processor with ISA Extensions Supporting Flexible Fixed-Point FormatsabstractAs the number of individuals experiencing hearing problems rises, research in this area is increasing. In particular, algorithms for enhancing sound quality are becoming more advanced. However, the energy consumption of hearing aids is restricted to a few milliwatts. New hearing aid processors and hardware must be developed to tackle this issue. Therefore, this paper presents and evaluates custom hardware units for hearing aids suitable for flexible fixed-point formats. The units are added to the instruction set architecture (ISA) of a Tensilica Fusion G6. This is one of two high-level programmable application-specific instruction-set processors (ASIP) integrated into the Smart Hearing Aid Processor (SmartHeaP), a hearing aid system-on-chip (SoC) fabricated in 22nm fully depleted silicon on insulator (FD-SOI) technology. As many audio algorithms operate in the frequency domain, complex-domain units like a complex multiply-accumulate (CMAC) unit are introduced first. Additionally, coordinate rotation digital computer (CORDIC) operations have been added to speed up nonlinear functions such as logarithms, another function frequently used in hearing aid applications. The implemented extensions are integrated into a MATLAB fixed-point framework to simplify access to the ISA extensions for algorithm developers. It automatically generates fixed-point C code with direct access to the added instructions, reducing development time and complexity. Integrating the seven proposed instructions with corresponding register files increases the core area by 20% or 0.065 mm2in a 22nm front-end synthesis. On the other hand, these extensions reduce the cycle count for two evaluated hearing aid algorithms, a beamformer and a loudness compensator, by 93% and 38%. A reduction of up to 84% is achieved for standalone mathematical functions, like logarithmic calculations. These performance improvements directly translate into power savings by allowing a reduced clock frequency. Jens Karrenbauer, Sven Schönewald, Simon C. Klein, Holger Blume |
ASAP | 4 |
| 2024 | Multi-Level Prototyping of a Vertical Vector AI Processing SystemabstractModern embedded systems must be designed carefully to cope with the complexity and real-time requirements of modern AI (Artificial Intelligence) driven automotive applications, such as Advanced Driver-Assistance Systems (ADAS). Despite increasing complexity, the time to market is decreasing. In this work, a SystemC-based Virtual Prototype of a neural network processing platform is exploited to bypass the limitations of standalone instruction set simulators (ISS) and FPGA prototyping. The processing platform under test is based on a novel massive parallel vector processor architecture coupled with a RISC- V control core that runs widely used convolutional neural networks (CNNs) for object detection. The paper discusses the variations and appropriateness of the three prototyping methods outlined, demonstrating how the Virtual Prototype can address the aforementioned constraints, resulting in a 2.07x increase in accuracy, 16x greater configurations, and more profound insights into the system compared to standalone and FPGA prototyping. Frederik Kautz, Sven Gesper, Gia Bao Thieu, Hans-Martin Blüthgen, Holger Blume, Guillermo Payá-Vayá |
ASAP | 5 |
| 2024 | Design Space Exploration of Semantic Segmentation CNN SalsaNext for Constrained ArchitecturesabstractThe growing use of LiDAR systems and constrained computing resources in the automotive sector require efficient LiDAR processing. SalsaNext, a convolutional neural network for semantic segmentation, is a promising candidate for deployment in that area. To extend the research regarding its quantization and investigate its adaptability to constrained resources, a design space exploration is performed. The design space, defined by model size, topology, and compute precision, is evaluated on a Jetson AGX Orin regarding classification accuracy, latency, and energy efficiency. The results display a trade-off between classification accuracy and runtime. The smallest model evaluated in INT8 on the GPU provides the smallest latency of 14.48 ms with a mloU score of 43.2%. A mloU score of 47.7% at a latency of 26.92 ms can be achieved with the medium-sized model and modified topology evaluated in INT8 on the DLA. The medium-sized model with modified topology provides good classification accuracy evaluated in FP32 on the GPU with a mloU score of 55.2% in 67.85 ms. Oliver Renke, Christoph Riggers, Jens Karrenbauer, Holger Blume |
ASAP | 4 |
| 2024 | PTP-Synchronized Tri-Level Sync Generation for Networked Multi-Sensor SystemsabstractSynchronization of sensor devices is crucial for concurrent data acquisition. Numerous protocols have emerged for this task, and for some multi-sensor setups to operate synchronized, a conversion between deployed protocols is needed. This paper presents a bare-metal implementation of a Tri- Level Sync signal generator on a microcontroller unit (MCU) synchronized to a master clock via the IEEE 1588 Precision Time Protocol (PTP). Cameras can be synchronized by locking their frame generators to the Tri-Level Sync signal. As this synchronization depends on a stable analog signal, a careful design of the signal generation based on a PTP-managed clock is required. The limited tolerance of a camera to clock frequency adjustments for continuous operations imposes rate-limits on the PTP-controller. Simulations using a software model demonstrate the resulting controller instabilities from rate-limiting. This problem is addressed by introducing a linear prediction mode to the controller, which estimates the realizable offset change during rate-limited frequency alignment. By adjusting the frequency in a timely manner, a large overshoot of the controller can be avoided. Additionally, a cascading controller design that decouples the PTP from the clock update rate proved to be advantageous to increase the camera’s tolerable frequency change. This paper demonstrates that a MCU is a viable platform to perform PTP-synchronized Tri-Level Sync generation. Our open source implementation is available for use by the research community at https://github.com/IMS-AS-LUH/t41-tri-sync-ptp. Christoph Riggers, Jens Schleusner, Oliver Renke, Holger Blume |
RTCSA | 4 |
| 2023 | Exploiting Subword Permutations to Maximize CNN Compute Performance and EfficiencyabstractNeural networks (NNs) are quantized to decrease their computational demands and reduce their memory foot-print. However, specialized hardware is required that supports computations with low bit widths to take advantage of such optimizations. In this work, we propose permutations on subword level that build on top of multi-bit-width multiply-accumulate operations to effectively support low bit width computations of quantized NNs. By applying this technique, we extend the data reuse and further improve compute performance for convolution operations compared to simple vectorization using SIMD (single-instruction-multiple-data). We perform a design space exploration using a cycle accurate simulation with MobileNet and VGG16 on a vector-based processor. The results show a speedup of up to$3.7\times$and a reduction of up to$1.9\times$for required data transfers. Additionally, the control overhead for orchestrating the computation is decreased by up to$3.9\times$. Michael Beyer, Sven Gesper, Andre Guntoro, Guillermo Payá-Vayá, Holger Blume |
ASAP | 5 |
| 2023 | The ZuSE-KI-Mobil AI Accelerator SoC: Overview and a Functional Safety PerspectiveabstractZuSE-KI-Mobil (ZuKIMo) is a nationally funded research project, currently in its intermediate stage. The goal of the ZuKIMo project is to develop a new System-on-Chip (SoC) platform and corresponding ecosystem to enable efficient Artificial Intelligence (AI) applications with specific requirements. With ZuKIMo, we specifically target applications from the mobility domain, i.e. autonomous vehicles and drones. The initial ecosystem is built by a consortium consisting of seven partners from German academia and industry. We develop the SoC platform and its ecosystem around a novel AI accelerator design. The customizable accelerator is conceived from scratch to fulfill the functional and non-functional requirements derived from the ambitious use cases. A tape-out in 22 nm FDX-technology is planned in 2023. Apart from the System-on-Chip hardware design itself, the ZuKIMo ecosystem has the objective of providing software tooling for easy deployment of new use cases and hardware-CNN co-design. Furthermore, AI accelerators in safety-critical applications like our mobility use cases, necessitate the fulfillment of safety requirements. Therefore, we investigate new design methodologies for fault analysis of Deep Neural Networks (DNNs) and introduce our new redundancy mechanism for AI accelerators. Fabian Kempf, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Nael Fasfous, Alexander Frickenstein, Hans-Jörg Vögel, Simon Friedrich, Robert Wittig, Emil Matús, Gerhard P. Fettweis, Matthias Lüders, Holger Blume, Jens Benndorf, Darius Grantz, Martin Zeller, Dietmar Engelke, Karl-Heinz Eickel |
DATE | 13 |
| 2023 | ZuSE Ki-Avf: Application-Specific AI Processor for Intelligent Sensor Signal Processing in Autonomous DrivingabstractModern and future AI-based automotive applications, such as autonomous driving, require the efficient real-time processing of huge amounts of data from different sensors, like camera, radar, and LiDAR. In the ZuSE-KI-AVF project, multiple university, and industry partners collaborate to develop a novel massive parallel processor architecture, based on a cus-tomized RISC-V host processor, and an efficient high-performance vertical vector coprocessor. In addition, a software development framework is also provided to efficiently program AI-based sensor processing applications. The proposed processor system was verified and evaluated on a state-of-the-art UltraScale+ FPGA board, reaching a processing performance of up to 126.9 FPS, while executing the YOLO-LITE CNN on 224x224 input images. Further optimizations of the FPGA design and the realization of the processor system on a 22nm FDSOI CMOS technology are planned. Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá, Christoph Riggers, Oliver Renke, Till Fiedler, Jakob Marten, Tobias Stuckenberg, Holger Blume, Christian Weis, Lukas Steiner, Chirag Sudarshan, Norbert Wehn, Lennart M. Reimann, Rainer Leupers, Michael Beyer, Daniel Köhler, Alisa Jauch, Jan Micha Borrmann, Setareh Jaberansari, Tim Berthold, Meinolf Blawat, Markus Kock, Gregor Schewior, Jens Benndorf, Frederik Kautz, Hans-Martin Blüthgen, Christian Sauer 0001 |
DATE | 9 |
| 2023 | Fault Detection on Multi COTS FPGA Systems for Physics Experiments on the International Space StationabstractField-programmable gate arrays (FPGAs) in space applications come with the drawback of radiation effects, which inevitably will occur in devices of small process size. This also applies to the electronics of the Bose Einstein Condensate and Cold Atom Laboratory (BECCAL) apparatus, which is planned to operate on the International Space Station for several years. A total of more than 100 FPGAs distributed in the setup will be used for high-precision control of specialized sensors and actuators at nanosecond scale. Due to the large amount of devices in BECCAL, commercial off-the-shelf (COTS) FPGAs are used which are not radiation hardened. In this work, we detect and mitigate radiation effects in an application specific COTS-FPGA-based communication network. For that redundancy is integrated into the design while the firmware is optimized to stay within the FPGA's resource constraints. A redundant integrity checker module is developed which can notify preceding network devices about data and configuration bit errors. The firmware is evaluated by injecting faults into data and configuration registers in simulation and real hardware. The FPGA resource usage of the firmware is cut down by more than half, enabling the use of double modular redundancy for the switching fabric. Together with the triple modular redundancy protected integrity checker, this combination fully prevents silent data corruptions in the design as shown in simulations and by injecting faults in hardware using the Intel Fault Injection FPGA IP Core while staying in the resource limitation of a COTS FPGA. Tim Oberschulte, Jakob Marten, Holger Blume |
FPGA | 3 |
| 2023 | Improved Multi-Scale Grid Rendering of Point Clouds for Radar Object Detection NetworksabstractArchitectures that first convert point clouds to a grid representation and then apply convolutional neural networks achieve good performance for radar-based object detection. However, the transfer from irregular point cloud data to a dense grid structure is often associated with a loss of information, due to the discretization and aggregation of points. In this paper, we propose a novel architecture, multi-scale KPPillarsBEV, that aims to mitigate the negative effects of grid rendering. Specifically, we propose a novel grid rendering method, KPBEV, which leverages the descriptive power of kernel point convolutions to improve the encoding of local point cloud contexts during grid rendering. In addition, we propose a general multi-scale grid rendering formulation to incorporate multi-scale feature maps into convolutional backbones of detection networks with arbitrary grid rendering methods. We perform extensive experiments on the nuScenes dataset and evaluate the methods in terms of detection performance and computational complexity. The proposed multi-scale KPPillarsBEV architecture outperforms the baseline by 5.37% and the previous state of the art by 2.88% in Car AP4.0 (average precision for a matching threshold of 4 meters) on the nuScenes validation set. Moreover, the proposed single-scale KPBEV grid rendering improves the Car AP4.0 by 2.90% over the baseline while maintaining the same inference speed. Daniel Köhler, Maurice Quach, Michael Ulrich, Frank Meinl, Bastian Bischoff, Holger Blume |
FUSION | 6 |
| 2023 | Online Quantization Adaptation for Fault-Tolerant Neural Network Inference
Michael Beyer, Jan Micha Borrmann, Andre Guntoro, Holger Blume |
SAFECOMP | 4 |
| 2023 | Dynamic Model-Based Safety Margins for High-Density Matrix Headlight SystemsabstractReal-time masking of vehicles in a dynamic road environment is a demanding task for adaptive driving beam systems of modern headlights. Next-generation high-density matrix headlights enable precise, high-resolution projections, while advanced driver assistance systems enable detection and tracking of objects with high update rates and low-latency estimation of the pose of the ego-vehicle. Accurate motion tracking and precise coverage of the masked vehicles are necessary to avoid glare while maintaining a high light throughput for good visibility. Safety margins are added around the mask to mitigate glare and flicker caused by the update rate and latency of the system. We provide a model to estimate the effects of spatial and temporal sampling on the safety margins for high- and low-density headlight resolutions and different update rates. The vertical motion of the ego-vehicle is simulated based on a dynamic model of a vehicle suspension system to model the impact of the motion-to-photon latency on the mask. Using our model, we evaluate the light throughput of an actual matrix headlight for the relevant corner cases of dynamic masking scenarios depending on pixel density, update rate, and system latency. We apply the masks provided by our model to a high beam light distribution to calculate the loss of luminous flux and compare the results to a light throughput approximation technique from the literature. Jens Schleusner, Holger Blume, Sebastian Lampe |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Self-Supervised Velocity Estimation for Automotive Radar Object Detection NetworksabstractThis paper presents a method to learn the Cartesian velocity of objects using an object detection network on automotive radar data. The proposed method is self-supervised in terms of generating its own training signal for the velocities. Labels are only required for single-frame, oriented bounding boxes (OBBs). Labels for the Cartesian velocities or contiguous sequences, which are expensive to obtain, are not required. The general idea is to pre-train an object detection network without velocities using single-frame OBB labels, and then exploit the network’s OBB predictions on unlabelled data for velocity training. In detail, the network’s OBB predictions of the unlabelled frames are updated to the timestamp of a labelled frame using the predicted velocities and the distances between the updated OBBs of the unlabelled frame and the OBB predictions of the labelled frame are used to generate a self-supervised training signal for the velocities. The detection network architecture is extended by a module to account for the temporal relation of multiple scans and a module to represent the radars’ radial velocity measurements explicitly. A twostep approach of first training only OBB detection, followed by training OBB detection and velocities is used. Further, a pre-training with pseudo-labels generated from radar radial velocity measurements bootstraps the self-supervised method of this paper. Experiments on the publicly available nuScenes dataset show that the proposed method almost reaches the velocity estimation performance of a fully supervised training, but does not require expensive velocity labels. Furthermore, we outperform a baseline method which uses only radial velocity measurements as labels. Daniel Niederlöhner, Michael Ulrich, Sascha Braun, Daniel Köhler, Florian Faion, Claudius Gläser, André Treptow, Holger Blume |
IV | 8 |
| 2022 | Predictive accuracy of CNN for cortical oscillatory activity in an acute rat model of parkinsonism
Ali Abdul Nabi Ali, Mesbah Alam, Simon C. Klein, Nicolai Behmann, Joachim K. Krauss, Theodor Doll, Holger Blume, Kerstin Schwabe |
Neural Networks | 7 |
| 2019 | Statistical Performance Prediction for Multicore Applications Based on Scalability CharacteristicsabstractMulticore processors serve as target platforms in a broad variety of applications ranging from high-performance computing to embedded mobile computing and automotive. But, the required parallel programming opens up a huge design space of parallelization strategies each with potential bottlenecks. Therefore, an early estimation of an application's performance is a desirable development tool. However, out-of-order execution, superscalar instruction pipelines, as well as communication costs and (shared-) cache effects essentially influence the performance of parallel programs. While offering a good modeling and simulation speed, analytic models provide moderate prediction results so far. Virtual prototyping requires a time-consuming simulation, but produces better accuracy. Furthermore, even existing statistical methods often require detailed knowledge of the hardware for characterization. In this work, we present a concept and its evaluation for a statistical approach for performance prediction based on abstract runtime parameters, which describe an application's scalability behavior and can be extracted from profiles without user input. These scalability parameters not only include information on the interference of software demands and hardware capabilities, but indicate bottlenecks as well. Depending on the database setup, we achieve a competitive accuracy of 20 % mean prediction error (11 % median), which we also proof in a case study. Oliver Jakob Arndt, Matthias Lüders, Holger Blume |
ASAP | 3 |
| 2019 | Probabilistic 3D Point Cloud Fusion on Graphics Processors for Automotive (Poster)
Nicolai Behmann, Yihan Cheng, Jens Schleusner, Holger Blume |
FUSION | 4 |
| 2019 | KAVUAKA: A Low Power Application Specific Hearing Aid ProcessorabstractThe integration of application specific instruction set processors (ASIPs) in hearing aids requires various architectural customizations and software-side optimizations in order to meet the stringent power consumption constraints and processing performance demands. This paper presents the KAVUAKA application specific hearing aid processor and its ASIC integration as a system on chip (SoC). The final system contains four KAVUAKA processor cores and ten co-processors. Each of these processors and co-processors were individually customized and differ in their data path width. The processors are organized in two clusters, which share memories, an audio interface, co-processors and a serial interface. With this system, different hearing aid systems are evaluated in terms of performance, power and area by activating different processor and co-processor combinations. A 40 nm low power technology was used to build this research hearing aid system. The die size is 3.6 mm2with less than 1 mm2per core. The measured average power consumption is less than 1 mW per core. Lukas Gerlach 0001, Guillermo Payá-Vayá, Holger Blume |
VLSI-SoC | 3 |
| 2019 | FLINT+: A runtime-configurable emulation-based stochastic timing analysis framework
Moritz Weißbrich, Lukas Gerlach 0001, Holger Blume, Ardalan Najafi, Alberto García Ortiz, Guillermo Payá-Vayá |
Integr. | 3 |
| 2019 | Dynamic self-reconfiguration of a MIPS-based soft-core processor architecture
Stephan Nolting, Guillermo Payá-Vayá, Florian Giesemann, Holger Blume, Sebastian Niemann, Christian Müller-Schloer |
J. Parallel Distributed Comput. | 4 |
| 2019 | Online stereo camera calibration for automotive vision based on HW-accelerated A-KAZE-feature extraction
Nico Mentzer, Jannik Mahr, Guillermo Payá-Vayá, Holger Blume |
J. Syst. Archit. | 4 |
| 2018 | A HOG-based Real-time and Multi-scale Pedestrian Detector Demonstration System on FPGAabstractPedestrian detection will play a major role in future driver assistance and autonomous driving. One powerful algorithm in this field uses HOG features to describe the specific properties of pedestrians in images. To determine their locations, features are extracted and classified window-wise from different scales of an input image. The results of the classification are finally merged to remove overlapping detections. The real-time execution of this method requires specific FPGA- or ASIC-architectures. Recent work focused on accelerating the feature extraction and classification. Although merging is an important step in the algorithm, it is only rarely considered in hardware implementations. A reason for that could be its complexity and irregularity that is not trivial to implement in hardware. In this paper, we present a new bottom-up FPGA architecture that maps the full HOG-based algorithm for pedestrian detection including feature extraction, SVM classification, and multi-scale processing in combination with merging. For that purpose, we also propose a new hardware-optimized merging method. The resulting architecture is highly efficient. Additionally, we present an FPGA-based full real-time and multi-scale pedestrian detection demonstration system. Jan Dürre, Dario Paradzik, Holger Blume |
FPGA | 3 |
| 2018 | Low-Cost Channel Sounder Design Based on Software-Defined Radio and OFDMabstractIn this paper we describe the design of a low-cost and portable channel sounding system. The baseband signal processing in both the transmitter and the receiver are done in software, while the up/down conversion are done by means of a Software Defined Radio (SDR) frontend. The modulation procedure is based on the Orthogonal Frequency Division Multiplexing (OFDM) technique due to its robustness against inter-symbol interference (ISI). The sounder is designed to overcome some well known important issues of any OFDM system. For example, it does not need extra resources neither to guarantee a ISI-free transmission by means of a guard interval, nor for the synchronization by means of a preamble. Moreover, the excitation signal is designed to have a peak-to-average power ratio (PAPR) close to 0 dB while maintaining a flat spectrum. The system is calibrated and validated under laboratory conditions prior to the measurement of an air-to-ground channel in the C frequency band, the results are presented and discussed. Yasser Samayoa, Markus Kock, Holger Blume, Jörn Ostermann |
VTC Fall | 3 |
| 2017 | Real-time implementation of a GMM-based binaural localization algorithm on a VLIW-SIMD processorabstractLocalization algorithms have become of considerable interest for robot audition, acoustic navigation, teleconferencing, speaker localization, and many other applications over the last decade. In this paper, we present a real-time implementation of a Gaussian mixture model (GMM) based probabilistic sound source localization algorithm for a low-power VLIW-SIMD processor for hearing devices. The algorithm has been proven to allow for robust localization of multiple sound sources simultaneously in reverberant and noisy environments. Real-time computation for audio frames of 512 samples at 16 kHz was achieved by introducing algorithmic optimizations and hardware customizations. To the best of our knowledge, this is the first real-time capable implementation of a computationally complex GMM-based sound source localization algorithm on a low-power processor. The resulting estimated core area without consideration of memory in 40nm low-power TSMC technology is 188,511 pm2. Christopher Seifert, Joachim Thiemann, Lukas Gerlach 0001, Tobias Volkmar, Guillermo Payá-Vayá, Holger Blume, Steven van de Par |
ICME | 6 |
| 2017 | FPGA emulation methodology for fast and accurate power estimation of embedded processors
Sebastian Hesselbarth, Gregor Schewior, Holger Blume |
J. Syst. Archit. | 3 |
| 2017 | Small footprint synthesizable temperature sensor for FPGA devices
Guillermo Payá-Vayá, Christopher Bartels, Holger Blume |
J. Syst. Archit. | 3 |
| 2016 | FPGA-based frequency estimation of a DFB laser using Rb spectroscopy for space missionsabstractStable laser light sources are necessary for atom interferometry based experiments on space platforms such as sounding rockets or satellites. Diode lasers, commonly used as light sources in these experiments, lack long term stability, therefore external frequency stabilization systems are required. Commonly used spectroscopy based systems require the correct optical atomic reference transition to be manually identified before analog circuits can start tracking this transition. For automated systems, latencies below 100 fis are required making the design of such systems challenging. In this paper a scalable architecture for an automated FPGAbased frequency estimation using correlation based matching algorithms is introduced. The architecture is variable in terms of the matching algorithm (SAD/SSD/CC) as well as the number of matching cores. Due to the selected algorithm a frequency estimation error below 0.9 MHz is reached. The scalable architecture allows the variation of the execution time between 25.5 ms and 60.1 μs depending on the number of cores between 1 and 512. To reach the latency constraint, 277 or more cores have to be used. The configuration featuring minimal latency and optimal accuracy (512 SAD cores) requires less than 60% logic elements, 35% registers and 10% memory bits of an Alters Cyclone IV (EP4CE115) FPGA. To investigate the latency as well as the FPGA resource and power consumption, a full design space exploration of the laser frequency estimation architecture is performed. Christian Spindeldreier, Thijs J. Wendrich, Ernst M. Rasel, Wolfgang Ertmer, Holger Blume |
ASAP | 5 |
| 2015 | FLINT: layout-oriented FPGA-based methodology for fault tolerant ASIC design
Rochus Nowosielski, Lukas Gerlach 0001, Stephan Bieband, Guillermo Payá-Vayá, Holger Blume |
DATE | 5 |
| 2014 | ASEV - Automatic situation assessment for event-driven video analysisabstractMany complex maneuvers involving aircraft, vehicles and persons are carried out at airport aprons. Manual video surveillance used for safety and security purposes is inefficient and privacy protection must be guaranteed. In this paper, we propose a system named ASEV that automatically assesses situations for airport surveillance. It combines four main components: a low-level image processing unit based on a new hardware implementation to extract features in real time, a high-level image processing unit for scene analysis, a real-time inference engine for scene understanding, and a data protection stage for log encryption. In addition, four often neglected aspects are successfully addressed: two-way communication between system and operator, power consumption, monitored people privacy and operator activity control. Extensive evaluation at a real airport shows that the proposed system improves the operator performance with sound and visual alerts based on the automatic assessment of various events. Michele Fenzi, Jörn Ostermann, Nico Mentzer, Guillermo Payá-Vayá, Holger Blume, Tu Ngoc Nguyen, Thomas Risse 0001 |
AVSS | 5 |
| 2012 | Evaluation of Inertial Sensor Fusion Algorithms in Grasping Tasks Using Real Input Data: Comparison of Computational Costs and Root Mean Square ErrorabstractSensor fusion is an important computation step for acquiring reliable orientation information from inertial sensors. These sensors are very attractive in order to achieve a mobile capturing of human movements, which is desired for application in sports or rehabilitation. Commercial inertial sensors with small form factors and low power consumption can be used for capturing without any interference. There are several common techniques for calculating orientation data based on RAW sensor data. This paper gives an overview of the computational effort and achievable accuracy of integration algorithms, vector observation algorithms and Kalman filter algorithms for inertial sensor fusion. The sensor data were compared against an optical motion capturing system. The considered application is the capturing of arm movements during grasping tasks in stroke rehabilitation. Therefore, the algorithms are evaluated based on corresponding real world input data. The provided benchmark compares the sensor fusion algorithms in terms of computational cost and orientation estimation error. Hans-Peter Brückner, Christian Spindeldreier, Holger Blume, E. Schoonderwaldt, Eckart Altenmüller |
BSN | 3 |
| 2011 | Instruction set extension for high throughput disparity estimation in stereo image processingabstractThis paper presents the implementation and evaluation of an application-specific instruction set for a customizable RISC-processor for very high throughput stereo image processing. Compared to the base processor the overall processing time is accelerated by a factor of over 130, while the processor silicon area requirement increases only by a factor of 2.9. The processor has been enhanced with algorithm-specific extensions (i.e. special functional units), as well as with extensions that are not restricted to a specific algorithm (e.g. single-instruction-multiple-data). Hereby, the special functional units account for 50% of the speed-up, but less than 14% of the processor silicon area requirement. The proposed processor extensions thereby sustain the full flexibility of a programmable processor while enabling disparity estimation of 640×480 stereo video sequences at 20 fps when running at a clock frequency of 373 MHz. Christian Banz, Carsten Dolar, Fabian Cholewa, Holger Blume |
ASAP | 4 |
| 2011 | A FPGA architecture for real-time processing of variable-length FFTSabstractA new FFT architecture for real-time implementation of large FFTs is presented. The architecture supports both, high throughput and variable-length processing capabilities. The implementation is configurable at run-time, in order to compute power-of-two length ranging from 16 to 2n. It supports efficient integration of data scaling techniques. A radix-23DIT FFT algorithm is derived, which minimizes the number of multipliers and supports simple reordering. Stefan Langemeyer, Peter Pirsch, Holger Blume |
ICASSP | 3 |
| 2010 | A forwarding-sensitive instruction scheduling approach to reduce register file constraints in VLIW architecturesabstractThis paper presents a forwarding-based approach to increase the code compaction and consequently the processing performance of VLIW media-processors that implement monolithic or partitioned register file (RF) organizations with reduced number of read/write ports. This approach exploits the forwarding mechanism implemented in common pipelined VLIW architectures to reduce the number of RF accesses, which is one of the main limiting factors of the code compaction process. This RF access reduction enables a higher instruction scheduling efficiency and eventually decreases the power consumption, without requiring extra hardware. A forwarding-sensitive code generation algorithm based on an enhanced list scheduling algorithm is described in detail. In addition, three case studies are presented, where the proposed scheduling algorithm leads to performance improvements of up to 8.4% when running common image and video codec tasks on a generic VLIW architecture. This is attractively close to the maximum performance improvement (11.4%) that can be achieved when investing in hardware by using a RF with twice the number of ports. Guillermo Payá-Vayá, Javier Martín-Langerwerf, Holger Blume, Peter Pirsch |
ASAP | 3 |
| 2008 | Design flow for embedded FPGAs based on a flexible architecture templateabstractModern digital signal processing applications have an increasing demand for computational power while needing to preserve low power dissipation and high flexibility. For many applications, the growth of algorithmic complexity is already faster than the growth of computational power provided by discrete general purpose processors. A typical approach to address this problem is the combination of a processor core with dedicated accelerators. Since changes in standards or algorithms can change the demands on the accelerators, an attractive alternative to highly customised VLSI- macros is the use of reconfigurable embedded FPGAs (eFPGAs). First commercial products combining a general purpose processor core and an embedded FPGA recently emerged (e.g. Stretch S6000 Menta eFPGA- augmented CPUs). For many digital signal processing applications, a significantly improved efficiency in terms of power dissipation, throughput and chip area can be achieved by tailoring both the processor core and the reconfigurable accelerator to the given application domain. In this work, a methodology to design highly customisable eFPGA-architectures starting from a high level description is presented. The design framework elaborated during this work enables a physically optimised VLSI-design of the specified eFPGA and aims to support simulation of the according eFPGA-macros both on a functional and netlist-level by providing an elementary configuration tool based on the same high level description as the eFPGA-architecture. Bernd Neumann, Thorsten von Sydow, Holger Blume, Tobias G. Noll |
DATE | 3 |
| 2008 | Design of a Pareto-optimization environment and its application to motion estimationabstractThe characteristics of modern video signal processing algorithms are significantly influenced by a multitude of different configuration parameters. Hence, the selection of an optimum parameter set becomes a difficult task, since often various quality metrics have to be regarded, leading to a so-called multi-objective optimization problem. Furthermore, the high computational effort typically restricts the maximum possible number of simulated parameter configurations. Therefore, a flexible environment for multi-objective optimization of configuration parameters has been elaborated which uses evolutionary algorithms to efficiently explore the design-space. This environment is applied here for a detailed analysis and Pareto-optimization of a complex state-of-the-art motion estimation algorithm being used for frame rate conversion. Jörg von Livonius, Holger Blume, Tobias G. Noll |
MMSP | 2 |
| 2008 | OpenMP-based parallelization on an MPCore multiprocessor platform - A performance and power analysis
Holger Blume, Jörg von Livonius, Lisa Rotenberg, Tobias G. Noll, Harald Bothe, Jörg Brakensiek |
J. Syst. Archit. | 1 |
| 2008 | A Scalable Packet Sorting Circuit for High-Speed WFQ Packet SchedulingabstractA novel implementation of a tag sorting circuit for a weighted fair queueing (WFQ) enabled Internet protocol (IP) packet scheduler is presented. The design consists of a search tree, matching circuitry, and a custom memory layout. It is implemented using 130-nm silicon technology and supports quality of service (QoS) on networks at line speeds of 40 Gb/s, enabling next generation IP services to be deployed. Kieran McLaughlin, Sakir Sezer, Holger Blume, Xin Yang 0010, Friederich Kupzog, Tobias G. Noll |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | A Power Estimation Model for an FPGA-based Softcore ProcessorabstractWe describe the application of a hybrid functional level power analysis (FLPA) and instruction level power analysis (ILPA) approach to a processor model implemented on an FPGA. This technique enables the estimation of the task specific power consumption of the modeled processor, in our case a LEON2, very early during a system design flow, based on the software which will run on it. The FLPA/ILPA model used during our work as well as the test scenarios and the measured results are described. Later, the function block separation and the power consumption modeling are discussed. Finally, the model is validated by benchmarking. The obtained model is promising in the sense that a) its estimations are close (4% on average) to the measured data, and b) the model structure is similar to that of hardcore processors which is not a trivial result. Peter Zipf, Heiko Hinkelmann, Manfred Glesner, Holger Blume, Tobias G. Noll |
FPL | 5 |
| 2007 | Hybrid functional- and instruction-level power modeling for embedded and heterogeneous processor architectures
Holger Blume, Daniel Becker 0001, Lisa Rotenberg, Martin Botteck, Jörg Brakensiek, Tobias G. Noll |
J. Syst. Archit. | 1 |
| 2007 | Application of deterministic and stochastic Petri-Nets for performance modeling of NoC architectures
Holger Blume, Thorsten von Sydow, Daniel Becker 0001, Tobias G. Noll |
J. Syst. Archit. | 1 |
| 2006 | Quantitative Analysis of Embedded FPGA-Architectures for ArithmeticabstractEmbedding FPGAs (eFPGAs) in modern SoCs provides a high amount of flexibility while high-throughput digital signal processing algorithms can be realised efficiently. An analysis of eFPGA architectures and corresponding structural elements is presented to determine the optimisation potential for eFPGAs tailored to an arithmetic oriented application domain. The applied design flow incorporating an automated layout generation approach and the utilised simulation environment is discussed. An eFPGA macro designed and realised for arithmetic oriented applications is quantitatively compared to an actual commercial FPGA in terms of area, power consumption and delay time. It can be shown that this optimised eFPGA macro outperforms a state of the art commercial device for a couple of arithmetic operators which are commonly applied in arithmetic datapaths Thorsten von Sydow, Bernd Neumann, Holger Blume, Tobias G. Noll |
ASAP | 3 |
| 2006 | Design and analysis of matching circuit architectures for a closest match lookupabstractThis paper investigates the implementation of a number of circuits used to perform a high speed closest value match lookup. The design is targeted particularly for use in a search trie, as used in various networking lookup applications, but can be applied to many other areas where such a match is required. A range of different designs have been considered and implemented on FPGA. A detailed description of the architectures investigated is followed by an analysis of the synthesis results Kieran McLaughlin, Friederich Kupzog, Holger Blume, Sakir Sezer, Tobias G. Noll, John V. McCanny |
IPDPS | 3 |
| 2004 | Segmentation in the loop: an iterative object-based algorithm for motion estimationabstractMotion estimation algorithms are a key component for multimedia systems and optimization of these algorithms is still a topic of current research. Promising approaches try to integrate into the motion estimation process besides pure grey level similarities further types of information, contained in the image. Due to the moderate quality of this additional information the integration has to be performed rather conservatively in order to reduce the risk of an even dramatic degradation of the vector field quality in some cases. Up to now there is no robust algorithm available, which yields a noticeable improvement for all types of motion and image scenes, without causing a loss of quality in critical situations. Within the scope of this contribution the application of high performance segmentation for the enhancement of motion vector fields is analyzed. Starting from these results a new iterative concept for object based motion estimation is developed, which combines the results of a classic motion estimation with the information of image segmentation and features a high robustness against segmentation errors. The results of this new algorithm are analyzed on the basis of different objective evaluation criterions and compared to classic motion estimation algorithms. Holger Blume, Jörg von Livonius, Tobias G. Noll |
VCIP | 1 |
| 2002 | Model-Based Exploration of the Design Space for Heterogeneous Systems on ChipabstractThe exploration of the design space for heterogeneous reconfigurable systems on chip (SoC) becomes more and more important. As modern SoCs include a variety of different architecture blocks, ensuring flexibility as well as highest performance, it is mandatory to prune the design space in an early stage of the design process in order to achieve short innovation cycles for new products. Therefore, the goal of this work is to provide estimations of implementation specific parameters like throughput rate, power dissipation and silicon area by means of cost functions. A concept for a model based exploration strategy supporting the design flow for heterogeneous SoCs is presented. In order to prove the feasibility of this exploration strategy, first of all operations were implemented on discrete components like DSPs, FPGAs or dedicated ASICs. Implementation parameters are provided for a variety of basic operations frequently required in digital signal processing. These implementation parameters serve as a basis for deriving models for the design space exploration concept. Holger Blume, H. Hübert, H. T. Feldkämper, Tobias G. Noll |
ASAP | 1 |
| 2000 | Integration of High-Performance ASICs into Reconfigurable Systems Providing Additional Multimedia FunctionalityabstractThe computational power of many future multimedia applications is beyond the capabilities of today's multimedia systems. Therefore, the integration of additional high-performance multimedia components is most decisive. This paper presents the integration of multimedia components into computer systems using reconfigurable coprocessor boards. The goal of these reconfigurable platforms which can be adapted to several applications and which include digital signal processors, controlling and memory devices as well as dedicated multimedia ASICs is worked out. On the way to such a platform four ASICs for image and text processing are presented. The integration of these components into a computing system using a CardBus-based coprocessor board is shown. Holger Blume, Hans-Martin Blüthgen, Christiane Henning, Patrick Osterloh |
ASAP | 1 |
| 1999 | Nonlinear vector error tolerant interpolation of intermediate video images by weighted medians
Holger Blume |
Signal Process. Image Commun. | 1 |
| 1998 | FIR-filter design with spatial and frequency design constraints using evolution strategies
Ortwin Franzen, Holger Blume, Hartmut Schröder |
Signal Process. | 2 |