Ulrich Rückert 0001

dblp:36/3893-1 · DBLP profile ↗
← Back
85ranked-venue papers
3as first author
14since 2021 · last 2025
0009-0000-0465-9441ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 3 first-author · 7 since 2021Systems, architecture and hardware · 37 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2Computer networks · 1
YearPublicationVenuePosition
2025 Energy-Based Optimization of Wire Paths in Free Space Using Discrete Elastic Rod Models
abstract
This paper explores a method for generating plausible cable routings in free space by combining curve-energy formulations from structural dynamics, knot theory, and robot path planning. The approach minimizes a composite energy functional—accounting for smoothness and collision avoidance—while encouraging a predefined cable length through an arc-length energy term. Preliminary results demonstrate the method’s ability to produce smooth, collision-free cable curves in simple example tasks, offering a physics-inspired foundation for early-stage design of electrical wiring in free space.
Ruben Lipperts, Christian Klarhorst, Marc Hesse, Ulrich Rückert 0001
ETFA4
2025 Energy Efficient Online Stream Classification under Concept Drift on FPGAs for Edge Computing
abstract
With the increasing availability of data collected by edge devices over time, efficient algorithms running remotely on low-energy devices such as FPGAs are required. This includes Machine Learning algorithms, which constitute a valuable tool when analyzing and processing vast amounts of data. To keep accurate models under distributional changes, commonly referred to as concept drift, adaptive online learning models are required. While first works proposed FPGA implementations of several machine learning algorithms, in this work, we will focus on online learning using the neighbor-based SAM-kNN model, which showed good performance under heterogenous drifts. We propose an efficient FPGA implementation that yields considerable speed and energy efficiency advantages while keeping a competitive accuracy over a range of artificial and real-world benchmarks.The implementation code is available on GitHub at https://github.com/jvaquet/SAMkNN-on-FPGA.
Jonas Vaquet, Florian Porrmann, Sarah Pilz, Valerie Vaquet, Jens Hagemeyer, Ulrich Rückert 0001, Barbara Hammer
IJCNN6
2024 A Graph Neural Network Assisted Evolutionary Algorithm for Expensive Multi-Objective Optimization
abstract
Surrogate-assisted evolutionary algorithms (SAEAs) have emerged as a promising approach to addressing expensive and black-box problems. Most existing SAEAs leverage regression models to predict the objective values, reducing the use of true objective functions. However, these methods focus on learning the mapping from the decision space to the objective space and may fail to reveal the relationship between solutions in the decision space. Recently, graph neural networks (GNNs) have attracted increased attention due to their powerful ability to expose sample interaction. In this paper, we propose employing a graph neural network for learning embeddings of solutions in the decision space, followed by a classification task aimed at predicting dominance relationships between solutions in the objective space and a regression task for obtaining the estimated fitness values. To this end, we generate a graph at each generation to represent the topology relationship between solutions in the decision space, where nodes represent solutions, and edges are added depending on the Euclidean distances between nodes. In addition, a new acquisition function that adaptively weights the predictions on objective values and dominance relationships is proposed to effectively identify new samples. The performance of the proposed method is examined by extensive empirical studies on a widely used test suite in comparison to its peer algorithms, and the results confirm the effectiveness of the proposed method.
Xiangyu Wang 0013, Xilu Wang 0001, Yaochu Jin, Ulrich Rückert 0001
CEC4
2024 FOG: A Unified Framework for Federated Combinatorial Optimization on Graphs
abstract
With the revolutionary advancements in deep learning technologies, neural combinatorial optimization (NCO) emerges as a promising field for solving complex combinatorial optimization problems (COPs) in real-world scenarios. By lever-aging the learning ability of deep neural networks, NCO trains network models to learn the mapping from problem instances to optimal solutions, offering advantages in terms of scalability and generalization to unseen instances. However, in spite of the effectiveness, existing NCO approaches have been developed under the assumption of a centralized setting, which requires all historical data to be stored and accessed centrally on a single device. In practical applications, a more common scenario is the distributed storage of data across multiple devices. Therefore, the centralized assumption in NCO is difficult to meet due to the necessity of privacy protection and data security. To tackle this problem, in this paper, we propose a privacy-preserving framework of federated combinatorial optimization on graphs, named FOG. We formulate COPs as graph optimization tasks and propose to train a graph neural network (GNN) model via data-driven federated learning approaches. We further introduce the edge-aware message passing mechanism in the GNN model, which enhances the model capability by integrating both the edge and node information during the inference process. The proposed framework is evaluated on the well-known TSP benchmarks with various scales, and the experimental results show that FOG achieves highly competitive performance while being able to protect data privacy.
Shiqing Liu, Ulrich Rückert 0001, Yaochu Jin
CEC2
2024 A Spike Vision Approach for Multi-object Detection and Generating Dataset Using Multi-core Architecture on Edge Device
Sanaullah, Shamini Koravuna, Ulrich Rückert 0001, Thorsten Jungeblut
EANN3
2024 A Digital Twin Implementation for the AMiRo
abstract
Klarhorst C, Quirin D, Hesse M, Rückert U. A Digital Twin Implementation for the AMiRo. In: 2024 IEEE 29th International Conference on Emerging Technologies and Factory Automation (ETFA). IEEE; 2024: 1-4.
Christian Klarhorst, Dennis Quirin, Marc Hesse, Ulrich Rückert 0001
ETFA4
2024 Poster: Selection of Optimal Neural Model using Spiking Neural Network for Edge Computing*
abstract
Spiking Neural Networks (SNNs), are inspired by the biological brain's complicated signaling mechanisms and possess unique characteristics that set them apart from traditional artificial neural networks. This research study explores the challenging domain of image classification, specifically utilizing the well-known MNIST dataset through the development and thorough evaluation of different neural models for edge computing. However, the primary contribution is the autonomous selection of the best-performing SNN model through various early stopping approaches and validation functions, allowing the models to autonomously adapt during training. In addition, this article presents the standalone AutoML-SNN model, which is the introduction of dynamic elements into selected SNN domains, enhancing their adaptability to complex patterns within the dataset. Furthermore, the early stopping methodologies are used to reduce overfitting hazards, and using the 3000-neuron set, the LIF appeared as the most proficient neural model.
Sanaullah, Ulrich Rückert 0001, Thorsten Jungeblut
ICDCS3
2023 TinyML optimization for activity classification on the resource-constrained body sensor BI-Vital
abstract
Addressing the demand for scalable and efficient deployment of machine learning (ML) models on mobile devices and especially microcontrollers, this research introduces a hardware-in-the-loop (HIL) deployment setup within the producer-consumer software architecture of the BI-Vital, a chest-mounted body sensor designed by our research group. The sensor provides real-time monitoring of various physiological and environmental parameters, enhanced by the integration of tiny machine learning (TinyML). Leveraging the UCI-HAR dataset for Human Activity Recognition (HAR), this study focuses on optimizing ML models’ hyperparameters to balance inference time, memory, accuracy, and power consumption. An efficiency score E is proposed to assess models in this unique context. The results illustrate the trade-offs between the models: Decision Trees offer reduced power usage, Multilayer Perceptrons ensure high accuracy with minimal memory requirements, while Convolutional Neural Networks present limitations due to extended inference times. The results emphasize the potential of TinyML in wearable physiological monitoring, especially for optimizing models for resource-limited devices in real-world applications.
Kevin Penner, Felix Wittenfeld, Bastian Steinhagen, Marc Hesse, Ulrich Rückert 0001
BSN5
2023 Streamlined Training of GCN for Node Classification with Automatic Loss Function and Optimizer Selection
Sanaullah, Shamini Koravuna, Ulrich Rückert 0001, Thorsten Jungeblut
EANN3
2023 Evaluation of Spiking Neural Nets-Based Image Classification Using the Runtime Simulator RAVSim
abstract
Spiking Neural Networks (SNNs) help achieve brain-like efficiency and functionality by building neurons and synapses that mimic the human brain’s transmission of electrical signals. However, optimal SNN implementation requires a precise balance of parametric values. To design such ubiquitous neural networks, a graphical tool for visualizing, analyzing, and explaining the internal behavior of spikes is crucial. Although some popular SNN simulators are available, these tools do not allow users to interact with the neural network during simulation. To this end, we have introduced the first runtime interactive simulator, called Runtime Analyzing and Visualization Simulator (RAVSim),adeveloped to analyze and dynamically visualize the behavior of SNNs, allowing end-users to interact, observe output concentration reactions, and make changes directly during the simulation. In this paper, we present RAVSim with the current implementation of runtime interaction using the LIF neural model with different connectivity schemes, an image classification model using SNNs, and a dataset creation feature. Our main objective is to primarily investigate binary classification using SNNs with RGB images. We created a feed-forward network using the LIF neural model for an image classification algorithm and evaluated it by using RAVSim. The algorithm classifies faces with and without masks, achieving an accuracy of 91.8% using 1000 neurons in a hidden layer, 0.0758 MSE, and an execution time of ∼10[Formula: see text]min on the CPU. The experimental results show that using RAVSim not only increases network design speed but also accelerates user learning capability.
Shamini Koravuna, Ulrich Rückert 0001, Thorsten Jungeblut
Int. J. Neural Syst.2
2022 VEDLIoT: Very Efficient Deep Learning in IoT
abstract
The VEDLIoT project targets the development of energy-efficient Deep Learning for distributed AIoT applications. A holistic approach is used to optimize algorithms while also dealing with safety and security challenges. The approach is based on a modular and scalable cognitive IoT hardware platform. Using modular microserver technology enables the user to configure the hardware to satisfy a wide range of applications. VEDLIoT offers a complete design flow for Next-Generation IoT devices required for collaboratively solving complex Deep Learning applications across distributed systems. The methods are tested on various use-cases ranging from Smart Home to Automotive and Industrial IoT appliances. VEDLIoT is an H2020 EU project which started in November 2020. It is currently in an intermediate stage with the first results available.
Martin Kaiser, René Griessl, Nils Kucza, Carola Haumann, Lennart Tigges, Kevin Mika, Jens Hagemeyer, Florian Porrmann, Ulrich Rückert 0001, Micha vor dem Berge, Stefan Krupop, Mario Porrmann, Marco Tassemeier, Pedro Trancoso, Fareed Qararyah, Stavroula Zouzoula, António Casimiro, Alysson Neves Bessani, José Cecílio, Stefan Andersson, Oliver Brunnegård, Olof Eriksson, Roland Weiss 0001, Franz Meierhöfer, Hans Salomonsson, Elaheh Malekzadeh, Daniel Ödman, Anum Khurshid, Pascal Felber, Marcelo Pasin, Valerio Schiavoni, Jämes Ménétrey, Karol Gugala, Piotr Zierhoffer, Eric Knauss, Hans-Martin Heyn
DATE9
2022 SNNs Model Analyzing and Visualizing Experimentation Using RAVSim
Sanaullah, Shamini Koravuna, Ulrich Rückert 0001, Thorsten Jungeblut
EANN3
2022 ML4ProFlow: A Framework for Low-Code Data Processing from Edge to Cloud in Industrial Production
abstract
One necessary part of Industry 4.0 is the availability and accessibility of data processing pipelines. This paper shows the ongoing development of ML4ProFlow, a framework that brings together the following parts: First, it provides the management of execution environments. Second, it specifies processing modules that focus on reusability and cross-platform usage. Third, it comes with a benchmarking automation to help developers implementing and analyzing modules and their combination. Those three integral parts of the framework are presented and the usability is shown.
Christian Klarhorst, Dennis Quirin, Marc Hesse, Ulrich Rückert 0001
ETFA4
2021 Simurgh: a fully decentralized and secure NVMM user space file system
abstract
The availability of non-volatile main memory (NVMM) has started a new era for storage systems and NVMM specific file systems can support extremely high data and metadata rates, which are required by many HPC and data-intensive applications. Scaling metadata performance within NVMM file systems is nevertheless often restricted by the Linux kernel storage stack, while simply moving metadata management to the user space can compromise security or flexibility.
Nafiseh Moti, Frederic Schimmelpfennig, Reza Salkhordeh, David Klopp, Toni Cortes, Ulrich Rückert 0001, André Brinkmann
SC6
2020 Benchmarking Deep Spiking Neural Networks on Neuromorphic Hardware
Christoph Ostrau, Jonas Dominik Homburg, Christian Klarhorst, Michael Thies, Ulrich Rückert 0001
ICANN (2)5
2019 Jointly Trained Variational Autoencoder for Multi-Modal Sensor Fusion
Timo Korthals, Marc Hesse, Jürgen Leitner, Andrew Melnik, Ulrich Rückert 0001
FUSION5
2019 Multi-Modal Generative Models for Learning Epistemic Active Sensing
abstract
We present a novel approach of multi-modal deep generative models and apply this to coordinated heterogeneous multi-agent active sensing. A major approach to achieve this objective is to train a multi-modal variational Auto Encoder (M2VAE) that integrates the information of different sensor modalities into a joint latent representation. Furthermore, we derive an objective from the M2VAE that enables the maximization of the evidence lower bound via selection of sensor modalities. Using this approach as a direct reward signal to a multi-modal and multi-agent deep reinforcement learning setup leads intuitively to an epistemic active sensing behavior that coordinately resolves the ambiguity of observations.
Timo Korthals, Daniel Rudolph, Jürgen Leitner, Marc Hesse, Ulrich Rückert 0001
ICRA5
2019 A Bidirectional Object Tracking and Navigation System using a True-Range Multilateration Method
abstract
In the past, several contributions and proposals for the implementation of Ultra-wideband (UWB)-based localization and positioning solutions on the system level were made. However, most of them are limited to a unidirectional approach, i.e. data communication is in one direction (from a transmitter to a receiver). This restricts the systems' use-case to either navigation or tracking. In this paper, we demonstrate an UWB-based bidirectional localization system which is capable of acting as both a navigation and tracking system in a single wireless platform. Regarding this, we proposed a complete set of such a system and outline the implementation process in the paper. A true-range multilateration method is used as a positioning algorithm in the implemented system, for which we proposed a novel Non-Line-of-Sight (NLOS) mitigation technique. The experimental evaluation of the proposed system was done in comparison with a commercially available UWB system. In the experiments, we used a Vicon camera system as a reference.
Cung Lian Sang, Michael Adams 0002, Timo Korthals, Timm Hörmann, Marc Hesse, Ulrich Rückert 0001
IPIN6
2019 Towards an SSVEP-BCI Controlled Smart Home
abstract
Brain-Computer Interfaces (BCIs) based on Steady-State Visually Evoked Potentials (SSVEPs) can be used as hand-free control device. To utilize this control method in a real life scenario, we created a system in which a smart home is controlled by BCI. Six devices in the smart home environment could be controlled with the BCI system: The entrance door, the wardrobe, the kitchens' worktop and drawers, the light system of all the rooms and a guide light. In the presented paper, the visual stimuli for the BCI were placed at multiple screens in the smart home (placed at different locations such as the kitchen and the living room). The processing was done on one computer, located in the living room. The placement of the visual stimuli corresponded to the actuators that were controlled, e.g. the kitchen drawers were linked to the stimuli displayed in the kitchen. An online experiment was conducted where participants went through a scenario consisting of thirteen SSVEP-BCI selections in total. Eight healthy participants took part in the experiments. For BCI signal acquisition, a mobile EEG amplifier was used. Participants walked freely around the rooms during the experiment. An average accuracy of 81 % was achieved, which suggests that the SSVEP-system is suitable to control the external devices in the smart home, and that the system can be expanded to involve more actuators.
Michael Adams 0002, Sadok Ben-Salem, Arne Vogelsang, Thorsten Jungeblut, Ulrich Rückert 0001, Ivan Volosyak, Mihaly Benda, Abdul Saboor, André Frank Krause, Aya Rezeika, Felix Gembler, Piotr Stawicki, Marc Hesse, Kai Essig
SMC6
2018 Resource-efficient Reconfigurable Computer-on-Module for Embedded Vision Applications
abstract
The paper proposes a novel architecture for a highly customisable FPGA-SoC-based Computer-on-Module (CoM) targeting embedded vision applications. Apart from a Xilinx Zynq SoC, the module integrates an Adapteva Epiphany floating point accelerator in a Toradex Apalis compliant form factor. The CoM has been successfully integrated into two robot platforms to enhance their vision processing capabilities. For evaluation, visually-guided collision avoidance and navigation has been implemented, mimicking the behaviour of insects. The hardware/software partitioning is presented together with a comparison to an HLS-based solution for the given application. The proposed stream-based FPGA implementation achieves a speedup of 721 and an increase in energy efficiency by a factor of 800 compared to an OpenCV-based implementation on one of the embedded ARM processors of the Zynq SoC.
Daniel Klimeck, Hanno Gerd Meyer, Jens Hagemeyer, Mario Porrmann, Ulrich Rückert 0001
ASAP5
2018 Generic Architecture for Modular Real-time Systems in Robotics
Thomas Schöpping, Timo Korthals, Marc Hesse, Ulrich Rückert 0001
ICINCO (2)4
2018 An Analytical Study of Time of Flight Error Estimation in Two-Way Ranging Methods
abstract
In absence of clock synchronization, Two-Way Ranging (TWR) is the most commonly used technique for measuring the distance between two wireless transceivers. The existing time-of-flight (TOF) error estimation model, the IEEE 802.15.4-2011 standard, is specifically based on clock drift error. However, it is insufficient when an in-depth comparative analysis of different TWR methods is required. In this paper, we propose an extended TOF error estimation model for TWR methods, based on the IEEE 802.15.4 standard. Using the proposed model, we perform an analytical study of TOF error estimation among different TWR methods. The model is validated with numerical simulation results. Moreover, we demonstrate the pitfalls of the symmetric double-sided TWR (SDS-TWR) method, which is commonly used to reduce the TOF error due to clock drifts.
Cung Lian Sang, Michael Adams 0002, Timm Hörmann, Marc Hesse, Mario Porrmann, Ulrich Rückert 0001
IPIN6
2018 CoreVA-MPSoC: A Many-Core Architecture with Tightly Coupled Shared and Local Data Memories
abstract
MPSoCs with hierarchical communication infrastructures are promising architectures for low power embedded systems. Multiple CPU clusters are coupled using an Network-on-Chip (NoC). Our CoreVA-MPSoC targets streaming applications in embedded systems, like signal and video processing. In this work we introduce a tightly coupled shared data memory to each CPU cluster, which can be accessed by all CPUs of a cluster and the NoC with low latency. The main focus is the comparison of different memory architectures and their connection to the NoC. We analyze memory architectures with local data memory only, shared data memory only, and a hybrid architecture integrating both. Implementation results are presented for a 28 nm FD-SOI standard cell technology. A CPU cluster with shared memory shows similar area requirements compared to the local memory architecture. We use post place and route simulations for precise analysis of energy consumption on both cluster and NoC level using the different memory architectures. An architecture with shared data memory shows best performance results in combination with a high resource efficiency. On average, the use of shared memory shows a 17.2 percent higher throughput for a benchmark suite of 10 applications compared to the use of local memory only.
Johannes Ax, Gregor Sievers, Julian Daberkow, Martin Flasskamp, Marten Vohrmann, Thorsten Jungeblut, Wayne Kelly, Mario Porrmann, Ulrich Rückert 0001
IEEE Trans. Parallel Distributed Syst.9
2017 FPGA-based multi-robot tracking
Arif Irwansyah, Omar W. Ibraheem, Jens Hagemeyer, Mario Porrmann, Ulrich Rückert 0001
J. Parallel Distributed Comput.5
2016 Towards a comprehensive power consumption model for wireless sensor nodes
abstract
Energy efficiency is the most outstanding design criterion for wireless sensor nodes and especially wireless body sensors. Because a detailed measurement of the system's power consumption is not possible during the design process and often too complex for already manufactured devices, the power consumption has to be estimated. This leads to the need for a comprehensive and modular model for the power consumption of WSNs, which is proposed in this work. Due to the modular structure of the model the user is able to get a first estimate in an early stage of the design process (e.g. choose components) and to get a more accurate estimation later in the design process by lowering the abstraction level. This tackles the demanding trade-off between accuracy and usability in modeling.
Marc Hesse, Michael Adams 0002, Timm Hörmann, Ulrich Rückert 0001
BSN4
2016 A software assistant for user-centric calibration of a wireless body sensor
abstract
Body sensors have a promising contribution to health promotion in many areas of daily life (telemedicine, corporate health care or recreational sports). However, the valid measurement of vital signs and kinematic data strongly depends on the signals' quality and the users' compliance (proper usage). Although, there is a lot of research work concerning accuracy and calibration of wireless body sensors the human user is typically not involved. Thus, in this work, we present a software assistant (wizard) that guides users during the process of attaching and setting up a wireless body sensor. Furthermore, insights of the implemented software as well as the utilized quality measures and calibration steps are given (ECG, respiration sensor and accelerometer). With the proposed software assistant, the users are instructed to correctly attach the body sensor and calibrate or verify the operability of the various sensor elements. The primary goal is to encourage compliance and the users' sense of control. In this way, we want to reduce faulty operation and ensure optimal signal quality.
Timm Hörmann, Marc Hesse, Michael Adams 0002, Ulrich Rückert 0001
BSN4
2016 Occupancy Grid Mapping with Highly Uncertain Range Sensors based on Inverse Particle Filters
abstract
A huge number of techniques for detecting and mapping obstacles based on LIDAR and SONAR exist, though not taking approximative sensors with high levels of uncertainty into consideration. The proposed mapping method in this article is undertaken by detecting surfaces and approximating objects by distance using sensors with high localization ambiguity. Detection is based on an Inverse Particle Filter, which uses readings from single or multiple sensors as well as a robot’s motion. This contribution describes the extension of the Sequential Importance Resampling filter to detect objects based on an analytical sensor model and embedding into Occupancy Grid Maps. The approach has been applied to the autonomous mini robot AMiRo in a distributed way. There were promising results for its low-power, low-cost proximity sensors in various real life mapping scenarios, which outperform the standard Inverse Sensor Model approach.
Timo Korthals, Marvin Barther, Thomas Schöpping, Stefan Herbrechtsmeier, Ulrich Rückert 0001
ICINCO (2)5
2015 Robust estimation of physical activity by adaptively fusing multiple parameters
abstract
Raising the awareness of being physically active by utilizing wearable body sensors has become a popular research topic. Recent approaches combine physical and physiological information to obtain a precise prediction of a person;s physical activity ratio. However, the error in the determination of physical activity due to invalid physiological values that are resulting from underlying signal disturbances, has so far not been considered. We therefore present a robust measure of activity that fuses accelerometer data, heart rate and other personalized features, and is adaptively responding to missing physiological sensor data. To set up the model, we make use of regression analysis (MARS). Our findings indicate the need for considering signal quality when estimating physical activity. The predictive model shows close agreement (R2= 0.97) to the reference from indirect calorimetry, even if the physiological information is partly corrupted.
Timm Hörmann, Peter Christ, Marc Hesse, Ulrich Rückert 0001
BSN4
2015 Evaluation of interconnect fabrics for an embedded MPSoC in 28 nm FD-SOI
abstract
Embedded many-core architectures contain dozens to hundreds of CPU cores that are connected via a highly scalable NoC interconnect. Our Multiprocessor-System-on-Chip CoreVA-MPSoC combines the advantages of tightly coupled bus-based communication with the scalability of NoC approaches by adding a CPU cluster as an additional level of hierarchy. In this work, we analyze different cluster interconnect implementations with 8 to 32 CPUs and compare them in terms of resource requirements and performance to hierarchical NoCs approaches. Using 28 nm FD-SOI technology the area requirement for 32 CPUs and AXI crossbar is 5.59 mm2including 23.61% for the interconnect at a clock frequency of 830 MHz. In comparison, a hierarchical MPSoC with 4 CPU cluster and 8 CPUs in each cluster requires only 4.83 mm2including 11.61% for the interconnect. To evaluate the performance, we use a compiler for streaming applications to map programs to the different MPSoC configurations. We use this approach for a design-space exploration to find the most efficient architecture and partitioning for an application.
Gregor Sievers, Johannes Ax, Nils Kucza, Martin Flasskamp, Thorsten Jungeblut, Wayne Kelly, Mario Porrmann, Ulrich Rückert 0001
ISCAS8
2014 CoreVA: A Configurable Resource-Efficient VLIW Processor Architecture
abstract
Mobile signal processing applications have a limited energy budget and require resource-efficient processing elements. General purpose VLIW CPUs offer a high energy efficiency and allow for the execution of a wide range of applications in this domain. In this work we present the configurable 32 bit VLIW processor architecture CoreVA. Besides the number of issue slots, it allows for a fine-grained configuration of the amount and characteristics of the processor's functional units (e.g., ALUs, MACs, or LD/ST units). A design-space exploration is performed to evaluate how these functional units impact area and power consumption. The basic configuration with one ALU, MAC, DIV, and LD/ST unit has a power consumption of 11.796 mW and an area of 0.142 mm2 at a clock frequency of 750 MHz in a 28 nm FD-SOI process. The maximum clock frequency in this process node is 833 MHz. To bear a relation of the hardware requirements to possible performance gains of the application, a signal processing algorithm is used as a benchmark to evaluate the energy consumption of different hardware configurations. The lowest energy consumption is observed with a configuration of 4 issue slots using 4 ALUs, 4 MACs, and 2 LD/ST units. This is an improvement by a factor of 1.68 compared to the single issue slot configuration.
Boris Hübener, Gregor Sievers, Thorsten Jungeblut, Mario Porrmann, Ulrich Rückert 0001
EUC5
2013 A reconfigurable neuroprocessor for self-organizing feature maps
Jan Lachmair, Erzsébet Merényi, Mario Porrmann, Ulrich Rückert 0001
Neurocomputing4
2013 A systematic approach for optimized bypass configurations for application-specific embedded processors
abstract
The diversity of today's mobile applications requires embedded processor cores with a high resource efficiency, that means, the devices should provide a high performance at low area requirements and power consumption. The fine-grained parallelism supported by multiple functional units of VLIW architectures offers a high throughput at reasonable low clock frequencies compared to single-core RISC processors. To efficiently utilize the processor pipeline, common system architectures have to cope with data hazards due to data dependencies between consecutive operations. On the one hand, such hazards can be resolved by complex forwarding circuits (i.e., a pipeline bypass) which forward intermediate results to a subsequent instruction. On the other hand, the pipeline bypass can strongly affect or even dominate the total resource requirements and degrade the maximum clock frequency. In this work the CoreVA VLIW architecture is used for the development and the analysis of application-specific bypass configurations. It is shown that many paths of a comprehensive bypass system are rarely used and may not be required for certain applications. For this reason, several strategies have been implemented to enhance the efficiency of the total system by introducing application-specific bypass configurations. The configuration can be carried out statically by only implementing required paths or at runtime by dynamically reconfiguring the hardware. An algorithm is proposed which derives an optimized configuration by iteratively disabling single bypass paths. The adaptation of these application-specific bypass configurations allows for a reduction of the critical path by 26%. As a result, the execution time and energy requirements could be reduced by up to 21.5%. Using Dynamic Frequency Scaling (DFS) and dynamic deactivation/reactivation of bypass paths allows for a runtime reconfiguration of the bypass system. This ensures the highest efficiency while processing varying applications.
Thorsten Jungeblut, Boris Hübener, Mario Porrmann, Ulrich Rückert 0001
ACM Trans. Embed. Comput. Syst.4
2012 Hardware accelerated real time classification of hyperspectral imaging data for coffee sorting
Andreas Backhaus, Jan Lachmair, Ulrich Rückert 0001, Udo Seiffert
ESANN3
2012 gNBXe - a Reconfigurable Neuroprocessor for Various Types of Self-Organizing Maps
Jan Lachmair, Erzsébet Merényi, Mario Porrmann, Ulrich Rückert 0001
ESANN4
2012 Parallel neural hardware: the time is right
Ulrich Rückert 0001, Erzsébet Merényi
ESANN1
2012 A TCMS-based architecture for GALS NoCs
abstract
In this work we propose a TCMS (Tightly Coupled Mesochronous Synchronizer)-based architecture of Globally-Asynchronous Locally-Synchronous (GALS) Network-on-Chips (NoC). The NoC is based on the GigaNoC approach, a scalable NoC featuring packet-switched wormhole routing. At a clock frequency of 750MHz a link bandwidth of up to 6 GByte/s is achieved. To provide a high computational performance, the processing engines (PEs) are based on the CoreVA VLIW architecture. The resource efficiency of mesochronous (TMCS-based) and asynchronous (FIFO-based) communication links is analyzed. In addition an asynchronous coupling of the PE to the switch boxes is evaluated. This allows for multi-voltage/multi-frequency scenarios, where the performance of each PE is adapted to the current performance requirements. Analyses have shown, that TCMS-based communication links and asynchronously coupled PEs allow for the high efficiency of GALS-based NoCs with moderate additional resource requirements.
Thorsten Jungeblut, Johannes Ax, Mario Porrmann, Ulrich Rückert 0001
ISCAS4
2011 Automatic HDL-Based Generation of Homogeneous Hard Macros for FPGAs
abstract
The regularity of resources found in FPGAs is a unique feature, which can be utilized in a number of applications, e.g., in timing critical applications or applications with a demand for homogeneous routing. Current synthesis tools do not support an automatic generation of homogeneous FPGA designs, such that a time-consuming hand-crafted design is required. We present a tool flow, which automatically generates homogeneous hard macros for Xilinx FPGAs starting from a high-level description, such as VHDL. Key functionalities of the tool flow are a homogeneous placer and a suitable routing algorithm, which aim at maintaining the homogeneity of the resulting hard macro. The place and route tools use a resource library that is automatically generated for the target FPGA family by extracting relevant information from the vendor tools. The tool chain is demonstrated for the design of hard macros for a time-to-digital converter and a tiled partially reconfigurable region. The resulting designs are evaluated with respect to resource requirements and timing constraints.
Sebastian Korf, Dario Cozzi, Markus Köster, Jens Hagemeyer, Mario Porrmann, Ulrich Rückert 0001, Marco D. Santambrogio
FCCM6
2011 Integrated circuit optimization by means of evolutionary multi-objective optimization
abstract
The design of resource efficient integrated circuits (ICs) requires solving a minimization problem which consists of more than one objective given as measures of the available resources. This multi-objective optimization problem (MOP) can be solved on the smallest unit of the IC, the standard cells, to improve the performance of the entire circuit. In this work, transistor sizing of an IC is approached via a multi-objective approach which includes the use of multi-objective evolutionary algorithms (MOEAs). We compare the performance of two MOEAs on a four-dimensional MOP of a particular standard cell. The results indicate that evolutionary strategies are suitable for the treatment of such problems and advantageous against other rather classical methods.
Matthias W. Blesken, Anouar Chebil, Ulrich Rückert 0001, Xavier Esquivel, Oliver Schütze 0001
GECCO3
2011 Applying dynamic reconfiguration in the mobile robotics domain: A case study on computer vision algorithms
abstract
Mobile robots are widely used in industrial environments and are expected to be widely available in human environments in the near future, for example, in the area of care and service robots. This article proposes an implementation for a highly customizable color recognition module based on Field Programmable Gate Array (FPGA) hardware to accomplish tasks like real-time frame processing for image streams. In comparison to a pure software solution on a CPU, an attached FPGA-based hardware accelerator enables real-time image processing and significantly reduces the required computing power of the CPU. Instead, the CPU can be used for tasks that cannot be efficiently implemented on FPGAs, for example, because of a large control overhead. We concentrate on a multirobot scenario where a group of robots follows a human team member by keeping a specific formation in order to support the human in exploration and object detection. Additionally, the robots provide a communication infrastructure to maintain a stable multihop communication network between the human and a base station recording all actions and evaluating the captured images and transmitted data. Depending on the current operating conditions, the robot system has to be able to execute a wide variety of different tasks. Since only a small number of tasks have to be executed concurrently, dynamic reconfiguration of the FPGA can be used to avoid the parallel implementation of all tasks on the FPGA. Within this context, this article discusses application fields where dynamic reconfiguration of FPGA-based coprocessors significantly reduces the CPU load and presents examples of how dynamic reconfiguration can be used in exploration.
Federico Nava, Donatella Sciuto, Marco D. Santambrogio, Stefan Herbrechtsmeier, Mario Porrmann, Ulf Witkowski, Ulrich Rückert 0001
ACM Trans. Reconfigurable Technol. Syst.7
2011 Design Optimizations for Tiled Partially Reconfigurable Systems
abstract
In partially reconfigurable architectures, system components can be dynamically loaded and unloaded allowing resources to be shared over time. Dynamic system components are represented by partial reconfiguration (PR) modules. In comparison to a static system, the design of a partially reconfigurable system requires additional design steps, such as partitioning the device resources into static and dynamic regions. We present the concept of tiled PR regions, which enables a flexible online-placement of PR modules. Dynamic reconfiguration requires a suitable communication infrastructure to interconnect the static and dynamic system components. We present an embedded communication macro, a communication infrastructure that interconnects PR modules in a tiled PR region. Efficient online-placement of PR modules depends not only on the placement algorithm, but also on design-time aspects such as the chosen synthesis regions of the PR modules. We propose a design method for selecting suitable synthesis regions for the PR modules aiming to optimize their placement at run-time.
Markus Köster, Wayne Luk, Jens Hagemeyer, Mario Porrmann, Ulrich Rückert 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2010 Multiobjective optimization for transistor sizing sub-threshold CMOS logic standard cells
abstract
Transistor sizing of sub-threshold standard cells for digital ultra-low power systems is a very challenging task because robustness has to be considered as an important design objective in addition to the competing resources power consumption and propagation delay. In this paper we regard this task as a multiobjective optimization problem (MOP) and show that the support of MOP algorithms is necessary and beneficial in the design process of sub-threshold CMOS logic standard cells. Optimization results are presented for an inverter, NAND gate, and NOR gate in a 65 nm process technology.
Matthias W. Blesken, Sven Lütkemeier, Ulrich Rückert 0001
ISCAS3
2010 High level specification of embedded listeners for monitoring of Network-on-Chips
abstract
Nowadays, the Network-on-Chip (NoC) paradigm has become more and more popular for building an on-chip communication infrastructure. Like in every traditional network, debugging and performance monitoring are also very important issues in NoC-based systems. Unfortunately, the design process of monitoring hardware is a time consuming activity. The work presented in this paper is based on a high level specification language, called SiLLis (Simplified Language for Listeners), for the convenient development of generic monitoring hardware. SiLLis allows the designer to define complex filter rules on a high abstraction level. In this way, the design time as well as the bandwidth requirements for monitoring data can be drastically reduced. To present the benefits of SiLLis, we define a performance monitor that is integrated into a NoC-based multiprocessor System-on-Chip and can be used both to analyze the performance of the system and to optimize the routing strategy at run-time. By using SiLLis, the performance monitor can be realized with a area overhead of only 0.58 % per NoC node.
Christoph Puttmann, Mario Porrmann, Paolo Roberto Grassi, Marco D. Santambrogio, Ulrich Rückert 0001
ISCAS5
2010 Design Space Exploration for Memory Subsystems of VLIW Architectures
abstract
In this work we present a design space exploration of the memory subsystem of our configurable CoreVA VLIW architecture. The development of resource efficient processor architectures is based on a two-stage tool flow using a high-level processor specification as a reference. We evaluate several memory configurations like one memory port or two memory ports, as well as different write-miss-allocation modes. Applications ranging from LTE protocol stack over baseband processing up to cryptography and multimedia are evaluated in terms of execution time and energy efficiency. Analyses have shown that the application specific configuration of the memory subsystem can improve energy by up to 25%. Our environment allows the rapid profiling and evaluation of algorithms to choose the most efficient configuration.
Thorsten Jungeblut, Gregor Sievers, Mario Porrmann, Ulrich Rückert 0001
NAS4
2010 Runtime Reconfiguration of Multiprocessors Based on Compile-Time Analysis
abstract
In multiprocessors, performance improvement is typically achieved by exploring parallelism with fixed granularities, such as instruction-level, task-level, or data-level parallelism. We introduce a new reconfiguration mechanism that facilitates variations in these granularities in order to optimize resource utilization in addition to performance improvements. Our reconfigurable multiprocessor QuadroCore combines the advantages of reconfigurability and parallel processing. In this article, a unified hardware-software approach for the design of our QuadroCore is presented. This design flow is enabled via compiler-driven reconfiguration which matches application-specific characteristics to a fixed set of architectural variations. A special reconfiguration mechanism has been developed that alters the architecture within a single clock cycle. The QuadroCore has been implemented on Xilinx XC2V6000 for functional validation and on UMC’s 90nm standard cell technology for performance estimation. A diverse set of applications have been mapped onto the reconfigurable multiprocessor to meet orthogonal performance characteristics in terms of time and power. Speedup measurements show a 2--11 times performance increase in comparison to a single processor. Additionally, the reconfiguration scheme has been applied to save power in data-parallel applications. Gate-level simulations have been performed to measure the power-performance trade-offs for two computationally complex applications. The power reports confirm that introducing this scheme of reconfiguration results in power savings in the range of 15--24%.
Madhura Purnaprajna, Mario Porrmann, Ulrich Rückert 0001, Michael Hussmann, Michael Thies, Uwe Kastens
ACM Trans. Reconfigurable Technol. Syst.3
2008 An automated platform for minirobots experiments
abstract
In this paper, a platform for managing and providing remote access to robots was developed and constructed. The system helps to schedule, perform and analyze experiments using minirobots. A solution for recharging the robots automatically has been included in the system in order to save the time needed for manual recharging. The system can automatically interrupt the experiments, charge robots and then resume experiments. To reach this level of autonomy, a positioning system, path planning technique along with video streaming have been developed and implemented.
Ulf Witkowski, Emad Monier, Ulrich Rückert 0001, Sally El Ghoul, M. S. El-Ghoniemy, Mohamed Saied Abdel-Wahab, A. Fouad, Ashraf S. Hussein, A. Kamal, M. Abdel-Meniem, W. Abo El Khair
ICARCV3
2007 GigaNoC - A Hierarchical Network-on-Chip for Scalable Chip-Multiprocessors
abstract
Due to the technological progress in the semiconductor industry, more and more components can be integrated on a single die forming a complex System-on-Chip. For enabling an efficient interaction between the various building blocks of today's SoCs, efficient communication structures become more and more essential. In this paper, we present the GigaNoC, a hierarchical Network-on-Chip that is especially suitable for scalable Chip-Multiprocessor architectures. The GigaNoC approach features a packet-switched wormhole routing on-chip network that provides the backbone of our multiprocessor architecture. In order to meet bandwidth requirements of different application domains, our Network-on-Chip is easily scalable and parameterizable in various aspects. This work highlights the communication protocol and shows a performance evaluation for different congestion scenarios. Furthermore, we present an FPGA-based prototypical realization and introduce a debugging and verification environment. Finally, implementation results for a standard cell technology are discussed.
Christoph Puttmann, Jörg-Christian Niemann, Mario Porrmann, Ulrich Rückert 0001
DSD4
2007 Controlling complexity of RBF networks by similarity
Ulrich Rückert 0001, Ralf Eickhoff
ESANN1
2007 A Digital Framework for Pulse Coded Neural Network Hardware with Bit-Serial Operation
abstract
This publication presents a digital framework for building up pulse coded neural networks with leaky integrate-and-fire neurons and static synapses as well as dynamic synapses. The system, including a novel communication infrastructure, is mainly focused on ASIC synthesis but also shows a small footprint on Virtex2(Pro) FPGAs. Its bit-serial operation has been verified by simulations.
Tim Kaulmann, Deniz Dikmen, Ulrich Rückert 0001
HIS3
2007 Impact of Shrinking Technologies on the Activation Function of Neurons
Ralf Eickhoff, Tim Kaulmann, Ulrich Rückert 0001
ICANN (1)3
2007 A Control Approach to a Biophysical Neuron Model
Tim Kaulmann, Axel Löffler, Ulrich Rückert 0001
ICANN (1)3
2007 Partial Dynamic Reconfiguration in a Multi-FPGA Clustered Architecture Based on Linux
abstract
Dynamically reconfigurable hardware allows for implementing systems that can be adapted at run-time according to the needs of the user. This paper presents an architecture that is composed of multiple FPGAs that are connected to an embedded processor. Thus, the architecture is referred to as a multi-FPGA clustered architecture (MFCA). All FPGAs can be partially and dynamically reconfigured to integrate user-defined IP-cores into the system at run-time. For the resource management and communication management we have implemented a Linux operating system on the embedded processor that can be used to control the reconfiguration of the FPGAs by means of simple function calls. Furthermore, the Linux OS completely hides the physical infrastructure of the MFCA from user applications, offering a consistent interface to utilize partial reconfiguration.
Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto, Boris Kettelhoit, Markus Köster, Mario Porrmann, Ulrich Rückert 0001
IPDPS7
2007 Robustness of radial basis functions
Ralf Eickhoff, Ulrich Rückert 0001
Neurocomputing2
2007 Resource efficiency of the GigaNetIC chip multiprocessor architecture
Jörg-Christian Niemann, Christoph Puttmann, Mario Porrmann, Ulrich Rückert 0001
J. Syst. Archit.4
2007 Characterization of Analog Local Cluster Neural Network Hardware for Control
abstract
The local cluster neural network (LCNN) was designed for analog realization especially suited to applications in control systems. It uses clusters of sigmoidal neurons to generate basis functions that are localized in multidimensional input space. Sigmoidal neurons are well suited to analog electronic realization. In this paper, we report the results of extensive measurements that characterize the computational capabilities of the first analog very large scale integration (VLSI) realization of the LCNN. Despite manufacturing fluctuations and the inherent low precision of analog electronics, the test results suggest that it may be suitable for use in feedback control systems.
Joaquin Sitte, Ulrich Rückert 0001
IEEE Trans. Neural Networks3
2006 Robust Local Cluster Neural Networks
Ralf Eickhoff, Joaquin Sitte, Ulrich Rückert 0001
ESANN3
2006 Pareto-optimal Noise and Approximation Properties of RBF Networks
Ralf Eickhoff, Ulrich Rückert 0001
ICANN (1)2
2006 SIRENS: A Simple Reconfigurable Neural Hardware Structure for artificial neural network implementations
abstract
Artificial neural networks are used in various applications and research areas. Mathematically inspired approaches use these types of networks to solve complex classification or function approximation tasks whereas biologically motivated models attempt to adapt desired properties from biology such as robustness or fault tolerance to technical systems and architectures. Therefore, a great variety of different models have been proposed in literature which can be separated in time-dependent and time-independent models. To verify these models and to accelerate simulations prototypes are often implemented in integrated circuits using digital or analog designs. In this work, a simple reconfigurable neural hardware structure (SIRENS) is introduced which is capable to represent several different models of neurons, time-independent and time-dependent models as well. Therefore, this system can be used for several applications (classification or simulation) and purposes (acceleration or operation). The underlying mathematical principles are presented and, furthermore, design considerations are given in this paper.
Ralf Eickhoff, Tim Kaulmann, Ulrich Rückert 0001
IJCNN3
2006 Enhancing Fault Tolerance of Radial Basis Functions
abstract
The challenge of future nanoelectronic applications, e.g. in quantum computing or in molecular computing, is to assure reliable computation facing a growing number of malfunctioning and failing computational units. Modeled on biology artificial neural networks are intended to be one preferred architecture for these applications because their architectures allow distributed information processing and, therefore, will result in tolerance to malfunctioning neurons and in robustness to noise. In this work, methods to enhance fault tolerance to permanently failing neurons of Radial Basis Function networks are investigated for function approximation applications. Therefore, a relevance measure is introduced which can be used to enhance the fault tolerance or, on the contrary, to control the network complexity if it is used for pruning.
Ralf Eickhoff, Ulrich Rückert 0001
IJCNN2
2006 Bio-inspired massively parallel architectures for nanotechnologies
abstract
Massively parallel single-chip multiprocessors (CMP) share a number of traits with biological systems such as neural networks. These biological systems have therefore inspired a number of concepts that may help to overcome some of the problems that will come up in future circuit technologies. In this work we present a first comparison of CMPs based on processor cores of different complexity and estimate the efficiency of CMPs with regards to overall performance and energy consumption. The analysis is based on an analytical model of chip multiprocessing that can help to estimate the runtime and energy consumption of different parallel algorithms. As in previous work we will use the GigaNetIC architecture as a basis for the different CMP architectures
Björn Jäger, Mario Porrmann, Ulrich Rückert 0001
ISCAS3
2005 Analytical approach to massively parallel architectures for nanotechnologies
abstract
In the emerging field of single-chip multiprocessors (CMP), analytical models of performance and power consumption are necessary for design space exploration and the analysis of existing architectures. In the light of ever decreasing structure sizes in microchips the scalability of proposed CMPs is of great interest to the developers. Looking even further into the future at the possibilities offered by, e. g., nanotechnology, a set of such models may help to identify promising architectures and possible bottlenecks even before the enabling technologies exist. In this paper, we present our current work in this area in the form of two models. The first and very basic model is based on Amdahl's law and gives a first promising outlook on chip multiprocessing. Based on the more complex BSP model, our second model takes the on-chip communication into account and thus allows a much more detailed look at the architecture. In both cases, basic laws of circuit technology have been combined with the underlying models that now take the effects of device scaling into account. Later, we also present the GigaNetIC architecture, a CMP developed by our research group. It was then analyzed by applying the BSP-based model.
Björn Jäger, Jörg-Christian Niemann, Ulrich Rückert 0001
ASAP3
2005 Tolerance of Radial Basis Functions Against Stuck-At-Faults
Ralf Eickhoff, Ulrich Rückert 0001
ICANN (2)2
2005 Low-cost Bluetooth Communication for the Autonomous Mobile Minirobot Khepera
abstract
This paper presents a low-cost Bluetooth communication for autonomous mobile minirobots. The interfacing of the Bluetooth hardware is done by a simple UART connection, which makes the approach easily portable. Using ASCII commands and events, which are exchanged over this serial link, the Bluetooth operations can be controlled and monitored. Since internal Bluetooth stack operations are concealed, no deeper knowledge of the Bluetooth technology is necessary to utilize this wireless communication. These features make the presented approach perfectly suited for the integration into minirobots like the Khepera, where computational power and spatial resources are strictly limited.
Michael Grosseschallau, Ulf Witkowski, Ulrich Rückert 0001
ICRA3
2005 CSD: cell-based service discovery in large-scale robot networks
abstract
If robots are deployed in large numbers in our environment in future, collaboration between the presumably specialized robots will be essential for a successful operation. The robots will set up mobile ad-hoc networks for communication and efficient routing and discovery protocols for such robot networks will provide the basic layer for a successful collaboration of the robots. In this paper we present a service discovery protocol that allows robots to efficiently discover available services in the network. It is specifically designed for large-scale robot networks. It takes into account the high dynamics of robot networks and exploits the position data of the robots to increase scalability and efficiency. A cell-based grid with master nodes in each cell forms the basic structure. Through proactive intra-cell communication and reactive inter-cell communication, scalability is ensured and the effects of node movements on the overall network are minimized. We implemented our solution in an example scenario.
Jia Lei Du, Ulf Witkowski, Ulrich Rückert 0001
IROS3
2005 A low complexity directional scheme for mobile ad hoc networks
abstract
In the context of single frequency band and omnidirectional communication, the throughput of wireless network is interference limited. Collision may occur if other signals impinge the destination while it is receiving the intended signal. To prevent collisions, CSMA/CA is applied to serialize the communication tasks. Recently, numerous directional schemes have been proposed to further increase the network throughput. In this paper, a practical scheme for directional communication based on a simplified switched beam technique is proposed on the physical layer. Additionally, the corresponding modifications on the CSMA/CA protocol are also studied to optimize the overall system performance. It is worthy of notice that this scheme is feasible for the mobile portable device under state-of-the-art technologies in terms of dimension, power consumption, hardware complexity and computing requirements. Preliminary simulations show that this scheme can increase the network throughput up to two times compared to an omni-directional system at the cost of only 10% higher power consumption
Matthias Grünewald, Ulrich Rückert 0001
PIMRC3
2005 Fault-tolerance of basis function networks using tensor product stabilizers
abstract
Neural networks are intended to be used in future nanoelectronics since these architectures seem to be fault-tolerant to malfunctioning elements and robust to noise. In this paper, the robustness to noise of basis function networks using tensor product stabilizers is analyzed and upper bounds of the mean square error under noise contaminated weights or inputs are determined. Furthermore, consequences of permanently malfunctioning neurons are investigated and their impact on the mean squared error is analyzed. To achieve a reliable operation of the neural network necessary restrictions are introduced. Finally, the impact of technical realizations is investigated and its complexity is compared to radial basis functions.
Ralf Eickhoff, Ulrich Rückert 0001
SMC2
2005 Defragmentation Algorithms for Partially Reconfigurable Hardware
abstract
Dynamic reconfiguration is a promising approach for resource efficient utilization of microelectronic systems. Standard platforms for partial dynamic reconfiguration are field-programmable gate arrays (FPGAs). Multiple hardware tasks can share the same FPGA resources over time, which increases the device utilization in comparison to non-reconfigurable systems. Although, similar resource management is already known in the area of operating systems, there is a requirement to adapt these concepts to the special needs of dynamically reconfigurable systems. Additionally, there is a lack of underlying mechanisms, e.g., to suspend hardware tasks and restart them at a different position within the FPGA. In this article we introduce a mechanism for task relocation that includes saving and restoring of state information of the task. Based on this approach we address the problem of defragmentation. We present defragmentation algorithms that minimize different types of costs. With the help of a detailed simulation model and a benchmark, we finally provide realistic simulation results and compare the different algorithms.
Markus Köster, Heiko Kalte, Mario Porrmann, Ulrich Rückert 0001
VLSI-SoC4
2004 A Mapping Strategy for Resource-Efficient Network Processing on Multiprocessor SoC
abstract
Hardware architectures based on a field of hardware-extended processors can provide flexible computing power for applications where parallelism can be exploited. For multiprocessors, the assignment of functionality to execution units can have a great impact on the performance. Additionally, finding the optimal mapping can be a time-consuming task. We present a multiprocessor architecture along with a suitable design method that includes an automated solution to the mapping problem. Our hardware architecture employs a network-on-chip (NoC) to achieve a high degree of scalability for the application and for the system in respect to future integration technologies. We also show how to reduce the packet buffer requirements with a proper scheduling strategy and present first estimates for the resource consumption of an application targeted for mobile networking.
Matthias Grünewald, Jörg-Christian Niemann, Mario Porrmann, Ulrich Rückert 0001
DATE4
2004 Hardware Support for Dynamic Reconfiguration in Reconfigurable SoC Architectures
Björn Griese, Erik Vonnahme, Mario Porrmann, Ulrich Rückert 0001
FPL4
2004 Study on column wise design compaction for reconfigurable systems
abstract
Some of currently available field programmable gate arrays (FPGAs) can be reconfigured partially, which makes it possible to build up dynamic systems that can be adapted to changing demands during runtime. One basic aspect of such a system is the way the dynamic hardware modules are placed on the FPGA. As most FPGAs offer partial reconfiguration in a column wise manner, a 1D placement of column wise implemented modules seems to be promising. Within This work we present a design study that determines the effects of a column wise module implementation on the resulting frequency and power consumption.
Heiko Kalte, Gareth Lee, Mario Porrmann, Ulrich Rückert 0001
FPT4
2004 gNBX - reconfigurable hardware acceleration of self-organizing maps
abstract
In this work a new FPGA based hardware accelerator (gNBX) for self-organizing maps is introduced. New principles for hardware acceleration of self-organizing maps, which increase the degree of parallelity and therefore the acceleration gain was presented. Our technology independent design description can be mapped on application specific integrated circuits if very high performance is required, as well as on field programmable gate arrays (FPGAs), which offer mid level performance (a speed up factor of up to 70 in comparison with PCs for typical datasets is achieved) at relatively low costs. Additionally, FPGAs offer the flexibility to adapt the hardware to the changing requirements of the application during runtime. Therefore, the hardware can be exploited optimally at all times during the simulation process. Several benchmark scenarios with well known datasets shows the performance of our system.
Christopher Pohl, Marc Franzmeier, Mario Porrmann, Ulrich Rückert 0001
FPT4
2004 System-on-Programmable-Chip Approach Enabling Online Fine-Grained 1D-Placement
abstract
Summary form only given. The increasing logic density of current FPGAs (field programmable gate arrays) enables the integration of whole systems on one programmable chip. Some of these FPGAs provide the additional feature of partial dynamic reconfiguration, which permits to change parts of the device while other parts keep working. Combining the features of system level density and partial dynamic reconfiguration enables the integration of dynamic systems that can be adopted to changing demands during runtime. A lot of theoretical work in this challenging research area has been done on efficiently placing and scheduling modules on the FPGA area. However, there is a lack of applied approaches that can be realized by existing tools and FPGAs. We present a new, realizable approach for the dynamic system integration on Xilinx Virtex FPGAs. In contrast to the existing approaches that consider fixed slots for the module placement, our approach enables the fine-grained placement of modules with variable width along a horizontal communication infrastructure.
Heiko Kalte, Mario Porrmann, Ulrich Rückert 0001
IPDPS3
2004 V: Drive - Costs and Benefits of an Out-of-Band Storage Virtualization System
André Brinkmann, Michael Heidebuer, Friedhelm Meyer auf der Heide, Ulrich Rückert 0001, Kay Salzwedel, Mario Vodisek
MSST4
2003 A holistic methodology for network processor design
abstract
The GigaNetIC project aims to develop high-speed components for networking applications based on massively parallel architectures. A central part of this project is the design, evaluation, and realization of a parameterizable network processing unit. In this paper we present a design methodology for network processors which encompasses the research areas from the application software down to the gate level of the chip. Key components of this holistic approach have been successfully applied to characteristic examples of architecture refinements.
Olaf Bonorden, Nikolaus Brüls, Uwe Kastens, Dinh Khoi Le, Friedhelm Meyer auf der Heide, Jörg-Christian Niemann, Mario Porrmann, Ulrich Rückert 0001, Adrian Slowik, Michael Thies
LCN8
2003 A massively parallel architecture for self-organizing feature maps
abstract
A hardware accelerator for self-organizing feature maps is presented. We have developed a massively parallel architecture that, on the one hand, allows a resource-efficient implementation of small or medium-sized maps for embedded applications, requiring only small areas of silicon. On the other hand, large maps can be simulated with systems that consist of several integrated circuits that work in parallel. Apart from the learning and recall of self-organizing feature maps, the hardware accelerates data pre- and postprocessing. For the verification of our architectural concepts in a real-world environment, we have implemented an ASIC that is integrated into our heterogeneous multiprocessor system for neural applications. The performance of our system is analyzed for various simulation parameters. Additionally, the performance that can be achieved with future microelectronic technologies is estimated.
Mario Porrmann, Ulf Witkowski, Ulrich Rückert 0001
IEEE Trans. Neural Networks3
2002 A reconfigurable SOM hardware accelerator
Mario Porrmann, Marc Franzmeier, Heiko Kalte, Ulf Witkowski, Ulrich Rückert 0001
ESANN5
2002 Dynamically Reconfigurable Hardware - A New Perspective for Neural Network Implementations
Mario Porrmann, Ulf Witkowski, Heiko Kalte, Ulrich Rückert 0001
FPL4
2002 A Direction Sensitive Network Based on a Biophysical Neurone Model
Burkhard Iske, Axel Löffler, Ulrich Rückert 0001
ICANN3
2002 Continuous Sonar Sensing for Mobile Mini-Robots
abstract
Ultrasonic sensors enable mobile autonomous systems to obtain information about obstacles in large environments. In the presented work, a 5 cm broad array of three piezo-ceramic ultrasonic transducers is employed for getting two-dimensional impressions of the surroundings. Deviating from the pulse echo measurement techniques used so far the time-continuous transmitting and receiving from modulated pseudo-random sequences are considered. Thus the narrow bandwidth of a piezo-ceramic transducer can be compensated by an increased measuring period. An advanced analysis of the correlated signals allows the rejection of phantoms caused by multiple reflections. Furthermore, a classification of objects such as wall, corner, log or cylinder is possible.
Jürgen Klahold, Jens Rautenberg, Ulrich Rückert 0001
ICRA3
2002 Simulation of spiking neural networks -- architectures and implementations
Martin Schäfer, Tim Schönauer, Carsten Wolff, Georg Hartmann, Heinrich Klar, Ulrich Rückert 0001
Neurocomputing6
2001 A methodology for behaviour design of autonomous systems
abstract
A new methodology of behaviour design and modelling is proposed, which is based on a tree structure. The tree allows a structured design and overview of autonomous systems behaviours. A behaviour of higher abstraction level consists of a combination of one or several behaviours of lower abstraction level. The advantage of the tree becomes clear when wanting to reuse already developed behaviours. In order to reuse a behaviour of higher level of abstraction several behaviours of lower level of abstraction are required, which can easily be identified when describing a behaviour in the proposed tree structure. Lower levels of behaviour are mostly system dependant. By replacing only the behaviours of low level of abstraction behaviours can easily be transferred to other systems. Additionally, the behaviour tree enables the estimation and evaluation of resource requirements of different behaviours of different levels of abstraction.
Burkhard Iske, Ulrich Rückert 0001
IROS2
2000 A Bootstrapping Method for Autonomous and in Site Learning of Generic Navigation Behavior
abstract
To understand the behaviour of natural autonomous systems, research is carried out on artificial autonomous agents. The paper focuses on how simple behaviours can be learnt autonomously using a bootstrapping method. Firstly, a two dimensional self-organising map is realised which provides the agent's sense of orientation. Once this relative positioning system has been established, the agent learns to navigate towards a target using the reinforcement learning technique of Q-learning. Since only neural network processing is used, this technique emulates the distributed and adaptive information processing found in natural autonomous systems. Furthermore, due to its generality, the neural implementation developed is transferable to other artificial autonomous agents with different sensors and effector suites.
Burkhard Iske, Ulrich Rückert 0001, Kurt Malmstrom, Joaquin Sitte
ICPR2
1998 SOM accelerator system
Stefan Rüping 0002, Mario Porrmann, Ulrich Rückert 0001
Neurocomputing3
1998 Local cluster neural net analog VLSI design
Joaquin Sitte, Tim Körner, Ulrich Rückert 0001
Neurocomputing3
1993 Acceleratorboard for neural associative memories
Ulrich Rückert 0001, Andreas Funke, Christof Pintaske
Neurocomputing1
1984 Intelligent memories in VLSI
Karl Goser, Carsten Foelster, Ulrich Rückert 0001
Inf. Sci.3