EDBT 2026 Demo / reviewers in the wild / expert
Andreas Herkersdorf
dblp:62/5988
· DBLP profile ↗
106ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-8886-5345ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 82 · 2 first-author · 15 since 2021Software engineering, systems software and programming languages · 27 · 1 first-author · 6 since 2021Computer networks · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Advancing European Semiconductor and Chiplet Innovation Through the Bavarian Chip Design CenterabstractEurope’s semiconductor industry relies heavily on Asian and US manufacturers. The EU Chips Act seeks to strengthen Europe’s capabilities across the semiconductor value chain. Aligned with this goal, the Bavarian Chip Design Center (BCDC) supports local chip design, manufacturing, and talent development, with a focus on RISC-V computing and heterogeneous integration. Within BCDC, the Technical University of Munich and Fraunhofer are developing a chiplet-based architecture optimized for low-power edge AI. The system integrates two chiplets, combining a security-enhanced RISC-V core and AI accelerators, connected via a chiplet-optimized serial interface that supports encrypted data. The chiplets are mounted on a custom interposer with low-capacitance wires for efficient data transmission. System-and component-level development is currently ongoing, with a tapeout in 22 nm FD-SOI planned for 2027. The overall goal is to deliver a proof of concept for a small-scale energy-efficient chiplet system that demonstrates Bavaria’s and Europe’s capability to drive innovation in novel chip design fields. Hussam Amrouch, Jehaan Joseph, Michael Schirmer, Johannes Geier, Ulf Schlichtmann, Michael Meidinger, Thomas Wild, Andreas Herkersdorf, Jens Nöpel, Georg Sigl, Carsten Trinitis, Aswathy Nedumpalli Sankaranarayanan, Martin Schulz 0001, Andreas Korb, Konrad Hohentanner |
DATE | 8 |
| 2026 | STEP: Spatial Footprint Prefetcher with Multi-Point Temporal Triggers
Yuanji Ye, Oliver Lenke, Thomas Wild, Andreas Herkersdorf |
ISCA | 4 |
| 2025 | HiPerNoC: A High-Performance Network-an-Chip for Flexible and Scalable FPGA-Based SmartNICsabstractA recent approach that the research community has proposed to address the steep growth of network traffic and the attendant rise in computing demands is in-network computing. This paradigm shift is bringing about an increase in the types of computations performed by network devices. Consequently, processing demands are becoming more varied, requiring flexible packet-processing architectures. State-of-the-art switch-based smart network interface cards (SmartNICs) provide high versatility without sacrificing performance but do not scale well concerning resource usage. In this paper, we introduce HiPerNoC-a flexible and scalable field-programmable gate array (FPGA)-based SmartNIC architecture deploying a 2D-mesh network-on-chip (NoC) with a novel router design to manage network traffic with diverse processing demands. The NoC can forward incoming network packets to the available processing engines in the required sequence at a traffic load of up to 91.1 Gbit/s (0.89 flit/node/cycle). Each router applies distributed switch allocation and avoids head-of-line blocking by deploying queues at the switch crosspoints of input-output connections used by the routing algorithm. It also prevents deadlocks by employing non-blocking virtual cut-through switching. We implemented a prototype of HiPerNoC as a 4x4 2D-mesh NoC in SystemVerilog and evaluated it with synthetic network traffic via cycle-accurate register-transfer level simulations in Vivado. The evaluation results show that HiPerNoC achieves up to 53 % higher saturation throughput, occupies 53 % fewer lookup tables and block RAMs, and consumes 16 % less power on an Alveo U55C than ProNoC-a state-of-the-art FPGA-based NoC. Klajd Zyla, Marco Liess, Thomas Wild, Andreas Herkersdorf |
DATE | 4 |
| 2025 | Rule-Based Reinforcement Learning on FPGA for QoS-Aware Dynamic Frequency ScalingabstractTo improve the system performance of multiprocessor system-on-chips (MPSoCs), modern processors have several built-in hardware features, such as prefetchers, which respond to short-term variations in processor load that occurs on a submillisecond scale. However, even the latest dynamic (voltage) frequency scaling governors using reinforcement learning (RL) are implemented in software and, thus, cannot take advantage of these variations. In this work, we propose a hardware RL agent, augmented with preemptive shielding and eligibility traces, to optimize the execution of deadline-bound quality-of-service (QoS) tasks in mixed-critical environments. We demonstrate the features of our algorithm in a hardware-in-the-loop simulation by running LLVM’s single-source benchmarks on SparcV8 processors. We also present our field-programmable gate array (FPGA) implementation with optimized resource usage and timing performance achieved through quantization and approximation. Florian Maurer 0003, Michael Meidinger, Matthias Schlemmer, Thomas Hallermeier, Anmol Surhonne, Thomas Wild, Andreas Herkersdorf |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2024 | Hardware Assist for Linux IPC on an FPGA PlatformabstractSpecialized hardware units often accelerate compute-intensive or memory-heavy functions. In previous publications, we proposed concepts to assist Linux with a hardware unit for managing waiting threads to improve blocking inter-process communication (IPC) mechanisms. This paper assesses the effectiveness of this hardware support on a Zynq platform. Although main memory accesses by our hardware unit are time-consuming, a consumer-producer application achieved an up to 220% increased message rate. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
CF | 6 |
| 2024 | HASIIL: Hardware-Assisted Scheduling to Improve IPC Latency in LinuxabstractInter-processes communication (IPC) is essential for multi-threaded applications to achieve efficient execution. Synchronization through IPC can become a bottleneck for these applications. The effectiveness of IPC is determined by both its latency and CPU utilization needed for the associated functions. Our research has revealed that for blocking IPC mechanisms, the thread scheduling functions within the Linux operating system significantly contribute to the notification latency. To address this issue, we propose a novel concept called HASIIL, which combines offloading IPC functionality with hardware-assisted scheduling to enhance IPC latency. Through this approach, we can improve the latency of blocking IPC mechanisms by up to 36% in Linux, while also improving CPU utilization by 40%. Tim Twardzik, Lars Nolte, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
CF | 6 |
| 2024 | EMDRIVE Architecture: Embedded Distributed Computing and Diagnostics from Sensor to EdgeabstractFuture automotive architectures are expected to transition from a network-centric to a domain-centered architecture featuring central compute units. Powerful domain controllers or smart sensors alleviate the load on these central units and communication systems. These controllers execute tasks with varying criticalities on heterogeneous multicore processors, and are ideally capable of dynamically balancing the computing load between the central unit and sensors. Here, Artificial Intelligence (AI) capabilities playa crucial role, as it is in high demand for such an automotive architecture. However, AI still requires specialized accelerators to improve their computation performance. Task-oriented distributed computing with criticalities up to ASIL-D necessitates the development and utilization of specialized methodologies, such as safety, through the isolation and abstraction of low-level hardware concepts. Meanwhile, online monitoring and diagnostics become vital features to detect errors during operation. The EMDRIVE architecture includes methods, components, and strategies to enhance the performance, safety, and security of such distributed computing platforms. The nationally funded EMDRIVE project connects its twelve partners from academia and industry and is currently in its intermediate stage. Patrick Schmidt 0003, Iuliia Topko, Matthias Stammler, Tanja Harbaum, Jürgen Becker 0001, Rico Berner, Omar Ahmed, Jakub Jagielski, Thomas Seidler, Markus Abel, Marius Kreutzer, Maximilian Kirschner, Victor Pazmino Betancourt, Robin Sehm, Lukas Groth, Andrija Neskovic, Rolf Meyer, Saleh Mulhem, Mladen Berekovic, Matthias Probst, Manuel Brosch, Georg Sigl, Thomas Wild, Matthias Ernst, Andreas Herkersdorf, Florian Aigner, Stefan Hommes, Sebastian Lauer, Maximilian Seidler, Thomas Raste, Gasper Skvarc Bozic, Ibai Irigoyen Ceberio, Albrecht Mayer |
DATE | 25 |
| 2024 | ecoNIC: Saving Energy Through SmartNIC-Based Load Balancing of Mixed-Critical Ethernet TrafficabstractIn next-generation automotive, industrial, data cen-ter, and other mixed-critical networks, Ethernet is expected to power the backbone interconnect among multi-core compute nodes. On attached Network Interface Cards (NICs) Receive Side Scaling (RSS) supports the CPU in balancing workloads across cores for reduced tail latencies. However, state-of-the-art solutions are primarily designed for performance and less for energy-efficiency which will play an equally important role. For this reason we present ecoNIC, an RSS-based hardware load balancer for SmartNICs, and an agile Dynamic Voltage and Frequency Scaling (DVFS) governor, for energy-saving network processing. ecoNI C efficiently pins flow priorities to CPU core clusters, reducing the workload of select cores in the process, and dynamically adjusts their clock speed to exploit freed-up capacities and save energy. Within a cluster, it proactively redirects packet bursts of priority-separated flow bundles among available cores, or offloads them to neighbor nodes, once local resources tend to become highly loaded. The per-core energy consumption this way is reduced at the expense of low priority packet latencies, while high priority service qualities are maintained. Experimental evaluations applying real-world network traces yield energy savings of up to 37.9 % at an increase from 559 µs to 3.06 ms in low priority end-to-end tail latency compared to an even workload distribution without frequency scaling. Franz Biersack, Marco Liess, Markus Absmann, Fabiana Lotter, Thomas Wild, Andreas Herkersdorf |
DSD | 6 |
| 2024 | FlexCross: High-Speed and Flexible Packet Processing via a Crosspoint-Queued CrossbarabstractThe fast pace at which new online services emerge leads to a rapid surge in the volume of network traffic. A recent approach that the research community has proposed to tackle this issue is in-network computing, which means that network devices perform more computations than before. As a result, processing demands become more varied, creating the need for flexible packet-processing architectures. State-of-the-art approaches provide a high degree of flexibility at the expense of performance for complex applications, or they ensure high performance but only for specific use cases. In order to address these limitations, we propose FlexCross. This flexible packet-processing design can process network traffic with diverse processing requirements at over 100 Gbit/s on FPGAs. Our design contains a crosspoint-queued crossbar that enables the execution of complex applications by forwarding incoming packets to the required processing engines in the specified sequence. The crossbar consists of distributed logic blocks that route incoming packets to the specified targets and resolve contentions for shared resources, as well as memory blocks for packet buffering. We implemented a prototype of FlexCross in Verilog and evaluated it via cycle-accurate register-transfer level simulations. We also conducted test runs with real-world network traffic on an FPGA. The evaluation results demonstrate that FlexCross outperforms state-of-the-art flexible packet-processing designs for different traffic loads and scenarios. The synthesis results show that our prototype consumes roughly 21% of the resources on a Virtex XCU55 UltraScale+ FPGA. Klajd Zyla, Marco Liess, Thomas Wild, Andreas Herkersdorf |
DSD | 4 |
| 2024 | FlexRoute: A Fast, Flexible and Priority-Aware Packet-Processing DesignabstractAs the world becomes more connected and new digital services emerge at a fast pace, the amount of network traffic increases rapidly. Consequently, processing requirements become more varied and drive the need for flexible packet-processing designs, especially as in-network computing gains traction. Traditional approaches deploy hardware accelerators in a pipeline in the sequence that the associated tasks are supposed to be executed. Hence, they do not accommodate flows with different processing requirements and provide no possibility to remap flows to task sequences in runtime. In order to address these limitations, we propose FlexRoute, a fast, flexible and priority-aware packet-processing design that can process network traffic at a rate of over 100 Gbit/s on FPGAs. Our design consists of a reconfigurable parser and several processing engines that are arranged in a pipeline. The processing engines are equipped with processing units that execute specific tasks, flexible forwarding logic and priority-aware queuing/scheduling logic. We implement a prototype of FlexRoute in Verilog and evaluate it via cycle-accurate register-transfer level simulations. We also synthesize and implement our design on the Alveo U55C High Performance Compute Card and show its resource usage. The evaluation results demonstrate that FlexRoute can process packets of arbitrary size with different processing requirements at a traffic rate of about 70 Gbit/s significantly faster than two state-of-the-art flexible packet-processing designs. Klajd Zyla, Marco Liess, Thomas Wild, Andreas Herkersdorf |
PDP | 4 |
| 2024 | HW-FUTEX: Hardware-Assisted Futex SyscallabstractEfficient thread synchronization primitives are crucial in modern computer systems for the performant execution of interdependent code segments. In Linux, the futex() syscall is used to construct blocking synchronization primitives such as mutexes or conditional variables. When using futex, the uncontended case is efficiently handled entirely in user space. In the event of contention, the kernel is called to put the waiting thread to sleep until the state of the primitive changes to uncontended. The kernel must be notified of this change by a futex() syscall to wake-up the sleeping thread. This syscall must be issued by the thread that changes the primitive, which is a significant burden on this thread. To remove this burden, we introduce HW-FUTEX to offload the futex wake functionality to a hardware unit (HW Unit) that asynchronously initiates wake-ups of the sleeping threads. This reduces the time required to issue the futex wake functionality by at least 90% to 350 cycles, with no additional overhead in the uncontended case. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2023 | HAWEN: Hardware Accelerator for Thread Wake-Ups in Linux Event NotificationabstractThe performance of multi-threaded applications relies on efficient inter-process communication. One common practice is putting a thread asleep while waiting for a certain condition. Exemplary Linux kernel mechanisms that use this practice include futex, sockets, epoll, eventfd and pipe. Once the condition is met, i.e., the associated event has occurred, the waiting thread is notified. Optimizations for event notification mechanisms in Linux mostly target the thread which receives events. Contrarily, we identified high potential in relieving the event-generating thread and propose HAWEN, a hardware accelerator for thread wake-up support. HAWEN has been integrated into Linux event notification in a minimally intrusive manner. Gem5-based multi-core architecture simulations revealed up to 80% faster thread wake-up times and a 53% shorter event-generating syscall. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Clara Kowalsky, Thomas Wild, Andreas Herkersdorf |
DAC | 8 |
| 2023 | Information Processing Factory 2.0 - Self-awareness for Autonomous Collaborative SystemsabstractThis paper summarizes the talks of a special session on the IPF 2.0 project, a collaborative German-US research project that leverages self-awareness principles for the self-management of distributed systems of autonomous multiprocessor systems-on-chip (MPSoCs). Nora Sperling, Alex Bendrick, Dominik Stöhrmann, Rolf Ernst, Bryan Donyanavard, Florian Maurer 0003, Oliver Lenke, Anmol Surhonne, Andreas Herkersdorf, Walaa Amer, Caio Batista de Melo, Ping-Xiang Chen, Quang Anh Hoang, Rachid Karami, Biswadip Maity, Paul Nikolian, Mariam Rakka, Dongjoo Seo, Saehanseul Yi, Minjun Seo, Nikil Dutt, Fadi J. Kurdahi |
DATE | 9 |
| 2023 | Priority-aware Inter-Server Receive Side ScalingabstractNext-generation automotive networks will be characterized by a high number of interconnected sensors, actuators and applications on electronic control units communicating with each other over a high-speed Ethernet backbone network. As these applications have various criticalities, high volumes of fluctuating traffic with different priorities will have to be processed in a reliable and efficient manner. To cope with these challenges, we present Priority-aware Inter-Server Receive Side Scaling (prioRSS), a new SmartNIC-based hardware accelerator designed for automotive compute nodes. prioRSS builds upon Receive Side Scaling and introduces priority-awareness into an intra- and inter-node load balancer. It uses a priority-partitioned indirection table within which flows of the same priority are bundled. Low-latency reconfigurations issued by a Network Health Monitoring software allow for adapting the table content to changing network conditions. Simulative evaluations and comparisons to a priority-unaware version of our design show that prioRSS enables per-priority resource assignments without degrading end-to-end packet latencies while using the same table memory space. Paired with a priority-aware scheduler, end-to-end latencies of high priority flows can be notably reduced compared to average packet latencies, at the expense of lowest priority traffic. The best results are acquired when partitioning the table proportionally to the associated traffic share. Franz Biersack, Kilian Holzinger, Henning Stubbe, Thomas Wild, Georg Carle, Andreas Herkersdorf |
PDP | 6 |
| 2023 | FlexPipe: Fast, Flexible and Scalable Packet Processing for High-Performance SmartNICsabstractData centers have been struggling to provide the necessary processing capacity to handle the surging rate of network traffic that is generated in an increasingly connected and service-oriented world. As a result, SmartNICs play an even more important role than before as they can offload various network applications and hence free CPU resources for application-layer processing, increase performance and reduce processing time. However, they often do not support flows with different offload requirements and cannot dynamically allocate offloads in run-time. In order to address these limitations, we propose FlexPipe, a fast, flexible and scalable packet-processing architecture for high-performance SmartNICs. Our design enables low-latency and runtime-reconfigurable packet forwarding at high traffic rates with minimal area overhead. Furthermore, it provides load-aware packet steering toward multiple offload units of the same type for low-bandwidth offloads. We implement a prototype of FlexPipe in Verilog and validate it via cycle-accurate register-transfer level simulations. Our evaluation results show that FlexPipe can process packets of arbitrary size with different offload requirements at line rate and on average 1.9x faster than a SmartNIC with a predefined sequence of offloads and 1.8x faster than PANIC, a flexible state-of-the-art SmartNIC. Klajd Zyla, Marco Liess, Thomas Wild, Andreas Herkersdorf |
VLSI-SoC | 4 |
| 2022 | SmartNIC-based Load Management and Network Health Monitoring for Time Sensitive ApplicationsabstractTime sensitive network applications, for example in Intra-Vehicular Networks, aim to give predictable end-to-end latency guarantees. As a consequence, processing resources of involved host systems remain partially unused, because they are reserved for rare worst cases. This circumstance provides the opportunity to reduce dimensioning overheads by managing the load on the nodes flexibly within the network. In our proposed approach, a SmartNIC involving an FPGA-based load balancer achieves dynamic routing of flows whilst preserving end-to-end latency guarantees. A flow-oriented online network measurement component continuously supervises network traffic with regards to compliance to flow specifications and constraints such as bounded one-way delay, absence of packet loss, and jitter. We use the supervisor to enhance forwarding decisions on the data plane. Initial evaluation yields a saving potential of around 30 %. We showcase quick dynamic reconfiguration of the FPGA when triggered by real-time measurement of the one-way delay using realistic automotive network traffic. Kilian Holzinger, Franz Biersack, Henning Stubbe, Angela Gonzalez Mariño, Abdoul Kane, Francesc Fons, Haigang Zhang, Thomas Wild, Andreas Herkersdorf, Georg Carle |
NOMS | 9 |
| 2021 | Precise real-time monitoring of time-critical flowsabstractEthernet is increasingly used in areas where time-critical and safety-relevant data are transported over the network along with best-effort flows, for example in intra vehicle networks or industrial networks. The resulting complex network architectures, time-sensitive networking configurations and system interactions are hard to foresee during the design phase. Therefore, it is hard to rule out any violations of flow specifications or timing and reliability requirements, especially in the presence of unpredictable failures. Kilian Holzinger, Henning Stubbe, Franz Biersack, Angela Gonzalez Mariño, Abdoul Kane, Francesc Fons, Haigang Zhang, Thomas Wild, Andreas Herkersdorf, Georg Carle |
CoNEXT | 9 |
| 2021 | Long Short-Term Memory Neural Network-based Power Forecasting of Multi-Core ProcessorsabstractWe propose a novel technique to forecast the power consumption of processor cores at run-time. Power consumption varies strongly with different running applications and within their execution phases. Accurately forecasting future power changes is highly relevant for proactive power/thermal management. While forecasting power is straightforward for known or periodic workloads, the challenge for general unknown workloads at different voltage/frequency (v/n-levels is still unsolved. Our technique is based on a long short-term memory (LSTM) recurrent neural network (RNN) to forecast the average power consumption for both the next 1ms and 10ms periods. The runtime inputs for the LSTM RNN are current and past power information as well as performance counter readings. An LSTM RNN enables this forecasting due to its ability to preserve the history of power and performance counters. Our LSTM RNN needs to be trained only once at design-time while adapting during run-time to different system behavior through its internal memory. We demonstrate that our approach accurately forecasts power for unseen applications at different v/f-levels. The experimental results shows that the forecasts of our LSTM RNN provide 43% lower worst case error for the 1ms forecasts and 38% for the 10ms forecasts. comnared to the state of the art. Mark Sagi, Martin Rapp, Heba Khdr, Yizhe Zhang 0005, Nael Fasfous, Nguyen Anh Vu Doan, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
DATE | 9 |
| 2021 | SEAMS: Self-Optimizing Runtime Manager for Approximate Memory HierarchiesabstractMemory approximation techniques are commonly limited in scope, targeting individual levels of the memory hierarchy. Existing approximation techniques for a full memory hierarchy determine optimal configurations at design-time provided a goal and application. Such policies are rigid: they cannot adapt to unknown workloads and must be redesigned for different memory configurations and technologies. We propose SEAMS: the first self-optimizing runtime manager for coordinating configurable approximation knobs across all levels of the memory hierarchy. SEAMS continuously updates and optimizes its approximation management policy throughout runtime for diverse workloads. SEAMS optimizes the approximate memory configuration to minimize energy consumption without compromising the quality threshold specified by application developers. SEAMS can (1) learn a policy at runtime to manage variable application quality of service ( QoS ) constraints, (2) automatically optimize for a target metric within those constraints, and (3) coordinate runtime decisions for interdependent knobs and subsystems. We demonstrate SEAMS’ ability to efficiently provide functions (1)–(3) on a RISC-V Linux platform with approximate memory segments in the on-chip cache and main memory. We demonstrate SEAMS’ ability to save up to 37% energy in the memory subsystem without any design-time overhead. We show SEAMS’ ability to reduce QoS violations by 75% with < 5% additional energy. Biswadip Maity, Bryan Donyanavard, Anmol Surhonne, Amir-Mohammad Rahmani, Andreas Herkersdorf, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2020 | Emergent Control of MPSoC Operation by a Hierarchical Supervisor / Reinforcement Learning ApproachabstractMPSoCs increasingly depend on adaptive resource management strategies at runtime for efficient utilization of resources when executing complex application workloads. In particular, conflicting demands for adequate computation performance and power-/energy-efficiency constraints make desired application goals hard to achieve. We present a hierarchical, cross-layer hardware/software resource manager capable of adapting to changing workloads and system dynamics with zero initial knowledge. The manager uses rule-based reinforcement learning classifier tables (LCTs) with an archive-based backup policy as leaf controllers. The LCTs directly manipulate and enforce MPSoC building block operation parameters in order to explore and optimize potentially conflicting system requirements (e.g., meeting a performance target while staying within the power constraint). A supervisor translates system requirements and application goals into per-LCT objective functions (e.g., core instructions-per-second (IPS). Thus, the supervisor manages the possibly emergent behavior of the low-level LCT controllers in response to 1) switching between operation strategies (e.g., maximize performance vs. minimize power; and 2) changing application requirements. This hierarchical manager leverages the dual benefits of a software supervisor (enabling flexibility), together with hardware learners (allowing quick and efficient optimization). Experiments on an FPGA prototype confirmed the ability of our approach to identify optimized MPSoC operation parameters at runtime while strictly obeying given power constraints. Florian Maurer 0003, Bryan Donyanavard, Amir-Mohammad Rahmani, Nikil Dutt, Andreas Herkersdorf |
DATE | 5 |
| 2020 | Inter-Server RSS: Extending Receive Side Scaling for Inter-Server Workload DistributionabstractNetwork Function Virtualization enables operators to schedule diverse network processing workloads on a general-purpose hardware infrastructure. However, short-lived processing peaks make an efficient dimensioning of processing resources under stringent tail latency constraints challenging. To reduce dimensioning overheads, several load balancing approaches, which either adaptively steer network traffic to a group of servers or to their internal CPU cores, have separately been investigated.In this paper, we present Inter-Server RSS (isRSS), a hardware mechanism built on top of Receive Side Scaling in the network interface card, which combines intra-and inter-server load balancing. In a first step, isRSS targets a balanced utilization of processing resources by steering packet bursts to CPU cores based on per-core load feedback. If all local CPU cores are highly loaded, isRSS avoids high queueing delays by redirecting newly arriving packet bursts to other servers, which execute the same network functions, exploiting that processing peaks are unlikely to occur at all servers at the same time. Our evaluation based on real-world network traces shows that compared to Receive Side Scaling, the joint intra-and inter-server load balancing approach is able to reduce the processing capacity dimensioned for network function execution by up to 38.95% and limit packet reordering to 0.0589% while maintaining tail latencies. Andreas Oeldemann, Franz Biersack, Thomas Wild, Andreas Herkersdorf |
PDP | 4 |
| 2020 | Power- and Cache-Aware Task Mapping with Dynamic Power Budgeting for Many-CoresabstractTwo factors primarily affect the performance of multi-threaded tasks on many-core processors with logically-shared and physically-distributed Last-Level Cache (LLC): the LLC latencies of threads running on different cores and the per-core power budgets that aim to guarantee thermally safe operation. Two knobs affect these factors: First, the mapping of threads to cores affects both the LLC latencies and the power budgets. Second, dynamic power budgeting refines the power budgets during task execution. A mapping that spatially distributes threads across the many-core increases the power budgets, but unfortunately also increases the LLC latencies. Contrarily, mapping all threads near the center of the many-core minimizes the LLC latencies, but unfortunately also decreases the power budgets. Consequently, both metrics cannot be simultaneously optimal, which leads to a Pareto-optimization for task mapping that has formerly not been exploited. Dynamic power budgeting reallocates the power budgets according to the tasks' execution phases. This results in a dynamically changing non-uniform power budget, which further increases the performance. We are the first to present a run-time algorithm PCGov combining task-agnostic task mapping and task-aware dynamic power budgeting for many-cores with shared distributed LLC. PCGov yields up to 21 percent lower response time and 13 percent lower energy consumption compared to the state-of-the-art, with a low overhead of less than 0.5 percent. Martin Rapp, Mark Sagi, Anuj Pathania, Andreas Herkersdorf, Jörg Henkel |
IEEE Trans. Computers | 4 |
| 2020 | A Lightweight Nonlinear Methodology to Accurately Model Multicore Processor PowerabstractMany power management algorithms demand accurate and fine-grained runtime estimations of dynamic core power. In the absence of fine-grained power sensors, model-based estimations are needed. Such power models commonly approximate the switching activity of logic gates using performance counters while assuming a linear performance counter/power relation at a fixed frequency and voltage. It has been shown that this relation cannot be captured accurately enough with purely linear models and that well-established nonlinear modeling techniques, e.g., polynomial modeling, easily overfit the underlying performance/power relations. Although neural-network-based modeling has shown to accurately capture nonlinear relations, it has a large training and inference overhead which is too high for fine-grained models on core-level and estimation rates in the range of 1-10 kHz. We propose a methodology for nonlinear transformation of specific performance counters to increase power modeling accuracy at constant frequency and voltage with a relatively low overhead for both model generation and run-time application over a linear model. Furthermore, we use least-angle regression (LARS) to determine a ranking of the performance counter inputs for use in linear and nonlinear modeling and show that the transformed performance counters are better suited for power modeling. The generated dynamic power model consisting of a nonlinear transformation block and a linear regression block reduces relative estimation error on average by 4% and in worst-case scenarios by 7% compared to state-of-the-art fine-grained linear power models. Compared to a state-of-the-art polynomial regression model our proposed approach reduces the relative estimation error by 10% in worst-case scenarios. Mark Sagi, Nguyen Anh Vu Doan, Martin Rapp, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | Self-aware Cyber-Physical SystemsabstractIn this article, we make the case for the new class of Self-aware Cyber-physical Systems. By bringing together the two established fields of cyber-physical systems and self-aware computing, we aim at creating systems with strongly increased yet managed autonomy, which is a main requirement for many emerging and future applications and technologies. Self-aware cyber-physical systems are situated in a physical environment and constrained in their resources, and they understand their own state and environment and, based on that understanding, are able to make decisions autonomously at runtime in a self-explanatory way. In an attempt to lay out a research agenda, we bring up and elaborate on five key challenges for future self-aware cyber-physical systems: (i) How can we build resource-sensitive yet self-aware systems? (ii) How to acknowledge situatedness and subjectivity? (iii) What are effective infrastructures for implementing self-awareness processes? (iv) How can we verify self-aware cyber-physical systems and, in particular, which guarantees can we give? (v) What novel development processes will be required to engineer self-aware cyber-physical systems? We review each of these challenges in some detail and emphasize that addressing all of them requires the system to make a comprehensive assessment of the situation and a continual introspection of its own state to sensibly balance diverse requirements, constraints, short-term and long-term objectives. Throughout, we draw on three examples of cyber-physical systems that may benefit from self-awareness: a multi-processor system-on-chip, a Mars rover, and an implanted insulin pump. These three very different systems nevertheless have similar characteristics: limited resources, complex unforeseeable environmental dynamics, high expectations on their reliability, and substantial levels of risk associated with malfunctioning. Using these examples, we discuss the potential role of self-awareness in both highly complex and rather more simple systems, and as a main conclusion we highlight the need for research on above listed topics. Kirstie L. Bellman, Christopher Landauer, Nikil Dutt, Lukas Esterle, Andreas Herkersdorf, Axel Jantsch, Nima Taherinejad, Peter R. Lewis 0001, Marco Platzner, Kalle Tammemäe |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2020 | Machine Learning Approaches for Efficient Design Space Exploration of Application-Specific NoCsabstractIn many Multi-Processor Systems-on-Chip (MPSoCs), traffic between cores is unbalanced. This motivates the use of an application-specific Network-on-Chip (NoC) that is customized and can provide a high performance at low cost in terms of power and area. However, finding an optimized application-specific NoC architecture is a challenging task due to the huge design space. This article proposes to apply machine learning approaches for this task. Using graph rewriting, the NoC Design Space Exploration (DSE) is modelled as a Markov Decision Process (MDP). Monte Carlo Tree Search (MCTS), a technique from reinforcement learning, is used as search heuristic. Our experimental results show that—with the same cost function and exploration budget—MCTS finds superior NoC architectures compared to Simulated Annealing (SA) and a Genetic Algorithm (GA). However, the NoC DSE process suffers from the high computation time due to expensive cycle-accurate SystemC simulations for latency estimation. This article therefore additionally proposes to replace latency simulation by fast latency estimation using a Recurrent Neural Network (RNN). The designed RNN is sufficiently general for latency estimation on arbitrary NoC architectures. Our experiments show that compared to SystemC simulation, the RNN-based latency estimation offers a similar speed-up as the widely used Queuing Theory (QT). Yet, in terms of estimation accuracy and fidelity, the RNN is superior to QT, especially for high-traffic scenarios. When replacing SystemC simulations with the RNN estimation, the obtained solution quality decreases only slightly, whereas it suffers significantly when QT is used. Marcel Mettler, Daniel Mueller-Gritschneder, Thomas Wild, Andreas Herkersdorf, Ulf Schlichtmann |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2020 | Combinatorial Auctions for Temperature-Constrained Resource Management in ManycoresabstractAlthough manycore processors have plenty of cores, not all of them may run simultaneously at full speed and even some of them might need to be power-gated in order to keep the chip within safe temperature limits. Hence, a resource management technique, that allocates cores to application aiming at maximizing the system performance, will not be able to achieve its goal without taking into account the on-chip temperature and its impact on the availability of the chip's resources. However, considering a temperature constraint by the resource management will further increase its complexity, especially in manycores, and thus implementing it in a centralized scheme might lead to a computation bottleneck and a single point of failure. To avoid such scenarios, it is inevitable to distribute the computation required by the resource management technique throughout the chip. In this article, we propose a distributed resource management technique that considers temperature as an essential factor in allocating cores to applications and determining the power states of these cores and their voltage/frequency levels, while taking into account the performance models of the applications in order to maximize the overall system performance under a temperature constraint. Our proposed technique employs, for the first time, combinatorial auctions within an agent system to achieve the targeted goal in a distributed manner. The experimental evaluations show that our proposed technique achieves significant performance improvements with an average of 41% compared to several distributed resource management techniques. Heba Khdr, Muhammad Shafique 0001, Santiago Pagani, Andreas Herkersdorf, Jörg Henkel |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Cryptographic Hashing in P4 Data PlanesabstractP4 introduces a standardized, universal way for data plane programming. Secure and resilient communication typically involves the processing of payload data and specialized cryptographic hash functions. We observe that current P4 targets lack the support for both. Therefore, applications and protocols, which require message authentication codes or hashing structures that are resilient against attacks such as denial-of-service, cannot be implemented. To enable authentication and resilience, we make the case for extending P4 targets with cryptographic hash functions. We propose an extension of the P4 Portable Switch Architecture for cryptographic hashes and discuss our prototype implementations for three different P4 target platforms: CPU, NPU, and FPGA. To assess the practical applicability, we conduct a performance evaluation and analyze the resource consumption. Our prototype implementations show that cryptographic hashing can be integrated efficiently. We cannot identify a single hash function delivering satisfying performance on all investigated platforms. Therefore, we recommend a set of hash functions to optimize target-specific performance. Dominik Scholz, Andreas Oeldemann, Fabien Geyer, Sebastian Gallenmüller, Henning Stubbe, Thomas Wild, Andreas Herkersdorf, Georg Carle |
ANCS | 7 |
| 2019 | SOSA: Self-Optimizing Learning with Self-Adaptive Control for Hierarchical System-on-Chip ManagementabstractResource management strategies for many-core systems dictate the sharing of resources among applications such as power, processing cores, and memory bandwidth in order to achieve system goals. System goals require consideration of both system constraints (e.g., power envelope) and user demands (e.g., response time, energy-efficiency). Existing approaches use heuristics, control theory, and machine learning for resource management. They all depend on static system models, requiring a priori knowledge of system dynamics, and are therefore too rigid to adapt to emerging workloads or changing system dynamics. Bryan Donyanavard, Tiago Rogério Mück, Amir-Mohammad Rahmani, Nikil Dutt, Armin Sadighi, Florian Maurer 0003, Andreas Herkersdorf |
MICRO | 7 |
| 2019 | Channel mapping strategies for effective protection switching in fail-operational hard real-time NoCsabstractWith Multi Processor System-on-Chips (MPSoC) scaling up to thousands of processing elements, bus-based solutions have been dropped in favor of Network-on-Chips (NoC) as proposed in [2]. However, MPSoCs are yet hesitantly adopted in safety-critical fields, mainly due to the difficulty of ensuring strict isolation between different applications running on a single MPSoC as well as providing communication with Guaranteed Service (GS) to critical applications. This is particularly difficult in the NoC as it constitutes a network of shared resources. Moreover, safety-critical applications require some degree of Fault-Tolerance (FT) to guarantee safe operation at all times. Max Koenen, Nguyen Anh Vu Doan, Thomas Wild, Andreas Herkersdorf |
NOCS | 4 |
| 2019 | APEC: improved acknowledgement prioritization through erasure coding in bufferless NoCsabstractBufferless NoCs have been proposed as they come with a decreased silicon area footprint and a reduced power consumption, when compared to buffered NoCs. However, while known for their inherent simplicity, they suffer from early saturation and depend on additional measures to ensure reliable packet delivery, such as control protocols based on ACKs or NACKs. In this paper, we propose APEC, a novel concept for bufferless NoCs that allows to prioritize ACKs and NACKs over single payload flits of colliding packets by discarding the latter. Lightweight heuristic erasure codes are used to compensate for discarded payload flits. By trading off the erasure code overhead for packet retransmissions, a more efficient network operation is achieved. For ACK-based networks, APEC saturates at 2.1x and 2.875x higher generation rates than a conventional ACK-based bufferless NoC for packets between 5 and 17 flits. For NACK-based networks, APEC does not require concepts such as deflection routing or circuit-switched overlay NACK-networks, as prior work does. Therefore, it can simplify the network implementation compared to prior work while achieving similar performance. Michael Vonbun, Adrian Schiechel, Nguyen Anh Vu Doan, Thomas Wild, Andreas Herkersdorf |
NOCS | 5 |
| 2018 | BiSME: A Hardware Coprocessor to Perform Signature Matching at Multi-Gigabit RatesabstractHardware acceleration of signature matching is essential to perform content aware networking at predictable rates in modern network processors. Existing hardware accelerators either cannot perform signature matching at predictable rates due to the storage organization of the signatures or do not compress the signatures effectively resulting in inefficient on-chip memory usage. Addressing these problems, a bitmap based signature matching engine called BiSME is proposed in this paper, which is a flexible, programmable and scalable hardware coprocessor to perform signature matching at fixed, but guaranteed rates. The storage architectures proposed as part of BiSME, allows to efficiently store the compressed signatures in a flexible and programmable manner in on-chip memories. Each BiSME instance is fine-tuned to perform signature matching at 9.3 Gbps, with multiple instances capable of supporting increasing signature counts as well as increasing throughput. The BiSME was synthesized on a commercial 28nm technology library and only occupies 1.43 mm2of silicon area and consumes 155mW of power. The BiSME hardware implementation was thoroughly verified on the Cadence Palladium platform. Over 2GB of network traffic was injected simultaneously into BiSME and a software based signature matching solution and the identical signature matching results further validated the correctness of the design. Subramanian Shiva Shankar, Pinxing Lin, Andreas Herkersdorf, Thomas Wild |
ASAP | 3 |
| 2018 | Design methodologies for enabling self-awareness in autonomous systemsabstractThis paper deals with challenges and possible solutions for incorporating self-awareness principles in EDA design flows for autonomous systems. We present a holistic approach that enables self-awareness across the software/hardware stack, from systems-on-chip to systems-of-systems (autonomous car) contexts. We use the Information Processing Factory (IPF) metaphor as an exemplar to show how self-awareness can be achieved across multiple abstraction levels, and discuss new research challenges. The IPF approach represents a paradigm shift in platform design by envisioning the move towards a consequent platform-centric design in which the combination of self-organizing learning and formal reactive methods guarantee the applicability of such cyber-physical systems in safety-critical and high-availability applications. Armin Sadighi, Bryan Donyanavard, Thawra Kadeed, Kasra Moazzemi, Tiago Rogério Mück, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Thomas Wild, Nikil Dutt, Rolf Ernst, Andreas Herkersdorf, Fadi J. Kurdahi |
DATE | 11 |
| 2018 | FlueNT10G: A Programmable FPGA-based Network Tester for Multi-10-Gigabit EthernetabstractWe present FlueNT10G, an open-source FPGA-based network tester for precise replay of network traces, as well as for accurate packet capture and round-trip latency measurements. FlueNT10G streams replay and capture data between the host system and the FPGA board during active tests. It enables continuous measurements without being constrained by the memory capacity of the FPGA board. FlueNT10G is able to concurrently replay and capture traffic on three 10 Gbit/s network interfaces for all packet sizes. When operated exclusively in replay or capture mode, throughput increases to 4x 10 Gbit/s. Our design yields a temporal resolution of 6.4 ns for precise traffic pattern generation, as well as for accurate arrival timestamping and latency measurements. On the software-side, FlueNT10G is complemented by an API enabling the programmable execution of reproducible network measurements. Targeting the automated performance evaluation of different virtualized network function configurations, the API further integrates access to a bidirectional side-band channel for device-under-test reconfiguration and status feedback. FlueNT10G has been implemented on the NetFPGA-SUME platform (Xilinx Virtex-7 XC7VX690T) with an FPGA resource utilization of no more than 25%, which leaves sufficient capacity available for future design extensions. Andreas Oeldemann, Thomas Wild, Andreas Herkersdorf |
FPL | 3 |
| 2018 | Platform-Centric Self-Awareness as a Key Enabler for Controlling Changes in CPSabstractFuture cyber-physical systems will host a large number of coexisting distributed applications on hardware platforms with thousands to millions of networked components communicating over open networks. These applications and networks are subject to continuous change. The current separation of design process and operation in the field will be superseded by a life-long design process of adaptation, infield integration, and update. Continuous change and evolution, application interference, environment dynamics and uncertainty lead to complex effects which must be controlled to serve a growing set of platform and application needs. Self-adaptation based on self-awareness and self-configuration has been proposed as a basis for such a continuous in-field process. Research is needed to develop automated in-field design methods and tools with the required safety, availability, and security guarantees. The paper shows two complementary use cases of self-awareness in architectures, methods, and tools for cyber-physical systems. The first use case focuses on safety and availability guarantees in self-aware vehicle platforms. It combines contracting mechanisms, tool based self-analysis and self-configuration. A software architecture and a runtime environment executing these tools and mechanisms autonomously are presented including aspects of self-protection against failures and security threats. The second use case addresses variability and long term evolution in networked MPSoC integrating hardware and software mechanisms of surveillance, monitoring, and continuous adaptation. The approach resembles the logistics and operation principles of manufacturing plants which gave rise to the metaphoric term of an Information Processing Factory that relies on incremental changes and feedback control. Both use cases are investigated by larger research groups. Despite their different approaches, both use cases face similar design and design automation challenges which will be summarized in the end. We will argue that seemingly unrelated research challenges, such as in machine learning and security, could also profit from the methods and superior modeling capabilities of self-aware systems. Mischa Möstl, Johannes Schlatow, Rolf Ernst, Nikil Dutt, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Fadi J. Kurdahi, Thomas Wild, Armin Sadighi, Andreas Herkersdorf |
Proc. IEEE | 10 |
| 2017 | A non-intrusive, operating system independent spinlock profiler for embedded multicore systemsabstractLocks are widely used as a synchronization method to guarantee the mutual exclusion for accesses to shared resources in multi-core embedded systems. They have been studied for years to improve performance, fairness, predictability etc. and a variety of lock implementations optimized for different scenarios have been proposed. In practice, applying an appropriate lock type to a specific scenario is usually based on the developer's hypothesis, which could mismatch the actual situation. A wrong lock type applied may result in lower performance and unfairness. Thus, a lock profiling tool is needed to increase the system transparency and guarantee the proper lock usage. In this paper, an operating-system-independent lock profiling approach is proposed as there are many different operating systems in the embedded field. This approach detects lock acquisition and lock releasing using hardware tracing based on hardware-level spinlock characteristics instead of specific libraries or APIs. The spinlocks are identified automatically; lock profiling statistics can be measured and performance-harmful lock behaviors are detected. With this information, the lock usage can be improved by the software developer. A prototype as a Java tool was implemented to conduct hardware tracing and analyze locks inside applications running on the Infineon AURIX microcontrollers. Lin Li 0046, Philipp Wagner 0001, Albrecht Mayer, Thomas Wild, Andreas Herkersdorf |
DATE | 5 |
| 2017 | Self-awareness in autonomous automotive systemsabstractSelf-awareness has been used in many research fields in order to add autonomy to computing systems. In automotive systems, we face several system layers that must be enriched with self-awareness to build truly autonomous vehicles. This includes functional aspects like autonomous driving itself, its integration on the hardware/software platform, and among others dependability, real-time, and security aspects. However, self-awareness mechanisms of all layers must be considered in combination in order to build a coherent vehicle self-awareness that does not cause conflicting decisions or even catastrophic effects. In this paper, we summarize current approaches for establishing self-awareness on those layers and elaborate why self-awareness needs to be addressed as a cross-layer problem, which we illustrate by practical examples. Johannes Schlatow, Mischa Möstl, Rolf Ernst, Marcus Nolte, Inga Jatzkowski, Markus Maurer, Christian Herber, Andreas Herkersdorf |
DATE | 8 |
| 2017 | Adaptive Reliability for Fault Tolerant Multicore SystemsabstractIn an era of continuously shrinking technology and escalating power density, Multiprocessor System on Chips (MPSoCs) suffer from a growing prominence of device defects and increase of dependability-related issues. This paper tackles the dependability challenge by suggesting an adaptive reliability enhancement strategy for multicore systems. We dynamically adapt the reliability enhancement to the actual tasks requirements as well as cores runtime operating conditions. As reliability improvement may adversely affect the parameters of embedded systems, we suggest a runtime recovery method. In fact, we implement a 3-mode mapping technique to limit redundancy overheads through judicious task migrating and dropping. Our experiments show promising results in terms of error mitigation with controllable power and thermal overheads. Ihsen Alouani, Thomas Wild, Andreas Herkersdorf, Smaïl Niar |
DSD | 3 |
| 2017 | A Divide and Conquer State Grouping Method for Bitmap Based Transition CompressionabstractMember State Bitmask Technique (MSBT) is a hardware oriented transition compression technique which can compress the redundant transitions in a finite automaton. While the compressed automaton is stored in on-chip memories; a dedicated hardware accelerator performs signature matching by comparing the network streams against the compressed automaton at line rate. The MSBT consists of three functional steps which include the intra-state transition compression, state grouping and the inter-state transition compression. The state grouping algorithm which is currently used in MSBT is not compression aware and results in sub-optimal transition compression. To address this weakness, a compression aware Divide and Conquer state grouping method is proposed in this paper, which can efficiently group states that improves the transition compression in MSBT. Experimental evaluation of the proposed state grouping method, results in a reduced on-chip memory usage of the order of 10-30%. The reduction in the memory usage allows to accommodate more signatures in on-chip memories and perform signature matching with them at line rate. Subramanian Shiva Shankar, Pinxing Lin, Andreas Herkersdorf, Thomas Wild |
PDCAT | 3 |
| 2017 | DiaSys: Improving SoC insight through on-chip diagnosis
Philipp Wagner 0001, Thomas Wild, Andreas Herkersdorf |
J. Syst. Archit. | 3 |
| 2017 | Efficient task spawning for shared memory and message passing in many-core architectures
Aurang Zaib, Thomas Wild, Andreas Herkersdorf, Jan Heisswolf, Jürgen Becker 0001, Andreas Weichslgartner, Jürgen Teich |
J. Syst. Archit. | 3 |
| 2016 | Disaggregated FPGAs: Network Performance Comparison against Bare-Metal Servers, Virtual Machines and Linux ContainersabstractFPGAs (Field Programmable Gate Arrays) are making their way into data centers (DC). They are used as accelerators to boost the compute power of individual server nodes and to improve the overall power efficiency. Meanwhile, DC infrastructures are being redesigned to pack ever more compute capacity into the same volume and power envelops. This redesign leads to the disaggregation of the server and its resources into a collection of standalone computing, memory, and storage modules. To embrace this evolution, we propose an architecture that decouples the FPGA from the CPU of the server by connecting the FPGA directly to the DC network. This proposal turns the FPGA into a network-attached computing resource that can be incorporated with disaggregated servers into these emerging data centers. We implemented a prototype and compared its network performance with that obtained from bare metal servers (Native), virtual machines (VM), and containers (CT). The results show that standalone network-attached FPGAs outperform them in terms of network latency and throughput by a factor of up to 35x and 73x, respectively. We also observed that the proposed architecture consumes only 14% of the total FPGA resources. Jagath Weerasinghe, François Abel, Christoph Hagleitner, Andreas Herkersdorf |
CloudCom | 4 |
| 2016 | Resolving Performance Interference in SR-IOV Setups with PCIe Quality-of-Service ExtensionsabstractPCI Express (PCIe) Single Root I/O Virtualization (SR-IOV) enables low latency and high performance virtualization of I/O devices. It has been embraced in cloud computing and is considered a promising foundation for sharing I/O in future multi-core embedded and mixed-criticality systems. Unfortunately, SR-IOV is vulnerable to Denial-of-Service (DoS) attacks, which cause performance interference. For cloud computing, an approach that mitigates ongoing attacks via software scheduling has been proposed. However, for embedded and mixed-criticality systems, solutions that go beyond mitigation are preferred. In this paper, we propose two integrated hardware architectures that completely prevent DoS attacks. As a foundation, we utilize optional Quality-of-Service (QoS) extensions from the PCIe specification. We determine which QoS extensions are needed, and show how virtualized multi-core CPUs need to implement and interface them (an aspect explicitly not covered in the PCIe specification) to enable DoS protection. The two proposed architectures are optimized for different goals, scheduling freedom or minimal hardware costs. As PCIe QoS is absent from current hardware, we evaluate our architectures with a QoS-enabled SystemC model of a real-world lab-setup. Results show that both architectures successfully prevent DoS attacks. To the best of our knowledge, we are the first to explore and evaluate feasibility of PCIe QoS for SR-IOV DoS prevention. Andre Oliver Richter, Christian Herber, Thomas Wild, Andreas Herkersdorf |
DSD | 4 |
| 2016 | Linux apps-usage-driven power dissipation-aware schedulerabstractIn modern symmetrical chip multiprocessor (CMP) architecture, problems in cache coherence, context switch overheads and serialized code bottleneck are major causes of excessive computing power dissipation in the application of simultaneous multithreading (SMT) technique. This research models and manages above-mentioned problems based on user application usage patterns identified in a mobile computing platform. A novel scheduler has been developed to realize power management schemes based on the Linux kernel (version. 3.0.1) and deployed in Android 4.0 ICS. The scheduler monitors multiple system performance metrics and predicts power dissipation based on the historical user application usage values as well as the content of the scheduler run queue. The length of the time slices and the variables of process control blocks are adjusted to optimize power dissipation according to the prediction. The proposed scheduler module has achieved a power dissipation reduction of 13 to 24% in a GEM5 simulated environment. Hou Zhao Qi Rex, Ching-Chuen Jong, Andreas Herkersdorf |
ISCAS | 3 |
| 2016 | A Rule-based Methodology for Hardware Configuration Validation in Embedded SystemsabstractAs the complexity of multicore SoCs increases, more potential system issues are arising. Hardware-related configuration issues are becoming more complicated owing to the introduction of more cores and various complex peripherals. Considering the complexity of multicore programming, consultation of the main source of guidance, i.e. the user manual, is not an efficient approach to identify such problems. Improper hardware-related configurations could lead to either functional or performance issues. Some of these issues are even subtle and hard to detect. Therefore, a rule-based validation methodology is proposed to deal with hardware-related configuration issues in an efficient and reliable way. Hardware trace is applied in this methodology to detect issues even before symptoms appear. The method directly observes the register accesses and detects bugs based on trace data. It is independent of the application as long as they are run on the given platform, which means the same method implementation could be applied to any applications on the same platform. In this paper, an initial proof-of-concept for the proposed methodology has been implemented and demonstrated on the Infineon TC29 device. Lin Li 0046, Philipp Wagner 0001, Ramesh Ramaswamy, Albrecht Mayer, Thomas Wild, Andreas Herkersdorf |
SCOPES | 6 |
| 2015 | A Hardware/Software Approach for Mitigating Performance Interference Effects in Virtualized Environments Using SR-IOVabstractSingle Root I/O Virtualization (SR-IOV) is an extension to the PCI Express (PCIe) standard that allows virtual machines (VMs) to directly access shared I/O devices without host involvement. This enabled SR-IOV to become the best-performing solution for virtual I/O to date, which lead to its commercial adoption, e.g., In the Amazon EC2. On the downside, a malicious VM can exploit the direct access to an SR-IOV device by flooding it with PCIe packets. This results in a congestion on the PCIe interconnect, which leads to performance interference effects between the malicious VM, concurrent VMs and even the host. In this paper, we present a hardware/software approach that detects and mitigates such Denial-of-Service (DoS) attacks. On the hardware side, we propose monitoring extensions within SR-IOV devices that distinguish legal device use from malicious device use by observing the rate of incoming PCIe transactions at VM granularity. Malicious VMs are reported to the host via interrupts. On the software side, performance interference effects can then be mitigated by dynamically adjusting the host's scheduling of the malicious VM or even shutting it down. We implement a prototype with a commercial off-the-shelf SR-IOV Ethernet controller and an FPGA board. On it, we demonstrate that appropriate scheduling of malicious VMs successfully mitigates interference effects for three cloud-relevant benchmarks. For example, Memcached is restored to 99.4% of baseline performance (compared to 61.8% without our extensions). In contrast to QoS features proposed in the PCIe 3.0 standard, our solution is more flexible. Additionally, it can be realized as an add-on to existing misuse detection hardware like the Intel Malicious Driver Detection (MDD). Andre Oliver Richter, Christian Herber, Stefan Wallentowitz, Thomas Wild, Andreas Herkersdorf |
CLOUD | 5 |
| 2015 | Real-time capable CAN to AVB ethernet gateway using frame aggregation and scheduling
Christian Herber, Andre Oliver Richter, Thomas Wild, Andreas Herkersdorf |
DATE | 4 |
| 2015 | MPIOV: scaling hardware-based I/O virtualization for mixed-criticality embedded real-time systems using non transparent bridges to (multi-core) multi-processor systems
Daniel Münch, Michael Paulitsch, Oliver Hanka, Andreas Herkersdorf |
DATE | 4 |
| 2015 | An Analytic Approach on End-to-End Packet Error Rate Estimation for Network-on-ChipabstractNetwork-on-Chip (NoC) are well-established for scalable on-chip communication, but technology generations of 22~nm and below, as well as aggressive voltage scaling to reduce NoC power consumption, introduce new variability challenges resulting in errors on wires and registers. Based on the probabilities of single bit flips, this paper focuses on the expected end-to-end packet error probabilities in NoC. We investigate the influence of individual bit error probabilities, the number of hops between communication partners, as well as the packet size. To evaluate these parameters, we propose an analytic approach which abstracts technology details of NoC data transport entities, such as links and buffers, and models each entity as a binary symmetric channel (BSC). The proposed probabilistic approach obtains equations for system-level NoC reliability estimates which allow an evaluation without the necessity to deploy time-consuming simulations. Michael Vonbun, Stefan Wallentowitz, Andreas Oeldemann, Andreas Herkersdorf |
DSD | 4 |
| 2015 | A hardware-based multi-objective thread mapper for tiled manycore architecturesabstractThread mapping is typically performed as an integral part of cooperative or pre-emptive operating system (OS) scheduling in order to share the processor core(s) among competing applications. Schedulers usually follow a single-objective performance optimization, such as maximizing core utilization or satisfying deadlines by the prioritization of threads. Meeting multiple orthogonal objectives, like performance vs. power or thermal resilience, in the era of manycore processors is a challenge because of the associated scalability and thread management overhead. We tackle these challenges by employing a two stage thread management strategy. In the first stage (not covered in this short paper), threads are assigned to regions or compute tiles. For the second stage we introduce in this paper the TCU (Thread Control Unit), a configurable, low latency, low overhead hardware thread mapper that takes various runtime sensor parameters into account. It can map threads within a small and bounded number of clock cycles in round robin, single or multi-objective manner. TCU is designed to consider not just load balancing or performance criteria but also physical constraints like power budgets, temperature limits and reliability aspects. TCU macro achieves 150K thread mappings per second on a tiled MPSoC FPGA prototype while operating at moderate 50 Mz. Evaluations of different mapping policies show that multi-objective thread mapping provides about 10 to 40% less mapping latency for periodic and bursty traffic compared to single-objective or round robin schemes. FPGA and ASIC syntheses reveal a 9% hardware overhead for the TCU on a four core compute tile. Ravi Kumar Pujari, Thomas Wild, Andreas Herkersdorf |
ICCD | 3 |
| 2015 | Introduction to the Special Issue on Testing, prototyping, and debugging of multi-core architectures
Frank Hannig, Andreas Herkersdorf |
J. Syst. Archit. | 2 |
| 2015 | Denial-of-Service attacks on PCI passthrough devices: Demonstrating the impact on network- and storage-I/O performance
Andre Oliver Richter, Christian Herber, Thomas Wild, Andreas Herkersdorf |
J. Syst. Archit. | 4 |
| 2014 | CAP: Communication Aware ProgrammingabstractNetworks on Chip (NoC) come along with increased complexity from the implementation and management perspective. This leads to higher energy consumption and programming complexity of NoC architectures. Jan Heisswolf, Aurang Zaib, Andreas Zwinkau, Sebastian Kobbe, Andreas Weichslgartner, Jürgen Teich, Jörg Henkel, Gregor Snelting, Andreas Herkersdorf, Jürgen Becker 0001 |
DAC | 9 |
| 2014 | Distributed cooperative shared last-level caching in tiled multiprocessor system on chipabstractIn a shared-memory based tiled many-core system-on-chip architecture, memory accesses present a huge performance bottleneck in terms of access latency as well as bandwidth requirements. The best practice approach to address this issue is to provide a multi-level cache hierarchy and a suitable cache-coherency mechanism. This paper presents a method to increase the memory access performance in distributed-directory-coherency-protocol based tiled many-core systems. The proposed method introduces an alternate design for the system-wide shared last-level caches (LLC) placed between the memory and the node private caches (NPC). The proposed system-wide shared LLC layer is distributed over the entire network and it interacts with the home directories of specific cache lines. Results from simulating SPEC2000 benchmark applications executed on a SystemC model of the proposed design show a minimum performance improvement of 20-25% when compared to a model without the shared cache layer at the expense of an additional 2% of the total cache memory space (NPC + LLC memory). In addition, the proposed design shows a minimum 7-15% and an average 14-15% improvement in performance in comparison to centralized system-wide shared LLC of equivalent size and dynamic mapped distributed LLC of equivalent size respectively. Preethi P. Damodaran, Stefan Wallentowitz, Andreas Herkersdorf |
DATE | 3 |
| 2014 | System integration - The bridge between More than Moore and More MooreabstractSystem Integration using 3D technology is a very promising way to cope with current and future requirements for electronic systems. Since the pure shrinking of devices (known as “More Moore”) will come to an end due to physical and economic restrictions, the integration of systems (e.g. by stacking dies, or by adding sensor functions) shows a way to maintain the growth in complexity as well as in diversity which is necessary for future applications. This so called “More than Moore” approach complements the conventional SoC product engineering. This paper gives insights in System Integration design challenges from different perspectives, ranging from design technology over MEMS product engineering and 3D interconnect to automotive cyber physical systems. Andy Heinig, Manfred Dietrich, Andreas Herkersdorf, Felix Miller, Thomas Wild, Kai Hahn, Armin Grünewald, Rainer Brück 0001, Steffen Krohnert, Jochen Reisinger |
DATE | 3 |
| 2014 | Hardware virtualization support for shared resources in mixed-criticality multicore systemsabstractElectric/Electronic architectures in modern automobiles evolve towards an hierarchical approach where functionalities from several ECUs are consolidated into few domain computers. Performance requirements directly lead to multicore solutions but also to a combination of very different requirements on such ECUs. Using virtualization in addition is one promising way of achieving segregation in time and space of shared resources. Based on examples taken from the automotive domain several concepts for efficient hardware extensions of coprocessors and I/O devices are shown in this contribution. These provide mechanisms to ensure quality of service (QoS) levels in terms of execution time, throughput and latency. The resulting infotainment architecture is a feasibility study and is integrated into a vehicle demonstrator as centralized infotainment platform (VCT). Oliver Sander, Timo Sandmann, Viet Vu Duy, Steffen Bähr, Falco Bapp, Jürgen Becker 0001, Hans-Ulrich Michel, Dirk Kaule, Daniel Adam, Enno Lübbers, Jürgen Hairbucher, Andre Oliver Richter, Christian Herber, Andreas Herkersdorf |
DATE | 14 |
| 2014 | Connecting different worlds - Technology abstraction for reliability-aware design and TestabstractThe rapid shrinking of device geometries in the nanometer regime requires new technology-aware design methodologies. These must be able to evaluate the resilience of the circuit throughout all System on Chip (SoC) abstraction levels. To successfully guide design decisions at the system level, reliability models, which abstract technology information, are required to identify those parts of the system where additional protection in the form of hardware or software coun-termeasures is most effective. Interfaces such as the presented Resilience Articulation Point (RAP) or the Reliability Interchange Information Format (RIIF) are required to enable EDA-assisted analysis and propagation of reliability information. The models are discussed from different perspectives, such as design and test. Ulf Schlichtmann, Veit Kleeberger, Jacob A. Abraham, Adrian Evans, Christina Gimmler-Dumont, Michael Glaß, Andreas Herkersdorf, Sani R. Nassif, Norbert Wehn |
DATE | 7 |
| 2014 | Iterative FPGA Implementation Easing Safety Certification for Mixed-Criticality Embedded Real-Time SystemsabstractThe design and operation of an aircraft, a railway, and a nuclear power station that include either safety-critical or safety-related systems require a proof that its safety is assured. The process providing this proof is called certification. This paper suggests an iterative FPGA implementation and iterative certification concept for FPGA-based systems to provide design-time adaptability while the complexity is still kept low to ease certification. The practical evaluation of this concept demonstrates that reuse at implementation level of a previously implemented part is to 100% usable for iterative certification. Regarding the resource utilization and complexity, the evaluation shows that there are potential savings in resource utilization and complexity compared to conventional run-time configurable designs. Iterative certification reduces the recertification of a whole design to a recertification of the changed part only and a verification tool qualification. It is shown that tool qualification can be accomplished with relatively moderate effort. Therefore, the presented concept substantially eases the certification process when using modular design and building block reuse. Daniel Münch, Michael Paulitsch, Michael Honold, Wolfgang Schlecker, Andreas Herkersdorf |
DSD | 5 |
| 2014 | Dependable task and communication migration in tiled manycore system-on-chipabstractPower densities and thermal hotspots are a major concern for the dependability of future multi-processor systemon- chip. They can lead to transient faults affecting the functionality in the short term and can cause permanent damage of a device. The dependability problem can be tackled on different layers such as technology hardening or application awareness. This work is based on an approach that addresses the issue for tile-based manycore system-on-chip on software and architecture layer. An agent-based system management employs task migration to react to thermal hotspots and pro-actively avoid them. The inter-task communication plays an important role as communication channels need to be migrated accordingly. The presented work focuses on the issue of communication migration and is based on the idea of handling it transparently to the task migration. Network-on-chip protection switching techniques have been introduced before and in this paper we evaluate the potential and bottlenecks of such methods in a realistic platform. Stefan Wallentowitz, Stefan Rosch, Thomas Wild, Andreas Herkersdorf, Volker Wenzel, Jörg Henkel |
FDL | 4 |
| 2014 | A network virtualization approach for performance isolation in controller area network (CAN)abstractAn important trend in automotive CPS is the shift from federated to integrated IT architectures, where multiple functions are consolidated on shared electronic resources instead of distributed electronic control units (ECUs). It is driven by increasing complexity, cost and installation space requirements of todays architectures. However, side-by-side integration of mixed-criticality functions poses new challenges with respect to safety and security. To achieve isolated performance for multiple integrated partitions with different criticalities, an efficient separation within computing and communication resources is required. This paper introduces a network virtualization approach for CAN, which enables concurrence of mixed-criticality communication on a single physical CAN bus through a strict performance isolation. We present a design concept as well as a prototypical implementation. The feasibility of our approach is demonstrated by an analytic evaluation of message latencies and through experimental case studies. Christian Herber, Andre Oliver Richter, Thomas Wild, Andreas Herkersdorf |
RTAS | 4 |
| 2013 | AUTO-GS: Self-Optimization of NoC Traffic through Hardware Managed Virtual ConnectionsabstractNetworks-on-Chip have shown their scalability for future many-core systems on chip. In real world scenarios, where multiple applications are being executed over a shared NoC based platform, efficient utilization of Networks-on-Chip resources becomes challenging. Methodologies are required to ensure better utilization of NoC, especially in the scenarios, where the communication patterns of NoC traffic are difficult to predict before run-time. In this paper, we propose a self-optimization mechanism which detects frequent communication by monitoring communication patterns at run-time and uses this information to establish virtual connections autonomously. Communication monitoring and connection establishment are realized in hardware. Hardware managed virtual connections lead to better utilization of NoC resources and reduce the communication latencies suffered by applications. In addition, energy consumption by the communication infrastructure is reduced. The proposed concept is investigated through simulation of real world application scenarios. The simulation results highlight the performance improvement and synthesis results show the low area overhead of the proposed hardware implementation. Aurang Zaib, Jan Heisswolf, Andreas Weichslgartner, Thomas Wild, Jürgen Teich, Jürgen Becker 0001, Andreas Herkersdorf |
DSD | 7 |
| 2013 | A Design Space Exploration Framework For Automotive Embedded Systems And Their Power ManagementabstractThe E/E (electric/electronic) architecture of a modern vehicle is a complex distributed system, where up to 80 electronic control units (ECUs), interconnected by several communication buses, need to collaborate with each other in order to implement the various comfort and safety features. The presented E/E design space exploration framework supports engineers during the development process of new E/E architectures by providing a graphical modeling and simulation environment with an interactive visualization of simulation-based results. Furthermore, it contains an advanced business logic which administrates the modeling, storing, retrieving and cloning of evaluation sessions consisting of complex experiments. Due to a high-level modeling approach, future architectures can be evaluated in respect to power consumption and performance values already in an early stage of the design process and design alternatives can be easily compared with each other. Gregor Walla, Zaur Molotnikov, Hans-Ulrich Michel, Walter Stechele, Andreas Barthels, Andreas Herkersdorf |
ECMS | 6 |
| 2013 | Introduction to the special section on multiprocessor system-on-chip for cyber-physical systemsabstractNo abstract available. Michael Hübner 0001, Andreas Herkersdorf |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | Virtual networks - distributed communication resource management
Jan Heisswolf, Aurang Zaib, Andreas Weichslgartner, Ralf König 0001, Thomas Wild, Jürgen Teich, Andreas Herkersdorf, Jürgen Becker 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2012 | Invasive manycore architecturesabstractThis paper introduces a scalable hardware and software platform applicable for demonstrating the benefits of the invasive computing paradigm. The hardware architecture consists of a heterogeneous, tile-based manycore structure while the software architecture comprises a multi-agent management layer underpinned by distributed runtime and OS services. The necessity for invasive-specific hardware assist functions is analytically shown and their integration into the overall manycore environment is described. Jörg Henkel, Andreas Herkersdorf, Lars Bauer, Thomas Wild, Michael Hübner 0001, Ravi Kumar Pujari, Artjom Grudnitsky, Jan Heisswolf, Aurang Zaib, Benjamin Vogel, Vahid Lari, Sebastian Kobbe |
ASP-DAC | 2 |
| 2012 | Virtual platforms: Breaking new groundsabstractThe case for developing and using virtual platforms (VPs) has now been made. If developers of complex HW/SW systems are not using VPs for their current design, complexity of next generation designs demands for their adoption. In addition, the users of these complex systems are asking either for virtual or real platforms in order to develop and validate the software that runs on them, in context with the hardware that is used to deliver some of the functionality. Debugging the erroneous interactions of events and state in a modern platform when things go wrong is hard enough on a VP; on a real platform (such as an emulator or FPGA-based prototype) it can become impossible unless a new level of sophistication is offered. The priority now is to ensure that the capabilities of these platforms meet the requirements of every application domain for electronics and software-based product design. And to ensure that all the use cases are satisfied. A key requirement is to keep pace with Moore's Law and the ever increasing embedded SW complexity by providing novel simulation technologies in every product release. This paper summarizes a special session focused on the latest applications and latest use cases for VPs. It gives an overview of where this technology is going and the impact on complex system design and verification. Rainer Leupers, Grant Martin, Roman Plyaskin, Andreas Herkersdorf, Frank Schirrmeister, Tim Kogel, Martin Vaupel |
DATE | 4 |
| 2012 | A low-overhead monitoring ring interconnect for MPSoC parameter optimizationabstractMPSoCs need to integrate self-x properties in order to get rid of the worst-case design style which is no longer affordable in large SoCs. Integrating self-x properties in SoCs is possible through a monitoring interconnect which carries monitor information to evaluators that decide on actions that will tune the SoC operation mode. We have designed a customized interconnect for SoC monitoring/actuation. We have implemented it in VHDL and tested it in FPGA. The prototype proved that this customized interconnect provides good results regarding latency and area overheads and is a key component in enabling self-optimization in our FPGA MPSoC prototype. Abdelmajid Bouajila, Abdallah Lakhtel, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf |
DDECS | 5 |
| 2012 | Analytical Design Space Exploration Based on Statistically Refined Runtime and Logic Estimation for Software Defined RadiosabstractThe exploration of the design space for complex hardware-software systems requires accurate models for the system components, which are often not available in early design phases, resulting in error-prone resource estimations. For a HWSW system with a finite set of design points, we present an analytical approach to evaluate the quality of a distinctive design point choice. Our approach enables the designer to gain a measure for statistical confidence whether an application with realtime requirements can be successfully implemented on a chosen set of processors and reconfigurable logic. By a statistical evaluation of runtime, latency, logic resources and memory requirements, a probability metric for each realization alternative in the system is derived, that gives a realization probability for different mappings and different combinations of chips. We apply our principles to an FPGA/DSP digital radio receiver system and evaluate the realization probabilities for a different combination of chip sizes and mappings. Finally, we compare our approach against conventional estimation techniques, such as worst-case evaluation. Matthias Ihmig, Michael Feilen, Andreas Herkersdorf |
DSD | 3 |
| 2012 | TSV-virtualization for Multi-protocol-Interconnect in 3D-ICsabstractThrough Silicon Vias (TSVs) are the method of choice to realize vertical connections between different chip layers in three dimensional Integrated Circuits (3D-ICs). These TSVs offer a fast connection and due to their short wire length, only a small capacitive load to the driving circuitry. On the other hand TSVs consume a relative large amount of chip area and as TSV-count increases the overall yield generally drops due to TSV manufacturing difficulties. As a result of the low capacitance, TSVs can be clocked much higher than conventional intra-layer links. To fully utilize the TSV-based vertical bandwidth we propose using them in a multiplexed manner and share them between several virtual links. On top of that we propose using TSVs to stretch state-of-the art interconnects like busses, crossbars or NoCs to other silicon layers in the 3D stack. This reduces TSV count and gives designers the opportunity to easily migrate from 2D to 3D designs and to largely benefit from reuse of existing IP blocks and interconnection schemes. Felix Miller, Thomas Wild, Andreas Herkersdorf |
DSD | 3 |
| 2012 | Dependable embedded systems: The German research foundation DFG priority program SPP 1500abstractWhen migrating to future technology nodes, dependability becomes a major design problem as variability, aging and susceptibility to soft errors increase. The purpose of this program is to research cross-layer solutions that address the physical problems at system-level i.e. at hardware-level, operating system level, application level etc. The goals and an overview of the DFG SPP 1500 research program are presented. Jörg Henkel, Oliver Bringmann 0001, Andreas Herkersdorf, Wolfgang Rosenstiel, Norbert Wehn |
ETS | 3 |
| 2012 | An integrated simulation framework for invasive computing
Michael Gerndt, Frank Hannig, Andreas Herkersdorf, Andreas Hollmann, Marcel Meyer, Sascha Roloff, Josef Weidendorfer, Thomas Wild, Aurang Zaib |
FDL | 3 |
| 2012 | A framework for Open Tiled Manycore System-On-ChipabstractTiled manycore architectures have become dominant for the integration of tens or even a hundred processor cores on a chip. While commercial products are increasingly available, research on the hardware of such platforms and especially prototyping often rely on building such a platform from scratch or is bound to abstract simulation. In this paper we present the Open Tiled Manycore System-on-Chip (Op-TiMSoC) which is a library-based tool flow that helps generating a tiled manycore platform based on a library of open standard components. OpTiMSoC allows for research and prototyping of both shared memory and distributed memory platforms. It includes LISNoC which is a flexible NoC implementation. An OpTiMSoC system can easily be generated based on the publicly available repository and prototyped on an FPGA. As exemplary targets we evaluated the usage of different FPGA boards and an emulation platform. Stefan Wallentowitz, Andreas Lankes, Aurang Zaib, Thomas Wild, Andreas Herkersdorf |
FPL | 5 |
| 2012 | Benefits of selective packet discard in networks-on-chipabstractToday, Network on Chip concepts principally assume inherent lossless operation. Considering that future nanometer CMOS technologies will witness increased sensitivity to all forms of manufacturing and environmental variations (e.g., IR drop, soft errors due to radiation, transient temperature induced timing problems, device aging), efforts to cope with data corruption or packet loss will be unavoidable. Possible counter measures against packet loss are the extension of flits with ECC or the introduction of error detection with retransmission. We propose to make use of the perceived deficiency of packet loss as a feature. By selectively discarding stuck packets in the NoC, a proven practice in computer networks, all types of deadlocks can be resolved. This is especially advantageous for solving the problem of message-dependent deadlocks, which otherwise leads to high costs either in terms of throughput or chip area. Strict ordering, the most popular approach to this problem, results in a significant buffer overhead and a more complex router architecture. In addition, we will show that eliminating local network congestions by selectively discarding individual packets also can improve the effective throughput of the network. The end-to-end retransmission mechanism required for the reliable communication, then also provides lossless communication for the cores. Andreas Lankes, Thomas Wild, Stefan Wallentowitz, Andreas Herkersdorf |
ACM Trans. Archit. Code Optim. | 4 |
| 2011 | An approach to improve accuracy of source-level TLMs of embedded softwareabstractVirtual Prototypes (VPs) based on Transaction Level Models (TLMs) have become a de-facto standard for design space exploration and validation of complex software-centric multicore or multiprocessor systems. The most popular method to get timed software TLMs is to annotate timing information at the basic-block level granularity back into application source code, called source code instrumentation (SCI). The existing SCI approaches realize the back-annotation of timing information based on mapping between source code and binary code. However, optimizing compilation has a large impact on the code mapping and will lower the accuracy of the generated source-level TLMs. In this paper, we present an efficient approach to tackle this problem. We propose to use mapping between source-level and binary-level control flows as the basis for timing annotation instead of code mapping. Software TLMs generated by our approach allow for accurate evaluation of multiprocessor systems at a very high speed. This has been proven by our experiments with a set of benchmark programs and a case study. Zhonglei Wang, Kun Lu 0005, Andreas Herkersdorf |
DATE | 3 |
| 2011 | An architecture and an FPGA prototype of a reliable processor pipeline towards multiple soft- and timing errorsabstractThis paper presents a reliable processor pipeline architecture resilient to multiple soft- and timing errors. It also presents a probabilistic quantification of its performance overheads. This reliable processor pipeline architecture has been implemented in the Leon3 VHDL open source processor. An FPGA prototype running under random fault injection has also been developed. This reliable processor pipeline has low performance overheads (relative CPI of 1.06 at an error injection rate of 3 %) and is therefore much better than techniques based on flushing. Abdelmajid Bouajila, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf |
DDECS | 4 |
| 2011 | Context-aware compiled simulation of out-of-order processor behavior based on atomic tracesabstractAdvanced out-of-order processors exhibit complex dynamic behavior. Therefore, they are difficult to model at abstraction levels higher than cycle-accurate instruction set simulators (ISS's). Conventional compiled simulation techniques have been widely used for fast performance estimation. However, they assume static time intervals between memory accesses and do not consider diverse behavior of out-of-order processors. In this paper, we introduce context-aware compiled simulation, in which the timing of basic blocks is defined dynamically, depending on the previously executed basic blocks. We extend binary-level compiled simulation correspondingly and show that consideration of contexts can significantly increase the accuracy of performance estimation. With the proposed technique, we could reduce the average error of timing estimation to 0,47% at average speedup of 45× compared to sim-outorder simulator from the SimpleScalar tool set. Roman Plyaskin, Andreas Herkersdorf |
VLSI-SoC | 2 |
| 2010 | A folded pipeline network processor architecture for 100 Gbit/s networksabstractEthernet, although initially conceived as a Local Area Network technology, has been steadily making inroads into access and core networks. This has led to a need for higher link speeds, which are now reaching 100 Gbit/s. Packet processing at this rate represents a significant challenge, that needs to be met efficiently, while minimizing power consumption and chip area. This level of throughput favours a pipelined approach, thus this paper takes a traditional pipeline and breaks it down to mini-pipelines, which can perform coarse-grained processing (like process an MPLS label to completion). These mini-pipelines are then parellelized and used to construct a folded pipeline architecture, which augments the traditional approach by significantly reducing power consumption, a key problem in future routers. The paper compares the two approaches, discusses their advantages and disadvantages and demonstrates by quantitative measures that the folded pipeline architecture is the better solution for 100 Gbit/s processing. Kimon Karras, Thomas Wild, Andreas Herkersdorf |
ANCS | 3 |
| 2010 | A rapid prototyping system for error-resilient multi-processor systems-on-chipabstractStatic and dynamic variations, which have negative impact on the reliability of microelectronic systems, increase with smaller CMOS technology. Thus, further downscaling is only profitable if the costs in terms of area, energy and delay for reliability keep within limits. Therefore, the traditional worst case design methodology will become infeasible. Future architectures have to be error resilient, i.e., the hardware architecture has to tolerate autonomously transient errors. In this paper, we present an FPGA based rapid prototyping system for multi-processor systems-on-chip composed of autonomous hardware units for error-resilient processing and interconnect. This platform allows the fast architectural exploration of various error protection techniques under different failure rates on the microarchitectural level while keeping track of the system behavior. We demonstrate its applicability on a concrete wireless communication system. Matthias May 0001, Norbert Wehn, Abdelmajid Bouajila, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf, Daniel Ziener, Jürgen Teich |
DATE | 6 |
| 2010 | Architectural Vulnerability Factor Estimation with Backwards AnalysisabstractSingle-Event-Upsets in synchronous register-based designs are a severe problem for safety-critical applications. Exact and detailed error rate estimations are needed to determine a system's level of reliability. Available methods for estimation consider only special effects, use special reliability models or are computationally intensive. We present an innovative method that is able to calculate the architectural vulnerability factor (AVF)of any RT-level circuit description by applying time-reversed stimulus values. This method, which we call Backwards Analysis, considers all major masking effects (logic masking, information lifetime, timing derating, transitive masking) in a single algorithm and delivers results in several levels of detail from average AVF through sensitivity waveforms. The results show the critical parts and states of a design, which could be used for reliability assessment and selective hardening of the circuit to reach a target failure rate. Robert Hartl, Andreas J. Rohatschek, Walter Stechele, Andreas Herkersdorf |
DSD | 4 |
| 2010 | An Application-Aware Load Balancing Strategy for Network Processors
Rainer Ohlendorf, Michael Meitinger, Thomas Wild, Andreas Herkersdorf |
HiPEAC | 4 |
| 2010 | Comparison of Deadlock Recovery and Avoidance Mechanisms to Approach Message Dependent Deadlocks in On-chip NetworksabstractWith the transition from buses to on-chip networks in SoCs the problem of deadlocks in on-chip interconnects arises. Deadlocks can be caused by routing cycles in the network, or by message dependencies, even if the network itself is actually free of routing cycles. Two basic approaches to counter message dependent deadlocks exist: deadlock avoidance, which is most popular in NoCs, and deadlock recovery, which has until now only been used in parallel computer networks. Deadlock recovery promises low buffer space requirements and does not impose restrictions on connections between individual communication partners. For this study, we have adapted a deadlock recovery scheme for the use in NoCs and compared it to strict ordering as a representative of deadlock avoidance in terms of throughput and buffer space. The results show significant buffer space savings for deadlock recovery, however, at the cost of reduced data throughput. Andreas Lankes, Thomas Wild, Andreas Herkersdorf, Sören Sonntag, Helmut Reinig |
NOCS | 3 |
| 2010 | High-level timing analysis of concurrent applications on MPSoC platforms using memory-aware trace-driven simulationsabstractDue to the growing complexity of multiprocessor systems-on-chip (MPSoCs), there is an increasing demand on efficient design space exploration techniques. In addition to the analysis of diverse hardware architectures, these techniques should assist the designer in the flexible evaluation of various scheduling policies and application mappings while taking effects of the shared on-chip communication infrastructure into account. Most available simulation approaches are either unable to cover all these aspects jointly or have poor simulation performance. In this paper, we present a framework for timing analysis of MPSoC architectures using abstract and yet accurate traces. The traces capture both precise processing latencies and memory access patterns and represent application- and OS-related workload. Performance estimation is performed by an interleaved execution of the traces on a highly configurable multiprocessor platform modeled in our trace-driven SystemC TLM simulator. Using the flexible scheduler model presented in this paper, various mappings and scheduling policies can be rapidly evaluated while considering on-chip interconnect contention and usage of shared resources. Due to the abstraction of the trace-driven simulations, the proposed framework allows for both fast and accurate explorations of MPSoC design alternatives. Roman Plyaskin, Alejandro Masrur, Martin Geier 0001, Samarjit Chakraborty, Andreas Herkersdorf |
VLSI-SoC | 5 |
| 2010 | Software performance simulation strategies for high-level embedded system design
Zhonglei Wang, Andreas Herkersdorf |
Perform. Evaluation | 2 |
| 2009 | An efficient approach for system-level timing simulation of compiler-optimized embedded softwareabstractSoftware accounts for more than 80% of embedded system development efforts, so software performance estimation is a very important issue in system design. Recently, source level simulation (SLS) has become a state-of-the-art approach for software simulation in system level design. However, the simulation accuracy relies on the mapping between source code and binary code, which can be destroyed by compiler optimizations. This drawback strongly limits the usability of this technique in practical system design. We introduce an approach to overcome this limitation by converting source code to a low level representation, called intermediate source code (ISC). ISC has accounted for most compiler optimizations and has a structure close to binary code, so it allows for accurate back-annotation of timing information from the binary level. To show the benefits of our approach, we present a quantitative comparison of the related techniques with the proposed one, using a set of benchmarks. Zhonglei Wang, Andreas Herkersdorf |
DAC | 2 |
| 2009 | SysCOLA: a framework for co-development of automotive software and system platformabstractA modeling language with formal semantics is able to capture a system's functionality unambiguously, without concerning implementation details. Such a formal language is well-suited for a design process that employs formal techniques and supports hardware/software synthesis. On the other hand, SystemC is a widely used system level design language with hardware-oriented modeling features. It provides a desirable simulation framework for system architecture design and exploration. This paper presents a design framework, called SysCOLA, that makes use of the unique advantages of both a new formal modeling language, COLA, and SystemC, and allows for parallel development of application software and system platform. In SysCOLA, function design and architecture exploration are done in the COLA based modeling environment and the SystemC based virtual prototyping environment, respectively. Our concepts of abstract platform and virtual platform abstraction layer facilitate the orthogonalization of functionality and architecture by means of mapping and integration in the respective environments. As SysCOLA is targeted at the automotive domain, the whole design approach is showcased using a case study of designing an automotive system. Zhonglei Wang, Andreas Herkersdorf, Wolfgang Haberl, Martin Wechs |
DAC | 2 |
| 2009 | Hierarchical NoCs for Optimized Access to Shared Memory and IO ResourcesabstractThe concept of on-chip networks (NoCs) has been developed to cope with the increasing communication requirements in systems-on-chip (SoCs) consisting of an ever-growing number of cores. Proposals of NoC architectures are often made assuming evenly distributed traffic, where all tiles receive and produce the same amount of traffic. However, in real systems specific, communication centric tiles, for example off-chip memory controllers or other data IO interfaces, consume and generate a significant part of the overall traffic. In this paper we propose hierarchical NoC topologies to improve access to this type of shared resources. The hierarchical networks may be built from different types of sub-networks, e.g. meshes, rings, crossbars and buses. We investigate different hierarchical network architectures and compare them to the popular 2D mesh topology in terms of hop count, latency and network throughput. Our results show that the proposed approach allows to significantly reduce network latencies to these communication centric tiles. Andreas Lankes, Thomas Wild, Andreas Herkersdorf |
DSD | 3 |
| 2009 | An Efficient Hardware Architecture for Packet Re-sequencing in Network Processors MPSoCsabstractDue to the multi-processor nature of Network Processors (NP), data packets entering the system are processed in parallel and might be transmitted out-of-order at the output leading to a significant degradation in network performance. In this paper we propose a new well-structured, area-efficient, and high speed hardware architecture for packet re-sequencing. For this purpose, several buffering techniques were investigated and analyzed in terms of complexity and memory requirements, taking into consideration the networking application and the impact of the number of processing elements (PE) on packet reordering. The proposed architecture, based on the appropriate buffering mechanism, is then demonstrated and implemented on our FPGA-based prototyping platform. In contrast to other solutions, our results showed 80% more efficient resource utilization while being capable to achieve 10% higher data rate of 3.2 Gbit/s. Shadi Traboulsi, Michael Meitinger, Rainer Ohlendorf, Andreas Herkersdorf |
DSD | 4 |
| 2009 | Flow Analysis on Intermediate Source Code for WCET Estimation of Compiler-Optimized ProgramsabstractMany WCET analysis tools developed in academia integrate WCET analysis into program compilation, either to transform flow information extracted from the source code level to the object code level, or to perform flow analysis on a special intermediate representation. This integration increases analysis complexity, forces software developers to use a special compiler, and thus, strongly limits the usability of the tools in practice. Motivated by this limitation in the existing flow analysis approaches, this paper presents a more efficient approach, that performs flow analysis on the intermediate source code (ISC), transformed from the original source code. ISC retains the functional behavior and executability of the original source code but has a structure close to the object code. This low level structure facilitates the transformation of the flow facts, extracted from the ISC, down to the object code level. In the whole approach, no modification of standard tools is needed. Zhonglei Wang, Andreas Herkersdorf |
RTCSA | 2 |
| 2008 | Buffer allocation for advanced packet segmentation in Network ProcessorsabstractIn current network processors, incoming variable-length packets are sliced using only one small segment size and then stored in the buffer. Inconveniently, short data bursts are inadequate for accessing SDRAM, commonly used for packet buffers, due to high activation and pre-charging latencies. Using large segment sizes is not optimal either because though it increases memory bandwidth, the benefit comes at the price of a heavy reduction in storing efficiency. A good solution to achieve simultaneously high performance and memory utilization consists in storing a single packet segmented using multiple segment sizes. In this paper, we study how to allocate memory for these different-sized segments in an efficient way. First we analyze the appropriate segment pool size for a multitude of traffic scenarios. Our experiments show that simple static buffer allocation does not always suffice as different segment pools may be exhausted depending on traffic. Hence we introduce a method for handling multiple segment pools not only in a static but also in a dynamic way, taking advantage of a new set of control structures based on a combination of bitmaps and linked lists. We demonstrate that our method achieves a huge reduction in control buffer size requirements in comparison to state-of-the-art control structures, together with decreasing the average number of accesses to control data. Daniel Llorente, Kimon Karras, Thomas Wild, Andreas Herkersdorf |
ASAP | 4 |
| 2008 | Design Flows, Communication Based Design and Architectures in Automotive Electronic SystemsabstractSummary form only given. The complete presentation was not made available for publication as part of the conference proceedings. A steadily increasing number of microprocessors and electronic components with the heavy demand of computation performance in automotive electronic systems affect substantially the design of networked ECUs in today as well as future cars. Novel approaches, based on heterogeneous hardware (Coarse- fine Grained reconfigurable Hardware, Microprocessors) could be a solution to handle the computation intensive tasks, e.g. for driver-assistance systems. The challenge here is to find an optimal trade-off between power consumption, cost, performance and flexibility which leads to the question which technology and which distribution (automotive function centralisation - decentralisation trade-offs!) will be targeted in future car electronics. Introducing novel architecture topologies and corresponding tool flows with standardised specification and verification are here severe challenges. A first approach to meet these challenges is the AUTOSAR development partnership, which aims at a standardisation of automotive software architecture. The purpose of this tutorial is to evaluate and discuss new concepts for communication based design of automotive electronic and car network systems, as well as to discuss and envisage future system design in automotive electronics. Both aspects, hardware / software design and tool-integration will be discussed. The main emphasis in this session is design-flow, tool-development, applications and system design. The tutorial is addressed to hardware and system engineers as well as to researchers. A set of presentations intended to set the stage for the discussion, will be followed by a panel where selected world-wide specialists in the field of automotive electronics will discuss the demands and interests of industry on novel technologies and systems and research activities for future automotive systems. Jürgen Becker 0001, Michael Hübner 0001, Robert Esser, Andreas Herkersdorf, Walter Stechele, Vera Lauer |
DATE | 4 |
| 2008 | A Model Driven Development Approach for Implementing Reactive Systems in HardwareabstractTo deal with the increasing complexity of digital systems, the model driven development approach has proven to be beneficial. This paper presents a model driven hardware design process that is dedicated to reactive embedded systems. The approach is based on the component language (COLA), a synchronous data flow language with formal semantics. COLA follows the hypothesis of perfect synchrony. Models thus do not assume specific timing properties and remain deterministic as long as data flow requirements are retained. This is an essential feature for modeling safety-critical systems. Further, the well-defined semantics not only allows that the resulting models can be formally reasoned about, but is also the key to translation to domain-specific languages. This paper describes the approach of translating the models to VHDL descriptions from their graphical representations. As COLA is well-adapted to both data flow description and control automata, the generated VHDL code can be synthesized to very efficient FPGA circuits, comparable to that synthesized from hand-written VHDL code according to our case study. Zhonglei Wang, Andreas Herkersdorf, Stefano Merenda, Michael Tautschnig |
FDL | 2 |
| 2008 | Fine grain reconfigurable architecturesabstractIn this booth on fine grain reconfigurable architectures, several research groups demonstrate their joint work on operating concepts for managing dynamic and partial reconfiguration, visualization of bitstreams and routing, presenting an application applying dynamic reconfiguration for video engines as well as work on minimization of reconfiguration data. Unique is that all the above four projects present their work using the same reconfigurable FPGA-based fabric called Erlangen slot machine that has also been built within one project just the purpose of experimenting with dynamic fine grain reconfiguration as an interdisciplinary platform. Josef Angermeier, Mateusz Majer, Jürgen Teich, Lars Braun, Tobias Schwalb, Philipp Graf, Michael Hübner 0001, Jürgen Becker 0001, Enno Lübbers, Marco Platzner, Christopher Claus, Walter Stechele, Andreas Herkersdorf, Markus Rullmann, Renate Merker |
FPL | 13 |
| 2008 | Network processorsabstractTraditional design of network processors is complicated by two conflicting demands, flexibility and performance. On the one side, network processors should be flexible enough to adapt to changing protocols and varying traffic profiles, on the other side they have to cope with increasing data rates of network links. This demonstrator shows that runtime reconfigurable systems have the potential to optimise both criteria without affecting each other negatively. The demonstrator addresses edge router applications and consists of two independently developed subsystems, the FlexPath NP architecture designed at the TU Munchen and the Dyna-CORE architecture designed at the University of Lubeck. Thilo Pionteck, Roman Koch, Carsten Albrecht, Erik Maehle, Michael Meitinger, Rainer Ohlendorf, Thomas Wild, Andreas Herkersdorf |
FPL | 8 |
| 2008 | A Simulation Approach for Performance Validation during Embedded Systems Design
Zhonglei Wang, Wolfgang Haberl, Andreas Herkersdorf, Martin Wechs |
ISoLA | 3 |
| 2008 | A Processing Path Dispatcher in Network Processor MPSoCsabstractMulti-field packet classification problems discussed in the literature are typically constrained to the Internet five-tuple and primarily address the problem of network quality-of-service (QoS) support and access control. In this paper, we present a solution for a classification problem that is used for optimized packet assignment to different data paths within a network processor system-on-chip (SoC). In contrast to the five-tuple-based rules discussed in the prior art, our problem has rules that consider a larger set of fields from the packet header. However, for each individual rule a different sub-set of fields is relevant and the number of rules is smaller. Based on a specification of the usage case for our classifier we derive heterogeneous decision graph algorithm (HDGA), a heuristic approach to construct a decision tree classifier that integrates external lookup results for certain types of rules. We evaluate various parameters for optimizing the proposed decision tree and present simulation results to show the scalability of HDGA for typical problem sizes. This paper is concluded with the results of an implementation on our field-programmable gate-array (FPGA)-based prototyping platform. Rainer Ohlendorf, Michael Meitinger, Thomas Wild, Andreas Herkersdorf |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2007 | Simulated and measured performance evaluation of RISC-based SoC platforms in network processing applications
Rainer Ohlendorf, Thomas Wild, Michael Meitinger, Holm Rauchfuss, Andreas Herkersdorf |
J. Syst. Archit. | 5 |
| 2006 | AutoVision: flexible processor architecture for video-assisted drivingabstractSummary form only given. Future automotive security systems will benefit from visual scene analysis based on a fusion of video, infrared, and radar images. Today we have already functions like lane departure warning and automatic cruise control (ACC) for pretty well defined driving environments, such as highways and primary roads. Recent research activities concentrate on more complex environments, such as city traffic with a wide variety of traffic participants moving in an unpredictable manner, e.g. bikes, pedestrians, children, and even animals, and under changing weather and lighting conditions. The ITRS semiconductor roadmap for microelectronics forecasts a continued doubling of transistor capacity per chip every 2 to 2.5 years enabling billion transistor ASIC designs in the near future. Multi processor system on chip (MPSoC) solutions with 8, 16 or even more standard RISC CPU cores, mega-bytes of fast (ns access latencies) on-chip SRAM memories, giga-byte per second interconnect buses or NoC (network on chip) meshes, high-speed serial I/Os and, last but not least, million gate equivalent dedicated hardware accelerator functions in eFPGA (embedded field programmable gate array) logic are becoming reality on a single silicon substrate. Examples of current research projects shall illustrate our perception on how this tremendous increase in functionality and computational performance per chip area may impact automotive control unit (ACU) architectures for driver assistance applications. The AutoVision processor is a dynamically reconfigurable MPSoC prototype where video-specific pixel processing engines are on-the-fly loaded or exchanged without interrupting regular system operations. For the time being, pixel processing engines cover functions such as object edge detection or luminance segmentation, and are implemented as dedicated hardware accelerators to ensure real-time frame processing capabilities of the AutoVision processor. Dynamic replacement of processing engines ensures an automatic and area efficient adaptation to various driving conditions. Segmented objects are, in a subsequent step, characterized by means of standard MPEG-7 descriptors and entered as search criteria into traffic scene analysis databases. Goal is to obtain a clean distinction between passenger cars, trucks, and big rectangular traffic signs, and to identify pedestrians or bikers in complex traffic situations. The AutoVision processor project is supported by the German Research Foundation (DFG) in the special emphasis research programme "reconfigurable computing" Andreas Herkersdorf, Walter Stechele |
DATE | 1 |
| 2006 | Performance evaluation for system-on-chip architectures using trace-based transaction level simulationabstractThe ever increasing complexity and heterogeneity of modern system-on-chip (SoC) architectures make an early and systematic exploration of alternative solutions mandatory. Efficient performance evaluation methods are of highest importance for a broad search in the solution space. In this paper we present an approach that captures the SoC functionality for each architecture resource as sequences of trace primitives. These primitives are translated at simulation runtime into transactions and superposed on the system architecture. The method uses SystemC as modeling language, requires low modeling effort and yet provides accurate results within reasonable turnaround times. A concluding application example demonstrates the effectiveness of our approach Thomas Wild, Andreas Herkersdorf, Rainer Ohlendorf |
DATE | 2 |
| 2006 | Organic Computing at the System on Chip LevelabstractThe evolution of CMOS technologies leads to integrated circuits with ever smaller device sizes, lower supply voltage, higher clock frequency and more process variability. Intermittent faults effecting logic and timing are becoming a major challenge for future integrated circuit designs. This paper presents an organic computing inspired SoC architecture which applies self-organization and self-calibration concepts to build reliable SoCs with lower overheads and a broader fault coverage than classical fault-tolerance techniques. We demonstrate the feasibility of this approach by example on the processing pipeline of a public-domain RISC CPU core Abdelmajid Bouajila, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf, Andreas Bernauer, Oliver Bringmann 0001, Wolfgang Rosenstiel |
VLSI-SoC | 4 |
| 2005 | Reduction of CMOS Power Consumption and Signal Integrity Issues by Routing OptimizationabstractThis paper suggests a methodology to decrease the power of a static CMOS standard cell design at layout level by focusing on switched capacitance. The term switched is the key: if a capacitance is not switched often, it may be high. If it is frequently switched, it should be minimized in order to reduce power consumption. This can be done by an algorithm based on forces that automatically optimizes the position and length of every single wire segment in a routed design. The forces are proportional to the toggle activities derived from a gate level simulation. The novelty is that this allows us to iteratively find a new topology for the wire segments. Our algorithm takes as input an already given, grid routed layout. Paul Zuber, Armin Windschiegl, Raúl Medina Beltrán de Otálora, Walter Stechele, Andreas Herkersdorf |
DATE | 5 |
| 2005 | Robust header compression (ROHC) in next-generation network processorsabstractRobust Header Compression (ROHC) provides for more efficient use of radio links for wireless communication in a packet switched network. Due to its potential advantages in the wireless access area and the proliferation of network processors in access infrastructure, there exists a need to understand the resource requirements and architectural implications of implementing ROHC in this environment. We present an analysis of the primary functional blocks of ROHC and extract the architectural implications on next-generation network processor design for wireless access. The discussion focuses on memory space and bandwidth dimensioning as well as processing resource budgets. We conclude with an examination of resource consumption and potential performance gains achievable by offloading computationally intensive ROHC functions to application specific hardware assists. We explore the design tradeoffs for hardware assists in the form of reconfigurable hardware, Application-Specific Instruction-set Processors (ASIPs), and Application-Specific Integrated Circuits (ASICs). David E. Taylor, Andreas Herkersdorf, Andreas C. Döring, Gero Dittmann |
IEEE/ACM Trans. Netw. | 2 |
| 2004 | Buffer Schemes for Runtime Reconfiguration of Function Variants in Communication SystemsabstractThis contribution is an extension of our work, which introduced distributed buffer schemes for runtime reconfiguration in adaptive processing architectures, e.g., for streaming media applications. We propose a reconfiguration control protocol and depict results for a reference system implementation. With dynamic reconfiguration, area-cost of field-programmable logic (FPL) can be reduced by reuse, and potential for adaptive signal processing techniques can be enabled. The challenge with runtime reconfiguration is the reconfiguration latency. Given the limitation regarding reconfiguration latency with traditional approaches, we proposed distributed buffer schemes. The simulation results show that our approach enables potential for runtime reconfiguration for adaptive signal processing under real-time constraints. The proposed control protocol enables regular structures for modular based dynamic reconfiguration handling. Finally, we present implementation details and results of an architectural system implementation. Dirk Eilers, Helmut Steckenbiller, Andreas Herkersdorf |
FCCM | 3 |
| 2003 | Performance evaluation of network processor architectures: combining simulation with analytical estimation
Samarjit Chakraborty, Simon Künzli 0001, Lothar Thiele, Andreas Herkersdorf, Patricia Sagmeister |
Comput. Networks | 4 |
| 2003 | Design methodology for a modular service-driven network processor architecture
Maria Gabrani, Gero Dittmann, Andreas C. Döring, Andreas Herkersdorf, Patricia Sagmeister, Jan van Lunteren |
Comput. Networks | 4 |
| 1999 | A scalable modular architecture for SDH/SONET technologyabstractA scalable architecture for integrated SDH/SONET framers is presented. It exploits the fact that not only those framer functions that are obviously suited for parallel processing but all SDH/SONET overhead processing and most of the payload processing functions can be implemented as distributed algorithms. Thereby an M/spl times/STM-N framer also handles STM-M/spl times/N and STM-(M/spl times/N)c frames. These distributed algorithms cover frame scrambling, overhead byte processing, ATM and PPP payload handling, and interface implementations. A series of modular building blocks for this architecture and first complete framers is now available. Rolf Clauberg, Andreas Herkersdorf, Wolfram W. Lemppenau, Hans R. Schindler |
ICCCN | 2 |
| 1995 | Route Discovery for Multistage Fabrics in ATM Switching Nodes
Andreas Herkersdorf, L. Heusler, Erik Maehle |
Perform. Evaluation | 1 |
| 1993 | Fast Connection Establishment in Large-Scale NetworksabstractA hierarchical decomposition of network nodes which permits the networkwide topology database and algorithms to be independent of node internals, yet allows (implicit) access to special intranodal features when required is described. A complementary procedure for the establishment of bandwidth-reserved connections is outlined that makes efficient use of this node structure to support network-level optimization through use of node-level features. Based on a significantly increased concurrency of path computation and bandwidth reservation, the execution time of the setup procedure is independent of the network's complexity (such as number of nodes and links) and is essentially bounded from above by the round-trip delay only.> Willibald A. Doeringer, Doug Dykeman, Antonius P. J. Engbersen, Roch Guérin, Andreas Herkersdorf, L. Heusler |
INFOCOM | 5 |