VLDB 2026 Research / reviewers in the wild / expert
Thomas Wild
dblp:15/4766
· DBLP profile ↗
60ranked-venue papers
1as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 1 first-author · 13 since 2021Software engineering, systems software and programming languages · 12 · 1 first-author · 5 since 2021Computer networks · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Advancing European Semiconductor and Chiplet Innovation Through the Bavarian Chip Design CenterabstractEurope’s semiconductor industry relies heavily on Asian and US manufacturers. The EU Chips Act seeks to strengthen Europe’s capabilities across the semiconductor value chain. Aligned with this goal, the Bavarian Chip Design Center (BCDC) supports local chip design, manufacturing, and talent development, with a focus on RISC-V computing and heterogeneous integration. Within BCDC, the Technical University of Munich and Fraunhofer are developing a chiplet-based architecture optimized for low-power edge AI. The system integrates two chiplets, combining a security-enhanced RISC-V core and AI accelerators, connected via a chiplet-optimized serial interface that supports encrypted data. The chiplets are mounted on a custom interposer with low-capacitance wires for efficient data transmission. System-and component-level development is currently ongoing, with a tapeout in 22 nm FD-SOI planned for 2027. The overall goal is to deliver a proof of concept for a small-scale energy-efficient chiplet system that demonstrates Bavaria’s and Europe’s capability to drive innovation in novel chip design fields. Hussam Amrouch, Jehaan Joseph, Michael Schirmer, Johannes Geier, Ulf Schlichtmann, Michael Meidinger, Thomas Wild, Andreas Herkersdorf, Jens Nöpel, Georg Sigl, Carsten Trinitis, Aswathy Nedumpalli Sankaranarayanan, Martin Schulz 0001, Andreas Korb, Konrad Hohentanner |
DATE | 7 |
| 2026 | STEP: Spatial Footprint Prefetcher with Multi-Point Temporal Triggers
Yuanji Ye, Oliver Lenke, Thomas Wild, Andreas Herkersdorf |
ISCA | 3 |
| 2025 | HiPerNoC: A High-Performance Network-an-Chip for Flexible and Scalable FPGA-Based SmartNICsabstractA recent approach that the research community has proposed to address the steep growth of network traffic and the attendant rise in computing demands is in-network computing. This paradigm shift is bringing about an increase in the types of computations performed by network devices. Consequently, processing demands are becoming more varied, requiring flexible packet-processing architectures. State-of-the-art switch-based smart network interface cards (SmartNICs) provide high versatility without sacrificing performance but do not scale well concerning resource usage. In this paper, we introduce HiPerNoC-a flexible and scalable field-programmable gate array (FPGA)-based SmartNIC architecture deploying a 2D-mesh network-on-chip (NoC) with a novel router design to manage network traffic with diverse processing demands. The NoC can forward incoming network packets to the available processing engines in the required sequence at a traffic load of up to 91.1 Gbit/s (0.89 flit/node/cycle). Each router applies distributed switch allocation and avoids head-of-line blocking by deploying queues at the switch crosspoints of input-output connections used by the routing algorithm. It also prevents deadlocks by employing non-blocking virtual cut-through switching. We implemented a prototype of HiPerNoC as a 4x4 2D-mesh NoC in SystemVerilog and evaluated it with synthetic network traffic via cycle-accurate register-transfer level simulations in Vivado. The evaluation results show that HiPerNoC achieves up to 53 % higher saturation throughput, occupies 53 % fewer lookup tables and block RAMs, and consumes 16 % less power on an Alveo U55C than ProNoC-a state-of-the-art FPGA-based NoC. Klajd Zyla, Marco Liess, Thomas Wild, Andreas Herkersdorf |
DATE | 3 |
| 2025 | Rule-Based Reinforcement Learning on FPGA for QoS-Aware Dynamic Frequency ScalingabstractTo improve the system performance of multiprocessor system-on-chips (MPSoCs), modern processors have several built-in hardware features, such as prefetchers, which respond to short-term variations in processor load that occurs on a submillisecond scale. However, even the latest dynamic (voltage) frequency scaling governors using reinforcement learning (RL) are implemented in software and, thus, cannot take advantage of these variations. In this work, we propose a hardware RL agent, augmented with preemptive shielding and eligibility traces, to optimize the execution of deadline-bound quality-of-service (QoS) tasks in mixed-critical environments. We demonstrate the features of our algorithm in a hardware-in-the-loop simulation by running LLVM’s single-source benchmarks on SparcV8 processors. We also present our field-programmable gate array (FPGA) implementation with optimized resource usage and timing performance achieved through quantization and approximation. Florian Maurer 0003, Michael Meidinger, Matthias Schlemmer, Thomas Hallermeier, Anmol Surhonne, Thomas Wild, Andreas Herkersdorf |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | Hardware Assist for Linux IPC on an FPGA PlatformabstractSpecialized hardware units often accelerate compute-intensive or memory-heavy functions. In previous publications, we proposed concepts to assist Linux with a hardware unit for managing waiting threads to improve blocking inter-process communication (IPC) mechanisms. This paper assesses the effectiveness of this hardware support on a Zynq platform. Although main memory accesses by our hardware unit are time-consuming, a consumer-producer application achieved an up to 220% increased message rate. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
CF | 5 |
| 2024 | HASIIL: Hardware-Assisted Scheduling to Improve IPC Latency in LinuxabstractInter-processes communication (IPC) is essential for multi-threaded applications to achieve efficient execution. Synchronization through IPC can become a bottleneck for these applications. The effectiveness of IPC is determined by both its latency and CPU utilization needed for the associated functions. Our research has revealed that for blocking IPC mechanisms, the thread scheduling functions within the Linux operating system significantly contribute to the notification latency. To address this issue, we propose a novel concept called HASIIL, which combines offloading IPC functionality with hardware-assisted scheduling to enhance IPC latency. Through this approach, we can improve the latency of blocking IPC mechanisms by up to 36% in Linux, while also improving CPU utilization by 40%. Tim Twardzik, Lars Nolte, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
CF | 5 |
| 2024 | EMDRIVE Architecture: Embedded Distributed Computing and Diagnostics from Sensor to EdgeabstractFuture automotive architectures are expected to transition from a network-centric to a domain-centered architecture featuring central compute units. Powerful domain controllers or smart sensors alleviate the load on these central units and communication systems. These controllers execute tasks with varying criticalities on heterogeneous multicore processors, and are ideally capable of dynamically balancing the computing load between the central unit and sensors. Here, Artificial Intelligence (AI) capabilities playa crucial role, as it is in high demand for such an automotive architecture. However, AI still requires specialized accelerators to improve their computation performance. Task-oriented distributed computing with criticalities up to ASIL-D necessitates the development and utilization of specialized methodologies, such as safety, through the isolation and abstraction of low-level hardware concepts. Meanwhile, online monitoring and diagnostics become vital features to detect errors during operation. The EMDRIVE architecture includes methods, components, and strategies to enhance the performance, safety, and security of such distributed computing platforms. The nationally funded EMDRIVE project connects its twelve partners from academia and industry and is currently in its intermediate stage. Patrick Schmidt 0003, Iuliia Topko, Matthias Stammler, Tanja Harbaum, Jürgen Becker 0001, Rico Berner, Omar Ahmed, Jakub Jagielski, Thomas Seidler, Markus Abel, Marius Kreutzer, Maximilian Kirschner, Victor Pazmino Betancourt, Robin Sehm, Lukas Groth, Andrija Neskovic, Rolf Meyer, Saleh Mulhem, Mladen Berekovic, Matthias Probst, Manuel Brosch, Georg Sigl, Thomas Wild, Matthias Ernst, Andreas Herkersdorf, Florian Aigner, Stefan Hommes, Sebastian Lauer, Maximilian Seidler, Thomas Raste, Gasper Skvarc Bozic, Ibai Irigoyen Ceberio, Albrecht Mayer |
DATE | 23 |
| 2024 | ecoNIC: Saving Energy Through SmartNIC-Based Load Balancing of Mixed-Critical Ethernet TrafficabstractIn next-generation automotive, industrial, data cen-ter, and other mixed-critical networks, Ethernet is expected to power the backbone interconnect among multi-core compute nodes. On attached Network Interface Cards (NICs) Receive Side Scaling (RSS) supports the CPU in balancing workloads across cores for reduced tail latencies. However, state-of-the-art solutions are primarily designed for performance and less for energy-efficiency which will play an equally important role. For this reason we present ecoNIC, an RSS-based hardware load balancer for SmartNICs, and an agile Dynamic Voltage and Frequency Scaling (DVFS) governor, for energy-saving network processing. ecoNI C efficiently pins flow priorities to CPU core clusters, reducing the workload of select cores in the process, and dynamically adjusts their clock speed to exploit freed-up capacities and save energy. Within a cluster, it proactively redirects packet bursts of priority-separated flow bundles among available cores, or offloads them to neighbor nodes, once local resources tend to become highly loaded. The per-core energy consumption this way is reduced at the expense of low priority packet latencies, while high priority service qualities are maintained. Experimental evaluations applying real-world network traces yield energy savings of up to 37.9 % at an increase from 559 µs to 3.06 ms in low priority end-to-end tail latency compared to an even workload distribution without frequency scaling. Franz Biersack, Marco Liess, Markus Absmann, Fabiana Lotter, Thomas Wild, Andreas Herkersdorf |
DSD | 5 |
| 2024 | FlexCross: High-Speed and Flexible Packet Processing via a Crosspoint-Queued CrossbarabstractThe fast pace at which new online services emerge leads to a rapid surge in the volume of network traffic. A recent approach that the research community has proposed to tackle this issue is in-network computing, which means that network devices perform more computations than before. As a result, processing demands become more varied, creating the need for flexible packet-processing architectures. State-of-the-art approaches provide a high degree of flexibility at the expense of performance for complex applications, or they ensure high performance but only for specific use cases. In order to address these limitations, we propose FlexCross. This flexible packet-processing design can process network traffic with diverse processing requirements at over 100 Gbit/s on FPGAs. Our design contains a crosspoint-queued crossbar that enables the execution of complex applications by forwarding incoming packets to the required processing engines in the specified sequence. The crossbar consists of distributed logic blocks that route incoming packets to the specified targets and resolve contentions for shared resources, as well as memory blocks for packet buffering. We implemented a prototype of FlexCross in Verilog and evaluated it via cycle-accurate register-transfer level simulations. We also conducted test runs with real-world network traffic on an FPGA. The evaluation results demonstrate that FlexCross outperforms state-of-the-art flexible packet-processing designs for different traffic loads and scenarios. The synthesis results show that our prototype consumes roughly 21% of the resources on a Virtex XCU55 UltraScale+ FPGA. Klajd Zyla, Marco Liess, Thomas Wild, Andreas Herkersdorf |
DSD | 3 |
| 2024 | FlexRoute: A Fast, Flexible and Priority-Aware Packet-Processing DesignabstractAs the world becomes more connected and new digital services emerge at a fast pace, the amount of network traffic increases rapidly. Consequently, processing requirements become more varied and drive the need for flexible packet-processing designs, especially as in-network computing gains traction. Traditional approaches deploy hardware accelerators in a pipeline in the sequence that the associated tasks are supposed to be executed. Hence, they do not accommodate flows with different processing requirements and provide no possibility to remap flows to task sequences in runtime. In order to address these limitations, we propose FlexRoute, a fast, flexible and priority-aware packet-processing design that can process network traffic at a rate of over 100 Gbit/s on FPGAs. Our design consists of a reconfigurable parser and several processing engines that are arranged in a pipeline. The processing engines are equipped with processing units that execute specific tasks, flexible forwarding logic and priority-aware queuing/scheduling logic. We implement a prototype of FlexRoute in Verilog and evaluate it via cycle-accurate register-transfer level simulations. We also synthesize and implement our design on the Alveo U55C High Performance Compute Card and show its resource usage. The evaluation results demonstrate that FlexRoute can process packets of arbitrary size with different processing requirements at a traffic rate of about 70 Gbit/s significantly faster than two state-of-the-art flexible packet-processing designs. Klajd Zyla, Marco Liess, Thomas Wild, Andreas Herkersdorf |
PDP | 3 |
| 2024 | HW-FUTEX: Hardware-Assisted Futex SyscallabstractEfficient thread synchronization primitives are crucial in modern computer systems for the performant execution of interdependent code segments. In Linux, the futex() syscall is used to construct blocking synchronization primitives such as mutexes or conditional variables. When using futex, the uncontended case is efficiently handled entirely in user space. In the event of contention, the kernel is called to put the waiting thread to sleep until the state of the primitive changes to uncontended. The kernel must be notified of this change by a futex() syscall to wake-up the sleeping thread. This syscall must be issued by the thread that changes the primitive, which is a significant burden on this thread. To remove this burden, we introduce HW-FUTEX to offload the futex wake functionality to a hardware unit (HW Unit) that asynchronously initiates wake-ups of the sleeping threads. This reduces the time required to issue the futex wake functionality by at least 90% to 350 cycles, with no additional overhead in the uncontended case. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | HAWEN: Hardware Accelerator for Thread Wake-Ups in Linux Event NotificationabstractThe performance of multi-threaded applications relies on efficient inter-process communication. One common practice is putting a thread asleep while waiting for a certain condition. Exemplary Linux kernel mechanisms that use this practice include futex, sockets, epoll, eventfd and pipe. Once the condition is met, i.e., the associated event has occurred, the waiting thread is notified. Optimizations for event notification mechanisms in Linux mostly target the thread which receives events. Contrarily, we identified high potential in relieving the event-generating thread and propose HAWEN, a hardware accelerator for thread wake-up support. HAWEN has been integrated into Linux event notification in a minimally intrusive manner. Gem5-based multi-core architecture simulations revealed up to 80% faster thread wake-up times and a 53% shorter event-generating syscall. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Clara Kowalsky, Thomas Wild, Andreas Herkersdorf |
DAC | 7 |
| 2023 | Priority-aware Inter-Server Receive Side ScalingabstractNext-generation automotive networks will be characterized by a high number of interconnected sensors, actuators and applications on electronic control units communicating with each other over a high-speed Ethernet backbone network. As these applications have various criticalities, high volumes of fluctuating traffic with different priorities will have to be processed in a reliable and efficient manner. To cope with these challenges, we present Priority-aware Inter-Server Receive Side Scaling (prioRSS), a new SmartNIC-based hardware accelerator designed for automotive compute nodes. prioRSS builds upon Receive Side Scaling and introduces priority-awareness into an intra- and inter-node load balancer. It uses a priority-partitioned indirection table within which flows of the same priority are bundled. Low-latency reconfigurations issued by a Network Health Monitoring software allow for adapting the table content to changing network conditions. Simulative evaluations and comparisons to a priority-unaware version of our design show that prioRSS enables per-priority resource assignments without degrading end-to-end packet latencies while using the same table memory space. Paired with a priority-aware scheduler, end-to-end latencies of high priority flows can be notably reduced compared to average packet latencies, at the expense of lowest priority traffic. The best results are acquired when partitioning the table proportionally to the associated traffic share. Franz Biersack, Kilian Holzinger, Henning Stubbe, Thomas Wild, Georg Carle, Andreas Herkersdorf |
PDP | 4 |
| 2023 | FlexPipe: Fast, Flexible and Scalable Packet Processing for High-Performance SmartNICsabstractData centers have been struggling to provide the necessary processing capacity to handle the surging rate of network traffic that is generated in an increasingly connected and service-oriented world. As a result, SmartNICs play an even more important role than before as they can offload various network applications and hence free CPU resources for application-layer processing, increase performance and reduce processing time. However, they often do not support flows with different offload requirements and cannot dynamically allocate offloads in run-time. In order to address these limitations, we propose FlexPipe, a fast, flexible and scalable packet-processing architecture for high-performance SmartNICs. Our design enables low-latency and runtime-reconfigurable packet forwarding at high traffic rates with minimal area overhead. Furthermore, it provides load-aware packet steering toward multiple offload units of the same type for low-bandwidth offloads. We implement a prototype of FlexPipe in Verilog and validate it via cycle-accurate register-transfer level simulations. Our evaluation results show that FlexPipe can process packets of arbitrary size with different offload requirements at line rate and on average 1.9x faster than a SmartNIC with a predefined sequence of offloads and 1.8x faster than PANIC, a flexible state-of-the-art SmartNIC. Klajd Zyla, Marco Liess, Thomas Wild, Andreas Herkersdorf |
VLSI-SoC | 3 |
| 2022 | SmartNIC-based Load Management and Network Health Monitoring for Time Sensitive ApplicationsabstractTime sensitive network applications, for example in Intra-Vehicular Networks, aim to give predictable end-to-end latency guarantees. As a consequence, processing resources of involved host systems remain partially unused, because they are reserved for rare worst cases. This circumstance provides the opportunity to reduce dimensioning overheads by managing the load on the nodes flexibly within the network. In our proposed approach, a SmartNIC involving an FPGA-based load balancer achieves dynamic routing of flows whilst preserving end-to-end latency guarantees. A flow-oriented online network measurement component continuously supervises network traffic with regards to compliance to flow specifications and constraints such as bounded one-way delay, absence of packet loss, and jitter. We use the supervisor to enhance forwarding decisions on the data plane. Initial evaluation yields a saving potential of around 30 %. We showcase quick dynamic reconfiguration of the FPGA when triggered by real-time measurement of the one-way delay using realistic automotive network traffic. Kilian Holzinger, Franz Biersack, Henning Stubbe, Angela Gonzalez Mariño, Abdoul Kane, Francesc Fons, Haigang Zhang, Thomas Wild, Andreas Herkersdorf, Georg Carle |
NOMS | 8 |
| 2021 | Precise real-time monitoring of time-critical flowsabstractEthernet is increasingly used in areas where time-critical and safety-relevant data are transported over the network along with best-effort flows, for example in intra vehicle networks or industrial networks. The resulting complex network architectures, time-sensitive networking configurations and system interactions are hard to foresee during the design phase. Therefore, it is hard to rule out any violations of flow specifications or timing and reliability requirements, especially in the presence of unpredictable failures. Kilian Holzinger, Henning Stubbe, Franz Biersack, Angela Gonzalez Mariño, Abdoul Kane, Francesc Fons, Haigang Zhang, Thomas Wild, Andreas Herkersdorf, Georg Carle |
CoNEXT | 8 |
| 2021 | Long Short-Term Memory Neural Network-based Power Forecasting of Multi-Core ProcessorsabstractWe propose a novel technique to forecast the power consumption of processor cores at run-time. Power consumption varies strongly with different running applications and within their execution phases. Accurately forecasting future power changes is highly relevant for proactive power/thermal management. While forecasting power is straightforward for known or periodic workloads, the challenge for general unknown workloads at different voltage/frequency (v/n-levels is still unsolved. Our technique is based on a long short-term memory (LSTM) recurrent neural network (RNN) to forecast the average power consumption for both the next 1ms and 10ms periods. The runtime inputs for the LSTM RNN are current and past power information as well as performance counter readings. An LSTM RNN enables this forecasting due to its ability to preserve the history of power and performance counters. Our LSTM RNN needs to be trained only once at design-time while adapting during run-time to different system behavior through its internal memory. We demonstrate that our approach accurately forecasts power for unseen applications at different v/f-levels. The experimental results shows that the forecasts of our LSTM RNN provide 43% lower worst case error for the 1ms forecasts and 38% for the 10ms forecasts. comnared to the state of the art. Mark Sagi, Martin Rapp, Heba Khdr, Yizhe Zhang 0005, Nael Fasfous, Nguyen Anh Vu Doan, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
DATE | 7 |
| 2020 | Inter-Server RSS: Extending Receive Side Scaling for Inter-Server Workload DistributionabstractNetwork Function Virtualization enables operators to schedule diverse network processing workloads on a general-purpose hardware infrastructure. However, short-lived processing peaks make an efficient dimensioning of processing resources under stringent tail latency constraints challenging. To reduce dimensioning overheads, several load balancing approaches, which either adaptively steer network traffic to a group of servers or to their internal CPU cores, have separately been investigated.In this paper, we present Inter-Server RSS (isRSS), a hardware mechanism built on top of Receive Side Scaling in the network interface card, which combines intra-and inter-server load balancing. In a first step, isRSS targets a balanced utilization of processing resources by steering packet bursts to CPU cores based on per-core load feedback. If all local CPU cores are highly loaded, isRSS avoids high queueing delays by redirecting newly arriving packet bursts to other servers, which execute the same network functions, exploiting that processing peaks are unlikely to occur at all servers at the same time. Our evaluation based on real-world network traces shows that compared to Receive Side Scaling, the joint intra-and inter-server load balancing approach is able to reduce the processing capacity dimensioned for network function execution by up to 38.95% and limit packet reordering to 0.0589% while maintaining tail latencies. Andreas Oeldemann, Franz Biersack, Thomas Wild, Andreas Herkersdorf |
PDP | 3 |
| 2020 | A Lightweight Nonlinear Methodology to Accurately Model Multicore Processor PowerabstractMany power management algorithms demand accurate and fine-grained runtime estimations of dynamic core power. In the absence of fine-grained power sensors, model-based estimations are needed. Such power models commonly approximate the switching activity of logic gates using performance counters while assuming a linear performance counter/power relation at a fixed frequency and voltage. It has been shown that this relation cannot be captured accurately enough with purely linear models and that well-established nonlinear modeling techniques, e.g., polynomial modeling, easily overfit the underlying performance/power relations. Although neural-network-based modeling has shown to accurately capture nonlinear relations, it has a large training and inference overhead which is too high for fine-grained models on core-level and estimation rates in the range of 1-10 kHz. We propose a methodology for nonlinear transformation of specific performance counters to increase power modeling accuracy at constant frequency and voltage with a relatively low overhead for both model generation and run-time application over a linear model. Furthermore, we use least-angle regression (LARS) to determine a ranking of the performance counter inputs for use in linear and nonlinear modeling and show that the transformed performance counters are better suited for power modeling. The generated dynamic power model consisting of a nonlinear transformation block and a linear regression block reduces relative estimation error on average by 4% and in worst-case scenarios by 7% compared to state-of-the-art fine-grained linear power models. Compared to a state-of-the-art polynomial regression model our proposed approach reduces the relative estimation error by 10% in worst-case scenarios. Mark Sagi, Nguyen Anh Vu Doan, Martin Rapp, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Machine Learning Approaches for Efficient Design Space Exploration of Application-Specific NoCsabstractIn many Multi-Processor Systems-on-Chip (MPSoCs), traffic between cores is unbalanced. This motivates the use of an application-specific Network-on-Chip (NoC) that is customized and can provide a high performance at low cost in terms of power and area. However, finding an optimized application-specific NoC architecture is a challenging task due to the huge design space. This article proposes to apply machine learning approaches for this task. Using graph rewriting, the NoC Design Space Exploration (DSE) is modelled as a Markov Decision Process (MDP). Monte Carlo Tree Search (MCTS), a technique from reinforcement learning, is used as search heuristic. Our experimental results show that—with the same cost function and exploration budget—MCTS finds superior NoC architectures compared to Simulated Annealing (SA) and a Genetic Algorithm (GA). However, the NoC DSE process suffers from the high computation time due to expensive cycle-accurate SystemC simulations for latency estimation. This article therefore additionally proposes to replace latency simulation by fast latency estimation using a Recurrent Neural Network (RNN). The designed RNN is sufficiently general for latency estimation on arbitrary NoC architectures. Our experiments show that compared to SystemC simulation, the RNN-based latency estimation offers a similar speed-up as the widely used Queuing Theory (QT). Yet, in terms of estimation accuracy and fidelity, the RNN is superior to QT, especially for high-traffic scenarios. When replacing SystemC simulations with the RNN estimation, the obtained solution quality decreases only slightly, whereas it suffers significantly when QT is used. Marcel Mettler, Daniel Mueller-Gritschneder, Thomas Wild, Andreas Herkersdorf, Ulf Schlichtmann |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2019 | Cryptographic Hashing in P4 Data PlanesabstractP4 introduces a standardized, universal way for data plane programming. Secure and resilient communication typically involves the processing of payload data and specialized cryptographic hash functions. We observe that current P4 targets lack the support for both. Therefore, applications and protocols, which require message authentication codes or hashing structures that are resilient against attacks such as denial-of-service, cannot be implemented. To enable authentication and resilience, we make the case for extending P4 targets with cryptographic hash functions. We propose an extension of the P4 Portable Switch Architecture for cryptographic hashes and discuss our prototype implementations for three different P4 target platforms: CPU, NPU, and FPGA. To assess the practical applicability, we conduct a performance evaluation and analyze the resource consumption. Our prototype implementations show that cryptographic hashing can be integrated efficiently. We cannot identify a single hash function delivering satisfying performance on all investigated platforms. Therefore, we recommend a set of hash functions to optimize target-specific performance. Dominik Scholz, Andreas Oeldemann, Fabien Geyer, Sebastian Gallenmüller, Henning Stubbe, Thomas Wild, Andreas Herkersdorf, Georg Carle |
ANCS | 6 |
| 2019 | Channel mapping strategies for effective protection switching in fail-operational hard real-time NoCsabstractWith Multi Processor System-on-Chips (MPSoC) scaling up to thousands of processing elements, bus-based solutions have been dropped in favor of Network-on-Chips (NoC) as proposed in [2]. However, MPSoCs are yet hesitantly adopted in safety-critical fields, mainly due to the difficulty of ensuring strict isolation between different applications running on a single MPSoC as well as providing communication with Guaranteed Service (GS) to critical applications. This is particularly difficult in the NoC as it constitutes a network of shared resources. Moreover, safety-critical applications require some degree of Fault-Tolerance (FT) to guarantee safe operation at all times. Max Koenen, Nguyen Anh Vu Doan, Thomas Wild, Andreas Herkersdorf |
NOCS | 3 |
| 2019 | APEC: improved acknowledgement prioritization through erasure coding in bufferless NoCsabstractBufferless NoCs have been proposed as they come with a decreased silicon area footprint and a reduced power consumption, when compared to buffered NoCs. However, while known for their inherent simplicity, they suffer from early saturation and depend on additional measures to ensure reliable packet delivery, such as control protocols based on ACKs or NACKs. In this paper, we propose APEC, a novel concept for bufferless NoCs that allows to prioritize ACKs and NACKs over single payload flits of colliding packets by discarding the latter. Lightweight heuristic erasure codes are used to compensate for discarded payload flits. By trading off the erasure code overhead for packet retransmissions, a more efficient network operation is achieved. For ACK-based networks, APEC saturates at 2.1x and 2.875x higher generation rates than a conventional ACK-based bufferless NoC for packets between 5 and 17 flits. For NACK-based networks, APEC does not require concepts such as deflection routing or circuit-switched overlay NACK-networks, as prior work does. Therefore, it can simplify the network implementation compared to prior work while achieving similar performance. Michael Vonbun, Adrian Schiechel, Nguyen Anh Vu Doan, Thomas Wild, Andreas Herkersdorf |
NOCS | 4 |
| 2018 | BiSME: A Hardware Coprocessor to Perform Signature Matching at Multi-Gigabit RatesabstractHardware acceleration of signature matching is essential to perform content aware networking at predictable rates in modern network processors. Existing hardware accelerators either cannot perform signature matching at predictable rates due to the storage organization of the signatures or do not compress the signatures effectively resulting in inefficient on-chip memory usage. Addressing these problems, a bitmap based signature matching engine called BiSME is proposed in this paper, which is a flexible, programmable and scalable hardware coprocessor to perform signature matching at fixed, but guaranteed rates. The storage architectures proposed as part of BiSME, allows to efficiently store the compressed signatures in a flexible and programmable manner in on-chip memories. Each BiSME instance is fine-tuned to perform signature matching at 9.3 Gbps, with multiple instances capable of supporting increasing signature counts as well as increasing throughput. The BiSME was synthesized on a commercial 28nm technology library and only occupies 1.43 mm2of silicon area and consumes 155mW of power. The BiSME hardware implementation was thoroughly verified on the Cadence Palladium platform. Over 2GB of network traffic was injected simultaneously into BiSME and a software based signature matching solution and the identical signature matching results further validated the correctness of the design. Subramanian Shiva Shankar, Pinxing Lin, Andreas Herkersdorf, Thomas Wild |
ASAP | 4 |
| 2018 | Design methodologies for enabling self-awareness in autonomous systemsabstractThis paper deals with challenges and possible solutions for incorporating self-awareness principles in EDA design flows for autonomous systems. We present a holistic approach that enables self-awareness across the software/hardware stack, from systems-on-chip to systems-of-systems (autonomous car) contexts. We use the Information Processing Factory (IPF) metaphor as an exemplar to show how self-awareness can be achieved across multiple abstraction levels, and discuss new research challenges. The IPF approach represents a paradigm shift in platform design by envisioning the move towards a consequent platform-centric design in which the combination of self-organizing learning and formal reactive methods guarantee the applicability of such cyber-physical systems in safety-critical and high-availability applications. Armin Sadighi, Bryan Donyanavard, Thawra Kadeed, Kasra Moazzemi, Tiago Rogério Mück, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Thomas Wild, Nikil Dutt, Rolf Ernst, Andreas Herkersdorf, Fadi J. Kurdahi |
DATE | 8 |
| 2018 | FlueNT10G: A Programmable FPGA-based Network Tester for Multi-10-Gigabit EthernetabstractWe present FlueNT10G, an open-source FPGA-based network tester for precise replay of network traces, as well as for accurate packet capture and round-trip latency measurements. FlueNT10G streams replay and capture data between the host system and the FPGA board during active tests. It enables continuous measurements without being constrained by the memory capacity of the FPGA board. FlueNT10G is able to concurrently replay and capture traffic on three 10 Gbit/s network interfaces for all packet sizes. When operated exclusively in replay or capture mode, throughput increases to 4x 10 Gbit/s. Our design yields a temporal resolution of 6.4 ns for precise traffic pattern generation, as well as for accurate arrival timestamping and latency measurements. On the software-side, FlueNT10G is complemented by an API enabling the programmable execution of reproducible network measurements. Targeting the automated performance evaluation of different virtualized network function configurations, the API further integrates access to a bidirectional side-band channel for device-under-test reconfiguration and status feedback. FlueNT10G has been implemented on the NetFPGA-SUME platform (Xilinx Virtex-7 XC7VX690T) with an FPGA resource utilization of no more than 25%, which leaves sufficient capacity available for future design extensions. Andreas Oeldemann, Thomas Wild, Andreas Herkersdorf |
FPL | 2 |
| 2018 | Platform-Centric Self-Awareness as a Key Enabler for Controlling Changes in CPSabstractFuture cyber-physical systems will host a large number of coexisting distributed applications on hardware platforms with thousands to millions of networked components communicating over open networks. These applications and networks are subject to continuous change. The current separation of design process and operation in the field will be superseded by a life-long design process of adaptation, infield integration, and update. Continuous change and evolution, application interference, environment dynamics and uncertainty lead to complex effects which must be controlled to serve a growing set of platform and application needs. Self-adaptation based on self-awareness and self-configuration has been proposed as a basis for such a continuous in-field process. Research is needed to develop automated in-field design methods and tools with the required safety, availability, and security guarantees. The paper shows two complementary use cases of self-awareness in architectures, methods, and tools for cyber-physical systems. The first use case focuses on safety and availability guarantees in self-aware vehicle platforms. It combines contracting mechanisms, tool based self-analysis and self-configuration. A software architecture and a runtime environment executing these tools and mechanisms autonomously are presented including aspects of self-protection against failures and security threats. The second use case addresses variability and long term evolution in networked MPSoC integrating hardware and software mechanisms of surveillance, monitoring, and continuous adaptation. The approach resembles the logistics and operation principles of manufacturing plants which gave rise to the metaphoric term of an Information Processing Factory that relies on incremental changes and feedback control. Both use cases are investigated by larger research groups. Despite their different approaches, both use cases face similar design and design automation challenges which will be summarized in the end. We will argue that seemingly unrelated research challenges, such as in machine learning and security, could also profit from the methods and superior modeling capabilities of self-aware systems. Mischa Möstl, Johannes Schlatow, Rolf Ernst, Nikil Dutt, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Fadi J. Kurdahi, Thomas Wild, Armin Sadighi, Andreas Herkersdorf |
Proc. IEEE | 8 |
| 2017 | A non-intrusive, operating system independent spinlock profiler for embedded multicore systemsabstractLocks are widely used as a synchronization method to guarantee the mutual exclusion for accesses to shared resources in multi-core embedded systems. They have been studied for years to improve performance, fairness, predictability etc. and a variety of lock implementations optimized for different scenarios have been proposed. In practice, applying an appropriate lock type to a specific scenario is usually based on the developer's hypothesis, which could mismatch the actual situation. A wrong lock type applied may result in lower performance and unfairness. Thus, a lock profiling tool is needed to increase the system transparency and guarantee the proper lock usage. In this paper, an operating-system-independent lock profiling approach is proposed as there are many different operating systems in the embedded field. This approach detects lock acquisition and lock releasing using hardware tracing based on hardware-level spinlock characteristics instead of specific libraries or APIs. The spinlocks are identified automatically; lock profiling statistics can be measured and performance-harmful lock behaviors are detected. With this information, the lock usage can be improved by the software developer. A prototype as a Java tool was implemented to conduct hardware tracing and analyze locks inside applications running on the Infineon AURIX microcontrollers. Lin Li 0046, Philipp Wagner 0001, Albrecht Mayer, Thomas Wild, Andreas Herkersdorf |
DATE | 4 |
| 2017 | Adaptive Reliability for Fault Tolerant Multicore SystemsabstractIn an era of continuously shrinking technology and escalating power density, Multiprocessor System on Chips (MPSoCs) suffer from a growing prominence of device defects and increase of dependability-related issues. This paper tackles the dependability challenge by suggesting an adaptive reliability enhancement strategy for multicore systems. We dynamically adapt the reliability enhancement to the actual tasks requirements as well as cores runtime operating conditions. As reliability improvement may adversely affect the parameters of embedded systems, we suggest a runtime recovery method. In fact, we implement a 3-mode mapping technique to limit redundancy overheads through judicious task migrating and dropping. Our experiments show promising results in terms of error mitigation with controllable power and thermal overheads. Ihsen Alouani, Thomas Wild, Andreas Herkersdorf, Smaïl Niar |
DSD | 2 |
| 2017 | A Divide and Conquer State Grouping Method for Bitmap Based Transition CompressionabstractMember State Bitmask Technique (MSBT) is a hardware oriented transition compression technique which can compress the redundant transitions in a finite automaton. While the compressed automaton is stored in on-chip memories; a dedicated hardware accelerator performs signature matching by comparing the network streams against the compressed automaton at line rate. The MSBT consists of three functional steps which include the intra-state transition compression, state grouping and the inter-state transition compression. The state grouping algorithm which is currently used in MSBT is not compression aware and results in sub-optimal transition compression. To address this weakness, a compression aware Divide and Conquer state grouping method is proposed in this paper, which can efficiently group states that improves the transition compression in MSBT. Experimental evaluation of the proposed state grouping method, results in a reduced on-chip memory usage of the order of 10-30%. The reduction in the memory usage allows to accommodate more signatures in on-chip memories and perform signature matching with them at line rate. Subramanian Shiva Shankar, Pinxing Lin, Andreas Herkersdorf, Thomas Wild |
PDCAT | 4 |
| 2017 | DiaSys: Improving SoC insight through on-chip diagnosis
Philipp Wagner 0001, Thomas Wild, Andreas Herkersdorf |
J. Syst. Archit. | 2 |
| 2017 | Efficient task spawning for shared memory and message passing in many-core architectures
Aurang Zaib, Thomas Wild, Andreas Herkersdorf, Jan Heisswolf, Jürgen Becker 0001, Andreas Weichslgartner, Jürgen Teich |
J. Syst. Archit. | 2 |
| 2016 | Resolving Performance Interference in SR-IOV Setups with PCIe Quality-of-Service ExtensionsabstractPCI Express (PCIe) Single Root I/O Virtualization (SR-IOV) enables low latency and high performance virtualization of I/O devices. It has been embraced in cloud computing and is considered a promising foundation for sharing I/O in future multi-core embedded and mixed-criticality systems. Unfortunately, SR-IOV is vulnerable to Denial-of-Service (DoS) attacks, which cause performance interference. For cloud computing, an approach that mitigates ongoing attacks via software scheduling has been proposed. However, for embedded and mixed-criticality systems, solutions that go beyond mitigation are preferred. In this paper, we propose two integrated hardware architectures that completely prevent DoS attacks. As a foundation, we utilize optional Quality-of-Service (QoS) extensions from the PCIe specification. We determine which QoS extensions are needed, and show how virtualized multi-core CPUs need to implement and interface them (an aspect explicitly not covered in the PCIe specification) to enable DoS protection. The two proposed architectures are optimized for different goals, scheduling freedom or minimal hardware costs. As PCIe QoS is absent from current hardware, we evaluate our architectures with a QoS-enabled SystemC model of a real-world lab-setup. Results show that both architectures successfully prevent DoS attacks. To the best of our knowledge, we are the first to explore and evaluate feasibility of PCIe QoS for SR-IOV DoS prevention. Andre Oliver Richter, Christian Herber, Thomas Wild, Andreas Herkersdorf |
DSD | 3 |
| 2016 | A Rule-based Methodology for Hardware Configuration Validation in Embedded SystemsabstractAs the complexity of multicore SoCs increases, more potential system issues are arising. Hardware-related configuration issues are becoming more complicated owing to the introduction of more cores and various complex peripherals. Considering the complexity of multicore programming, consultation of the main source of guidance, i.e. the user manual, is not an efficient approach to identify such problems. Improper hardware-related configurations could lead to either functional or performance issues. Some of these issues are even subtle and hard to detect. Therefore, a rule-based validation methodology is proposed to deal with hardware-related configuration issues in an efficient and reliable way. Hardware trace is applied in this methodology to detect issues even before symptoms appear. The method directly observes the register accesses and detects bugs based on trace data. It is independent of the application as long as they are run on the given platform, which means the same method implementation could be applied to any applications on the same platform. In this paper, an initial proof-of-concept for the proposed methodology has been implemented and demonstrated on the Infineon TC29 device. Lin Li 0046, Philipp Wagner 0001, Ramesh Ramaswamy, Albrecht Mayer, Thomas Wild, Andreas Herkersdorf |
SCOPES | 5 |
| 2015 | A Hardware/Software Approach for Mitigating Performance Interference Effects in Virtualized Environments Using SR-IOVabstractSingle Root I/O Virtualization (SR-IOV) is an extension to the PCI Express (PCIe) standard that allows virtual machines (VMs) to directly access shared I/O devices without host involvement. This enabled SR-IOV to become the best-performing solution for virtual I/O to date, which lead to its commercial adoption, e.g., In the Amazon EC2. On the downside, a malicious VM can exploit the direct access to an SR-IOV device by flooding it with PCIe packets. This results in a congestion on the PCIe interconnect, which leads to performance interference effects between the malicious VM, concurrent VMs and even the host. In this paper, we present a hardware/software approach that detects and mitigates such Denial-of-Service (DoS) attacks. On the hardware side, we propose monitoring extensions within SR-IOV devices that distinguish legal device use from malicious device use by observing the rate of incoming PCIe transactions at VM granularity. Malicious VMs are reported to the host via interrupts. On the software side, performance interference effects can then be mitigated by dynamically adjusting the host's scheduling of the malicious VM or even shutting it down. We implement a prototype with a commercial off-the-shelf SR-IOV Ethernet controller and an FPGA board. On it, we demonstrate that appropriate scheduling of malicious VMs successfully mitigates interference effects for three cloud-relevant benchmarks. For example, Memcached is restored to 99.4% of baseline performance (compared to 61.8% without our extensions). In contrast to QoS features proposed in the PCIe 3.0 standard, our solution is more flexible. Additionally, it can be realized as an add-on to existing misuse detection hardware like the Intel Malicious Driver Detection (MDD). Andre Oliver Richter, Christian Herber, Stefan Wallentowitz, Thomas Wild, Andreas Herkersdorf |
CLOUD | 4 |
| 2015 | Real-time capable CAN to AVB ethernet gateway using frame aggregation and scheduling
Christian Herber, Andre Oliver Richter, Thomas Wild, Andreas Herkersdorf |
DATE | 3 |
| 2015 | A hardware-based multi-objective thread mapper for tiled manycore architecturesabstractThread mapping is typically performed as an integral part of cooperative or pre-emptive operating system (OS) scheduling in order to share the processor core(s) among competing applications. Schedulers usually follow a single-objective performance optimization, such as maximizing core utilization or satisfying deadlines by the prioritization of threads. Meeting multiple orthogonal objectives, like performance vs. power or thermal resilience, in the era of manycore processors is a challenge because of the associated scalability and thread management overhead. We tackle these challenges by employing a two stage thread management strategy. In the first stage (not covered in this short paper), threads are assigned to regions or compute tiles. For the second stage we introduce in this paper the TCU (Thread Control Unit), a configurable, low latency, low overhead hardware thread mapper that takes various runtime sensor parameters into account. It can map threads within a small and bounded number of clock cycles in round robin, single or multi-objective manner. TCU is designed to consider not just load balancing or performance criteria but also physical constraints like power budgets, temperature limits and reliability aspects. TCU macro achieves 150K thread mappings per second on a tiled MPSoC FPGA prototype while operating at moderate 50 Mz. Evaluations of different mapping policies show that multi-objective thread mapping provides about 10 to 40% less mapping latency for periodic and bursty traffic compared to single-objective or round robin schemes. FPGA and ASIC syntheses reveal a 9% hardware overhead for the TCU on a four core compute tile. Ravi Kumar Pujari, Thomas Wild, Andreas Herkersdorf |
ICCD | 2 |
| 2015 | Denial-of-Service attacks on PCI passthrough devices: Demonstrating the impact on network- and storage-I/O performance
Andre Oliver Richter, Christian Herber, Thomas Wild, Andreas Herkersdorf |
J. Syst. Archit. | 3 |
| 2014 | System integration - The bridge between More than Moore and More MooreabstractSystem Integration using 3D technology is a very promising way to cope with current and future requirements for electronic systems. Since the pure shrinking of devices (known as “More Moore”) will come to an end due to physical and economic restrictions, the integration of systems (e.g. by stacking dies, or by adding sensor functions) shows a way to maintain the growth in complexity as well as in diversity which is necessary for future applications. This so called “More than Moore” approach complements the conventional SoC product engineering. This paper gives insights in System Integration design challenges from different perspectives, ranging from design technology over MEMS product engineering and 3D interconnect to automotive cyber physical systems. Andy Heinig, Manfred Dietrich, Andreas Herkersdorf, Felix Miller, Thomas Wild, Kai Hahn, Armin Grünewald, Rainer Brück 0001, Steffen Krohnert, Jochen Reisinger |
DATE | 5 |
| 2014 | Dependable task and communication migration in tiled manycore system-on-chipabstractPower densities and thermal hotspots are a major concern for the dependability of future multi-processor systemon- chip. They can lead to transient faults affecting the functionality in the short term and can cause permanent damage of a device. The dependability problem can be tackled on different layers such as technology hardening or application awareness. This work is based on an approach that addresses the issue for tile-based manycore system-on-chip on software and architecture layer. An agent-based system management employs task migration to react to thermal hotspots and pro-actively avoid them. The inter-task communication plays an important role as communication channels need to be migrated accordingly. The presented work focuses on the issue of communication migration and is based on the idea of handling it transparently to the task migration. Network-on-chip protection switching techniques have been introduced before and in this paper we evaluate the potential and bottlenecks of such methods in a realistic platform. Stefan Wallentowitz, Stefan Rosch, Thomas Wild, Andreas Herkersdorf, Volker Wenzel, Jörg Henkel |
FDL | 3 |
| 2014 | A network virtualization approach for performance isolation in controller area network (CAN)abstractAn important trend in automotive CPS is the shift from federated to integrated IT architectures, where multiple functions are consolidated on shared electronic resources instead of distributed electronic control units (ECUs). It is driven by increasing complexity, cost and installation space requirements of todays architectures. However, side-by-side integration of mixed-criticality functions poses new challenges with respect to safety and security. To achieve isolated performance for multiple integrated partitions with different criticalities, an efficient separation within computing and communication resources is required. This paper introduces a network virtualization approach for CAN, which enables concurrence of mixed-criticality communication on a single physical CAN bus through a strict performance isolation. We present a design concept as well as a prototypical implementation. The feasibility of our approach is demonstrated by an analytic evaluation of message latencies and through experimental case studies. Christian Herber, Andre Oliver Richter, Thomas Wild, Andreas Herkersdorf |
RTAS | 3 |
| 2013 | AUTO-GS: Self-Optimization of NoC Traffic through Hardware Managed Virtual ConnectionsabstractNetworks-on-Chip have shown their scalability for future many-core systems on chip. In real world scenarios, where multiple applications are being executed over a shared NoC based platform, efficient utilization of Networks-on-Chip resources becomes challenging. Methodologies are required to ensure better utilization of NoC, especially in the scenarios, where the communication patterns of NoC traffic are difficult to predict before run-time. In this paper, we propose a self-optimization mechanism which detects frequent communication by monitoring communication patterns at run-time and uses this information to establish virtual connections autonomously. Communication monitoring and connection establishment are realized in hardware. Hardware managed virtual connections lead to better utilization of NoC resources and reduce the communication latencies suffered by applications. In addition, energy consumption by the communication infrastructure is reduced. The proposed concept is investigated through simulation of real world application scenarios. The simulation results highlight the performance improvement and synthesis results show the low area overhead of the proposed hardware implementation. Aurang Zaib, Jan Heisswolf, Andreas Weichslgartner, Thomas Wild, Jürgen Teich, Jürgen Becker 0001, Andreas Herkersdorf |
DSD | 4 |
| 2013 | Virtual networks - distributed communication resource management
Jan Heisswolf, Aurang Zaib, Andreas Weichslgartner, Ralf König 0001, Thomas Wild, Jürgen Teich, Andreas Herkersdorf, Jürgen Becker 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2012 | Invasive manycore architecturesabstractThis paper introduces a scalable hardware and software platform applicable for demonstrating the benefits of the invasive computing paradigm. The hardware architecture consists of a heterogeneous, tile-based manycore structure while the software architecture comprises a multi-agent management layer underpinned by distributed runtime and OS services. The necessity for invasive-specific hardware assist functions is analytically shown and their integration into the overall manycore environment is described. Jörg Henkel, Andreas Herkersdorf, Lars Bauer, Thomas Wild, Michael Hübner 0001, Ravi Kumar Pujari, Artjom Grudnitsky, Jan Heisswolf, Aurang Zaib, Benjamin Vogel, Vahid Lari, Sebastian Kobbe |
ASP-DAC | 4 |
| 2012 | TSV-virtualization for Multi-protocol-Interconnect in 3D-ICsabstractThrough Silicon Vias (TSVs) are the method of choice to realize vertical connections between different chip layers in three dimensional Integrated Circuits (3D-ICs). These TSVs offer a fast connection and due to their short wire length, only a small capacitive load to the driving circuitry. On the other hand TSVs consume a relative large amount of chip area and as TSV-count increases the overall yield generally drops due to TSV manufacturing difficulties. As a result of the low capacitance, TSVs can be clocked much higher than conventional intra-layer links. To fully utilize the TSV-based vertical bandwidth we propose using them in a multiplexed manner and share them between several virtual links. On top of that we propose using TSVs to stretch state-of-the art interconnects like busses, crossbars or NoCs to other silicon layers in the 3D stack. This reduces TSV count and gives designers the opportunity to easily migrate from 2D to 3D designs and to largely benefit from reuse of existing IP blocks and interconnection schemes. Felix Miller, Thomas Wild, Andreas Herkersdorf |
DSD | 2 |
| 2012 | An integrated simulation framework for invasive computing
Michael Gerndt, Frank Hannig, Andreas Herkersdorf, Andreas Hollmann, Marcel Meyer, Sascha Roloff, Josef Weidendorfer, Thomas Wild, Aurang Zaib |
FDL | 8 |
| 2012 | A framework for Open Tiled Manycore System-On-ChipabstractTiled manycore architectures have become dominant for the integration of tens or even a hundred processor cores on a chip. While commercial products are increasingly available, research on the hardware of such platforms and especially prototyping often rely on building such a platform from scratch or is bound to abstract simulation. In this paper we present the Open Tiled Manycore System-on-Chip (Op-TiMSoC) which is a library-based tool flow that helps generating a tiled manycore platform based on a library of open standard components. OpTiMSoC allows for research and prototyping of both shared memory and distributed memory platforms. It includes LISNoC which is a flexible NoC implementation. An OpTiMSoC system can easily be generated based on the publicly available repository and prototyped on an FPGA. As exemplary targets we evaluated the usage of different FPGA boards and an emulation platform. Stefan Wallentowitz, Andreas Lankes, Aurang Zaib, Thomas Wild, Andreas Herkersdorf |
FPL | 4 |
| 2012 | Benefits of selective packet discard in networks-on-chipabstractToday, Network on Chip concepts principally assume inherent lossless operation. Considering that future nanometer CMOS technologies will witness increased sensitivity to all forms of manufacturing and environmental variations (e.g., IR drop, soft errors due to radiation, transient temperature induced timing problems, device aging), efforts to cope with data corruption or packet loss will be unavoidable. Possible counter measures against packet loss are the extension of flits with ECC or the introduction of error detection with retransmission. We propose to make use of the perceived deficiency of packet loss as a feature. By selectively discarding stuck packets in the NoC, a proven practice in computer networks, all types of deadlocks can be resolved. This is especially advantageous for solving the problem of message-dependent deadlocks, which otherwise leads to high costs either in terms of throughput or chip area. Strict ordering, the most popular approach to this problem, results in a significant buffer overhead and a more complex router architecture. In addition, we will show that eliminating local network congestions by selectively discarding individual packets also can improve the effective throughput of the network. The end-to-end retransmission mechanism required for the reliable communication, then also provides lossless communication for the cores. Andreas Lankes, Thomas Wild, Stefan Wallentowitz, Andreas Herkersdorf |
ACM Trans. Archit. Code Optim. | 2 |
| 2010 | A folded pipeline network processor architecture for 100 Gbit/s networksabstractEthernet, although initially conceived as a Local Area Network technology, has been steadily making inroads into access and core networks. This has led to a need for higher link speeds, which are now reaching 100 Gbit/s. Packet processing at this rate represents a significant challenge, that needs to be met efficiently, while minimizing power consumption and chip area. This level of throughput favours a pipelined approach, thus this paper takes a traditional pipeline and breaks it down to mini-pipelines, which can perform coarse-grained processing (like process an MPLS label to completion). These mini-pipelines are then parellelized and used to construct a folded pipeline architecture, which augments the traditional approach by significantly reducing power consumption, a key problem in future routers. The paper compares the two approaches, discusses their advantages and disadvantages and demonstrates by quantitative measures that the folded pipeline architecture is the better solution for 100 Gbit/s processing. Kimon Karras, Thomas Wild, Andreas Herkersdorf |
ANCS | 2 |
| 2010 | An Application-Aware Load Balancing Strategy for Network Processors
Rainer Ohlendorf, Michael Meitinger, Thomas Wild, Andreas Herkersdorf |
HiPEAC | 3 |
| 2010 | Comparison of Deadlock Recovery and Avoidance Mechanisms to Approach Message Dependent Deadlocks in On-chip NetworksabstractWith the transition from buses to on-chip networks in SoCs the problem of deadlocks in on-chip interconnects arises. Deadlocks can be caused by routing cycles in the network, or by message dependencies, even if the network itself is actually free of routing cycles. Two basic approaches to counter message dependent deadlocks exist: deadlock avoidance, which is most popular in NoCs, and deadlock recovery, which has until now only been used in parallel computer networks. Deadlock recovery promises low buffer space requirements and does not impose restrictions on connections between individual communication partners. For this study, we have adapted a deadlock recovery scheme for the use in NoCs and compared it to strict ordering as a representative of deadlock avoidance in terms of throughput and buffer space. The results show significant buffer space savings for deadlock recovery, however, at the cost of reduced data throughput. Andreas Lankes, Thomas Wild, Andreas Herkersdorf, Sören Sonntag, Helmut Reinig |
NOCS | 2 |
| 2009 | Hierarchical NoCs for Optimized Access to Shared Memory and IO ResourcesabstractThe concept of on-chip networks (NoCs) has been developed to cope with the increasing communication requirements in systems-on-chip (SoCs) consisting of an ever-growing number of cores. Proposals of NoC architectures are often made assuming evenly distributed traffic, where all tiles receive and produce the same amount of traffic. However, in real systems specific, communication centric tiles, for example off-chip memory controllers or other data IO interfaces, consume and generate a significant part of the overall traffic. In this paper we propose hierarchical NoC topologies to improve access to this type of shared resources. The hierarchical networks may be built from different types of sub-networks, e.g. meshes, rings, crossbars and buses. We investigate different hierarchical network architectures and compare them to the popular 2D mesh topology in terms of hop count, latency and network throughput. Our results show that the proposed approach allows to significantly reduce network latencies to these communication centric tiles. Andreas Lankes, Thomas Wild, Andreas Herkersdorf |
DSD | 2 |
| 2008 | Buffer allocation for advanced packet segmentation in Network ProcessorsabstractIn current network processors, incoming variable-length packets are sliced using only one small segment size and then stored in the buffer. Inconveniently, short data bursts are inadequate for accessing SDRAM, commonly used for packet buffers, due to high activation and pre-charging latencies. Using large segment sizes is not optimal either because though it increases memory bandwidth, the benefit comes at the price of a heavy reduction in storing efficiency. A good solution to achieve simultaneously high performance and memory utilization consists in storing a single packet segmented using multiple segment sizes. In this paper, we study how to allocate memory for these different-sized segments in an efficient way. First we analyze the appropriate segment pool size for a multitude of traffic scenarios. Our experiments show that simple static buffer allocation does not always suffice as different segment pools may be exhausted depending on traffic. Hence we introduce a method for handling multiple segment pools not only in a static but also in a dynamic way, taking advantage of a new set of control structures based on a combination of bitmaps and linked lists. We demonstrate that our method achieves a huge reduction in control buffer size requirements in comparison to state-of-the-art control structures, together with decreasing the average number of accesses to control data. Daniel Llorente, Kimon Karras, Thomas Wild, Andreas Herkersdorf |
ASAP | 3 |
| 2008 | Network processorsabstractTraditional design of network processors is complicated by two conflicting demands, flexibility and performance. On the one side, network processors should be flexible enough to adapt to changing protocols and varying traffic profiles, on the other side they have to cope with increasing data rates of network links. This demonstrator shows that runtime reconfigurable systems have the potential to optimise both criteria without affecting each other negatively. The demonstrator addresses edge router applications and consists of two independently developed subsystems, the FlexPath NP architecture designed at the TU Munchen and the Dyna-CORE architecture designed at the University of Lubeck. Thilo Pionteck, Roman Koch, Carsten Albrecht, Erik Maehle, Michael Meitinger, Rainer Ohlendorf, Thomas Wild, Andreas Herkersdorf |
FPL | 7 |
| 2008 | A Processing Path Dispatcher in Network Processor MPSoCsabstractMulti-field packet classification problems discussed in the literature are typically constrained to the Internet five-tuple and primarily address the problem of network quality-of-service (QoS) support and access control. In this paper, we present a solution for a classification problem that is used for optimized packet assignment to different data paths within a network processor system-on-chip (SoC). In contrast to the five-tuple-based rules discussed in the prior art, our problem has rules that consider a larger set of fields from the packet header. However, for each individual rule a different sub-set of fields is relevant and the number of rules is smaller. Based on a specification of the usage case for our classifier we derive heterogeneous decision graph algorithm (HDGA), a heuristic approach to construct a decision tree classifier that integrates external lookup results for certain types of rules. We evaluate various parameters for optimizing the proposed decision tree and present simulation results to show the scalability of HDGA for typical problem sizes. This paper is concluded with the results of an implementation on our field-programmable gate-array (FPGA)-based prototyping platform. Rainer Ohlendorf, Michael Meitinger, Thomas Wild, Andreas Herkersdorf |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | Power Estimation of Time Variant SoCs with TAPESabstractDuring the design process of modern SoCs (systems on chip), design tools and methods are required for the exploration of promising solutions. Evaluation criteria in this process are performance and often also power consumption. The design space is expanded by a trend towards time variant SoCs, which adapt their behaviour at run time to improve reliability or power consumption. This paper presents an extension to the TAPES system simulator in order to enable not only the exploration of architectures but also the investigation of power minimization strategies. The usefulness of the simulator is demonstrated in an architecture exploration of a network processor. Andreas Lankes, Thomas Wild, Johannes Zeppenfeld |
DSD | 2 |
| 2007 | Simulated and measured performance evaluation of RISC-based SoC platforms in network processing applications
Rainer Ohlendorf, Thomas Wild, Michael Meitinger, Holm Rauchfuss, Andreas Herkersdorf |
J. Syst. Archit. | 2 |
| 2006 | Performance evaluation for system-on-chip architectures using trace-based transaction level simulationabstractThe ever increasing complexity and heterogeneity of modern system-on-chip (SoC) architectures make an early and systematic exploration of alternative solutions mandatory. Efficient performance evaluation methods are of highest importance for a broad search in the solution space. In this paper we present an approach that captures the SoC functionality for each architecture resource as sequences of trace primitives. These primitives are translated at simulation runtime into transactions and superposed on the system architecture. The method uses SystemC as modeling language, requires low modeling effort and yet provides accurate results within reasonable turnaround times. A concluding application example demonstrates the effectiveness of our approach Thomas Wild, Andreas Herkersdorf, Rainer Ohlendorf |
DATE | 1 |
| 2003 | A Constructive Algorithm with Look-Ahead for Mapping and Scheduling of Task Graphs with Conditional EdgesabstractConstructive algorithms for mapping and scheduling take advantage of short execution times. However, since decisions for the mapping have to be made at a time when not all information of dynamic effects is available, unfavorable situations can arise which result in a degraded performance. In this paper, an enhancement for a constructive algorithm is shown to be effective for real-world applications. Improvements of the performance can be achieved by considering additional information, such as a look ahead of mandatory transfers. Additionally, an algorithm to determine mutual exclusion for arbitrary connected nodes is shown. Winthir Brunnbauer, Thomas Wild, Jürgen Foag, Nuria Pazos |
DSD | 2 |
| 2002 | Predictive methodology for high-performance networkingabstractNetworking devices have to offer short processing latencies and flexibility concerning the supported networking protocols and applications. The input packet-processing flow in conventional networking node devices follows a serial or a layer-specific pseudo-parallel method, serialized by the data dependencies of the inherent layer protocol types. This article describes a new methodology, which is based on the principles of branch prediction and speculative execution in microprocessors, for latency reduced input packet processing in a networking device. Through the use of protocol stack prediction, in combination with speculative protocol layer processing, a networking system may realize a mean system processing time reduction of up to 40 percent in a real networking environment in comparison to conventional processing methodologies. Jürgen Foag, Thomas Wild, Nuria Pazos, Winthir Brunnbauer |
ISCC | 2 |