EDBT 2026 Demo / reviewers in the wild / expert
Kees Goossens
dblp:g/KeesGoossens · also Kees G. W. Goossens
· DBLP profile ↗
132ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0001-7536-4050ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 102 · 8 first-author · 9 since 2021Software engineering, systems software and programming languages · 31 · 3 first-author · 1 since 2021Computer networks · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Time-Aware Shaper Scheduling for In-Vehicle Networks via Deep Reinforcement LearningabstractModern vehicles increasingly rely on distributed computing platforms that exchange large volumes of sensor and control data with strict timing requirements. Ensuring that this traffic meets its deadlines over Ethernet-based in-vehicle networks requires Time-Sensitive Networking (TSN) and, in particular, effective configuration of the Time-Aware Shaper (TAS). However, generating and updating TAS schedules that remain valid as traffic patterns evolve is an NP-hard problem that traditional optimization or heuristic methods address only partially. This paper introduces a Deep Reinforcement Learning (DRL) scheduler that learns to configure TAS schedules directly from network state while preserving standard compliance through analytical validation. The proposed DRL scheduler encodes the scenario (network topology and workload) of the in-vehicle network using a Graph Neural Network (GNN) and learns scheduling policies that balance deadline satisfaction, latency, and resource utilization. Evaluation on a comprehensive benchmark shows that the proposed approach consistently outperforms state-of-the-art heuristics and a topology-specific DRL baseline, achieving higher success rate and lower delay while maintaining efficient bandwidth use. Once trained, it can adapt to new traffic scenarios within milliseconds, demonstrating the potential of the DRL-based scheduler as a foundation for adaptive and reliable communication in next-generation software-defined vehicles. Mohammad Parsa Karimi, Majid Nabi, Andrew Nelson 0001, Kees Goossens, Twan Basten |
IEEE Internet Things J. | 4 |
| 2025 | Deep-Reinforcement-Learning-Based Scheduler for Time-Aware Shaper in In-Vehicle NetworksabstractAs vehicles develop into software-defined platforms with powerful automated driving capabilities and driver support systems, their in-vehicle networks become significantly more complicated. A key technique for ensuring deterministic, low-latency connectivity for crucial data traffic in such settings is Time-Sensitive Networking (TSN), and specifically the Time-Aware Shaper (TAS). However, current TAS scheduling techniques have difficulty adjusting schedules to dynamically shifting traffic patterns and changing operating conditions. This paper presents an adaptive scheduler using Deep Reinforcement Learning (DRL), which aims to meet strict deadlines, reducing latency and providing near-ideal resource usage. Experimental results for different vehicle scenarios show that our DRL-based scheduler performs better in terms of success rate, low latency, and overall network performance than state-of-the-art heuristic algorithms such as earliest deadline first (EDF) scheduling. Mohammadparsa Karimi, Majid Nabi, Andrew Nelson 0001, Kees Goossens, Twan Basten |
VTC2025-Spring | 4 |
| 2025 | INSIM: A Modular Simulation Platform for TSN-based In-Vehicle NetworksabstractIn-vehicle networks (IVNs) are rapidly evolving to support increasingly complex automotive applications, demanding higher bandwidth and deterministic timing bounds. Time-Sensitive Networking (TSN) has emerged as a promising Ethernet-based technology that addresses these stringent requirements. However, evaluating TSN-based IVN strategies remains a challenge due to the lack of standardized benchmarks and simulation tools. This paper introduces INSIM, a modular simulation platform specifically designed for TSN-based IVNs, providing an intuitive graphical interface, an extensible plug-in architecture, and integrated benchmarking features. INSIM integrates analytical performance models and discrete-event simulations (as plug-ins), enhancing the workflow for engineers by refining topology design, adjusting parameters, conducting simulations, and assessing performance, while providing researchers with a flexible platform to plug in, analyze, and compare custom network resource managers or analytical performance models. Mohammadparsa Karimi, Majid Nabi, Andrew Nelson 0001, Kees Goossens, Twan Basten |
VTC2025-Fall | 4 |
| 2025 | Hardware Implementation of Projection-Aggregation Decoders for Reed-Muller CodesabstractThis paper presents the hardware architecture and implementation of two variants of projection-aggregation-based decoding of Reed-Muller (RM) codes, namely unique projection aggregation (UPA) and collapsed projection aggregation (CPA). Through thorough analysis and experimentation, we observe that the hardware implementation of UPA exhibits superior resource usage and reduced energy consumption compared to CPA for the iterative projection aggregation (IPA) decoder. This finding underscores a critical insight: reducing computational cost, in isolation, may not necessarily translate into hardware cost effectiveness. Marzieh Hashemipour-Nazari, Andrea Nardi-Dei, Kees Goossens, Alexios Balatsoukas-Stimming |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | NPTSN: RL-Based Network Planning with Guaranteed Reliability for In-Vehicle TSSDNabstractTo achieve strict reliability goals with lower redundancy cost, Time-Sensitive Software-Defined Networking (TSSDN) enables run-time recovery for future in-vehicle networks. While the recovery mechanisms rely on network planning to establish reliability guarantees, existing network planning solutions are not suitable for TSSDN due to its domain-specific scheduling and reliability concerns. The sparse solution space and expensive reliability verification further complicate the problem. We propose NPTSN, a TSSDN planning solution based on deep Reinforcement Learning (RL). It represents the domain-specific concerns with the RL environment and constructs solutions with an intelligent network generator. The network generator iteratively proposes TSSDN solutions based on a failure analysis and trains a decision-making neural network using a modified actor-critic algorithm. Extensive performance evaluations show that NPTSN guarantees reliability for more test cases and shortens the decision trajectory compared to state-of-the-art solutions. It reduces the network cost by up to 6.8x in the performed experiments. Weijiang Kong, Majid Nabi, Kees Goossens |
DSN | 3 |
| 2023 | Recursive/Iterative Unique Projection-Aggregation Decoding of Reed-Muller CodesabstractWe describe recursive unique projection-aggregation (RUPA) decoding and iterative unique projection-aggregation (IUPA) decoding of Reed-Muller (RM) codes, which remove non-unique projections from the recursive projection-aggregation (RPA) and iterative projection-aggregation (IPA) algorithms respectively. We show that these algorithms have competitive error-correcting performance while requiring up to 95% projections lower than the baseline RPA algorithm. Marzieh Hashemipour-Nazari, Renate Debets, Kees Goossens, Alexios Balatsoukas-Stimming |
ICASSP | 3 |
| 2023 | Decentralized Configuration of TSCH-Based IoT Networks for Distinctive QoS: A Deep Reinforcement Learning ApproachabstractThe IEEE 802.15.4 time-slotted channel hopping (TSCH) is widely used as a reliable, low-power, and low-cost communication technology for many industrial Internet of Things (IoT) networks. In many applications, Quality-of-Service (QoS) requirements are different for heterogeneous nodes, necessitating nonequal parameter settings per node. This results in a very large configuration space making space exploration complex and time consuming. Moreover, network state and QoS requirements may change over time. Thus, run-time configuration mechanisms are needed for making decisions about proper node settings to consistently satisfy diverse and dynamic QoS requirements. In this article, we propose a run-time decentralized self-optimization framework based on deep reinforcement learning (DRL) for parameter configuration of a multihop TSCH network. DRL adopts neural networks as approximate functions to speed up the process of converging to QoS-satisfying configurations. Simulation results show that our proposed framework enables the network to use the right configuration settings according to the diverse QoS demands of different nodes. Moreover, it is shown that the convergence time of the learning framework is in the order of a few minutes which is acceptable for many IoT applications. Hamideh Hajizadeh, Majid Nabi, Kees Goossens |
IEEE Internet Things J. | 3 |
| 2023 | Pipelined Architecture for Soft-Decision Iterative Projection Aggregation Decoding for RM CodesabstractThe recently proposed recursive projection-aggregation (RPA) decoding algorithm for Reed-Muller codes has received significant attention as it provides near-ML decoding performance at reasonable complexity for short codes. However, its complicated structure makes it unsuitable for hardware implementation. Iterative projection-aggregation (IPA) decoding is a modified version of RPA decoding that simplifies the hardware implementation. In this work, we present a flexible hardware architecture for the IPA decoder that can be configured from fully-sequential to fully-parallel, thus making it suitable for a wide range of applications with different constraints and resource budgets. Our simulation and implementation results show that the IPA decoder has 41% lower area consumption, 44% lower latency, four times higher throughput, but currently seven times higher power consumption for a code with block length of 128 and information length of 29 compared to a state-of-the-art polar successive cancellation list (SCL) decoder with comparable decoding performance. Marzieh Hashemipour-Nazari, Yuqing Ren, Kees Goossens, Alexios Balatsoukas-Stimming |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | An Evaluation Framework for Vision-in-the-Loop Motion Control SystemsabstractIndustrial applications and processes such as quality inspections, pick and place operations, and semiconductor manufacturing require accurate positioning control for achieving the high throughput of the assembly machines. Vision-based sensing is considered to be a potential means to achieve robust positioning control which is referred to as a vision-in-the-loop (VIL) system. In such motion systems, the point-of-control and the point-of-interest are often different due to several physical factors. In this case, validation of a system is done only when a machine prototype is available. A physical prototype is often expensive and infeasible in real-life. This paper proposes an evaluation framework for VIL systems targeting a predictable multi-core embedded platform. The presented framework offers model-in-the-loop (MIL), software-in-the-loop (SIL), and processor-in-the-loop (PIL) simulation features for evaluating the closed-loop performance of industrial motion control systems. As a deployment platform, we consider a predictable embedded platform CompSOC. The predictable nature of the CompSOC platform guarantees periodic and deterministic execution of the control applications and allows verification of the timing properties and performance of the VIL system. Additionally, the framework offers automatic code generation feature targeting the CompSOC platform. Closed-loop simulation setup models the system dynamics and camera position in the CoppeliaSim physics simulation engine and simulates the system software in C and MATLAB. CoppeliaSim runs as a server and MATLAB as a client in synchronous mode. We show the effectiveness of our framework using a vision-based motion control example. Chaitanya Jugade, Daniel Hartgers, Phan Dúc Anh, Sajid Mohamed, Mojtaba Haghi, Dip Goswami, Andrew Nelson 0001, Gijs van der Veen, Kees Goossens |
ETFA | 9 |
| 2022 | SynAVB: Route and Slope Synthesis Ensuring Guaranteed Service in Ethernet AVBabstractIn the automotive domain, Ethernet Audio Video Bridging (AVB) provides service guarantees for real-time traffic. Its configuration synthesis requires routing the flows and allocating the bandwidth for Credit Based Shaping (CBS). Current approaches typically employ deadline-oblivious bandwidth allocation and rely on routing to establish deadline guarantees. But the worst-case delay of AVB flows requires complex analysis. Thus, routing integrated with the delay analysis either supports limited Stream Reservation (SR) classes or suffers significant timing overhead. To enable efficient run-time flow setup, we propose SynAVB, a tool to synthesize the configuration of Ethernet AVB, which establishes deadline awareness in bandwidth allocation. SynAVB supports an arbitrary number of SR classes and processes them one by one. For the flows in an SR class, it first performs a deadline-oblivious routing based on Mixed Integer Linear Programming (MILP) to guarantee the necessary bandwidth required by the flows. Then, a deadline-aware slope allocation algorithm computes the bandwidth required on every link to satisfy the deadline of all flows. Our experiment demonstrates that, given the same routes, the deadline-aware bandwidth allocation results in fewer flows violating the deadline compared with other allocation strategies. Moreover, SynAVB can guarantee the deadline of more flows while its run-time is up to 14.3x faster than the state-of-the-art approaches. Weijiang Kong, Majid Nabi, Kees Goossens |
ETFA | 3 |
| 2022 | Run-time Per-Class Routing of AVB Flows in In-Vehicle TSN via Composable Delay AnalysisabstractTime-Sensitive Networking (TSN) is a promising solution for the next generation in-vehicle networks. To enable innovations in adaptability, there is a strong trend to integrate TSN with Software-Defined Networking (SDN). Software-defined TSN requires fast routing algorithms that guarantee the Worst-Case End-to-End Delay (WCED) of the Audio Video Bridging (AVB) flows. However, current AVB flow routing algorithms execute WCED analysis for an exponential amount of route candidates, which is not feasible at run-time. Current WCED analysis of the AVB flows depends on the specific flow setup of other traffic classes. When high-priority flows change, low-priority flows must be adjusted as well to maintain deadline guarantees. In this paper, we propose a per-class flow management scheme to enable run-time routing of the AVB flows. Our composable analysis computes a formal WCED bound of the AVB flows. It does not rely on the flow setup of other traffic classes. So, the WCED and routes can be computed for each traffic class independently and in parallel. Based on the analysis, we develop a fast heuristic algorithm to incrementally route the AVB flows with guaranteed deadlines. Our experiments demonstrate that, compared to the existing solutions, the composable analysis is 2.6x faster. The cost is that the WCED bound can be up to 31% higher. When given a similar run-time as existing solutions, the proposed routing algorithm can set up 1.6x more flows. The average time to set up one flow is within 0.14s. Weijiang Kong, Majid Nabi, Kees Goossens |
VTC Spring | 3 |
| 2021 | Modeling, implementation, and analysis of XRCE-DDS applications in distributed multi-processor real-time embedded systemsabstractThe Publish-Subscribe paradigm is a design pattern for transparent communication in many recent distributed applications. Data Distribution Service (DDS) is a machine-to-machine communication standard that aims to provide reliable, highperformance, inter-operable, and real-time data exchange based on publish-subscribe paradigm. However, the high resource requirement of DDS limits its usage in low-cost embedded systems. XRCE-DDS is a Client-Agent based standard to enable resource-constrained small embedded systems to connect to the DDS global data space. Current XRCE-DDS implementations suffer from dependencies with host operating systems, target only single processing units, and lack performance analysis methods. In this paper, we present a bare-metal implementation of XRCE-DDS standard on the CompSOC platform as an instance of Multi-Processor System on Chip (MPSoC). The proposed framework includes a hard real-time side hosting the XRCE-DDS Client, and a soft real-time side hosting the XRCE-DDS Agent. A Scenario Aware Data Flow (SADF) model is proposed to capture the dynamism of the system behavior in terms of different execution scenarios. We analyze the long-term expected value for throughput by capturing the probabilistic scenario switching using a proposed Markov model which is experimentally validated. Saeid Dehnavi, Dip Goswami, Martijn Koedam, Andrew Nelson 0001, Kees Goossens |
DATE | 5 |
| 2021 | A Deployment Framework for Quality-Sensitive Applications in Resource-Constrained Dynamic EnvironmentsabstractTraditional embedded systems and recent platforms used in emerging computing paradigms (e.g., fog computing) have resource limits and require their applications and services to be dynamically added (i.e., deployed) and removed at run-time. These applications often have non-functional (quality) requirements (e.g., end-to-end latency) which are only satisfied when sufficient resources are allocated to them. Hence, a run-time decision-maker is needed to optimize the deployments, in terms of resource budgets that are allocated to applications. Additionally, computing platforms have become heterogeneous in terms of their resources and the applications they execute. However, the existing deployment solutions are limited to specific resources and services. In this paper, we propose a run-time deployment framework that is more flexible in defining constraints and optimization goals and works with more heterogeneous resources and resource models than existing solutions. The framework is implemented on an embedded platform as a proof of concept. Shayan Tabatabaei Nikkhah, Marc Geilen, Dip Goswami, Martijn Koedam, Andrew Nelson 0001, Kees Goossens |
DSD | 6 |
| 2021 | Hardware Implementation of Iterative Projection-Aggregation Decoding of Reed-Muller CodesabstractIn this work, we present a simplification and a corresponding hardware architecture for hard-decision recursive projection-aggregation (RPA) decoding of Reed-Muller (RM) codes. In particular, we transform the recursive structure of RPA decoding into a simpler and iterative structure with minimal error-correction degradation. Our simulation results for RM(7,3) show that the proposed simplification has a small error-correcting performance degradation (0.005 in terms of channel crossover probability) while reducing the average number of computations by up to 40%. In addition, we describe the first fully parallel hardware architecture for simplified RPA decoding. We present FPGA implementation results for an RM(6,3) code on a Xilinx Virtex-7 FPGA showing that our proposed architecture achieves a throughput of 171 Mbps at a frequency of 80 MHz. Marzieh Hashemipour-Nazari, Kees Goossens, Alexios Balatsoukas-Stimming |
ICASSP | 2 |
| 2021 | CompROS: A composable ROS2 based architecture for real-time embedded robotic developmentabstractRobot Operating System (ROS) is a de-facto standard robot middleware in many academic and industrial use cases. However, utilizing ROS/ROS2 in safety-critical embedded applications with real-time requirement is challenging because of C1) Non-real-time underlying hardware, C2) No control on the host OS scheduler, C3) Unpredictable dynamic memory allocation, C4) High resource requirement, and C5) Unpredictable execution model for ROS nodes. In this paper, we address these limiting factors by proposing a hardwaresoftware architecture -CompROS- for ROS2 based robotic development in a Multi-Processor System on Chip (MPSoC) platform. The proposed hardware architecture consists of a Hard Real-Time (HRT) RISC-V based subsystem implemented in the Programmable Logic (PL) part of the MPSoC platform, a Soft Real-Time (SRT) ARM-based subsystem in the Processing System (PS) part of the MPSoC platform, and a Non-Real-Time (NRT) PC. While the proposed hardware architecture along with a partitioning layer overcomes the first two limiting factors, the rest are managed by the proposed multi-layer software architecture. We make a bare-metal implementation of XRCE-DDS standard for PL-PS communication, while peer-to-peer PL-PL communication is done through a proposed real-time publish-subscribe approach. The reliable communication for PS-PL communication is done through utilizing C-HEAP protocol. Further, we integrate ROS2 software layers on top of the proposed hardware and software layers. Finally, with respect to C5, we present a real-time execution model of ROS2 nodes by a mapping of ROS2 entities to CompROS entities, which is validated through experimental results. We run ROS2 middleware with an executable size of less than 200 KB on an MPSoC platform. Saeid Dehnavi, Martijn Koedam, Andrew Nelson 0001, Dip Goswami, Kees Goossens |
IROS | 5 |
| 2021 | Isolation of redundant and mixed-critical automotive applications: effects on the system architectureabstractFuture automotive systems, with Advanced Driving Assistance Systems and Autonomous Driving functionalities, will require fail-operational electronic systems. To achieve that, redundancy is a necessary technique, like in many other fields such as aviation. Moreover the applications have different safety requirements, from safety-critical related applications, for example for the driver replacement domain, to QoS-oriented applications, for example for the infotainment domain. Redundancy in mixed-criticality systems can be solved by physically separating system resources or by using isolated virtualized environments with e.g. hypervisors. There are costs associated to both solutions. In this work we describe a novel model we use to characterize a mixed-criticality automotive system and the analysis steps to obtain quantified metrics. The quantified metrics include cost, failure probability, total functional and communication loads, and total cable length, to compare the different solutions from a system-level perspective. We analyse the same set of mixed-criticality applications that represent a simplified automotive system in four scenarios. The architecture topology is either domain-based or zone-based, and we use either physical separation or virtualization to provide isolation. The obtained results show how the model and the analysis allows us to understand the trade-offs between the different solutions in specific applications scenarios, and how to vary the metrics used in the analysis to adapt to a different applications scenario. Alessandro Frigerio, Bart Vermeulen, Kees Goossens |
VTC Spring | 3 |
| 2021 | Reducing Library Characterization Time for Cell-aware Test while Maintaining Test QualityabstractAbstract Cell-aware test (CAT) explicitly targets faults caused by defects inside library cells to improve test quality, compared with conventional automatic test pattern generation (ATPG) approaches, which target faults only at the boundaries of library cells. The CAT methodology consists of two stages. Stage 1, based on dedicated analog simulation, library characterization per cell identifies which cell-level test pattern detects which cell-internal defect; this detection information is encoded in a defect detection matrix (DDM). In Stage 2, with the DDMs as inputs, cell-aware ATPG generates chip-level test patterns per circuit design that is build up of interconnected instances of library cells. This paper focuses on Stage 1, library characterization, as both test quality and cost are determined by the set of cell-internal defects identified and simulated in the CAT tool flow. With the aim to achieve the best test quality, we first propose an approach to identify a comprehensive set, referred to as full set, of potential open- and short-defect locations based on cell layout. However, the full set of defects can be large even for a single cell, making the time cost of the defect simulation in Stage 1 unaffordable. Subsequently, to reduce the simulation time, we collapse the full set to a compact set of defects which serves as input of the defect simulation. The full set is stored for the diagnosis and failure analysis. With inspecting the simulation results, we propose a method to verify the test quality based on the compact set of defects and, if necessary, to compensate the test quality to the same level as that based on the full set of defects. For 351 combinational library cells in Cadence’s GPDK045 45nm library, we simulate only 5.4% defects from the full set to achieve the same test quality based on the full set of defects. In total, the simulation time, via linear extrapolation per cell, would be reduced by 96.4% compared with the time based on the full set of defects. Min-Chun Hu 0002, Santosh Malagi, Joe Swenton, Jos Huisken, Kees Goossens, Erik Jan Marinissen |
J. Electron. Test. | 6 |
| 2021 | Interface Modeling for Quality and Resource Management
Martijn Hendriks, Marc Geilen, Kees Goossens, Rob de Jong, Twan Basten |
Log. Methods Comput. Sci. | 3 |
| 2020 | A Distributed Safety Mechanism using Middleware and Hypervisors for Autonomous VehiclesabstractAutonomous vehicles use cyber-physical systems to provide comfort and safety to passengers. Design of safety mechanisms for such systems is hindered by the growing quantity and complexity of SoCs (System-on-a-Chip) and software stacks required for autonomous operation. Our study tackles two challenges: (1) fault handling in an autonomous driving system distributed across multiple processing cores and SoCs, and (2) isolation of multiple software modules consolidated in one SoC. To address the first challenge, we extend the state-of-the-art E-Gas layered monitoring concept. Similar to E-Gas, our safety mechanism has function, controller and vehicle layers. We propose to distribute these safety layers on processors with different ASILs (Automotive Safety Integrity Level). Besides, we implement seif-test, fault injection and challenge-response protocols to detect faults at runtime in the safety mechanism itself. To facilitate distributed operation, our mechanism is built on top of the DDS (Data Distribution Service) software middleware for safety-critical embedded applications, as well as DDS-XRCE (eXtremely Resource Constrained Environment) for resource- constrained processor cores of the highest ASIL. To address the second challenge, our safety mechanism employs hardware- assisted hypervisors to isolate software modules and implement fail-silent behavior of faulty software stacks. We validate our safety mechanism on the NXP BiueBox hardware platform using the LG SVL simulator, Baidu Apollo software framework for autonomous driving, and Xen hypervisor. Our fault injection experiments demonstrate that the distributed safety mechanism successfully detects faults in an autonomous system and safely stops the vehicle when necessary. Tjerk Bijlsma, Andrii Buriachevskyi, Alessandro Frigerio, Yuting Fu, Kees Goossens, Ali Ors, Pieter J. van der Perk, Andrei Sergeevich Terechko, Bart Vermeulen |
DATE | 5 |
| 2020 | Parallel Implementation of Iterative Learning Controllers on Multi-core Platforms
Mojtaba Haghi, Yusheng Yao, Dip Goswami, Kees Goossens |
DATE | 4 |
| 2020 | A Performance Analysis Framework for Real-Time Systems Sharing Multiple ResourcesabstractTiming properties of applications strongly depend on resources that are allocated to them. Applications often have multiple resource requirements, all of which must be met for them to proceed. Performance analysis of event-based systems has been widely studied in the literature. However, the proposed works consider only one resource requirement for each application task. Additionally, they mainly focus on the rate at which resources serve applications (e.g., power, instructions or bits per second), but another aspect of resources, which is their provided capacity (e.g., energy, memory ranges, FPGA regions), has been ignored. In this work, we propose a mathematical framework to describe the provisioning rate and capacity of various types of resource. Additionally, we consider the simultaneous use of multiple resources. Conservative bounds on response times of events and their backlog are computed. We prove that the bounds are monotone in event arrivals and in required and provided rate and capacity, which enables verification of real-time application performance based on worst-case characterizations. The applicability of our framework is shown in a case study. Shayan Tabatabaei Nikkhah, Marc Geilen, Dip Goswami, Kees Goossens |
DATE | 4 |
| 2020 | Tightening the Mesh Size of the Cell-Aware ATPG Net for Catching All Detectable Weakest FaultsabstractCell-aware test (CAT) explicitly targets faults caused by cell-internal short and open defects and has been shown to significantly reduce test escape rates. CAT library cell characterization is typically done for only two defect resistance values: one representing hard opens and another one representing hard shorts. In this paper, similar to fishermen tightening the mesh size of their nets to catch small fish, we perform library characterization as efficiently as possible for a set of resistances representing increasingly weaker defects, and then adjust our ATPG flow to explicitly target faults caused by the weakest still-detectable variant of each potential defect. We implemented this novel approach in an experimental ATPG tool flow script, using functions of Cadence's Modus as building blocks. To assess the effectiveness of our approach, we formulate a new dedicated test metric: the weakest fault coverage wfc. Compared to conventional CAT targeting hard defects only, experimental results show that our new approach enhances detection of weakest faults and significantly reduces wfc escapes =1-wfc, while maintaining its original (hard-defect) fault coverage fc, of course at the expense of (acceptable) increases in the required number of test patterns and associated test generation time. Min-Chun Hu 0002, Santosh Malagi, Joe Swenton, Jos Huisken, Kees Goossens, Cheng-Wen Wu, Erik Jan Marinissen |
ETS | 6 |
| 2020 | Approximated Pareto Analysis for Fast Optimization of Large IEEE 802.15.4 TSCH NetworksabstractThe IEEE 802.15.4 Time-Slotted Channel Hopping (TSCH) is a widely used standard technology for industrial Wireless Sensor Networks (WSNs). The applications of such networks have diverse Quality-of-Service (QoS) demands that should be satisfied. Optimal configuration of the network parameters based on their QoS requirements and run-time adaptation to continuously meet the QoS requirements given all network dynamics is a great challenge for large-scale networks. The configuration space is very large for large-scale networks resulting in space exploration to be complex and extremely time-consuming. Moreover, such space exploration needs to be performed multiple times at design-time to give insight into the worst and best case performance of the mechanisms and network topology. Yet, it is needed for the run-time reconfiguration upon changes in the network. To address the stated challenges, we propose a fast and accurate enough algorithm based on Pareto algebra to extract a subset of Pareto configurations for large TSCH networks in a very short time. A proper configuration can be then picked from this set. Having such a fast optimization algorithm, the network can react to the changes in the link quality, routing topology, and reconfigure itself optimally at an appropriate time. The performance of the proposed technique is extensively evaluated, and compared with other approaches such as basic incremental Pareto analysis, and genetic algorithm, showing its superior execution time and quality of configurations in terms of accuracy and diversity. Hamideh Hajizadeh, Rasool Tavakoli, Majid Nabi, Kees Goossens |
PIMRC | 4 |
| 2019 | Chip Health Tracking Using Dynamic In-Situ Delay MonitoringabstractTracking the gradual effect of silicon aging requires fine-grain slack monitoring. Conventional slack monitoring techniques intend to measure worst-case static slack, i.e. the slack of longest timing path. In sharp contrast to the conventional techniques, we propose a novel technique that is based on dynamic excitation of in-situ delay monitors, i.e. dynamic excitation of the timing paths that are monitored. As the delays degrade, the path delays increase and the monitors are excited more frequently. With the proposed technique, a fine-grained signature of the delay degradation is extracted from the excitation rate of monitors. Hadi Ahmadi Balef, Kees Goossens, José Pineda de Gyvez |
DATE | 2 |
| 2019 | Model-Based Processor-in-the-Loop Framework for Composable Multi-core PlatformsabstractFrom model-based design to implementation on an embedded platform requires target-specific code generation, compilation, and execution. Processor-in-the-loop (PIL) simulation is an intermediate step meant for detailed testing and debugging in the development process. This paper presents a PIL simulation framework targeting multi-core FPGA-based embedded platforms. The presented framework allows for a fully automated process of performing PIL simulations on an FPGA-based embedded platform - CompSOC - starting from a Simulink model. The framework includes two PIL configurations - one configuration executes only the controller code on the target platform while other configuration executes both the controller and the plant code on the target platform. It considers scheduling of multiple applications and interference-free execution on the target platform under the PIL configurations. Further, the framework allows for logging various measurements of parameters such as execution time, memory usage and so on in the PIL configurations which can be used for testing and debugging purposes. Mojtaba Haghi, Martijn Koedam, Dip Goswami, Kees Goossens |
DSD | 4 |
| 2019 | Optimization of Cell-Aware ATPG Results by Manipulating Library Cells' Defect Detection MatricesabstractCell-aware test (CAT) explicitly targets defects inside library cells and therefore significantly reduces the number of test escapes compared to conventional automatic test pattern generation (ATPG) approaches that cover cell-internal defects only serendipitously. CAT consists of two steps, viz. (1) library characterization and (2) cell-aware ATPG. Defect detection matrices (DDMs) are used as the interface between both CAT steps; they record which cell-internal defects are detected by which cell-level test patterns. This paper proposes two algorithms that manipulate DDMs to optimize cell-aware ATPG results with respect to fault coverage, test pattern count, and compute time. Algorithm 1 identifies don't-care bits in cell patterns, such that the ATPG tool can exploit these during cell-to-chip expansion to increase fault coverage and reduce test-pattern count. Algorithm 2 selects, at cell level, a subset of preferential patterns that jointly provides maximal fault coverage at a minimized stimulus care-bit sum. To keep the ATPG compute time under control, we run cell-aware ATPG with the preferential patterns first, and a second ATPG run with the remaining patterns only if necessary. Selecting the preferential patterns maps onto a well-known NP-hard problem, for which we derive an innovative heuristic that outperforms solutions in the literature. Experimental results on twelve circuits show average reductions of 43% of non-covered faults and 10% in chip-pattern count. Min-Chun Hu 0002, Joe Swenton, Santosh Malagi, Jos Huisken, Kees Goossens, Erik Jan Marinissen |
ITC-Asia | 6 |
| 2019 | Application of Cell-Aware Test on an Advanced 3nm CMOS Technology LibraryabstractAdvanced technology nodes employ a large number of innovations. In addition, they require `scaling boosters' in the design of standard-cell libraries to be able to offer the scaling benefits in area, performance, and power that we have grown accustomed to. Consequently, sub-10nm standard cells are significantly more complex than their predecessors. Cell-aware test (CAT) explicitly targets cell-internal resistive open and short defects identified through extensive characterization of the library cells. This paper is (to the best of our knowledge) the first to report on the application of CAT library characterization on a sub-10nm technology node. We used Cadence's CAT tool flow on an experimental 114-cell-library in IMEC's 3nm CMOS technology iN5. Despite the increased cell complexity, we show that the CAT flow still works, and that compared with functionally-comparable library cells in a 45nm technology, the number of potential non-equivalent defect locations, cell-level test patterns, and defect coverage did not change drastically. Santosh Malagi, Min-Chun Hu 0002, Joe Swenton, Rogier Baert, Jos Huisken, Bilal Chehab, Kees Goossens, Erik Jan Marinissen |
ITC | 8 |
| 2019 | A Scalable and Fast Model for Performance Analysis of IEEE 802.15.4 TSCH NetworksabstractThe IEEE 802.15.4 Time-Slotted Channel Hopping (TSCH) protocol has received considerable attention in many industrial applications. However, analytical models for fast performance estimation of TSCH-based networks by considering the interaction between Medium Access Control (MAC) and Physical (PHY) layers is an open problem. In this paper, we propose a stochastic model for performance analysis of TSCH-based networks including dedicated and shared links with non-ideal wireless link properties. The proposed model is scalable and is able to evaluate the MAC performance of a large-scale network quickly. The developed model is verified by simulations and real-world experiments. The results confirm the accuracy of the proposed model for large-scale networks with orders of magnitude faster execution compared to the existing model in the literature. This confirms the speed and scalability of the model, which makes it a perfect tool for network design and optimization. Hamideh Hajizadeh, Majid Nabi, Rasool Tavakoli, Kees Goossens |
PIMRC | 4 |
| 2019 | Topology Management and TSCH Scheduling for Low-Latency Convergecast in In-Vehicle WSNsabstractWireless sensor networks (WSNs) are considered as a promising solution in intravehicle networking to reduce wiring and production costs. This application requires reliable and real-time data delivery, while the network is very dense. The time-slotted channel hopping (TSCH) mode of the IEEE 802.15.4 standard provides a reliable solution for low-power networks through guaranteed medium access and channel diversity. However, satisfying the stringent requirements of in-vehicle networks is challenging and demands for special consideration in network formation and TSCH scheduling. This paper targets convergecast in dense in-vehicle WSNs, in which all nodes can potentially directly reach the sink node. A cross-layer low-latency topology management and TSCH scheduling (LLTT) technique is proposed that provides a very high timeslot utilization for the TSCH schedule and minimizes communication latency. It first picks a topology for the network that increases the potential of parallel TSCH communications. Then, by using an optimized graph isomorphism algorithm, it extracts a proper match in the physical connectivity graph of the network for the selected topology. This network topology is used by a lightweight TSCH schedule generator to provide low data delivery latency. Two techniques, namely grouped retransmission and periodic aggregation, are exploited to increase the performance of the TSCH communications. The experimental results show that LLTT reduces the end-to-end communication latency compared to other approaches, while keeping the communications reliable by using dedicated links and grouped retransmissions. Rasool Tavakoli, Majid Nabi, Twan Basten, Kees Goossens |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Comparing Platform-aware Control Design Flows for Composable and Predictable TDM-based Execution PlatformsabstractWe compare three platform-aware feedback control design flows that are tailored for a composable and predictable Time Division Multiplexing (TDM)-based execution platform. The platform allows for independent execution of multiple applications. Using the precise timing knowledge of the platform execution, we accurately characterise the execution of the control application (i.e., sensing, computing, and actuating operations) to design efficient feedback controllers with high control performance in terms of settling time. The design flows are derived for Single-Rate (SR) and Multi-Rate (MR) sampling schemes. We show the applicability of the design flows based on two design considerations and their trade-off: control performance and resource utilisation. The design flows are validated by means of MATLAB and Hardware-in-the-Loop (HIL) experiments for a motion control application. Juan Valencia, Dip Goswami, Kees Goossens |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2019 | Timing Speculation With Optimal In Situ Monitoring Placement and Within-Cycle Error PreventionabstractIn this paper, a timing speculation technique with low-overhead in situ delay monitors placed along critical paths is presented. The proposed insertion of monitors enables timing error prevention within the same clock cycle. Compared to other techniques, the design cost per monitor in our technique is low because no additional gates for the guard banding, inspection window generation, and short path extension are required. We benchmarked our approach on an ARM Cortex M0. The insertion strategy reduces the number of monitors by up to ~23×, power by ~5.5×, and area by ~2.8× compared to the traditional in situ monitoring techniques that insert monitors at the flipflops. The timing error correction uses a global clock stretching unit to prevent errors within one cycle. With the proposed error prevention technique, -.22% more delay variation is tolerated with a negligible energy overhead of less than ~1%. Hadi Ahmadi Balef, Hamed Fatemi, Kees Goossens, José Pineda de Gyvez |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Fault-Tolerant Deployment of Dataflow Applications Using Virtual ProcessorsabstractMulti-processors are suited to host a dynamic mix of real-time dataflow applications, but are increasingly subject to faults because of the decreasing feature size. Applications can start and stop as needed if they execute on a private set of Virtual Processors (VPs) that are deployed on the physical processors. This allows online software updates, but makes it impossible to predict the deployment. If a fault renders a processor unusable, the free resources on other processors may be too fragmented to allow its VPs to be re-deployed. We show that mapping an application to more VPs reduces the maximum VP size. This increases the probability of successfully dealing with faults, at the cost of an increase of the total size. Such a mapping can either be run from the start, or we can split the VPs only when a fault occurs. Experiments confirm the feasibility of our approach, and show a trade-off between improved fault-tolerance and resource usage for both strategies. J. Reinier van Kampenhout, Sander Stuijk, Kees Goossens |
DSD | 3 |
| 2018 | Effective In-Situ Chip Health Monitoring with Selective Monitor Insertion Along Timing PathsabstractIn-situ delay monitoring is an advanced technique to monitor the robustness of digital circuits. Conventionally, in-situ delay monitors are inserted at the end-points of timing paths. To reduce the number of monitors and to increase their observability, intermediate points have been considered. In sharp contrast to these works, we propose a low overhead technique where the insertion points are selected along the timing paths such that timing violations can be predicted without false negative detections. With our approach, the number of required monitors is reduced by up to 11X compared to end-point insertion techniques. The observability to delay degradation is 8X better with our approach, compared to techniques with straight monitor placement at intermediate points. Hadi Ahmadi Balef, Hamed Fatemi, Kees Goossens, José Pineda de Gyvez |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | A Unified Programming Model for Time- and Data-Driven Embedded ApplicationsabstractModern embedded systems encompass a fast increasing range of applications, spanning from automotive to multimedia, and industrial automation. To tackle the increasing design complexity, the model-based design paradigm promotes the use of Models of Computation (MoCs) to capture the essential application properties. Existing MoCs are split between the event/time-triggered paradigm and the data-driven paradigm. However, time and data are two inter-related dimensions that are essential for defining the correct application behavior. In this paper we advocate a unified MoC that integrates the notions of time and data while accounting for imperfect clocks. We present the formal properties of our model and show how the Synchronous Data Flow (SDF) MoC can be used to analyze the time performance guarantees. Gabriela Breaban, Sander Stuijk, Kees Goossens |
PDP | 3 |
| 2018 | Hybrid Timeslot Design for IEEE 802.15.4 TSCH to Support Heterogeneous WSNsabstractThe IEEE 802.15.4 Time-Slotted Channel Hopping (TSCH) protocol defines two types of timeslots for communications, namely dedicated and shared timeslots. An upper layer in the protocol stack uses these timeslots to design a communication schedule for the network links, based on the required bandwidth for each link. Considering a network with time-varying data traffic generation by each node, the bandwidth requirements are changing over time for each link. This leads to poor efficiency of a predefined schedule when there is no data traffic for the dedicated timeslots, or there is too much data traffic injected to the shared timeslots. In this paper, we propose a new type of timeslot, called hybrid timeslot. A hybrid timeslot acts as a dedicated timeslot for a specific link, when there are packets available to be transmitted on that link. Otherwise, it acts as a shared timeslot that can be accessed by other links, using a contention-based mechanism. The hybrid timeslot has backward compatibility with the TSCH protocol and is functional with a few adaptations in the parameter setup of the TSCH protocol. Experimental and simulation results show that for heterogeneous networks using hybrid timeslots improves communication latency without reliability penalty. Rasool Tavakoli, Majid Nabi, Twan Basten, Kees Goossens |
PIMRC | 4 |
| 2018 | A Generic Method for a Bottom-Up ASIL Decomposition
Alessandro Frigerio, Bart Vermeulen, Kees Goossens |
SAFECOMP | 3 |
| 2018 | Guard-Time Design for Symmetric Synchronization in IEEE 802.15.4 Time-Slotted Channel HoppingabstractTime-Slotted Channel Hopping (TSCH) is considered as one of the most reliable MAC solutions for low- power wireless networking. In order to establish time-slotted communications, this technique requires all nodes to remain synchronized. The synchronization is continuously done through normal communications to compensate the clock drift between different nodes. In this paper, we present a detailed look into the behavior of the IEEE 802.15.4 PHY and MAC in terms of the synchronization task. We show that the relation between timeslot offsets provided by the standard leads to different synchronization error margins for positive and negative relative clock drifts. This is due to the time required for detection of ongoing transmissions at receivers. This may lead to the situation that two nodes are able to communicate in only one direction. Depending on which node is the source node, the available margin to compensate the relative clock drift is different. Accordingly, we provide new values for timeslot offsets to compensate positive and negative relative clock drifts equally. Simulation results confirm that the standard offsets reduce the performance of TSCH due to asymmetric synchronization error handling. The results also show that this negative effect is mitigated by using the new offsets provided in this paper. Rasool Tavakoli, Majid Nabi, Twan Basten, Kees Goossens |
VTC Spring | 4 |
| 2018 | Dependable Interference-Aware Time-Slotted Channel Hopping for Wireless Sensor NetworksabstractIEEE 802.15.4 Time-Slotted Channel Hopping (TSCH) aims to improve communication reliability in Wireless Sensor Networks (WSNs) by reducing the impact of the medium access contention, multipath fading, and blocking of wireless links. While TSCH outperforms single-channel communications, cross-technology interference on the license-free ISM bands may affect the performance of TSCH-based WSNs. For applications such as in-vehicle networks for which interference is dynamic over time, it leads to non-guaranteed reliability of the communications over time. This article proposes an Enhanced version of the TSCH protocol together with a Distributed Channel Sensing technique (ETSCH+DCS) that dynamically detects good quality channels to be used for communication. The quality of channels is extracted using a combination of a central and a distributed channel-quality estimation technique. The central technique uses Non-Intrusive Channel-quality Estimation (NICE) technique that proactively performs energy detections in the idle part of each timeslot at the coordinator of the network. NICE enables ETSCH to follow dynamic interference, while it does not reduce throughput of the network. The distributed channel quality estimation technique is executed by all the nodes in the network, based on their communication history, to detect interference sources that are hidden from the coordinator. We did two sets of lab experiments with controlled interferers and a number of simulations using real-world interference datasets to evaluate ETSCH. Experimental and simulation results show that ETSCH improves reliability of network communications, compared to basic TSCH and the state-of-the-art solution. In some experimental scenarios NICE itself has been able to increase the average packet reception ratio by 22% and shorten the length of burst packet losses by half, compared to the plain TSCH protocol. Further experiments show that DCS can reduce the effect of hidden interference (which is not detectable by NICE) on the packet reception ratio of the affected links by 50%. Rasool Tavakoli, Majid Nabi, Twan Basten, Kees Goossens |
ACM Trans. Sens. Networks | 4 |
| 2017 | Efficient synchronization methods for LET-based applications on a Multi-Processor System on ChipabstractDistributed control applications cover a wide range of areas such as automotive, avionics, and automation. The Logical Execution Time (LET) Model of Computation (MoC) was proposed as a formal method to describe the functional and timing behavior of such applications. However, modern Multi-Processor Systems on Chip (MPSOC) do not have a shared notion of time between processors, due to their use of Globally Asynchronous Locally Synchronous (GALS) architecture. In this paper we propose two methods (based on FIFO channels and barriers) to implement time and data synchronization on a MPSOC. While a barrier synchronizes the execution flows of tasks at predefined points in their executions, a FIFO is an asynchronous data communication method between two tasks. First, they are used to implement LET applications. Next, we show how dataflow applications and mixed LET-dataflow applications are supported too. We implemented both methods on a MPSOC prototyped on a FPGA, and show that the data synchronization outperforms the related work by 67% in terms of software overhead. Gabriela Breaban, Sander Stuijk, Kees Goossens |
DATE | 3 |
| 2017 | Programming and analysing scenario-aware dataflow on a multi-processor platformabstractThe FSM-SADF model of computation is especially suitable for analysing real-time applications with input-dependent behaviour such as different modes, variable execution times and scalable parallelism. Although FSM-SADF specifies which scenario transitions are possible, it does not specify how and when they are decided at runtime. Multiple actors of a scenario, e.g. video stream header parsing, may have to fire before it is known which scenario the application is in. We solve this causality dilemma with a concept for executing a sequence of scenarios, and demonstrate an implementation on multiple processors with rolling static-order scheduling. We furthermore present a platform-aware analysis model that covers concept and implementation, and integrate the contributions in a toolflow. A proof-of-concept confirms the low overhead of the implementation and the exact timing analysis of our model. J. Reinier van Kampenhout, Sander Stuijk, Kees Goossens |
DATE | 3 |
| 2017 | A Globally Arbitrated Memory Tree for Mixed-Time-Criticality SystemsabstractEmbedded systems are increasingly based on multi-core platforms to accommodate a growing number of applications, some of which have real-time requirements. Resources, such as off-chip DRAM, are typically shared between the applications using memory interconnects with different arbitration polices to cater to diverse bandwidth and latency requirements. However, traditional centralized interconnects are not scalable as the number of clients increase. Similarly, current distributed interconnects either cannot satisfy the diverse requirements or have decoupled arbitration stages, resulting in larger area, power and worst-case latency. The four main contributions of this article are: 1) a Globally Arbitrated Memory Tree (GAMT) with a distributed architecture that scales well with the number of cores, 2) an RTL-level implementation that can be configured with five arbitration policies (three distinct and two as special cases), 3) the concept of mixed arbitration policies that allows the policy to be selected individually per core, and 4) a worst-case analysis for a mixed arbitration policy that combines TDM and FBSP arbitration.We compare the performance of GAMT with centralized implementations and show that it can run up to four times faster and have over 51 and 37 percent reduction in area and power consumption, respectively, for a given bandwidth. Manil Dev Gomony, Jamie Garside, Benny Akesson, Neil C. Audsley, Kees Goossens |
IEEE Trans. Computers | 5 |
| 2016 | Resource utilization and Quality-of-Control trade-off for a composable platform
Juan Valencia, Eelco P. van Horssen, Dip Goswami, W. P. M. H. Heemels, Kees Goossens |
DATE | 5 |
| 2016 | An Experimental Study of Cross-Technology Interference in In-Vehicle Wireless Sensor NetworksabstractWireless in-vehicle networks are considered as a flexible and cost-efficient solution for the new generation of cars. One of the candidate wireless technologies for these wireless sensor networks is the IEEE 802.15.4 standard which operates in the 2.4 GHz ISM band. This is while the number of wireless devices that operate in this band is ever increasing. This broad usage of the same RF band may cause considerable performance degradation of wireless networks due to interference. There is some work on the coexistence of the IEEE 802.15.4 protocol and other standard technologies such as IEEE 802.11 (Wi-Fi) and IEEE 802.15.1 (Bluetooth), but none of it considers the highly dynamic conditions of in-vehicle networks. In this paper, we investigate the interference behavior in in-vehicle environments using real-world experiments. We consider different scenarios and measure the interference on all the 16 channels of IEEE 802.15.4 in the 2.4 GHz band.The measurement data set is available to the public. This real-world data set can be used for realistic and accurate network simulation. To study the effect of interference on in-vehicle networks, we use this data set to evaluate the performance of an IEEE 802.15.4e TSCH link. The simulation results show that the packet error rate for some interference scenarios is considerably high and dynamic over time. This shows the value of the data set and reveals the importance of using adaptive interference mitigation techniques to improve the reliability of wireless in-vehicle networks. Rasool Tavakoli, Majid Nabi, Twan Basten, Kees Goossens |
MSWiM | 4 |
| 2016 | Modeling and Verification of Dynamic Command Scheduling for Real-Time Memory ControllersabstractIn modern multi-core systems with multiple real-time (RT) applications, memory traffic accessing the shared SDRAM is increasingly diverse, e.g., transactions have variable sizes. RT memory controllers with dynamic command scheduling can efficiently address the diversity by issuing appropriate commands subject to the SDRAM timing constraints. However, the scheduling dependencies between commands make it challenging to derive tight bounds for the worst-case response time (WCRT) and the worst-case bandwidth (WCBW) of a memory controller. Existing modeling and analysis techniques either do not provide tight WCRT and WCBW bounds for diverse memory traffic with variable transaction sizes or are difficult to adapt to different RT memory controllers. This paper models a memory controller using Timed Automata (TA), where model checking is applied for analysis. Our TA model is modular and accurately captures the behavior of a RT memory controller with dynamic command scheduling. We obtain WCRT and WCBW bounds, which are validated by simulating the worst- case transaction traces obtained by model checking with a cycle-accurate model of the memory controller. Our method outperforms three state-of-the-art analysis techniques. We reduce WCRT bound by up to 20%, while the average improvement is 7.7%, and increase the WCBW bound by up to 25% with an average improvement of 13.6%. In addition, our modeling is generic enough to extend to memory controllers with different mechanisms. Yonghui Li 0002, Benny Akesson, Kai Lampka, Kees Goossens |
RTAS | 4 |
| 2016 | Architecture and analysis of a dynamically-scheduled real-time memory controllerabstractMemory controller design is challenging as mixed time-criticality embedded systems feature an increasing diversity of real-time (RT) and non-real-time (NRT) applications with variable transaction sizes. To satisfy the requirements of the applications, tight bounds on the worst-case response time (WCRT) of memory transactions must be provided to RT applications, while the lowest possible average response time must be given to the remaining applications. Existing real-time memory controllers cannot efficiently achieve this goal as they either bound the WCRT by sacrificing the average response time, or cannot efficiently support variable transaction sizes. In this article, we propose to use dynamic command scheduling, which is capable of efficiently dealing with transactions with variable sizes. The three main contributions of this article are: (1) a memory controller architecture consisting of a front-end and a back-end, where the former uses a TDM arbiter with a new work-conserving policy and the latter has a dynamic command scheduling algorithm that is independent of the front-end, (2) a formalization of the timings of the memory transactions for the proposed algorithm and architecture, and (3) an analysis of WCRT for transactions to capture the behavior of both the front-end and the back-end. This WCRT analysis supports variable transaction sizes and different degrees of bank parallelism. The critical part of the WCRT is the worst-case execution time (WCET) of a transaction, which is the time spent on command scheduling in the back-end. The WCET is bounded by two techniques applied to both fixed and variable transaction sizes, respectively. We experimentally evaluate the proposed memory controller and compare to an existing semi-static approach. The results demonstrate that dynamic command scheduling significantly outperforms the semi-static approach in the average case, while it performs equally well or better in the worst-case with only a few exceptions. The former reduces the average response time for NRT applications, and the latter pertains the WCRT for RT applications. Yonghui Li 0002, Benny Akesson, Kees Goossens |
Real Time Syst. | 3 |
| 2016 | Power/Performance Trade-Offs in Real-Time SDRAM Command SchedulingabstractReal-time safety-critical systems should provide hard bounds on an applications' performance. SDRAM controllers used in this domain should therefore have a bounded worst-case bandwidth, response time, and power consumption. Existing works on real-time SDRAM controllers only consider a narrow range of memory devices, and do not evaluate how their schedulers' performance varies across memory generations, nor how the scheduling algorithm influences power usage. The extent to which the number of banks used in parallel to serve a request impacts performance is also unexplored, and hence there are gaps in the tool set of a memory subsystem designer, in terms of both performance analysis, and configuration options. This article introduces a generalized close-page memory command scheduling algorithm that uses a variable number of banks in parallel to serve a request. To reduce the schedule length for DDR4 memories, we exploit bank grouping through a pairwise bank-group interleaving scheme. The algorithm is evaluated using an ILP formulation, and provides schedules of optimal length for most of the considered LPDDR, DDR2, DDR3, LPDDR2, LPDDR3 and DDR4 devices. We derive the worst-case bandwidth, power and execution time for the same set of devices, and discuss the observed trade-offs and trends in the scheduler-configuration design space based on these metrics, across memory generations. Sven Goossens, Karthik Chandrasekar 0001, Benny Akesson, Kees Goossens |
IEEE Trans. Computers | 4 |
| 2016 | Argo: A Real-Time Network-on-Chip Architecture With an Efficient GALS ImplementationabstractIn this paper, we present an area-efficient, globally asynchronous, locally synchronous network-on-chip (NoC) architecture for a hard real-time multiprocessor platform. The NoC implements message-passing communication between processor cores. It uses statically scheduled time-division multiplexing (TDM) to control the communication over a structure of routers, links, and network interfaces (NIs) to offer real-time guarantees. The area-efficient design is a result of two contributions: 1) asynchronous routers combined with TDM scheduling and 2) a novel NI microarchitecture. Together they result in a design in which data are transferred in a pipelined fashion, from the local memory of the sending core to the local memory of the receiving core, without any dynamic arbitration, buffering, and clock synchronization. The routers use two-phase bundled-data handshake latches based on the Mousetrap latch controller and are extended with a clock gating mechanism to reduce the energy consumption. The NIs integrate the direct memory access functionality and the TDM schedule, and use dual-ported local memories to avoid buffering, flow-control, and synchronization. To verify the design, we have implemented a 4 × 4 bitorus NoC in 65-nm CMOS technology and we present results on area, speed, and energy consumption for the router, NI, NoC, and post layout. Evangelia Kasapaki, Martin Schoeberl, Rasmus Bo Sørensen, Christoph Thomas Muller, Kees Goossens, Jens Sparsø |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | A generic, scalable and globally arbitrated memory tree for shared DRAM access in real-time systems
Manil Dev Gomony, Jamie Garside, Benny Akesson, Neil C. Audsley, Kees Goossens |
DATE | 5 |
| 2015 | A Scenario-Aware Dataflow Programming ModelabstractThe FSM-SADF model of computation allows to find a tight bound on the throughput of firm real-time applications by capturing dynamic variations in scenarios. We explore an FSM-SADF programming model, and propose three different alternatives for scenario switching. The best candidate for our CompSOC platform was implemented, and experiments confirm that the tight throughput bound results in a reduced resource budget. This comes at the cost of a predictable overhead at run-time as well as increased communication and memory budgets. We show that design choices offer interesting trade-offs between run-time cost and resource budgets. J. Reinier van Kampenhout, Sander Stuijk, Kees Goossens |
DSD | 3 |
| 2015 | Composable Platform-Aware Embedded Control Systems on a Multi-core ArchitectureabstractIn this work, we propose a design flow for efficient implementation of embedded feedback control systems targeted for multi-core platforms. We consider a composable tile-based architecture as an implementation platform and realise the proposed design flow onto one instance of this architecture. The proposed design flow implements the feedback loops in a data-driven fashion leading to time-varying sampling periods with short average sampling period. Our design flow is composed of two phases: (i) representing the timing behaviour imposed by the platform by a finite and known set of sampling periods, which is achieved exploiting the composability of the platform, and (ii) a linear matrix inequality (LMI) based platform-aware control algorithm that explicitly takes the derived platform timing characteristics and the shorter average sampling period into account. Our results show that the platform-aware implementation outperforms traditional control design flows (i.e., almost 2 times) in terms of quality of control (QoC). Juan Valencia, Dip Goswami, Kees Goossens |
DSD | 3 |
| 2015 | Distributed power management of real-time applications on a GALS multiprocessor SOCabstractIt is generally desirable to reduce the power consumption of embedded systems. Dynamic Voltage and Frequency Scaling (DVFS) is a commonly applied technique to achieve power reduction at the cost of computational performance. Multiprocessor System on Chips (MPSoCs) can have multiple voltage and frequency domains, e.g. per-core. When DVFS is applied to real-time applications, the effects must be accounted for in the associated formal timing model. In this work, we contribute our distributed multi-core run-time power-management technique for real-time dataflow applications that uses per-core lookup-tables to select low-power DVFS operating points that meet the application's timing requirement. We describe in detail how timing slack is observed locally at run-time on each core and is used to select a local DVFS operating point that meets the application's timing requirement. We further describe our static off-line formal analysis technique to generate these per-core lookup-tables that link timing slack to low-power DVFS operating points. We provide an experimental analysis of our proposed technique using an H.263 decoder application that is mapped onto an FPGA prototyped hardware platform. Andrew Nelson 0001, Kees Goossens |
EMSOFT | 2 |
| 2015 | Enhanced Time-Slotted Channel Hopping in WSNs Using Non-intrusive Channel-Quality EstimationabstractCross-technology interference on the license-free ISM bands has a major negative effect on the performance of Wireless Sensor Networks (WSNs). Channel hopping has been adopted in the Time-Slotted Channel Hopping (TSCH) mode of IEEE 802.15.4e to eliminate blocking of wireless links caused by external interference on some frequency channels. This paper proposes an Enhanced version of the TSCH protocol (ETSCH) which restricts the used channels for hopping to the channels that are measured to be of good quality. The quality of channels is extracted using a new Non-Intrusive Channel-quality Estimation (NICE) technique by performing energy detections in selected idle periods every timeslot. NICE enables ETSCH to follow dynamic interference well, while it does not reduce throughput of the network. It also does not change the protocol, and does not require non-standard hardware. ETSCH uses a small Enhanced Beacon hopping Sequence List (EBSL) to broadcast periodic Enhanced Beacons (EB) in the network to synchronize nodes at the start of timeslots. Experimental results show that ETSCH improves reliability of network communication, compared to basic TSCH and a more advanced mechanism ATSCH. It provides higher packet reception ratios and reduces the maximum length of burst packet losses. Rasool Tavakoli, Majid Nabi, Twan Basten, Kees Goossens |
MASS | 4 |
| 2015 | Dataflow formalisation of real-time streaming applications on a Composable and Predictable Multi-Processor SOC
Andrew Nelson 0001, Kees Goossens, Benny Akesson |
J. Syst. Archit. | 2 |
| 2015 | T-CREST: Time-predictable multi-core architecture for embedded systemsabstractReal-time systems need time-predictable platforms to allow static analysis of the worst-case execution time (WCET). Standard multi-core processors are optimized for the average case and are hardly analyzable. Within the T-CREST project we propose novel solutions for time-predictable multi-core architectures that are optimized for the WCET instead of the average-case execution time. The resulting time-predictable resources (processors, interconnect, memory arbiter, and memory controller) and tools (compiler, WCET analysis) are designed to ease WCET analysis and to optimize WCET performance. Compared to other processors the WCET performance is outstanding. The T-CREST platform is evaluated with two industrial use cases. An application from the avionic domain demonstrates that tasks executing on different cores do not interfere with respect to their WCET. A signal processing application from the railway domain shows that the WCET can be reduced for computation-intensive tasks when distributing the tasks on several cores and using the network-on-chip for communication. With three cores the WCET is improved by a factor of 1.8 and with 15 cores by a factor of 5.7. The T-CREST project is the result of a collaborative research and development project executed by eight partners from academia and industry. The European Commission funded T-CREST. Martin Schoeberl, Sahar Abbaspour, Benny Akesson, Neil C. Audsley, Raffaele Capasso, Jamie Garside, Kees Goossens, Sven Goossens, Scott Hansen, Reinhold Heckmann, Stefan Hepp, Benedikt Huber, Alexander Jordan, Evangelia Kasapaki, Jens Knoop, Yonghui Li 0002, Daniel Wiltsche-Prokesch, Wolfgang Puffitsch, Peter P. Puschner, André Rocha, Cláudio Silva 0002, Jens Sparsø, Alessandro Tocchi |
J. Syst. Archit. | 7 |
| 2015 | A Real-Time Multichannel Memory Controller and Optimal Mapping of Memory Clients to Memory ChannelsabstractEver-increasing demands for main memory bandwidth and memory speed/power tradeoff led to the introduction of memories with multiple memory channels, such as Wide IO DRAM. Efficient utilization of a multichannel memory as a shared resource in multiprocessor real-time systems depends on mapping of the memory clients to the memory channels according to their requirements on latency, bandwidth, communication, and memory capacity. However, there is currently no real-time memory controller for multichannel memories, and there is no methodology to optimally configure multichannel memories in real-time systems. As a first work toward this direction, we present two main contributions in this article: (1) a configurable real-time multichannel memory controller architecture with a novel method for logical-to-physical address translation and (2) two design-time methods to map memory clients to the memory channels, one an optimal algorithm based on an integer programming formulation of the mapping problem, and the other a fast heuristic algorithm. We demonstrate the real-time guarantees on bandwidth and latency provided by our multichannel memory controller architecture by experimental evaluation. Furthermore, we compare the performance of the mapping problem formulation in a solver and the heuristic algorithm against two existing mapping algorithms in terms of computation time and mapping success ratio. We show that an optimal solution can be found in 2 hours using the solver and in less than 1 second with less than 7% mapping failure using the heuristic for realistically sized problems. Finally, we demonstrate configuring a Wide IO DRAM in a high-definition (HD) video and graphics processing system to emphasize the practical applicability and effectiveness of this work. Manil Dev Gomony, Benny Akesson, Kees Goossens |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Maximizing the Number of Good Dies for Streaming Applications in NoC-Based MPSoCs Under Process VariationabstractScaling CMOS technology into nanometer feature-size nodes has made it practically impossible to precisely control the manufacturing process. This results in variation in the speed and power consumption of a circuit. As a solution to process-induced variations, circuits are conventionally implemented with conservative design margins to guarantee the target frequency of each hardware component in manufactured multiprocessor chips. This approach, referred to as worst-case design, results in a considerable circuit upsizing, in turn reducing the number of dies on a wafer. This work deals with the design of real-time systems for streaming applications (e.g., video decoders) constrained by a throughput requirement (e.g., frames per second) with reduced design margins, referred to as better-than-worst-case design . To this end, the first contribution of this work is a complete modeling framework that captures a streaming application mapped to an NoC-based multiprocessor system with voltage-frequency islands under process-induced die-to-die and within-die frequency variations . The framework is used to analyze the impact of variations in the frequency of hardware components on application throughput at the system level . The second contribution of this work is a methodology to use the proposed framework and estimate the impact of reducing circuit design margins on the number of good dies that satisfy the throughput requirement of a real-time streaming application . We show on both synthetic and real applications that the proposed better-than-worst-case design approach can increase the number of good dies by up to 9.6% and 18.8% for designs with and without fixed SRAM and IO blocks, respectively. Davit Mirzoyan, Benny Akesson, Sander Stuijk, Kees Goossens |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2014 | Exploiting expendable process-margins in DRAMs for run-time performance optimizationabstractManufacturing-time process (P) variations and runtime voltage (V) and temperature (T) variations can affect a DRAM's performance severely. To counter these effects, DRAM vendors provide substantial design-time PVT timing margins to guarantee correct DRAM functionality under worst-case operating conditions. Unfortunately, with technology scaling these timing margins have become large and very pessimistic for a majority of the manufactured DRAMs. While run-time variations are specific to operating conditions and as a result, their margins difficult to optimize, process variations are manufacturing-time effects and excessive process-margins can be reduced at run-time, on a per-device basis, if properly identified. In this paper, we propose a generic post-manufacturing performance characterization methodology for DRAMs that identifies this excess in process-margins for any given DRAM device at runtime, while retaining the requisite margins for voltage (noise) and temperature variations. By doing so, the methodology ascertains the actual impact of process-variations on the particular DRAM device and optimizes its access latencies (timings), thereby improving its overall performance. We evaluate this methodology on 48 DDR3 devices (from 12 DIMMs) and verify the derived timings under worst-case operating conditions, showing up to 33.3% and 25.9% reduction in DRAM read and write latencies, respectively. Karthik Chandrasekar 0001, Sven Goossens, Christian Weis, Martijn Koedam, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 7 |
| 2014 | Coupling TDM NoC and DRAM controller for cost and performance optimization of real-time systemsabstractExisting memory subsystems and TDM NoCs for real-time systems are optimized independently in terms of cost and performance by configuring their arbiters according to the bandwidth and/or latency requirements of their clients. However, when they are used in conjunction, and run in different clock domains, i.e. they are decoupled, there exists no structured methodology to select the NoC interface width and operating frequency for minimizing area and/or power consumption. Moreover, the multiple arbitration points, one in the NoC and the other in the memory subsystem, introduce additional overhead in the worst-case guaranteed latency. These makes it hard to design cost-efficient real-time systems. The three main contributions in this paper are: (1) We present a novel methodology to couple any existing TDM NoC with a realtime memory controller and compute the different NoC interface width and operating frequency combinations for minimal area and/or power consumption. (2) For two different TDM NoC types, one a packet-switched and the other circuit-switched, we show the trade-off between area and power consumption with the different NoC configurations, for different DRAM generations. (3) We compare the coupled and decoupled architectures with the two NoCs, in terms of guaranteed worst-case latency, area and power consumption by synthesizing the designs in 40 nm technology. Our experiments show that using a coupled architecture in a system consisting of 16 clients results in savings of over 44% in guaranteed latency, 18% and 17% in area, 19% and 11% in power consumption for a packet-switched and a circuit-switched TDM NoC, respectively, with different DRAM types. Manil Dev Gomony, Benny Akesson, Kees Goossens |
DATE | 3 |
| 2014 | CoMik: A predictable and cycle-accurately composable real-time microkernelabstractThe functionality of embedded systems is ever increasing. This has lead to mixed time-criticality systems, where applications with a variety of real-time requirements co-exist on the same platform and share resources. Due to inter-application interference, verifying the real-time requirements of such systems is generally non trivial. In this paper, we present the CoMik microkernel that provides temporally predictable and composable processor virtualisation. CoMik's virtual processors are cycle-accurately composable, i.e. their timing cannot affect the timing of co-existing virtual processors by even a single cycle. Real-time applications executing on dedicated virtual processors can therefore be verified and executed in isolation, simplifying the verification of mixed time-criticality systems. We demonstrate these properties through experimentation on an FPGA prototyped hardware platform. Andrew Nelson 0001, Ashkan Beyranvand Nejad, Anca Mariana Molnos, Martijn Koedam, Kees Goossens |
DATE | 5 |
| 2014 | Composable and Predictable Dynamic Loading for Time-Critical Partitioned SystemsabstractIn time-critical systems such as in avionics, for safety and timing guarantees, applications are isolated from each other. Resources are partitioned in time and space creating a partition per application. Such isolation allows fault containment and independent development, testing and verification of applications. Current partitioned systems do not allow dynamically adding applications. Applications are statically loaded in their respective partitions. However dynamic loading can be useful or even necessary for scenarios such as on-board software updates, dynamic reconfiguration or re-loading applications in case of a fault. In this paper we propose a software architecture to dynamically create and manage partitions and a method for compostable dynamic loading which ensures that loading applications do not affect the running applications and vice versa. Furthermore the loading time is also predictable i.e. the loading time can be bounded a priori. We achieve this by splitting the loading process into parts, wherein only a small part which reserves minimum required resources is executed in the system partition and the other parts are executed in the allocated application partition which ensures isolation from other applications. We implement the software architecture for a SoC prototype on an FPGA board and demonstrate its composability and predictability properties. Shubhendu Sinha, Martijn Koedam, Rob van Wijk, Andrew Nelson 0001, Ashkan Beyranvand Nejad, Marc Geilen, Kees Goossens |
DSD | 7 |
| 2014 | Dynamic Command Scheduling for Real-Time Memory ControllersabstractMemory controller design is challenging as real-time embedded systems feature an increasing diversity of real-time and non-real-time applications with variable transaction sizes. To satisfy the requirements of the applications, tight bounds on the worst-case execution time (WCET) of memory transactions must be provided to real-time applications, while the lowest possible average execution time must be given to the rest. Existing real-time memory controllers cannot efficiently achieve this goal as they either bound the WCET by sacrificing the average execution time, or are not scalable to directly support variable transaction sizes, or both. In this paper, we propose to use dynamic command scheduling, which is capable of efficiently dealing with transactions with variable sizes. The three main contributions of this paper are: 1) a back-end architecture for a real-time memory controller with a dynamic command scheduling algorithm, 2) a formalization of the timings of the memory transactions for the proposed architecture and algorithm, and 3) two techniques to bound the WCET of transactions with both fixed and variable sizes, respectively. We experimentally evaluate the proposed memory controller and compare both the worst-case and average-case execution times of transactions to a state-of-the-art semi-static approach. The results demonstrate that dynamic command scheduling outperforms the semi-static approach by 33.4% in the average case and performs at least equally well in the worst case. We also show the WCET is tight for transactions with fixed and variable sizes, respectively. Yonghui Li 0002, Benny Akesson, Kees Goossens |
ECRTS | 3 |
| 2014 | dAElite: A TDM NoC Supporting QoS, Multicast, and Fast Connection Set-UpabstractNetworks-on-Chip (NoC) are seen as promising interconnect solutions, offering the advantages of scalability and high-frequency operation which the traditional bus interconnects lack. Several NoC implementations have been presented in the literature, some of them having mature tool-flows. The main differentiating factor between the various implementations is the set of services and communication patterns they offer to the end-user. In this paper we present dAElite, a TDM Network-on-Chip that offers a unique combination of features, namely, guaranteed bandwidth and latency per connection, built-in support for multicast, and a short connection set-up time. While our NoC was designed from the ground up, we leverage on existing tools for network dimensioning, analysis, and instantiation. We have implemented and tested our proposal in hardware and we compared it to Æthereal, a state-of-the-art NoC with similar features, but no multicast. We find that the connection set-up time is reduced by a factor of 10 and the network traversal latency is decreased by 33 percent. Moreover, considering realistic values of the network parameters, dAElite has a lower hardware area when synthesized in 65 nm technology. Radu Andrei Stefan, Anca Mariana Molnos, Kees Goossens |
IEEE Trans. Computers | 3 |
| 2014 | Process-variation-aware mapping of best-effort and real-time streaming applications to MPSoCsabstractAs technology scales, the impact of process variation on the maximum supported frequency (FMAX) of individual cores in a multiprocessor system-on-chip (MPSoC) becomes more pronounced. Task allocation without variation-aware performance analysis can greatly compromise performance and lead to a significant loss in yield , defined as the percentage of manufactured chips satisfying the application timing requirement . We propose variation-aware task allocation for best-effort and real-time streaming applications modeled as task graphs. Our solutions are primarily based on the throughput requirement, which is the most important timing requirement in many real-time streaming applications. The four main contributions of this work are (1) distinguishing best-effort firm real-time and soft real-time application classes, which require different optimization criteria, (2) using dataflow graphs, which are well suited for modeling and analysis of streaming applications, we explicitly model task execution both in terms of clock cycles (which is independent of variation) and seconds (which does depend on the variation of the resource), which we connect by an explicit binding, (3) we present two optimization approaches, which give different improvement results at different costs, (4) we present both exhaustive and heuristic algorithms that implement the optimization approaches. Our variation-aware mapping algorithms are tested on models of seven real applications and are compared to mapping methods that are unaware of hardware variation. Our results demonstrate (1) improvements in the average performance (3% on average) for best-effort applications, and (2) for firm real-time and soft real-time applications, yield improvements of up to 27% with an average of 15%, showing the effectiveness of our approaches. Davit Mirzoyan, Benny Akesson, Kees Goossens |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2013 | Towards variation-aware system-level power estimation of DRAMs: an empirical approachabstractDRAM vendors provide pessimistic current measures in memory datasheets to account for worst-case impact of process variations and to improve their production yield, leading to unrealistic power consumption estimates. In this paper, we first demonstrate the possible effects of process variations on DRAM performance and power consumption by performing Monte-Carlo simulations on a detailed DRAM cross-section. We then propose a methodology to empirically determine the actual impact for any given DRAM memory by assessing its performance characteristics during the DRAM calibration phase at system boot-time, thereby enabling its optimal use at run-time. We further employ our analysis on Micron's 2Gb DDR3-1600-x16 memory and show considerable over-estimation in the datasheet measures and the energy estimates (up to 28%), by using realistic current measures for a set of MediaBench applications. Karthik Chandrasekar 0001, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DAC | 5 |
| 2013 | System and circuit level power modeling of energy-efficient 3D-stacked wide I/O DRAMsabstractJEDEC recently introduced its new standard for 3D-stacked Wide I/O DRAM memories, which defines their architecture, design, features and timing behavior. With improved performance/power trade-offs over previous generation DRAMs, Wide I/O DRAMs provide an extremely energy-efficient green memory solution required for next-generation embedded and high-performance computing systems. With both industry and academia pushing to evaluate and employ these highly anticipated memories, there is an urgent need for an accurate power model targeting Wide I/O DRAMs that enables their efficient integration and energy management in DRAM stacked SoC architectures. In this paper, we present the first system-level power model of 3D-stacked Wide I/O DRAM memories that is almost as accurate as detailed circuit-level power models of 3D-DRAMs. To verify its accuracy, we experimentally compare its power and energy estimates for different memory workloads and operations against those of a circuit-level 3D-DRAM power model and show less than 2% difference between the two sets of estimates. Karthik Chandrasekar 0001, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 5 |
| 2013 | Architecture and optimal configuration of a real-time multi-channel memory controllerabstractOptimal utilization of a multi-channel memory, such as Wide IO DRAM, as shared memory in multi-processor platforms depends on the mapping of memory clients to the memory channels, the granularity at which the memory requests are interleaved in each channel, and the bandwidth and memory capacity allocated to each memory client in each channel. Firm real-time applications in such platforms impose strict requirements on shared memory bandwidth and latency, which must be guaranteed at design-time to reduce verification effort. However, there is currently no real-time memory controller for multichannel memories, and there is no methodology to optimally configure multi-channel memories in real-time systems. This paper has four key contributions: (1) A real-time multi-channel memory controller architecture with a new programmable Multi-Channel Interleaver unit. (2) A novel method for logical-to-physical address translation that enables inter-leaving memory requests across multiple memory channels at different granularities. (3) An optimal algorithm based on an Integer Linear Program (ILP) formulation to map memory clients to memory channels considering their communication dependencies, and to configure the memory controller for minimum bandwidth utilization. (4) We experimentally evaluate the run-time of the algorithm and show that an optimal solution can be found within 15 minutes for realistically sized problems. We also demonstrate configuring a multi-channel Wide IO DRAM in a High-Definition (HD) video and graphics processing system to emphasize the effectiveness of our approach. Manil Dev Gomony, Benny Akesson, Kees Goossens |
DATE | 3 |
| 2013 | Conservative open-page policy for mixed time-criticality memory controllersabstractComplex Systems-on-Chips (SoC) are mixed time-criticality systems that have to support firm real-time (FRT) and soft real-time (SRT) applications running in parallel. This is challenging for critical SoC components, such as memory controllers. Existing memory controllers focus on either firm real-time or soft real-time applications. FRT controllers use a close-page policy that maximizes worst-case performance and ignore opportunities to exploit locality, since it cannot be guaranteed. Conversely, SRT controllers try to reduce latency and consequently processor stalling by speculating on locality. They often use an open-page policy that sacrifices guaranteed performance, but is beneficial in the average case. This paper proposes a conservative open-page policy that improves average-case performance of a FRT controller in terms of bandwidth and latency without sacrificing real-time guarantees. As a result, the memory controller efficiently handles both FRT and SRT applications. The policy keeps pages open as long as possible without sacrificing guarantees and captures locality in this window. Experimental results show that on average 70% of the locality is captured for applications in the CHStone benchmark, reducing the execution time by 17% compared to a close-page policy. The effectiveness of the policy is also evaluated in a multi-application use-case, and we show that the overall average-case performance improves if there is at least one FRT or SRT application that exploits locality. Sven Goossens, Benny Akesson, Kees Goossens |
DATE | 3 |
| 2013 | A General Framework for Average-Case Performance Analysis of Shared ResourcesabstractContemporary embedded systems are based on complex heterogeneous multi-core platforms to cater to the increasing number of applications, some of which have (soft) real-time requirements. To reduce cost, resources are shared using diverse arbitration mechanisms, such as Time-Division Multiplexing (TDM), Static-Priority (SP), and Round-Robin (RR), depending on application and resource requirements. However, resource sharing results in interference between sharing applications making it difficult to estimate if the average latency is sufficient to satisfy their real-time requirements. Existing work proposes isolated models that either fail to address the diversity of arbitration mechanisms or cannot capture the dynamic arrival and service processes of applications and resources in multi-core platforms. This paper addresses this problem by proposing a general framework for average-case performance analysis of shared resources in multi-core platforms. The two main contributions are: 1) a general model for resource sharing based on queuing theory that can be used with different arbiters and that captures architectural features of the shared resource, such as pipelining and arbitration delay, and 2) three arbiter models for TDM, SP, and RR, respectively that assume general distributions (G/G/1) and fits within the framework. Sahar Foroutan, Benny Akesson, Kees Goossens, Frédéric Pétrot |
DSD | 3 |
| 2013 | Router Designs for an Asynchronous Time-Division-Multiplexed Network-on-ChipabstractIn this paper we explore the design of an asynchronous router for a time-division-multiplexed (TDM) network-on-chip (NOC) that is being developed for a multi-processor platform for hard real-time systems. TDM inherently requires a common time reference, and existing TDM-based NOC designs are either synchronous or mesochronous, but both approaches have their limitations: a globally synchronous NOC is no longer feasible in today's sub micron technologies and a mesochronous NOC requires special FIFO-based synchronizers in all input ports of all routers in order to accommodate for clock phase differences. This adds hardware complexity and increases area and power consumption. We propose to use asynchronous routers in order to achieve a simpler, more robust and globally-asynchronous NOC, and this represents an unexplored point in the design space. The paper presents a range of alternative router designs. All routers have been synthesized for a 65nm CMOS technology, and the paper reports post-layout figures for area, speed and energy and compares the asynchronous designs with an existing mesochronous clocked router. The results show that an asynchronous router is 2 times smaller, marginally slower and with roughly the same energy consumption, while offering a robust solution to the clock distribution problem. The paper further explores "clock-gating" of the individual pipeline stages in the asynchronous routers, and shows that this can lead to significant power savings. Evangelia Kasapaki, Jens Sparsø, Rasmus Bo Sørensen, Kees Goossens |
DSD | 4 |
| 2013 | Run-Time Slack Distribution for Real-Time Data-Flow Applications on Embedded MPSoCabstractLow energy consumption is crucial for embedded systems, including the ones that employ tiled Multiprocessor Systems-on-Chip(MPSoC). Such systems often execute real-time applications consisting of several tasks synchronized in a data-flow manner and mapped over different MPSoC tiles. Energy can be saved by lowering the processor voltage and frequency, hence extending the application execution over periods of time otherwise left idle, i.e., exploiting slack. In this paper we propose a framework to distribute slack information at run-time, intra-and inter-tile, to enable accurate and conservative slack calculation within each tile. The slack is transferred along with the existing inter-task synchronization and as a result it is distributed across the MPSoC with low overhead. In each tile, we add a hardware block that calculates the slack received during inter-tile communication and a software library to program this hardware. We integrate this framework into an existing MPSoC platform and we prototype an entire system with two tiles on an Xilinx ML605 FPGA board. We demonstrate the effectiveness of our proposal with a simple, conservative, DVFS management policy applied to an H.264 decoder application. The experimental results suggest that our framework reduces %the average processors frequency with 56% and the energy consumption with 53%, the total energy consumption of tiles with 27%, when compared to a state-of-the-art intra-tile approach that uses a similar management policy. Our proposal introduces only a minor software overhead of up to 4% over the application execution time and negligible additional FPGA chip utilization of 0.002%. Pavel G. Zaykov, Georgi Kuzmanov, Anca Mariana Molnos, Kees Goossens |
DSD | 4 |
| 2013 | A software-based technique enabling composable hierarchical preemptive scheduling for time-triggered applicationsabstractMany embedded real-time applications are typically time-triggered and preemptive schedulers are used to execute tasks of such applications. Orthogonally, composable partitioned embedded platforms use preemptive time-division multiplexing mechanism to isolate applications. Existing composable systems that support two-level scheduling are restricted to cooperative intra-application schedulers, and thus cannot execute the time-triggered applications. In this work, we introduce a framework that allows concurrent, composable execution of such applications on temporally-partitioned systems. The framework is composed of an execution platform and a method for timing analysis of applications running on the platform. The platform realizes a software-based timed-interrupt virtualization technique on an existing composable system. Multiple time-triggered applications may run concurrently using different intra-application preemptive scheduling policies, e.g., fixed-priority and rate-monotonic. The analysis method formalizes the available processing time for executing each application on a processor in order to enable schedulability tests for different policies. Finally, these concepts are demonstrated by executing a number of applications, first on an FPGA prototype and second on a Matlab simulation of the platform. The results indicate a composable and concurrent execution of multiple time-triggered applications using our proposed framework. Furthermore, the implementation of the technique has low cost in terms of memory footprint and execution overhead. Ashkan Beyranvand Nejad, Anca Mariana Molnos, Kees Goossens |
RTCSA | 3 |
| 2013 | A unified execution model for multiple computation models of streaming applications on a composable MPSoC
Ashkan Beyranvand Nejad, Anca Mariana Molnos, Kees Goossens |
J. Syst. Archit. | 3 |
| 2013 | A hardware/software platform for QoS bridging over multi-chip NoC-based systems
Ashkan Beyranvand Nejad, Anca Mariana Molnos, Matias Escudero Martinez, Kees Goossens |
Parallel Comput. | 4 |
| 2012 | Run-time power-down strategies for real-time SDRAM memory controllersabstractPowering down SDRAMs at run-time reduces memory energy consumption significantly, but often at the cost of performance. If employed speculatively with real-time memory controllers, power-down mechanisms could impact both the guaranteed bandwidth and the memory latency bounds. This calls for power-down strategies that can hide or bound the performance loss, making run-time memory power-down feasible for real-time applications. Karthik Chandrasekar 0001, Benny Akesson, Kees Goossens |
DAC | 3 |
| 2012 | DRAM selection and configuration for real-time mobile systemsabstractThe performance and power consumption of mobile DRAMs (LPDDRs) depend on the configuration of system-level parameters, such as operating frequency, interface width, request size, and memory map. In mobile systems running both real-time and non-real-time applications, the memory configuration must satisfy bandwidth requirements of real-time applications, meet the power consumption budget, and offer the best average-case execution time to the non-real-time applications. There is currently no well-defined methodology for selecting a suitable memory configuration for real-time mobile systems. The worst-case bandwidth, average-case execution time, and power consumption of mobile DRAMs across generations have furthermore not been investigated. This paper has two main contributions. 1) We analyze the worst-case bandwidth, average-case execution time, and power consumption of mobile DRAMs across three generations: LPDDR, LPDDR2 and Wide-IO-based 3D-stacked DRAM. 2) Based on our analysis, we propose a methodology for selecting memory configurations in real-time mobile systems.We show that LPDDR (32-bit IO), LPDDR2 (32-bit IO) and 3D-DRAM (128-bit IO) provide worst-case bandwidth up to 0.75 GB/s, 1.6 GB/s and 3.1 GB/s, respectively. We furthermore show for an H.263 decoder that LPDDR2 and 3D-DRAM reduce power consumption with up to 25% and 67%, respectively, compared to LPDDR, and reduce the execution time with up to 18% and 25%. Manil Dev Gomony, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 5 |
| 2012 | Memory-map selection for firm real-time SDRAM controllersabstractA modern real-time embedded system must support multiple concurrently running applications. To reduce costs, critical SoC components like SDRAM memories are often shared between applications with a variety of firm real-time requirements. To guarantee that the system works as intended, the memory controller must be configured such that all the real-time requirements of all sharing applications are satisfied. The attainable worst-case bandwidth, latency, and power of the memory depend largely on memory map configuration. Sharing SDRAM amongst multiple applications is challenging, since their requirements might call for different memory maps. This paper presents an exploration of the memory-map design space. Two contributions improve the memory-map selection procedure. The first contribution reduces the minimum access granularity by interleaving requests over a configurable number of banks instead of all banks. This technique is beneficial for worst-case performance in terms of bandwidth, latency and power. As a second contribution, we present a methodology to derive a memory-map configuration, i.e. the access granularity and number of interleaved banks, from a specification of the real-time application requirements and an overall memory power budget. Sven Goossens, Tim Kouters, Benny Akesson, Kees Goossens |
DATE | 4 |
| 2012 | A TDM NoC supporting QoS, multicast, and fast connection set-upabstractNetworks-on-Chip are seen as promising interconnect solutions, offering the advantages of scalability and high frequency operation which the traditional bus interconnects lack. Several NoC implementations have been presented in the literature, some of them having mature tool-flows and ecosystems. The main differentiating factor between the various implementations are the services and communication patters they offer to the end-user. In this paper we present dAElite, a TDM Network-on-Chip that offers a unique combinations of features, namely guaranteed bandwidth and latency per connection, built-in support for multicast, and a short connection set-up time. While our NoC was designed from the ground up, we leverage on existing tools for network dimensioning, analysis and instantiation. We have implemented and tested our proposal in hardware and we found it to compare favorably to the other NoCs in terms of hardware area. Compared with aelite, which is closest in terms of offered services our network offers connection set-up times faster by a factor of 10 network, traversal latencies decreased by 33%, and improved bandwidth. Radu Andrei Stefan, Anca Mariana Molnos, Jude Angelo Ambrose, Kees Goossens |
DATE | 4 |
| 2012 | Composable Virtual Memory for an Embedded SoCabstractSystems on a Chip concurrently execute multiple applications that may start and stop at run-time, creating many use-cases. Composability reduces the verifcation effort, by making the functional and temporal behaviours of an application independent of other applications. Existing approaches link applications to static address ranges that cannot be reused between applications that are not simultaneously active, wasting resources. In this paper we propose a composable virtual memory scheme that enables dynamic binding and relocation of applications. Our virtual memory is also predictable, for applications with real-time constraints. We integrated the virtual memory on, CompSOC, an existing composable SoC prototyped in FPGA. The implementation indicates that virtual memory is in general expensive, because it incurs a performance loss around 39% due to address translation latency. On top of this, composability adds to virtual memory an insigni cant extra performance penalty, below 1%. Cor Meenderinck, Anca Mariana Molnos, Kees Goossens |
DSD | 3 |
| 2012 | A Predictor-Based Power-Saving Policy for DRAM MemoriesabstractReducing power/energy consumption is an important goal for all computer systems, from servers to battery-driven hand-held devices. To achieve this goal, the energy consumption of all system components needs to be reduced. One of the most power-hungry components is the off-chip DRAM, even when it is idle. DRAMs support different power-saving modes, such as self-refresh and power-down, but employing them every time the DRAM is idle, reduces performance due to their power-up latencies. The self-refresh mode offers large power savings, but incurs a long power-up latency. The power-down mode, on the other hand, has a shorter power-up latency, but provides lower power savings. In this paper, we propose and evaluate a novel power-saving policy that combines the best of both power-saving modes in order to achieve significant power reductions with a marginal performance penalty. To accomplish this, we use a history-based predictor to forecast the duration of an idle period and then either employ self-refresh, or power-down, or a combination of both power saving modes. Significant refinements are made to the predictor to maximize the energy savings and minimize the performance penalty. The presented policy is evaluated using several applications from the multimedia domain and the experimental results show that it reduces the total DRAM energy consumption between 68.8% and 79.9% at a negligible performance penalty between 0.3% and 2.2%. Gervin Thomas, Karthik Chandrasekar 0001, Benny Akesson, Ben H. H. Juurlink, Kees Goossens |
DSD | 5 |
| 2012 | Architecture and design flow for a debug event distribution interconnectabstractIn this paper, we describe and analyze the architecture of the proposed Debug Event Distribution Interconnect (EDI). The EDI transmits debug events, which are 1-bit signals, between debug entities in different areas of the Network-on-Chip based Multi-Processor System-on-Chip. The EDI replicates the NoC topology with an EDI node instantiated for each underlying NoC data module. Contention in the EDI node is handled by replicating the EDI in layers. The EDI generation is automatic, and uses as input the cross-triggering patterns that are not required to follow the communication patterns in the NoC. The generation and routing tool is also presented in this paper. The EDI is evaluated with four different implementations varying complexity and handling of contention. The area of a single EDI Layer is around 0.9% of the area occupied by the tested NoCs, using the lower area implementation. These results show that the proposed implementation of the EDI incurs low cost on the overall system. Arnaldo Azevedo, Bart Vermeulen, Kees Goossens |
ICCD | 3 |
| 2011 | Architectures and modeling of predictable memory controllers for improved system integrationabstractDesigning multi-processor systems-on-chips becomes increasingly complex, as more applications with realtime requirements execute in parallel. System resources, such as memories, are shared between applications to reduce cost, causing their timing behavior to become inter-dependent. Using conventional simulation-based verification, this requires all concurrently executing applications to be verified together, resulting in a rapidly increasing verification complexity. Predictable and composable systems have been proposed to address this problem. Predictable systems provide bounds on performance, enabling formal analysis to be used as an alternative to simulation. Composable systems isolate applications, enabling them to be verified independently. Predictable and composable systems are built from predictable and composable resources. This paper presents three general techniques to implement and model predictable and composable resources, and demonstrates their applicability in the context of a memory controller. The architecture of the memory controller is general and supports both SRAM and DDR2/DDR3 SDRAM and a wide range of arbiters, making it suitable for many predictable and composable systems. The modeling approach is based on a shared-resource abstraction that covers any combination of supported memory and arbiter and enables system-level performance analysis with a variety of well-known frameworks, such as network calculus or data-flow analysis. Benny Akesson, Kees Goossens |
DATE | 2 |
| 2011 | An FPGA bridge preserving traffic quality of service for on-chip network-based systemsabstractFPGA prototyping of recent large Systems on Chip (SoCs) is very challenging due to the resource limitation of a single FPGA. Moreover, having external access to SoCs for verification and debug purposes is essential. In this paper, we suggest to partition a network-on-chip (NoC) based system into smaller sub-systems each with their own NoC, and each of which is implemented on a separate FPGA board. Multiple SoC ASICs can be bridged in the same way. The scheme that interconnects the sub-systems should offer the application connections the required quality of service (QoS). In this paper, we investigate bridging schemes at different levels of the NoC protocol stack. Comparing the distinct design criteria for the proposed schemes, a bridge is designed. The bridge experiments show that it provides QoS in terms of bandwith and latency. Ashkan Beyranvand Nejad, Matias Escudero Martinez, Kees Goossens |
DATE | 3 |
| 2011 | Optimal scheduling of switched FlexRay networksabstractThis paper introduces the concept of switched FlexRay networks and proposes two algorithms to schedule data communication for this new type of network. Switched FlexRay networks use an intelligent star coupler, called a switch, to temporarily decouple network branches, thereby increasing the effective network bandwidth. Although scheduling for basic FlexRay networks is not new, prior work in this domain does not utilize the branch parallelism that is available when a FlexRay switch is used. In addition to the novel exploitation of branch parallelism, the scheduling algorithms proposed in this paper also support all slot multiplexing options as defined in the FlexRay v 3.0 protocol specification. This includes support for the newly added repetition rates and support for multiplexing frames from different sending nodes in the same slot. Our first algorithm quickly produces a schedule given the communication requirements, network topology and FlexRay parameters, but cannot guarantee an optimal schedule in terms of the bandwidth efficiency and extensibility. Therefore, a second, branch-and-price algorithm is introduced that does find optimal schedules. Thijs Schenkelaars, Bart Vermeulen, Kees Goossens |
DATE | 3 |
| 2011 | Improved Power Modeling of DDR SDRAMsabstractPower modeling and estimation has become one of the most defining aspects in designing modern embedded systems. In this context, DDR SDRAM memories contribute significantly to system power consumption, but lack accurate and generic power models. The most popular SDRAM power model provided by Micron, is found to be inaccurate or insufficient for several reasons. First, it does not consider the power consumed when transitioning to power-down and self-refresh modes. Second, it employs the minimal timing constraints between commands from the SDRAM datasheets and not the actual duration between the commands as issued by an SDRAM memory controller. Finally, without adaptations, it can only be applied to a memory controller that employs a close-page policy and accesses a single SDRAM bank at a time. These critical issues with Micron's power model impact the accuracy and the validity of the power values reported by it and resolving them, forms the focus of our work. In this paper, we propose an improved SDRAM power model that estimates power consumption during the state transitions to power-saving states, employs an SDRAM command trace to get the actual timings between the commands issued and is generic and applicable to all DDRx SDRAMs and all memory controller policies and all degrees of bank interleaving. We quantitatively compare the proposed model against the unmodified Micron model on power and energy for DDR3-800. We show differences of up to 60% in energy-savings for the precharge power-down mode for a power-down duration of 14 cycles and up to 80% for the self-refresh mode for a self-refresh duration of 560 cycles. Karthik Chandrasekar 0001, Benny Akesson, Kees Goossens |
DSD | 3 |
| 2011 | A Unified Execution Model for Data-Driven Applications on a Composable MPSoCabstractMulti-processor Systems on Chip (MPSoCs) execute multiple applications concurrently. These applications may belong to different domains, i.e., may have firm-, soft-, or non-real time requirements. A composable system simplifies system design, integration, and verification by avoiding the inter-application interference. Existing work demonstrates composability for applications expressed using a single model of computation. For example, Kahn Process Network (KPN) and dataflow are two common data-driven parallel models of computation, each with different properties and suited for different application domains. This paper extends existing work with support for concurrent, composable execution of KPN and dataflow applications on the same MPSoC platform. We formalize a unified execution model by defining its operations that implement the different models of computation on the MPSoC, and discuss the trade-offs involved. Our experiments indicate that multiple applications modeled in KPN and dataflow run composably on an MPSoC platform. Ashkan Beyranvand Nejad, Anca Mariana Molnos, Kees Goossens |
DSD | 3 |
| 2011 | Power Minimisation for Real-Time Dataflow ApplicationsabstractEnergy efficient execution of applications is important for many reasons, e.g. time between battery charges, device temperature. Voltage and Frequency Scaling (VFS) enables applications to be run at lower frequencies on hardware resources thereby consuming less power. Real-time applications have deadlines that must be met otherwise their output is devalued. Dataflow modelling of real-time applications enables off-line verification of the application's temporal requirements. In this paper we describe a method to reduce the combined static and dynamic energy consumption using a Dynamic VFS (DVFS) technique for dataflow modelled real-time applications that may be mapped onto multiple hardware resources. We achieve this by using an application's static slack in order to perform DVFS while still satisfying the application's temporal requirements. We show that by formulating a dataflow modelled application and its mapping as a convex optimisation problem, with energy consumption as the objective function, the problem can be solved with a generic convex optimisation solver, producing an energy optimal constant frequency per application task. Our method allows task frequencies to be constrained such that, e.g. one frequency per application or per processor may be achieved. Andrew Nelson 0001, Orlando Moreira, Anca Mariana Molnos, Sander Stuijk, Ba Thang Nguyen, Kees Goossens |
DSD | 6 |
| 2011 | PUMA: Placement Unification with Mapping and Guaranteed Throughput Allocation on an FPGA Using a Hardwired NoCabstractPlatform-based Field Programmable Gate Arrays (FPGAs) have gained popularity for implementing multiprocessor system on chips (MPSoCs). The applications in an MPSoC can have high complexities and stringent Quality-of-Service (QoS) demands. Consequently, the problem of binding an application on an FPGA has become more challenging. An application requires logic and communication resources for computing and transporting data among its IPs. This in turn divides an FPGA into two virtual planes, i.e., logic and communication. Therefore, the available resources in both the FPGA planes should be taken into account by an application binding solution. Our proposed scheme performs placement unification with mapping and allocation (PUMA). This means PUMA accounts for the required (application) to the available (FPGA) resources in both the logic plane and the communication plane, simultaneously. A hardwired Network on Chip (HWNoC) serves as the communication plane for our FPGA, because of its scalable and isolated nature. Moreover, PUMA ensures that a successful binding solution fulfills an application QoS constraints. PUMA is implemented by using cycle-accurate transaction-level SystemC. PUMA performance and scalability is evaluated by using a number of synthetic applications. The PUMA application binding success rate exists in between 35% and 90%. Additionally, the cost of PUMA is evaluated against a real-world H.264 encoder. Muhammad Aqeel Wahlah, Kees Goossens |
DSD | 2 |
| 2011 | A Non-Intrusive Online FPGA Test Scheme Using a Hardwired Network on ChipabstractModern Field Programmable Gate Arrays (FPGAs) posses small features, and have gained popularity in mission-critical systems. However, due to small FPGA features and harsh external conditions that can be faced by a mission-critical system, an FPGA chip can suffer from faults. This in turn raises the need to test an FPGA to ensure a reliable system performance. However, a mission-critical system requires that the test process should be non-intrusive, i.e., applications & FPGA regions that are not being tested remain unaffected. Hence, an online test scheme is required that not only verifies the correctness of an FPGA, but also does not degrade the performance of other, running FPGA applications. In this paper, we propose a Hardwired Network on Chip (HWNoC) as the Test Access Mechanism (TAM). Our online test scheme uses a HWNoC, to perform real-time streaming of test data that is non-intrusive to other communication traffic. Additionally, our online test scheme exhibits approx. 18 and 29 times lower spatial and temporal overheads as compared to existing schemes, respectively. Muhammad Aqeel Wahlah, Kees Goossens |
DSD | 2 |
| 2011 | Time-predictable and composable architectures for dependable embedded systemsabstractEmbedded systems must interact with their real-time environment in a timely and dependable fashion. Most embedded-systems architectures and design processes consider "non-functional" properties such as time, energy, and reliability as an afterthought, when functional correctness has (hopefully) been achieved. As a result, embedded systems are often fragile in their real-time behaviour, and take longer to design and test than planned. Several techniques have been proposed to make real-time embedded systems more robust, and to ease the process of designing embedded systems: Saddek Bensalem, Kees Goossens, Christoph M. Kirsch, Roman Obermaisser, Edward A. Lee, Joseph Sifakis |
EMSOFT | 2 |
| 2011 | Automatic Generation of Efficient Predictable Memory PatternsabstractVerifying firm real-time requirements gets increasingly complex, as the number of applications in embedded systems grows. Predictable systems reduce the complexity by enabling formal verification. However, these systems require predictable software and hardware components, which is problematic for resources with highly variable execution times, such as SDRAM controllers. A predictable SDRAM controller has been proposed that addresses this problem using predictable memory patterns, which are precomputed sequences of SDRAM commands. However, the memory patterns are derived manually, which is a time-consuming and error-prone process that must be repeated for every memory device, and may result in inefficient use of scarce and expensive bandwidth. This paper addresses this issue by proposing three algorithms for automatic generation of efficient memory patterns that provide different trade-offs between run-time of the algorithm and the bandwidth guaranteed by the controller. We experimentally evaluate the algorithms for a number of DDR2/DDR3 memories and show that an appropriate choice of algorithm reduces run-time to less than a second and increases the guaranteed bandwidth by up to 10.2%. Benny Akesson, Williston Hayes Jr., Kees Goossens |
RTCSA (1) | 3 |
| 2011 | Resource-Efficient Real-Time Scheduling Using Credit-Controlled Static-Priority ArbitrationabstractA present-day System-on-Chip (SoC) runs a wide range of applications with diverse real-time requirements. Resources, such as processors, interconnects and memories, are shared between these applications to reduce cost. Resource sharing causes temporal interference, which must be bounded by a suitable resource arbiter. System-level analysis techniques use the service guarantee of the arbiter to ensure that real-time requirements of these applications are satisfied. A service guarantee that underestimates the minimum service provided by an arbiter results in more allocation of resources than needed to satisfy latency and throughput requirements. For instance, a linear service guarantee cannot accurately capture burst service provision by many priority-based schedulers, such as Credit-Controlled Static Priority (CCSP) and Priority-Budget Scheduling (PBS). As a result, the timing analysis of these arbiters becomes too pessimistic. This leads to unnecessary cost penalties since some SoC resources, such as SDRAM bandwidth, are scarce and expensive. This paper addresses this problem for the CCSP arbiter. The two main contributions are: (1) a piecewise linear service guarantee that accurately captures bursty service provisioning, and (2) an equivalent dataflow model of the new service guarantee, which is an essential component to integrate the arbiter with dataflow-based system-level design techniques that analyze the worst-case latency and throughput of real-time applications. The new service guarantee enables efficient resource utilization under CCSP arbitration. Experimental results of an H.263 video decoder application show that memory bandwidth savings from 26% up to 67% can be achieved by using the new service guarantee as compared to the existing linear service guarantee. Firew Siyoum, Benny Akesson, Sander Stuijk, Kees Goossens, Henk Corporaal |
RTCSA (1) | 4 |
| 2010 | The aethereal network on chip after ten years: goals, evolution, lessons, and futureabstractThe goals for the Æthereal network on silicon, as it was then called, were set in 2000 and its concepts were defined early 2001. Ten years on, what has been achieved? Did we meet the goals, and what is left of the concepts? In this paper we answer those questions, and evaluate different implementations, based on a new performance: cost analysis. We discuss and reflect on our experiences, and conclude with open issues and future directions. Kees Goossens, Andreas Hansson 0001 |
DAC | 1 |
| 2010 | Composable Dynamic Voltage and Frequency Scaling and Power Management for Dataflow ApplicationsabstractComposability means that the behaviour of an application, including its timing, is not affected by the absence or presence of other applications. It is required to be able to design, test, and verify applications independently. In this paper we define composable dynamic voltage and frequency scaling (DVFS) hardware, and composable power management. We ensure that the functional and temporal behaviours of an application are not affected by other applications, even when they are power managed. For dataflow applications with worst-case execution times per task, our power management is also predictable, i.e. guarantees end-to-end real-time requirements, even when the application is mapped on multiple processors that are power managed independently. Our method can be used with various DVFS architectures, such as on-chip and off-chip VF regulators. Our FPGA implementation models a system with multiple tiles, each containing a processor with local memory running a real-time operating system (RTOS) and power management. Tiles are interconnected by a network on chip, and communicate using shared memories. Experiments indicate energy savings of 68% w.r.t. no power management, and 40% w.r.t. power gating only. We also demonstrate composability and predictability on the platform in the presence of power management. Kees Goossens, Dongrui She, Aleksandar Milutinovic, Anca Mariana Molnos |
DSD | 1 |
| 2010 | A distributed architecture to check global properties for post-silicon debugabstractPost-silicon validation and debug, or ensuring that software executes correctly on the silicon of a multi-processor system-on-chip (MPSOC) is complicated, as it involves checking global properties that are distributed on the chip. In this paper we define an architecture to non-intrusively observe global properties at run time using distributed monitors. The architecture enables to perform actions when a property holds, such as stopping (part of) the system for inspection. We apply this architecture to the problem of software races that result in incorrect communication between concurrent tasks on different processors. In a case study, where we implemented monitors, event distribution, and instruments to stop communication between intellectual property (IP) blocks, we demonstrate that these races can be detected and classified as timing violations or as FIFO protocol violations. Erik Larsson, Bart Vermeulen, Kees Goossens |
ETS | 3 |
| 2010 | Classification and Analysis of Predictable Memory PatternsabstractThe verification complexity of real-time requirements in embedded systems grows exponentially with the number of applications, as resource sharing prevents independent verification using simulation-based approaches. Formal verification is a promising alternative, although its applicability is limited to systems with predictable hardware and software. SDRAM memories are common examples of essential hardware components with unpredictable timing behavior, typically preventing use of formal approaches. A predictable SDRAM controller has been proposed that provides guarantees on bandwidth and latency by dynamically scheduling memory patterns, which are statically computed sequences of SDRAM commands. However, the proposed patterns become increasingly inefficient as memories become faster, making them unsuitable for DDR3 SDRAM. This paper extends the memory pattern concept in two ways. Firstly, we introduce a burst count parameter that enables patterns to have multiple SDRAM bursts per bank, which is required for DDR3 memories to be used efficiently. Secondly, we present a classification of memory pattern sets into four categories based on the combination of patterns that cause worst-case bandwidth and latency to be provided. Bounds on bandwidth and latency are derived that apply to all pattern types and burst counts, as opposed to the single case covered by earlier work. Experimental results show that these extensions are required to support the most efficient pattern sets for many use-cases. We also demonstrate that the burst count parameter increases efficiency in presence of large requests and enables a wider range of real-time requirements to be satisfied. Benny Akesson, Williston Hayes Jr., Kees Goossens |
RTCSA | 3 |
| 2010 | Bandwidth Analysis of Functional Interconnects Used as Test Access MechanismabstractTest data travels through a System on Chip (SOC) from the chip pins to the Core-Under-Test (CUT) and vice versa via a Test Access Mechanism (TAM). Conventionally, a TAM is implemented using dedicated communication infrastructure. However, also existing functional interconnect, such as a bus or Network on Chip (NOC), can be reused as TAM; this will reduce the overall design effort and associated silicon area. For a given core, its test set, and maximal bandwidth that the functional interconnect can offer between test equipment and core-under-test, our approach instantiates a test wrapper for the core-under-test such that the test length is minimized. Unfortunately, it is unavoidable that along with the test data also unused (idle) bits are transported. This paper presents a holistic TAM bandwidth under-utilization analysis when functional interconnect is considered for test data transportation. We classify the idle bits into four types that refer to the root-cause of bandwidth under-utilization and pinpoint design improvement opportunities. Experimental results show an average bandwidth utilization of 80%, while the remaining 20% is consumed by the idle bits. Ardy van den Berg, Pengwei Ren, Erik Jan Marinissen, Georgi Gaydadjiev, Kees Goossens |
J. Electron. Test. | 5 |
| 2009 | A high-level debug environment for communication-centric debugabstractA large part of a modern SOC's debug complexity resides in the interaction between the main system components. Transaction-level debug moves the abstraction level of the debug process up from the bit and cycle level to the transactions between IP blocks. In this paper we raise the debug abstraction level further, by utilising structural and temporal abstraction techniques, combined with debug data interpretation and logical communication views. The combination of these techniques and views allow us, among others, to single-step and observe the operation of the network on a per-connection basis. As an example, we show how these higher-level abstractions have been implemented in the debug environment for the AEligthereal NOC architecture and present a generic debug API, which can be used to visualise an SOC's state at the logical communication level. Kees Goossens, Bart Vermeulen, Ashkan Beyranvand Nejad |
DATE | 1 |
| 2009 | Aelite: A flit-synchronous Network on Chip with composable and predictable servicesabstractTo accommodate the growing number of applications integrated on a single chip, Networks on Chip (NoC) must offer scalability not only on the architectural, but also on the physical and functional level. In addition, real-time applications require Guaranteed Services (GS), with latency and throughput bounds. Traditionally, NoC architectures only deliver scalability on two of the aforementioned three levels, or do not offer GS. In this paper we present the composable and predictable aelite NoC architecture, that offers only GS, based on flit-synchronous Time Division Multiplexing (TDM). In contrast to other TDM-based NoCs, scalability on the physical level is achieved by using mesochronous or asynchronous links. Functional scalability is accomplished by completely isolating applications, and by having a router architecture that does not limit the number of service levels or connections. We demonstrate how aelite delivers the requested service to hundreds of simultaneous connections, and does so with 5 times less area compared to a state-of-the-art NoC. Andreas Hansson 0001, Mahesh Subburaman, Kees Goossens |
DATE | 3 |
| 2009 | Composable Resource Sharing Based on Latency-Rate ServersabstractVerification of application requirements is becoming a bottleneck in system-on-chip design, as the number of applications grows. Traditionally, the verification complexity increases exponentially with the number of applications and must be repeated if an application is added, removed, or modified. Predictable systems offering lower bounds on performance have been proposed to manage the increasing verification complexity, although this approach is only applicable to a restricted set of applications and systems. Composable systems, on the other hand, completely isolate applications in both the value and time domains, allowing them to be independently verified. However, existing approaches to composable system design are either restricted to applications that can be statically scheduled, or share resources using time-division multiplexing, which cannot efficiently satisfy tight latency requirements. In this paper, we present an approach to composable resource sharing based on latency-rate servers that supports any arbiter belonging to the class, providing a larger solution space for a given set of requirements. The approach can be combined with formal performance analysis using a variety of well-known modeling frame works. We furthermore propose an architecture for a resource front end that implements our concepts and provides composable service for any resource with bounded service time. The architecture supports both systems with buffers dimensioned to prevent overflow and systems with smaller buffers, where overflow is prevented with flow control. Finally, we experimentally demonstrate the usefulness of our approach with a simple use case sharing an SRAM memory. Benny Akesson, Andreas Hansson 0001, Kees Goossens |
DSD | 3 |
| 2009 | Internet-Router Buffered Crossbars Based on Networks on ChipabstractThe scalability and performance of the Internet depends critically on the performance of its packet switches. Current packet switches are based on single-hop crossbar fabrics, with line cards that use virtual output-queueing to reduce head-of-line blocking. In this paper we propose to use a multi-hop network on a chip (NOC) as the crossbar fabric, with FIFO-queued line cards. The use of a multi-hop crossbar fabric has several advantages. (1) Speed-up, i.e. the crossbar fabric can operate faster because NOC inter-router wires are shorter than those in a single-hop crossbar, and because arbitration is distributed instead of centralised. (2) Load balancing because paths from different input-output port pairs share the same router buffers, unlike the internal buffers of buffered crossbar fabric that are dedicated to a single inputoutput pair. (3) Path diversity allows traffic from an input port to follow different paths to its destination output port. This results in further load balancing, especially for non-uniform traffic patterns. (4) Simpler line-card design: the use of FIFOs on the line cards simplifies both the line cards and the (interchip) flow control between the crossbar fabric and line cards, reducing the number of (expensive) chip pins required for flow control. (5) Scalability, in the sense that the crossbar speed is independent of the number of ports, which is not the case for single-hop crossbar fabrics. We analyzed the performance of our architecture both analytically and by simulation, and show that it performs well for a wide range of traffic conditions and switch sizes. Additionally we prototyped a 32 × 32 NOC-based crossbar fabric in a 65 nm CMOS technology. The unoptimised implementation operates at 413 MHz, achieving an aggregate throughput in excess of 1010ATM cells per second. Kees Goossens, Lotfi Mhamdi, Iria Varela Senin |
DSD | 1 |
| 2009 | Conservative Dynamic Energy Management for Real-Time Dataflow Applications Mapped on Multiple ProcessorsabstractVoltage-frequency scaling (VFS) trades a linear processor slowdown for a potentially quadratic reduction in energy consumption. Complex dependencies may exist between different tasks of an application. The impact of VFS on the end-to-end application performance is difficult to predict, especially when these tasks are mapped on multiple processors that are scaled independently. This is a problem for real-time (RT) applications that require guaranteed end-to-end performance. In this paper we first classify the slack existing in RT applications consisting of multiple dependent tasks mapped on multiple processors independently using VFS, resulting in static, work, and share slack. Then we concentrate on work and share slack as they can only be detected at run time, thus their conservative use is challenging. We propose SlackOS, a dynamic, dependency-aware, task scheduling that conservatively scales the voltage and frequency of each processor, to respect RT deadlines. When applied to a H.264 application, our method delivers 22% to 33% energy reduction, compared to dynamic RT scheduling that is not energy aware. Anca Mariana Molnos, Kees Goossens |
DSD | 2 |
| 2009 | Efficient Multicast Support in Buffered Crossbars using Networks on ChipabstractThe Internet growth coupled with the variety of its services is creating an increasing need for multicast traffic support by backbone routers and packet switches. Recently, buffered crossbar (CICQ) switches have shown high potential in efficiently handling multicast traffic. However, they were unable to deliver optimal performance despite their expensive and complex crossbar fabric. This paper proposes an enhanced CICQ switching architecture suitable for multicast traffic. Instead of a dedicated internal crosspoint buffer for every input-output pair of ports, the crossbar is designed as a multi-hop Network on Chip (NoC). Designing the crossbar as a NoC offers several advantages such as low latency, internal fabric load balancing and path diversity. It also obviates the requirement of the virtual output queuing by allowing simple FIFO structure without performance degradation. We designed appropriate routing for the NoC as well as on-chip router scheduling and tested its performance under a wide range of input multicast traffic. Simulations results showed that our proposal outperforms the CICQ architecture and offers a viable architectural alternative. We also studied the effect of various parameters such as the depth of the NoC as well as the speedup requirement for high-bandwidth multicast switching. Iria Varela Senin, Lotfi Mhamdi, Kees Goossens |
GLOBECOM | 3 |
| 2009 | Modeling reconfiguration in a FPGA with a hardwired network on chipabstractWe propose that FPGAs use a hardwired network on chip (HWNOC) as a unified interconnect for functional communications (data and control) as well as configuration (bitstreams for soft IP). In this paper we model such a platform. Using the HWNOC applications mapped on hard or soft IPs are set up and removed using memory-mapped communications. Peer-to-peer streaming data is used to communicate data between IPs, and also to transport configuration bitstreams. The composable nature of the HWNOC ensures that applications can be dynamically configured, programmed, and can operate, without affecting other running (real-time) applications. We describe this platform and the steps required for dynamic reconfiguration of IPs. We then model the hardware, i.e. HWNOC and hard and soft IPs, in cycle-accurate transaction-level SystemC. Next, we model its dynamic behavior, including bitstream loading, HWNOC programming, dynamic (re)configuration, clocking, reset, and computation. Muhammad Aqeel Wahlah, Kees Goossens |
IPDPS | 2 |
| 2009 | Efficient Service Allocation in Hardware Using Credit-Controlled Static-Priority ArbitrationabstractResources in contemporary systems-on-chip (SoC) are shared between applications to reduce cost. Access to shared resources is provided by arbiters that require a small hardware implementation and must run at high speed. To manage heavily loaded resources, such as memory channels, it is also important that the arbiter minimizes over allocation. A Credit-Controlled Static-Priority (CCSP) arbiter comprised of a rate regulator and a static-priority scheduler has been proposed for scheduling access to SoC resources. The proposed rate regulator, however, is not straight-forward to implement in hardware, and assumes that service is allocated with infinite precision. In this paper, we introduce a fast and small hardware implementation of the CCSP rate regulator and formally prove its correctness. We also show an efficient way of representing the allocated service in hardware with finite precision. Based on this representation, we define and evaluate two allocation strategies, and derive tight bounds on their respective over allocations. We show that increasing the precision of the implementation results in an exponential reduction in maximum over allocation at the cost of a linear increase in area. We demonstrate that the allocation strategy has a large impact on the allocation success rate for use cases with high load. Finally, we compare CCSP to traditional frame-based approaches and conclude that having a fine allocation granularity that is decoupled from latency is essential to manage highly loaded resources in real-time systems. Benny Akesson, Elisabeth F. M. Steffens, Kees Goossens |
RTCSA | 3 |
| 2009 | CoMPSoC: A template for composable and predictable multi-processor system on chipsabstractA growing number of applications, often with firm or soft real-time requirements, are integrated on the same System on Chip, in the form of either hardware or software intellectual property. The applications are started and stopped at run time, creating different use-cases. Resources, such as interconnects and memories, are shared between different applications, both within and between use-cases, to reduce silicon cost and power consumption. The functional and temporal behaviour of the applications is verified by simulation and formal methods. Traditionally, designers resort to monolithic verification of the system as whole, since the applications interfere in shared resources, and thus affect each other's behaviour. Due to interference between applications, the integration and verification complexity grows exponentially in the number of applications, and the task to verify correct behaviour of concurrent applications is on the system designer rather than the application designers. In this work, we propose a Composable and Predictable Multi-Processor System on Chip (CoMPSoC) platform template. This scalable hardware and software template removes all interference between applications through resource reservations. We demonstrate how this enables a divide-and-conquer design strategy, where all applications, potentially using different programming models and communication paradigms, are developed and verified independently of one another. Performance is analyzed per application, using state-of-the-art dataflow techniques or simulation, depending on the requirements of the application. These results still apply when the applications are integrated onto the platform, thus separating system-level design and application design. Andreas Hansson 0001, Kees Goossens, Marco Bekooij, Jos Huisken |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2008 | Bandwidth Analysis for Reusing Functional Interconnect as Test Access MechanismabstractTest data travels through a System-on-Chip (SOC) from the chip pins to the module-under-test and vice versa via a Test Access Mechanism (TAM). Conventionally, a TAM is implemented with dedicated wires. However, also existing functional interconnect, such as a bus or Network-on-Chip (NOC), can be reused as TAM. This will reduce the overall design effort and the silicon area. For a given module, its test set, and maximal bandwidth that the functional interconnect can offer between ATE and module-under-test, our approach designs a test wrapper for the module-under-test such that the test length is minimized. Unfortunately, it is unavoidable that with the test data also unused (idle) bits are transported. This paper presents a TAM bandwidth utilization analysis and techniques for idle bits reduction, to minimize the test length. We classify the idle bits into four types which explain the reason for bandwidth under-utilization and pinpoint design improvement opportunities. Experimental results show an average bandwidth utilization of 80%, while the remaining 20% is consumed by the idle bits. Ardy van den Berg, Pengwei Ren, Erik Jan Marinissen, Georgi Gaydadjiev, Kees Goossens |
ETS | 5 |
| 2008 | Hardwired Networks on Chip in FPGAs to Unify Functional and Con?guration Interconnects
Kees Goossens, Martijn T. Bennebroek, Jae Young Hur, Muhammad Aqeel Wahlah |
NOCS | 1 |
| 2008 | Applying Dataflow Analysis to Dimension Buffers for Guaranteed Performance in Networks on Chip
Andreas Hansson 0001, Maarten Wiggers, Arno Moonen, Kees Goossens, Marco Bekooij |
NOCS | 4 |
| 2008 | Debugging Distributed-Shared-Memory Communication at Multiple Granularities in Networks on Chip
Bart Vermeulen, Kees Goossens, Siddharth Umrani |
NOCS | 2 |
| 2008 | Real-Time Scheduling Using Credit-Controlled Static-Priority ArbitrationabstractThe convergence of application domains in new systems-on-chip (SoC) results in systems with many applications with a mix of soft and hard real-time requirements. To reduce cost, resources, such as memories and interconnect, are shared between applications. However, resource sharing introduces interference between the sharing applications, making it difficult to satisfy their real-time requirements. Existing arbiters do not efficiently satisfy the requirements of applications in SoCs, as they either couple rate or allocation granularity to latency, or cannot run at high speeds in hardware with a low-cost implementation. The contribution of this paper is an arbiter called credit- controlled static-priority (CCSP), consisting of a rate regulator and a static-priority scheduler. The rate regulator isolates applications by regulating the amount of provided service in a way that decouples allocation granularity and latency. The static-priority scheduler decouples latency and rate, such that low latency can be provided to any application, regardless of the allocated rate. We show that CCSP belongs to the class of latency-rate servers and guarantees the allocated rate within a maximum latency, as required by hard real-time applications. We present a hardware implementation of the arbiter in the context of a DDR2 SDRAM controller. An instance with six ports running at 200 MHz requires an area of 0.0223 mm2in a 90 nm CMOS process. Benny Akesson, Elisabeth F. M. Steffens, Eelke Strooisma, Kees Goossens |
RTCSA | 4 |
| 2008 | A monitoring-aware network-on-chip design flow
Calin Ciordas, Andreas Hansson 0001, Kees Goossens, Twan Basten |
J. Syst. Archit. | 3 |
| 2007 | Congestion-controlled best-effort communication for networks-on-chipabstractCongestion has negative effects on network performance. In this paper, a novel congestion control strategy is presented for networks-on-chip (NoC). For this purpose we introduce a new communication service, congestion-controlled best-effort (CCBE). The load offered to a CCBE connection is controlled based on congestion measurements in the NoC. Link utilization is monitored as a congestion measure, and transported to a model predictive controller (MPC). Guaranteed bandwidth and latency connections in the NoC are used for this, to assure progress of link utilization data in a congested NoC. We also present a simple but effective model for link utilization for the model-based predictions. Experimental results show that the presented strategy is effective and has reaction speeds of several microseconds which is considered acceptable for realtime embedded systems Jan Willem van den Brand, Calin Ciordas, Kees Goossens, Twan Basten |
DATE | 3 |
| 2007 | Undisrupted quality-of-service during reconfiguration of multiple applications in networks on chipabstractNetworks on chip (NoC) have emerged as the design paradigm for scalable system on chip (SoC) communication infrastructure. Due to convergence, a growing number of applications are integrated on the same chip. When combined, these applications result in use-cases with different communication requirements. The NoC is configured per use-case and traditionally all running applications are disrupted during use-case transitions, even those continuing operation. In this paper we present a model that enables partial reconfiguration of NoCs and a mapping algorithm that uses the model to map multiple applications onto a NoC with undisrupted quality-of-service during reconfiguration. The performance of the methodology is verified by comparison with existing solutions for several SoC designs. We apply the algorithm to a mobile phone SoC with telecom, multimedia and gaming applications, reducing NoC area by more than 17% and power consumption by 50% compared to a state-of-the-art approach Andreas Hansson 0001, Martijn Coenen, Kees Goossens |
DATE | 3 |
| 2007 | Communication-Centric SoC Debug Using TransactionsabstractThe growth in system-on-chip complexity puts pressure on system verification. Due to limitations in the pre-silicon verification process, errors in hardware and software slip through to the stage when silicon and the complete software stack are first brought together. Finding the remaining errors at this stage is becoming increasing difficult. We propose that debugging should be communication-centric at first and based on transactions. We combine run-time, on-chip abstraction of system data to the transaction level, with system-level debug control over the communication infrastructure. We prove our concepts and architecture with a gate-level implementation that includes a network-on-chip, breakpoint monitors, clock and reset control (all programmable through an IEEE 1149.1 TAP), and give a quantification of the associated hardware cost. Bart Vermeulen, Kees Goossens, Remco van Steeden, Martijn T. Bennebroek |
ETS | 2 |
| 2007 | Transaction-Based Communication-Centric DebugabstractThe behaviour of systems on chip (SOC) is complex because they contain multiple processors that interact through concurrent interconnects, such as networks on chip (NOC). Debugging such SOCs is hard. Based on a classification of debug scope and granularity, we propose that debugging should be communication-centric and based on transactions. Communication-centric debug focuses on the communication and the synchronisation between the IP blocks, which are implemented by the interconnect using transactions. We define and implement a modular debug architecture, based on NOC, monitors, and a dedicated high-speed event-distribution broadcast interconnect. The manufacturing-test scan chains and IEEE1149.1 test access ports (TAP) are re-used for configuration and debug data read-out. Our debug architecture requires only small changes to the functional architecture. The additional area cost is limited to the monitors and the event distribution interconnect, which are 4.5% of the NOC area, or less than 0.2% of the SOC area. The debug architecture runs at NOC functional speed and reacts very quickly to debug events to stop the SOC close in time to the condition that raised the event. The speed at which data is retrieved from the SOC after stopping using the TAP is 10 MHz. We prove our concepts and architecture with a gate-level implementation that includes the NOC, event distribution interconnect, and clock, reset, and TAP controllers. We include gate-level signal traces illustrating debug at message and transaction levels Kees Goossens, Bart Vermeulen, Remco van Steeden, Martijn T. Bennebroek |
NOCS | 1 |
| 2007 | Trade-offs in the Configuration of a Network on Chip for Multiple Use-CasesabstractSystems on chip (SoC) are becoming increasingly complex, with a large number of applications integrated on the same chip. Such a system often supports a large number of use-cases and is dynamically reconfigured when platform conditions or user requirements change. Networks on chip (NoC) offer the designer unsurpassed runtime flexibility. This flexibility stems from the programmability of the individual routers and network interfaces. When a change in use-case occurs, the application task graph and the network connections change. To mitigate the complexity in programming the many registers controlling the NoC, an abstraction in the form of a configuration library is needed. In addition, such a library must leave the modified system in a consistent state, from which normal operation can continue. In this paper we present the facilities for controlling change in a reconfigurable NoC. We show the architectural additions and the many trade-offs in the design of a run-time library for NoC reconfiguration. We qualitatively and quantitatively evaluate the performance, memory requirements, predictability and reusability of the different implementations Andreas Hansson 0001, Kees Goossens |
NOCS | 2 |
| 2006 | Mapping and configuration methods for multi-use-case networks on chipsabstractTo provide a scalable communication infrastructure for systems on chips (SoCs), networks on chips (NoCs), a communication centric design paradigm is needed. To be cost effective, SoCs are often programmable and integrate several different applications or use-cases on to the same chip. For the SoC platform to support the different use-cases, the NoC architecture should satisfy the performance constraints of each individual use-case. In this work we motivate the need to consider multiple use-cases during the NoC design process. We present a method to efficiently map the applications on to the NoC architecture, satisfying the design constraints of each individual use-case. We also present novel ways to dynamically reconfigure the network across the different use-cases and explore the possibility of integrating dynamic voltage and frequency scaling (DVS/DFS) techniques with the use-case centric NoC design methodology. We validate the performance of the design methodology on several SoC applications. The dynamic reconfiguration of the NoC integrated with DVS/DFS schemes results in large power savings for the resulting NoC systems Srinivasan Murali, Martijn Coenen, Andrei Radulescu, Kees Goossens, Giovanni De Micheli |
ASP-DAC | 4 |
| 2006 | A methodology for mapping multiple use-cases onto networks on chipsabstractA communication-centric design approach, networks on chips (NoCs), has emerged as the design paradigm for designing a scalable communication infrastructure for future systems on chips (SoCs). As technology advances, the number of applications or use-cases integrated on a single chip increases rapidly. The different use-cases of the SoC have different communication requirements (such as different bandwidth, latency constraints) and traffic patterns. The underlying NoC architecture has to satisfy the constraints of all the use-cases. In this work, we present a methodology to map multiple use-cases onto the NoC architecture, satisfying the constraints of each use-case. We present dynamic re-configuration mechanisms that match the NoC configuration to the communication characteristics of each use-case, also accounting for use-cases that can run in parallel. The methodology is applied to several real and synthetic SoC benchmarks, which result in a large reduction in NoC area (an average of 80%) and power consumption (an average of 54%) compared to traditional design approaches Srinivasan Murali, Martijn Coenen, Andrei Radulescu, Kees Goossens, Giovanni De Micheli |
DATE | 4 |
| 2006 | A Monitoring-Aware Network-on-Chip Design FlowabstractNetworks-on-chip (NoC) are a scalable interconnect solution for systems on chip and are rapidly becoming reality. Monitoring is a key enabler for debugging or performance analysis and quality-of-service techniques. The NoC design problem and the NoC monitoring problem cannot be treated in isolation. We propose a monitoring-aware NoC design flow able to take into account the monitoring requirements in general. We illustrate our flow with a debug driven monitoring case study of transaction monitoring. By treating the NoC design and monitoring problems in synergy, the area cost of monitoring can be limited to 3-20% in general Calin Ciordas, Andreas Hansson 0001, Kees Goossens, Twan Basten |
DSD | 3 |
| 2006 | Wrapper Design for the Reuse of Networks-on-Chip as Test Access MechanismabstractThis paper proposes a wrapper design for interconnects with guaranteed bandwidth and latency services and on-chip protocol. We demonstrate that these interconnects abstract the interconnect details and provide predictability in the data transfer, which are desirable not only for the functional domain but also for the test application. The proposed wrapper is implemented in VHDL and integrated to the Æthereal NoC. The results show the impact of of bandwidth in the core test time. The wrapper area and core test time are compared with a wrapper design for dedicated TAM. Alexandre M. Amory, Kees Goossens, Erik Jan Marinissen, Marcelo Lubaszewski, Fernando Gehm Moraes |
ETS | 2 |
| 2006 | NoC monitoring: impact on the design flowabstractNetworks-on-chip (NoCs) are a scalable interconnects solution to large scale multiprocessor systems on chip and are rapidly becoming reality. As the ratio of embedded cores per I/O pin increases, the run-time observability becomes a bottleneck. Run-time NoC monitoring can alleviate this problem. As NoCs are the result of sophisticated synthesis design flows, monitoring must be taken into account during this process. We present several scalable alternatives for NoC monitoring. The alternatives vary from using physically separated interconnects for user data and monitoring data, to a completely shared single interconnect. For each alternative we evaluate area cost, required design flow modifications, non-intrusiveness and reusability of monitoring resources for application communication traffic. An interesting trade-off is presented showing that what is area efficient requires efforts in modifying the NoC design flow and in achieving non-intrusiveness. All the experiments are done in the context of the /Ethereal NoC and design flow Calin Ciordas, Kees Goossens, Andrei Radulescu, Twan Basten |
ISCAS | 2 |
| 2006 | Comparison of An Æthereal Network on Chip and A Traditional Interconnect for A Multi-Processor DVB-T System on ChipabstractGrowing complexity of multiprocessor systems on chip (MP-SoC) requires future communication resources that can only be met by highly scalable architectures. Networks-on-Chip (NoCs) offer this scalability and other advantages like modularity, quality-of-service (QoS), possibly smaller area footprint and lower power dissipation. Although many papers describe the advantages of NoCs and describe techniques to apply NoCs on certain application domains, few have actually employed the complete design chain to make a netlist level implementation and area comparison (Steenhof et al., 2006) and (Angiolini et al., 2006). This paper describes the application of the AEligthereal NoC to an existing bus-based MP-SoC design and an area comparison with the original interconnects structure down to netlist level Chris Bartels, Jos Huisken, Kees Goossens, Patrick Groeneveld, Jef L. van Meerbergen |
VLSI-SoC | 3 |
| 2005 | A Design Flow for Application-Specific Networks on Chip with Guaranteed Performance to Accelerate SOC Design and VerificationabstractSystems on chip (SOC) are composed of intellectual property blocks (IP) and interconnect. While mature tooling exists to design the former, tooling for interconnect design is still a research area. In this paper we describe an operational design flow that generates and configures application-specific network on chip (NOC) instances, given application communication requirements. The NOC can be simulated in SystemC and RTL VHDL. An independent performance verification tool verifies analytically that the NOC instance (hardware) and its configuration (software) together meet the application performance requirements. The Æthereal NOC's guaranteed performance is essential to replace time-consuming simulation by fast analytical performance validation. As a result, application-specific NOCs that are guaranteed to meet the application's communication requirements are generated and verified in minutes, reducing the number of design iterations. A realistic MPEG SOC example substantiates our claims. Kees Goossens, John Dielissen, Om Prakash Gangwal, Santiago González Pestana, Andrei Radulescu, Edwin Rijpkema |
DATE | 1 |
| 2005 | An efficient on-chip NI offering guaranteed services, shared-memory abstraction, and flexible network configurationabstractWe present a network interface (NI) for an on-chip network. Our NI decouples computation from communication by offering a shared-memory abstraction, which is independent of the network implementation. We use a transaction-based protocol to achieve backward compatibility with existing bus protocols such as AXI, OCP, and DTL. Our NI has a modular architecture, which allows flexible instantiation. It provides both guaranteed and best-effort services via connections. These are configured via NI ports using the network itself, instead of a separate control interconnect. An example instance of this NI with four ports has an area of 0.25 mm/sup 2/ after layout in 0.13-/spl mu/m technology, and runs at 500 MHz. Andrei Radulescu, John Dielissen, Santiago González Pestana, Om Prakash Gangwal, Edwin Rijpkema, Paul Wielage, Kees Goossens |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2005 | An event-based monitoring service for networks on chipabstractNetworks on chip (NoCs) are a scalable interconnect solution for multiprocessor systems on chip. We propose a generic reconfigurable online event-based NoC monitoring service, based on hardware probes attached to NoC components, offering run-time observability of NoC behavior and supporting system-level debugging. We present a probe architecture, its programming model, traffic management strategies, and a cost analysis. We prove feasibility via a prototype implementation for the Æthereal NoC. Two MPEG NoC examples show that the monitoring service area, without advanced optimizations, is 17--24% of the NoC area. Two realistic monitoring examples show that monitoring traffic is several orders of magnitude lower than the 2GB/s/link raw bandwidth. Calin Ciordas, Twan Basten, Andrei Radulescu, Kees Goossens, Jef L. van Meerbergen |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2004 | Cost-Performance Trade-Offs in Networks on Chip: A Simulation-Based ApproachabstractA challenge facing designers of systems on chip (SoC) containing networks on chip (NoC) is to find NoC instances that balance the cost (e.g. area) and performance (e.g. latency and throughput). In this paper we present a simulation-based approach to address this problem. We use XML to instantiate network components (routers, network interfaces) and their composition. NoCs are evaluated in terms of cost and performance by sweeping over different parameters (e.g. network topology, network interface queue depth). We then show, how we can obtain trade-off plots by using the results obtained with our simulation environment. Finally, by means of two examples we illustrate how trade-off plots can help the NoC designers in selecting the right network based on a set of different constraints. Santiago González Pestana, Edwin Rijpkema, Andrei Radulescu, Kees Goossens, Om Prakash Gangwal |
DATE | 4 |
| 2004 | An Efficient On-Chip Network Interface Offering Guaranteed Services, Shared-Memory Abstraction, and Flexible Network ConfigurationabstractIn this paper we present a network interface for an on-chip network. Our network interface decouples computation from communication by offering a shared-memory abstraction, which is independent of the network implementation. We use a transaction-based protocol to achieve backward compatibility with existing bus protocols such as AXI, OCP and DTL. Our network interface has a modular architecture, which allows flexible instantiation. It provides both guaranteed and best-effort services via connections. These are configured via network interface ports using the network itself, instead of a separate control interconnect. An example instance of this network interface with 4 ports has an area of 0.143 mm/sup 2/ in a 0.13 /spl mu/m technology, and runs at 500 MHz. Andrei Radulescu, John Dielissen, Kees Goossens, Edwin Rijpkema, Paul Wielage |
DATE | 3 |
| 2003 | Trade Offs in the Design of a Router with Both Guaranteed and Best-Effort Services for Networks on Chip
Edwin Rijpkema, Kees Goossens, Andrei Radulescu, John Dielissen, Jef L. van Meerbergen, Paul Wielage, Erwin Waterlander |
DATE | 2 |
| 2002 | The Cost of Communication Protocols and Coordination Languages in Embedded Systems
Kees Goossens, Om Prakash Gangwal |
COORDINATION | 1 |
| 2002 | Networks on Silicon: Combining Best-Effort and Guaranteed ServicesabstractWe advocate a network on silicon (NOS) as a hardware architecture to implement communication between IP cores in future technologies, and as a software model in the form of a protocol stack to structure the programming of NOSs. We claim guaranteed services are essential. In the ETHEREAL NOS they pervade the NOS as a requirement for hardware design, and as foundation for software programming. Kees Goossens, Paul Wielage, Ad M. G. Peeters, Jef L. van Meerbergen |
DATE | 1 |
| 2002 | Networks on Silicon: Blessing or Nightmare?abstractContinuing VLSI technology scaling raises several deep submicron (DSM) problems like relatively slow interconnect, power dissipation and distribution, and signal integrity. Those problems are encountered particularly on long wires for global interconnect. As clock frequencies increase, scaled wires become relatively slower and on-chip communication will be the limiting performance factor of future chips. We explain why efficiently sharing of the wires for long distance communication is the solution to this problem. We introduce networks on silicon (NoS), that route packets over shared (semi)-global wires. NoS performance is expected to be high, but comes at a cost. Balancing the performance and cost of a NoS is a major challenge, and we believe busses still have a role to play. Paul Wielage, Kees Goossens |
DSD | 2 |
| 1998 | The petrol approach to high-level power estimationabstractHigh-level power estimation is essential for designing complex low-power ICs. However, the lack of flexibility, or restriction to synthesizable code of previously presented high-level power estimation approaches limits their use. In this paper we present a novel, more general and flexible high-level power estimation approach, that avoids these limitations. Petrol, as we call it, is not limited to specialized application domains, synthesizable VHDL, or data path parts of a design. We show that glitches can be usefully modeled at higher levels of abstraction. The Petrol approach shows good correlation with gate-level power estimates. It is currently used for commercial designs. 1 Introduction The huge integration capability of modern technologies (several millions of transistors) allows very complex systems on a single IC, such as MPEG2 encoders and digital audio channel decoders. For these systems power dissipation is a critical issue, since it is often the bottleneck for further inte... Rafael Peset Llopis, Kees Goossens |
ISLPED | 2 |