EDBT 2026 Demo / reviewers in the wild / expert
Jens Sparsø
dblp:42/2580
· DBLP profile ↗
41ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0002-0961-9438ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PeakEngine: A Deterministic On-the-Fly Pruning Neural Network Accelerator for Hearing InstrumentsabstractRecurrent neural networks (RNNs) are well-suited for sequential tasks such as speech enhancement (SE). However, their performance comes with high-computational complexity and latency. This impedes their deployment to battery-powered and resource-constrained hearing instruments (HIs) that need to operate for 16–18 h daily at only a few milliwatts (mW). In this article, we introduce PeakEngine, a configurable ASIC accelerator that decreases the amount of computation and memory accesses, and thus latency, in a gated recurrent unit (GRU) by means of adaptive inference. The reduction is achieved by on-the-fly pruning that selects the top$K$elements based on magnitudes of delta changes across timesteps from both input and hidden state sequences. Since$K$is constant, it results in a deterministic execution time. PeakEngine is synthesized in a 22-nm CMOS process, and the simulations show that it dissipates$11.83 ~\mu \text{J}$per inference for the baseline (unpruned) network and only 4.14–$5.04 ~\mu \text{J}$for the pruned networks, with maximum acceptable degradation to no degradation in the improvement in audio quality and intelligibility. Moreover, the inference is on average sped up 2.2–$2.97\times $, hence meeting the real-time requirements imposed by a HI application. To the best of our knowledge, PeakEngine is the first ASIC accelerator for deterministic and dynamic pruning in RNNs targeting HIs and SE. Zuzana Jelcicová, Evangelia Kasapaki, Oskar Andersson, Jens Sparsø |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Asynchronous Circuit Design in Chisel Using Phase-Decoupled Click ElementsabstractThis paper explores using the hardware construction language Chisel to describe, implement, and test asynchronous circuits. The Chisel language supports clocked registers and the use of multiple clocks. It is hence capable of describing asynchronous designs that are built from clocked flip-flops and logic, such as circuits using the asynchronous (phase-decoupled) click-element template that we target in this paper. Previous work has used VHDL to describe a library of asynchronous data-flow components as well as an example circuit demonstrating the use of the components. In this paper, we re-implement this work using Chisel and compare against VHDL. As we see, Chisel offers more efficient and elegant support for describing asynchronous handshake channels and for connecting components. Compared to the VHDL version, the code is reduced by up to 80 %. The example circuit computes the greatest common divisor of two integers, and it is synthesized and tested on an FPGA board. All code is open source giving those interested easy access to asynchronous design and prototyping. Kasper Juul Hesse Rasmussen, Tjark Petersen, Jens Sparsø |
DSD | 3 |
| 2023 | A Min-Heap-Based Accelerator for Deterministic On-the-Fly Pruning in Neural NetworksabstractThis paper addresses the design of an area and energy efficient hardware accelerator that supports on-the-fly pruning in neural networks. In a layer of N neurons, the accelerator selects the top K neurons in every timestep. As K is fixed, the runtime of the pruned network is deterministic, which is an important property in real-time systems such as hearing aids. As a first contribution, we propose to use a min-heap for the top K selection due to its efficient data structure and low time complexity. As a second contribution, we design and implement a hardware accelerator for dynamic pruning that is based on the min-heap algorithm. The heap memory storing the top K neurons and their index is realized as a 3-port standard cell-based memory implemented with latches. As a third contribution, we evaluate the energy savings from pruning of a gated recurrent unit used in a neural network for speech enhancement (regression task). Our experiments demonstrate energy savings of ~78% without degrading the SNR improvement, and up to ~93% while reducing the SNR improvement by 0.1 - 1.11 dB. Moreover, the overhead of the hardware accelerator constitutes negligible ~0.5% of the total energy. The accelerator is implemented in a 22nm CMOS process. Zuzana Jelcicová, Evangelia Kasapaki, Oskar Andersson, Jens Sparsø |
ISCAS | 4 |
| 2022 | Comparing timed-division multiplexing and best-effort networks-on-chipabstractBest-effort (BE) networks-on-chips (NOCs) are usually preferred over time-division multiplexed (TDM) NOCs in multi-core platforms because they are work-conserving and have lower (zero-load) latency. On the other hand, BE NOCs are significantly more expensive to implement than TDM NOCs because of their virtual channel buffers, allocators/arbiters, and (credit-based) flow control; functionality that a TDM NOC avoids altogether. The objective of this paper is to compare the performance of BE and TDM NOCs, taking hardware cost into consideration. The networks are compared using graphs showing average latency as a function of offered load. For the BE NOCs, we use the BookSim simulator, and for the TDM NOCs, we derive a queuing theory model and an associated TDM NOC simulator. Through experiments with both router architectures, packet length, link width, and different traffic patterns, we show that for the same hardware cost, a TDM NOC can provide higher bandwidth and comparable latency. We also show that the packet length is the most important factor affecting the TDM period, which again is the primary factor affecting latency. The best TDM NOC design for BE traffic uses single flit packets, wide links/flits, and a router with two pipeline stages: link and router traversal. Jens Sparsø, Hans Jakob Damsgaard, Dimitrios Katsamanis, Martin Schoeberl |
J. Syst. Archit. | 1 |
| 2021 | Evaluating a Time-Triggered Runtime System by Distributing a Flight ControllerabstractWith the recent advancements in the Industrial Internet of Things and Industry 4.0, cyber-physical systems have become increasingly inter-connected. It is becoming a challenge to maintain the same quality-of-control and time-predictability of computation and communication required by safety-critical hard real-time systems as previously achieved through non-distributed architectures. This paper examines the problem of implementing and distributing a closed-loop command-control system over an Ethernet network with guaranteed timing bounds. To achieve bounded communication and computation time, we use an open-source software framework running on the T-CREST platform combined with a TTEthernet network star topology. We evaluate its quality-of-control performance in our experimental setup and compare the results against single-core and multi-core implementations. The proposed distributed time-triggered runtime system executes with jitter below 10µs and can perform a stable flight scenario as verified by the benchmark implementation. Eleftherios Kyriakakis, Jens Sparsø, Martin Schoeberl |
ETFA | 2 |
| 2021 | Synchronizing Real-Time Tasks in Time-Triggered NetworksabstractIn order to guarantee end-to-end latency and minimal jitter in distributed real-time systems, it is necessary to provide tight synchronization between computation and communication. This requires time-predictable execution of tasks across all processing nodes, and the use of a network protocol that can provide a global time base and bounded communication latency. TTEthernet is one such industrial communication protocol. This paper investigates the synchronization of the task execution schedule with the underlying communication schedule, and we propose an open-source software framework for time-triggered end-systems. We present the implementation of a static cyclic task schedule, on a time-predictable platform that is integrated within a TTEthernet network and synchronized with the communication schedule. We evaluate the presented framework by developing a simple one-sensor, one-actuator industrial control example, distributed over three nodes that communicate over a single TTEthernet switch. The presented real-time system can exchange messages with minimal jitter as the distributed tasks are synchronized over the TTEthernet network with about 1.6 us precision. Due to the tight time synchronization, the system can operate stably with zero missed frames, using a single receiver and a single transmitter buffer. Eleftherios Kyriakakis, Jens Sparsø, Peter P. Puschner, Martin Schoeberl |
ISORC | 2 |
| 2020 | Synchronizing Real-Time Tasks in Time-Aware Networks: Work-in-ProgressabstractDistributed safety-critical systems require both time-predictable task execution and communication. On the processor, the execution of the tasks is dictated by a scheduling policy, while on the network, different industrial communication protocols can be deployed to guarantee bounded message latency. In this paper, we investigate the synchronization of the task execution with the underlying communication schedule, and we propose an open-source software framework. We implement a cyclic executive task scheduling policy on a time-predictable platform and synchronize the task execution with the underlying TTEth-ernet communication schedule. We evaluate our framework by developing a simple one-sensor, one-actuator industrial control example, distributed over three nodes. The presented real-time system can exchange messages with minimal jitter, and the distributed tasks synchronize to a precision of ≈ 1.6μs. Eleftherios Kyriakakis, Jens Sparsø, Peter P. Puschner, Martin Schoeberl |
EMSOFT | 2 |
| 2020 | A time-predictable open-source TTEthernet end-systemabstractCyber-physical systems deployed in areas like automotive, avionics, or industrial control are often distributed systems. The operation of such systems requires coordinated execution of the individual tasks with bounded communication network latency to guarantee quality-of-control. Both the time for computing and communication needs to be bounded and statically analyzable. To provide deterministic communication between end-systems, real-time networks can use a variety of industrial Ethernet standards typically based on time-division scheduling and enforced by real-time enabled network switches. For the computation, end-systems need time-predictable processors where the worst-case execution time of the application tasks can be analyzed statically. This paper presents a time-predictable end-system with support for deterministic communication using the open-source processor Patmos. The proposed architecture is deployed in a TTEthernet network, and the protocol software stack is implemented, and the worst-case execution time is statically analyzed. The developed end-system is evaluated in an experimental network setup composed of six TTEthernet nodes that exchange periodic frames over a TTEthernet switch. Eleftherios Kyriakakis, Maja Lund, Luca Pezzarossa, Jens Sparsø, Martin Schoeberl |
J. Syst. Archit. | 4 |
| 2019 | Scratchpad Memories with OwnershipabstractA multicore processor for real-time systems needs a time-predictable way to communicate data between different threads running on different cores. Standard multicore processors support data sharing with shared main memory backed up by caches and cache coherence protocol. This sharing solution is hardly time predictable nor does it scale to more than a few cores.This paper presents a shared scratchpad memory (SPM) for time-predictable communication between cores. The base architecture uses time-division multiplexing for the arbitration of the access to the shared SPM. This allows the timing of programs executing on different cores to be completely independent of each other. We extend this architecture by the notion of ownership. A core can own the SPM. Having exclusive access to the SPM reduces the access time to a single clock cycle. The ownership of the SPM can then be transferred to a different core, implementing low latency communication of bulk data. As an extension, we propose to organize this memory as a pool of SPMs that can be owned by different cores and transferred as needed. We evaluate the proposed architecture within the T-CREST multicore architecture. Martin Schoeberl, Tórur Biskopstø Strøm, Oktay Baris, Jens Sparsø |
DATE | 4 |
| 2019 | Demonstration of a Time-predictable Flight Controller on a Multicore ProcessorabstractUnmanned aerial vehicles, or drones, have drawn extensive attention during the last decade together with the maturity of the technology. Often, real-time requirements are needed when they are deployed for critical missions. The increasing computational demand for drones leads to a shift towards multicore architectures. However, the timing analysis of multicore systems is a challenging task due to the timing interference between processor cores. In this demonstration, we deploy a parallelized flight controller system on a multicore architecture. We provide timing analysis of the system, and test it with a processor-in-the-loop setup, which includes a flight simulator connected to the flight controller running on a time-predictable multicore platform. Oktay Baris, Shibarchi Majumder, Tórur Biskopstø Strøm, Anders la Cour-Harbo, Jens Sparsø, Thomas Bak, Martin Schoeberl |
ISORC | 5 |
| 2019 | A Time-predictable TTEthenet NodeabstractDistributed real-time systems need time-predictable computation and communication to facilitate static analysis of timing requirements and deadlines. This paper presents the implementation of a deterministic network protocol, TTEthernet, on the time-predictable Patmos processor. The implementation uses the existing Ethernet controller on the processor and we tested it with a TTEthernet system provided by TTTech Inc. Further testing showed that the controller could send time-triggered messages with bounded latency and a small jitter of approximately 4.5 us. We also provide worst-case execution time analysis of the network code, which demonstrates a time-predictable end-to-end solution. This work enables Patmos to communicate with other nodes in a deterministic way. Thus, extending the possible uses of Patmos. Maja Lund, Luca Pezzarossa, Jens Sparsø, Martin Schoeberl |
ISORC | 3 |
| 2019 | Hardlock: Real-time multicore locking
Tórur Biskopstø Strøm, Jens Sparsø, Martin Schoeberl |
J. Syst. Archit. | 2 |
| 2017 | A Controller for Dynamic Partial Reconfiguration in FPGA-Based Real-Time SystemsabstractIn real-time systems, the use of hardware accelerators can lead to a worst-case execution-time speed-up, to a simplification of its analysis, and to a reduction of its pessimism. When using FPGA technology, dynamic partial reconfiguration (DPR) can be used to minimize the area, by only loading those accelerators that are needed at any given point in time. The DPR controllers provided by the FPGA vendors satisfy a wide range of requirements and rely on software to manage the reconfiguration. This approach may lead to slow reconfiguration and unpredictable timing. This paper presents an open-source DPR controller specially developed for hard real-time systems and prototyped in connection with the open-source multi-core platform for real-time applications T-CREST. The controller enables a processor to perform reconfiguration in a time-predictable manner and supports different operating modes. The paper also presents a software tool for bitstream conversion, compression, and for reconfiguration time analysis. The DPR controller is evaluated in terms of hardware cost, operating frequency, speed, and bitstream compression ratio vs. reconfiguration time trade-off. A simple application example is also presented with the scope of showing the reconfiguration features of the controller. Luca Pezzarossa, Martin Schoeberl, Jens Sparsø |
ISORC | 3 |
| 2017 | A resource-efficient network interface supporting low latency reconfiguration of virtual circuits in time-division multiplexing networks-on-chip
Rasmus Bo Sørensen, Luca Pezzarossa, Martin Schoeberl, Jens Sparsø |
J. Syst. Archit. | 4 |
| 2016 | An area-efficient TDM NoC supporting reconfiguration for mode changesabstractThis paper presents an area-efficient time-division-multiplexing (TDM) network-on-chip (NoC) intended for use in a multicore platform for hard real-time systems. In such a platform, a mode change at the application level requires the tear-down and set-up of some virtual circuits without affecting the virtual circuits that persist across the mode change. Our NoC supports such reconfiguration in a very efficient way, using the same resources that are used for transmission of regular data. We evaluate the presented NoC in terms of worst-case reconfiguration time, hardware cost, and maximum operating frequency. The results show that the hardware cost for an FPGA implementation of our architecture is a factor of 2.2 to 3.9 times smaller than other NoCs with reconfiguration functionalities, and that the worst-case time for a reconfiguration is shorter or comparable to those NoCs. Rasmus Bo Sørensen, Luca Pezzarossa, Jens Sparsø |
NOCS | 3 |
| 2016 | Avionics Applications on a Time-Predictable Chip-MultiprocessorabstractAvionics applications need to be certified for the highest criticality standard. This certification includes schedulability analysis and worst-case execution time (WCET) analysis. WCET analysis is only possible when the software is written to be WCET analyzable and when the platform is time-predictable. In this paper we present prototype avionics applications that have been ported to the time-predictable T-CREST platform. The applications are WCET analyzable, and T-CREST is supported by the aiT WCET analyzer. This combination allows us to provide WCET bounds of avionic tasks, even when executing on a multicore processor. André Rocha, Cláudio Silva 0002, Rasmus Bo Sørensen, Jens Sparsø, Martin Schoeberl |
PDP | 4 |
| 2016 | Argo: A Real-Time Network-on-Chip Architecture With an Efficient GALS ImplementationabstractIn this paper, we present an area-efficient, globally asynchronous, locally synchronous network-on-chip (NoC) architecture for a hard real-time multiprocessor platform. The NoC implements message-passing communication between processor cores. It uses statically scheduled time-division multiplexing (TDM) to control the communication over a structure of routers, links, and network interfaces (NIs) to offer real-time guarantees. The area-efficient design is a result of two contributions: 1) asynchronous routers combined with TDM scheduling and 2) a novel NI microarchitecture. Together they result in a design in which data are transferred in a pipelined fashion, from the local memory of the sending core to the local memory of the receiving core, without any dynamic arbitration, buffering, and clock synchronization. The routers use two-phase bundled-data handshake latches based on the Mousetrap latch controller and are extended with a clock gating mechanism to reduce the energy consumption. The NIs integrate the direct memory access functionality and the TDM schedule, and use dual-ported local memories to avoid buffering, flow-control, and synchronization. To verify the design, we have implemented a 4 × 4 bitorus NoC in 65-nm CMOS technology and we present results on area, speed, and energy consumption for the router, NI, NoC, and post layout. Evangelia Kasapaki, Martin Schoeberl, Rasmus Bo Sørensen, Christoph Thomas Muller, Kees Goossens, Jens Sparsø |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2015 | Message Passing on a Time-predictable Multicore ProcessorabstractReal-time systems need time-predictable computing platforms. For a multicore processor to be time-predictable, communication between processor cores needs to be time-predictable as well. This paper presents a time-predictable message-passing library for such a platform. We show how to build up abstraction layers from a simple, time-division multiplexed hardware push channel. We develop these time-predictable abstractions and implement them in software. To prove the time-predictability of these functions we analyze their worst-case execution time (WCET) with the aiT WCET analysis tool. We combine these WCET numbers with the calculation of the network latency of a message and then provide a statically computed end-to-end latency for this core-to-core message. Rasmus Bo Sørensen, Wolfgang Puffitsch, Martin Schoeberl, Jens Sparsø |
ISORC | 4 |
| 2015 | T-CREST: Time-predictable multi-core architecture for embedded systemsabstractReal-time systems need time-predictable platforms to allow static analysis of the worst-case execution time (WCET). Standard multi-core processors are optimized for the average case and are hardly analyzable. Within the T-CREST project we propose novel solutions for time-predictable multi-core architectures that are optimized for the WCET instead of the average-case execution time. The resulting time-predictable resources (processors, interconnect, memory arbiter, and memory controller) and tools (compiler, WCET analysis) are designed to ease WCET analysis and to optimize WCET performance. Compared to other processors the WCET performance is outstanding. The T-CREST platform is evaluated with two industrial use cases. An application from the avionic domain demonstrates that tasks executing on different cores do not interfere with respect to their WCET. A signal processing application from the railway domain shows that the WCET can be reduced for computation-intensive tasks when distributing the tasks on several cores and using the network-on-chip for communication. With three cores the WCET is improved by a factor of 1.8 and with 15 cores by a factor of 5.7. The T-CREST project is the result of a collaborative research and development project executed by eight partners from academia and industry. The European Commission funded T-CREST. Martin Schoeberl, Sahar Abbaspour, Benny Akesson, Neil C. Audsley, Raffaele Capasso, Jamie Garside, Kees Goossens, Sven Goossens, Scott Hansen, Reinhold Heckmann, Stefan Hepp, Benedikt Huber, Alexander Jordan, Evangelia Kasapaki, Jens Knoop, Yonghui Li 0002, Daniel Wiltsche-Prokesch, Wolfgang Puffitsch, Peter P. Puschner, André Rocha, Cláudio Silva 0002, Jens Sparsø, Alessandro Tocchi |
J. Syst. Archit. | 22 |
| 2014 | A Metaheuristic Scheduler for Time Division Multiplexed Networks-on-ChipabstractThis paper presents a metaheuristic scheduler for inter-processor communication in multi-processor platforms using time division multiplexed (TDM) networks on chip (NOC). Compared to previous works, the scheduler handles a broader and more general class of platforms. Another contribution, which has significant practical implications, is the minimization of the TDM schedule period by over-provisioning bandwidth to connections with the smallest bandwidth requirements. Our results show that this is possible with only negligible impact on the schedule period. We evaluate the scheduler with seven different applications from the MCSL NOC benchmark suite. In the special case of all-to-all communication with equal bandwidths on all communication channels, we obtain schedules with a shorter period than reported in previous work. Rasmus Bo Sørensen, Jens Sparsø, Mark Ruvald Pedersen, Jaspur Hojgaard |
ISORC | 2 |
| 2014 | A loosely synchronizing asynchronous router for TDM-scheduled NOCsabstractThis paper presents an asynchronous router design for use in time-division-multiplexed (TDM) networks-on-chip. Unlike existing synchronous, mesochronous and asynchronous router designs with similar functionality, the router is able to silently skip over cycles/TDM-slots where no traffic is scheduled and hence avoid all switching activity in the idle links and router ports. In this way switching activity is reduced to the minimum possible amount. The fact that this relaxed synchronization is sufficient to implement TDM scheduling represents a contribution at the conceptual level. The idea can only be implemented using asynchronous circuit techniques. To this end, the paper explores the use of “click-element” templates. Click-element templates use only flip-flops and conventional gates, and this greatly simplifies the design process when using conventional EDA tools and standard cell libraries. Few papers, if any, have explored this. I. Kotleas, D. Humphreys, Rasmus Bo Sørensen, Evangelia Kasapaki, Florian Brandner, Jens Sparsø |
NOCS | 6 |
| 2013 | An area-efficient network interface for a TDM-based network-on-chipabstractNetwork interfaces (NIs) are used in multi-core systems where they connect processors, memories, and other IP-cores to a packet switched Network-on-Chip (NOC). The functionality of a NI is to bridge between the read/write transaction interfaces used by the cores and the packet-streaming interface used by the routers and links in the NOC. The paper addresses the design of a NI for a NOC that uses time division multiplexing (TDM). By keeping the essence of TDM in mind, we have developed a new area-efficient NI micro-architecture. The new design completely eliminates the need for FIFO buffers and credit based flow control - resources which are reported to account for 50–85% of the area in existing NI designs. The paper discusses the design considerations, presents the new NI micro-architecture, and reports area figures for a range of implementations. Jens Sparsø, Evangelia Kasapaki, Martin Schoeberl |
DATE | 1 |
| 2013 | Router Designs for an Asynchronous Time-Division-Multiplexed Network-on-ChipabstractIn this paper we explore the design of an asynchronous router for a time-division-multiplexed (TDM) network-on-chip (NOC) that is being developed for a multi-processor platform for hard real-time systems. TDM inherently requires a common time reference, and existing TDM-based NOC designs are either synchronous or mesochronous, but both approaches have their limitations: a globally synchronous NOC is no longer feasible in today's sub micron technologies and a mesochronous NOC requires special FIFO-based synchronizers in all input ports of all routers in order to accommodate for clock phase differences. This adds hardware complexity and increases area and power consumption. We propose to use asynchronous routers in order to achieve a simpler, more robust and globally-asynchronous NOC, and this represents an unexplored point in the design space. The paper presents a range of alternative router designs. All routers have been synthesized for a 65nm CMOS technology, and the paper reports post-layout figures for area, speed and energy and compares the asynchronous designs with an existing mesochronous clocked router. The results show that an asynchronous router is 2 times smaller, marginally slower and with roughly the same energy consumption, while offering a robust solution to the clock distribution problem. The paper further explores "clock-gating" of the individual pipeline stages in the asynchronous routers, and shows that this can lead to significant power savings. Evangelia Kasapaki, Jens Sparsø, Rasmus Bo Sørensen, Kees Goossens |
DSD | 2 |
| 2013 | A 65-nm CMOS area optimized de-synchronization flow for sub-VT designsabstractThis paper proposes a process independent post layout de-synchronization flow implemented in tool command language working on designs operating in the sub-VTregime. The overhead due to the self-timed operation is combated by introducing full-custom delay elements and latches for a standard 65-nm CMOS process. The flow offers the possibility to adjust granularity based on user requirements. Case studies with different reference designs manifested an average reduction of area and power overhead from 105% to 9% and 174% to 58% in comparison to a full standard cell de-synchronization approach. Christoph Thomas Muller, Steffen Malkowsky, Oskar Andersson, Jens Sparsø, Joachim Neves Rodrigues |
VLSI-SoC | 5 |
| 2012 | A Statically Scheduled Time-Division-Multiplexed Network-on-Chip for Real-Time SystemsabstractThis paper explores the design of a circuit-switched network-on-chip (NoC) based on time-division-multiplexing (TDM) for use in hard real-time systems. Previous work has primarily considered application-specific systems. The work presented here targets general-purpose hardware platforms. We consider a system with IP-cores, where the TDM-NoC must provide directed virtual circuits -- all with the same bandwidth -- between all nodes. This may not be a frequent scenario, but a general platform should provide this capability, and it is an interesting point in the design space to study. The paper presents an FPGA-friendly hardware design, which is simple, fast, and consumes minimal resources. Furthermore, an algorithm to find minimum-period schedules for all-to-all virtual circuits on top of typical physical NoC topologies like 2D-mesh, torus, bidirectional torus, tree, and fat-tree is presented. The static schedule makes the NoC time-predictable and enables worst-case execution time analysis of communicating real-time tasks. Martin Schoeberl, Florian Brandner, Jens Sparsø, Evangelia Kasapaki |
NOCS | 3 |
| 2011 | The ReNoC Reconfigurable Network-on-Chip: Architecture, Configuration Algorithms, and EvaluationabstractThis article presents a reconfigurable network-on-chip architecture called ReNoC, which is intended for use in general-purpose multiprocessor system-on-chip platforms, and which enables application-specific logical NoC topologies to be configured, thus providing both efficiency and flexibility. The article presents three novel algorithms that synthesize an application-specific NoC topology, map it onto the physical ReNoC architecture, and create deadlock-free, application-specific routing algorithms. We apply our algorithms to a mixture of real and synthetic applications and target three different physical architectures. Compared to a conventional NoC, ReNoC reduces power consumption by up to 58% on average. Matthias Bo Stuart, Mikkel Bystrup Stensgaard, Jens Sparsø |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2009 | Behavioral Synthesis of Asynchronous Circuits Using Syntax Directed Translation as BackendabstractThe current state-of-the art in high-level synthesis of asynchronous circuits is syntax directed translation, which performs a one-to-one mapping of an HDL-description into a corresponding circuit. This paper presents a method for behavioral synthesis of asynchronous circuits which builds on top of syntax directed translation, and which allows the designer to perform automatic design space exploration guided by area or speed constraints. This paper presents an asynchronous implementation template consisting of a data-path and a control unit and its implementation using the asynchronous hardware description language Balsa. This ldquoconventionalrdquo template architecture allows us to adapt traditional synchronous synthesis techniques for resource sharing, scheduling, binding, etc., to the domain of asynchronous circuits. A prototype tool has been implemented on top of the Balsa framework, and the method is illustrated through the implementation of a set of example circuits. The main contributions of this paper are the fundamental idea, the template architecture and its implementation using asynchronous handshake components, and the implementation of a prototype tool. Sune Fallgaard Nielsen, Jens Sparsø, Jan Madsen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | ReNoC: A Network-on-Chip Architecture with Reconfigurable Topology
Mikkel Bystrup Stensgaard, Jens Sparsø |
NOCS | 2 |
| 2008 | A Reactive and Cycle-True IP Emulator for MPSoC ExplorationabstractThe design of MultiProcessor Systems-on-Chip (MPSoC) emphasizes intellectual-property (IP)-based communication-centric approaches. Therefore, for the optimization of the MPSoC interconnect, the designer must develop traffic models that realistically capture the application behavior as executing on the IP core. In this paper, we introduce a Reactive IP Emulator (RIPE) that enables an effective emulation of the IP-core behavior in multiple environments, including bit- and cycle-true simulation. The RIPE is built as a multithreaded abstract instruction-set processor, and it can generate reactive traffic patterns. We compare the RIPE models with cycle-true functional simulation of complex application behavior (task-synchronization, multitasking, and input/output operations). Our results demonstrate high-accuracy and significant speedups. Furthermore, via a case study, we show the potential use of the RIPE in a design-space-exploration context. Shankar Mahadevan, Federico Angiolini, Jens Sparsø, Luca Benini, Jan Madsen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | A scalable, timing-safe, network-on-chip architecture with an integrated clock distribution methodabstractGrowing system sizes together with increasing performance variability are making globally synchronous operation hard to realize. Mesochronous clocking constitutes a possible solution to the problems faced. The most fundamental of problems faced when communicating between mesochronously clocked regions concerns the possibility of data corruption caused by metastability. This paper presents an integrated communication and mesochronous clocking strategy, which avoids timing related errors while maintaining a globally synchronous system perspective. The architecture is scalable as timing integrity is based purely on local observations. It is demonstrated with a 90 nm CMOS standard cell network-on-chip design which implements completely timing-safe, global communication in a modular system Tobias Bjerregaard, Mikkel Bystrup Stensgaard, Jens Sparsø |
DATE | 3 |
| 2006 | Packetizing OCP Transactions in the MANGO Network-on-ChipabstractThe scaling of CMOS technology causes a widening gap between the performance of on-chip communication and computation. This calls for a communication-centric design flow. The MANGO network-on-chip architecture enables globally asynchronous locally synchronous (GALS) system-on-chip design, while facilitating IP reuse by standard socket access points. Two types of services are available: connection-less best-effort routing and connection-oriented guaranteed service (GS) routing. This paper presents the core-centric programming model for establishing and using GS connections in MANGO. We show how OCP transactions are packetized and transmitted across the shared network, and illustrate how this affects the end-to-end performance. A high predictability of the latency of communication on shared links is shown in a MANGO-based demonstrator system Tobias Bjerregaard, Jens Sparsø |
DSD | 2 |
| 2006 | A Simple Clockless Network-on-Chip for a Commercial Audio DSP ChipabstractWe design a very small, packet-switched, clockless network-on-chip (NoC) as a replacement for the existing crossbar-based communication infrastructure in a commercial audio DSP chip. Both solutions are laid out in a 0.18 mum process, and compared in terms of area, power consumption and routing complexity. Even though the NoC turns out to be larger and more power consuming than the existing crossbar implementation, it still accounts for less than 1% of the total chip area and power consumption, and is justified by a long list of advantages: the NoC is modular, scalable and in contrast to the existing crossbar, it allows all blocks to communicate. The total wire length is decreased by 22% which eases the layout process and makes the design less prone to routing congestion. Not least, the communicating blocks are decoupled by means of the NoC, providing a globally-asynchronous, locally-synchronous (GALS) system where independent clocking of the individual blocks is enabled. This study shows that NoCs are feasible even for small systems Mikkel Bystrup Stensgaard, Tobias Bjerregaard, Jens Sparsø, Johnny Halkjær Pedersen |
DSD | 3 |
| 2005 | A Router Architecture for Connection-Oriented Service Guarantees in the MANGO Clockless Network-on-ChipabstractOn-chip networks for future system-on-chip designs need simple, high performance implementations. In order to promote system-level integrity, guaranteed services (GS) need to be provided. We propose a network-on-chip (NoC) router architecture to support this, and demonstrate with a CMOS standard cell design. Our implementation is based on clockless circuit techniques, and thus inherently supports a modular GALS-oriented design flow. Our router exploits virtual channels to provide connection-oriented GS, as well as connection-less best-effort (BE) routing. The architecture is highly flexible, in that support for different types of BE routing and GS arbitration can be easily plugged into the router. Tobias Bjerregaard, Jens Sparsø |
DATE | 2 |
| 2005 | A Network Traffic Generator Model for Fast Network-on-Chip SimulationabstractFor systems-on-chip (SoC) development, a predominant part of the design time is the simulation time. Performance evaluation and design space exploration of such systems in bit- and cycle-true fashion is becoming prohibitive. We propose a traffic generation (TG) model that provides a fast and effective network-on-chip (NoC) development and debugging environment. By capturing the type and the timestamp of communication events at the boundary of an IP core in a reference environment, the TG can subsequently emulate the core's communication behavior in different environments. Access patterns and resource contention in a system are dependent on the interconnect architecture, and our TG is designed to capture the resulting reactiveness. The regenerated traffic, which represents a realistic workload, can thus be used to undertake faster architectural exploration of interconnection alternatives, effectively decoupling simulation of IP cores and of interconnect fabrics. The results with the TG on an AMBA interconnect show a simulation time speedup above a factor of 2 over a complete system simulation, with close to 100 % accuracy. Shankar Mahadevan, Federico Angiolini, Michael Storgaard, Rasmus Grøndahl Olsen, Jens Sparsø, Jan Madsen |
DATE | 5 |
| 2005 | A Traffic Injection Methodology with Support for System-Level Synchronization
Shankar Mahadevan, Federico Angiolini, Jens Sparsø, Luca Benini, Jan Madsen |
VLSI-SoC | 3 |
| 2004 | Towards Behavioral Synthesis of Asynchronous Circuits - An Implementation Template Targeting Syntax Directed CompilationabstractThis paper presents a method for behavioral synthesis of asynchronous circuits. Our approach aims at providing a synthesis flow which is very similar to what is found in existing synchronous design tools. We adapt the synchronous behavioral synthesis abstraction into the asynchronous handshake domain by introducing a computation model, which resembles the synchronous datapath and control architecture, but which is completely asynchronous. The datapath and control architecture is then expressed in the Balsa-language, and using syntax directed compilation a corresponding handshake circuit implementation is produced. The paper also reports area, speed and power figures for a couple of benchmark circuits, which have been synthesized to layout. Sune Fallgaard Nielsen, Jens Sparsø, Jan Madsen |
DSD | 2 |
| 2002 | Future directions in clocking multi-ghz systemsabstractNo abstract available. Vojin G. Oklobdzija, Jens Sparsø |
ISLPED | 2 |
| 1999 | Designing asynchronous circuits for low power: an IFIR filter bank for a digital hearing aidabstractThis paper addresses the design of asynchronous circuits for low power through an example: a filter bank for a digital hearing aid. The asynchronous design re-implements an existing synchronous circuit which is used in a commercial product. For comparison, both designs have been fabricated in the same 0.7 /spl mu/m CMOS technology. When processing typical data (less than 50 dB sound pressure), the asynchronous control and data-path logic, an improved RAM design, and by a mechanism that adapts the number range to the actual need (exploiting the fact that typical audio signals are characterized by numerically small samples). Apart from the improved RAM design, these measures are only viable in an asynchronous design. The principles and techniques explained in this paper are of a general nature, and they apply to the design of asynchronous low-power digital signal-processing circuits in a broader perspective. In fact, this understanding is one of the contributions of the paper. Finally, the paper can be read as an example-driven introduction to asynchronous low-power design. Lars Skovby Nielsen, Jens Sparsø |
Proc. IEEE | 2 |
| 1994 | Low-power operation using self-timed circuits and adaptive scaling of the supply voltageabstractRecent research has demonstrated that for certain types of applications like sampled audio systems, self-timed circuits can achieve very low power consumption, because unused circuit parts automatically turn into a stand-by mode. Additional savings may be obtained by combining the self-timed circuits with a mechanism that adaptively adjusts the supply voltage to the smallest possible, while maintaining the performance requirements. This paper describes such a mechanism, analyzes the possible power savings, and presents a demonstrator chip that has been fabricated and tested. The idea of voltage scaling has been used previously in synchronous circuits, and the contributions of the present paper are: 1) the combination of supply scaling and self-timed circuitry which has some unique advantages, and 2) the thorough analysis of the power savings that are possible using this technique.> Lars Skovby Nielsen, Cees Niessen, Jens Sparsø, Kees van Berkel 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1993 | Delay-insensitive multi-ring structures
Jens Sparsø, Jørgen Staunstrup |
Integr. | 1 |
| 1991 | An area-efficient path memory structure for VLSI implementation of high speed Viterbi decoders
Erik Paaske, Steen Pedersen, Jens Sparsø |
Integr. | 3 |