EDBT 2026 Demo / reviewers in the wild / expert
Sander Stuijk
dblp:42/6513
· DBLP profile ↗
98ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0002-2518-6847ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 71 · 9 first-author · 17 since 2021Software engineering, systems software and programming languages · 22 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Theory of computation · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LOKI: a 0.266 pJ/SOP Digital SNN Accelerator with Multi-Cycle Clock-Gated SRAM in 22 nmabstractBio-inspired sensors like Dynamic Vision Sensors (DVS) and silicon cochleas are often combined with Spiking Neural Networks (SNNs), enabling efficient, event-driven processing similar to biological sensory systems. To realize the low-power constraints of the edge, the SNN should run on a hardware architecture that can exploit the sparse nature of the spikes. In this paper, we introduce LOKI, a digital architecture for Fully-Connected (FC) SNNs. By using Multi-Cycle Clock-Gated (MCCG) SRAMs, LOKI can operate at 0.59 V, while running at a clock frequency of 667 MHz. At full throughput, LOKI only consumes $0.266 \mathrm{pJ} /$ SOP. We evaluate LOKI on both the Neuromorphic MNIST (N-MNIST) and the Keyword Spotting (KWS) tasks, achieving 98.0 % accuracy at 119.8 nJ /inference and $93.0 \%$ accuracy at 546.5 nJ /inference respectively. Rick Luiken, Lorenzo Pes, Manil Dev Gomony, Sander Stuijk |
ASP-DAC | 4 |
| 2026 | Optimize edge AI processing through innovative compilation techniquesabstractHeterogeneous architectures became a compelling choice for edge processors executing complex DNN workloads, as they provide an ideal blend of openness, customization, energy-efficient heterogeneity, and scalable performance. Compiler optimization for DNNs on heterogeneous System-on-Chip (SoC) architectures however, must navigate complex hardware-software co-design, data movement minimization, aggressive parallelism exploitation, and advanced static/dynamic code transformations to deliver high performance and energy efficiency.This paper presents a novel compiler ecosystem for highly heterogeneous SoCs with multiple back-end targets, spanning from typical CPUs, to programmable RISC-V clusters and up to dedicated and reconfigurable accelerators. It puts together static analysis, optimization, and scheduling infrastructure to overcome the limitations of current state-of-the-art tools for heterogeneous edge AI processors. Our compilation pipeline introduces several innovative features: (1) an automatic end-to-end flow for RISC-V-based platforms, (2) efficient data layout remapping (reducing memory footprint by 35% on average) and recognition of complex ternary reductions for auto-vectorization, (3) code layout adaptation for hardware simplification, (4) a novel MLIR-based RISC-V backend supporting optimized matrix-multiplication micro-kernels that reach 90% of peak performance, (5) periodic scheduling capabilities for layer-fused CNNs, and (6) automated mapping and scheduling onto heterogeneous CGRA templates for advanced parallel kernel execution, delivering 33% higher energy efficiency than the scalar implementation and up to 3.6× higher performance. These advances enable hardware-aware compilation that reduces manual optimization effort, lowers energy consumption through memory and computation optimization, and minimizes memory footprint and data transfers. Shreya Alladi, Alexandre Lopoukhine, Georgios Alexandris, Andrea Nardi-Dei, Ravikiran Ravindranath Reddy, Christos P. Lamprakos, Panagiotis Chaidos, Alexis Maras, Alberto Ros 0001, Tobias Grosser, Sotirios Xydis, Dimitrios Soudris, Marc Geilen, Sander Stuijk, Henk Corporaal, Alexandra Jimborean |
DATE | 14 |
| 2026 | A Novel Depth-First Scheduling for Spatially Dynamic Neural Networks
Steven Colleman, Andrea Nardi-Dei, Marc Geilen, Sander Stuijk, Toon Goedemé |
ICAART (4) | 4 |
| 2026 | AURA: A Reconfigurable Asynchronous Spiking Processor for Low-Power Sensory SystemsabstractProcessing weak analog signals from biomedical, environmental, and IoT sensors in a low-power event-driven manner is a major challenge, as conventional synchronous digitization wastes energy and fails to capture the fine temporal dynamics of slow and sparse natural signals. Neuromorphic computing has emerged as a promising paradigm for edge sensing; however, most state-of-the-art neuromorphic chips lack dedicated analog-to-spike interfaces, omit crucial signal conditioning such as gain control, and offer only limited flexibility, restricting their use in real-world sensing tasks. To overcome these limitations, we present AURA, a reconfigurable mixed-signal asynchronous spiking neural network (SNN) system for ultra-low-power sensing and computing. AURA integrates analog soma and synapse with on-chip analog front-end that directly encode sensory signals into spikes, reducing conversion overhead and enhancing temporal fidelity. The processor supports inter(intra)-chip communication based on four-phase handshake protocol and Address-Event Representation (AER). Online learning is enabled via reconfigurable connectivity and synaptic weight through external PC. Designed in IHP 130nm CMOS process, AURA achieves ~0.4pJ per spike (operated within 500 Hz (incl. integ.)) in simulations. Shimeng Ye, Stijn Van Himste, Roel Jordans, Sander Stuijk, Federico Corradi |
ISCAS | 4 |
| 2025 | STEMS: Spatial-Temporal Mapping for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are event-driven bio-inspired neural networks. Recent research has trained SNN models with accuracy on par with Artificial Neural Networks (ANNs) on computer vision tasks. Due to their sparse, event-based computation, SNNs are particularly promising for energy-efficient processing, especially in event-based vision applications. However, neurons have internal states which evolve over time and keeping track of them can be costly. Hence, efficiently deploying them, especially on memory-constrained edge devices, requires careful mapping of their computation across both spatial and temporal dimensions.To address this issue, we introduce STEMS, Spatial-Temporal Mapping for SNNs. STEMS supports inter-layer mapping exploration, as well as loop tiling optimizations. By applying STEMS inter-layer exploration, we show up to 12× reduction in external memory traffic and up-to 5× reduction in energy consumption. Finally, we show that neuron states may not be needed in early SNN layers. By optimizing neuron states in one of our benchmarks, we reduced neuron states by 20x and improved energy performance by 1.4x saving without sacrificing accuracy. Sherif Eissa, Sander Stuijk, Floran de Putter, Andrea Nardi-Dei, Federico Corradi, Henk Corporaal |
IEEE Trans. Computers | 2 |
| 2023 | ReMeCo: Reliable Memristor-Based in-Memory Neuromorphic ComputationabstractMemristor-based in-memory neuromorphic computing systems promise a highly efficient implementation of vector-matrix multiplications, commonly used in artificial neural networks (ANNs). However, the immature fabrication process of memristors and circuit level limitations, i.e., stuck-at-fault (SAF), IR-drop, and device-to-device (D2D) variation, degrade the reliability of these platforms and thus impede their wide deployment. In this paper, we present ReMeCo, a redundancy-based reliability improvement framework. It addresses the non-idealities while constraining the induced overhead. It achieves this by performing a sensitivity analysis on ANN. With the acquired insight, ReMeCo avoids the redundant calculation of least sensitive neurons and layers. ReMeCo uses a heuristic approach to find the balance between recovered accuracy and imposed overhead. ReMeCo further decreases hardware redundancy by exploiting the bit-slicing technique. In addition, the framework employs the ensemble averaging method at the output of every ANN layer to incorporate the redundant neurons. The efficacy of the ReMeCo is assessed using two well-known ANN models, i.e., LeNet, and AlexNet, running the MNIST and CIFAR10 datasets. Our results show 98.5% accuracy recovery with roughly 4% redundancy which is more than 20× lower than the state-of-the-art. Ali BanaGozar, Seyed Hossein Hashemi Shadmehri, Sander Stuijk, Mehdi Kamal, Ali Afzali-Kusha, Henk Corporaal |
ASP-DAC | 3 |
| 2023 | PetaOps/W edge-AI $\mu$ Processors: Myth or reality?abstractWith the rise of deep learning (DL), our world braces for artificial intelligence (AI) in every edge device, creating an urgent need for edge-AI SoCs. This SoC hardware needs to support high throughput, reliable and secure AI processing at ultra-low power (ULP), with a very short time to market. With its strong legacy in edge solutions and open processing platforms, the EU is well-positioned to become a leader in this SoC market. However, this requires AI edge processing to become at least 100 times more energy-efficient, while offering sufficient flexibility and scalability to deal with AI as a fast-moving target. Since the design space of these complex SoCs is huge, advanced tooling is needed to make their design tractable. The CONVOLVE project (currently in Inital stage) addresses these roadblocks. It takes a holistic approach with innovations at all levels of the design hierarchy. Starting with an overview of SOTA DL processing support and our project methodology, this paper presents 8 important design choices largely impacting the energy efficiency and flexibility of DL hardware. Finding good solutions is key to making smart-edge computing a reality. Manil Dev Gomony, Floran de Putter, Anteneh Gebregiorgis, Gianna Paulin, Linyan Mei, Vikram Jain, Said Hamdioui, Victor Sanchez, Tobias Grosser, Marc Geilen, Marian Verhelst, Friedemann Zenke, Frank K. Gürkaynak, Barry de Bruin, Sander Stuijk, Simon Davidson, Sayandip De, Mounir Ghogho, Alexandra Jimborean, Sherif Eissa, Luca Benini, Dimitrios Soudris, Rajendra Bishnoi, Sam Ainsworth 0001, Federico Corradi, Ouassim Karrakchou, Tim Güneysu, Henk Corporaal |
DATE | 15 |
| 2023 | Vision-Based Multi-Size Object PositioningabstractAccurate object positioning is critical in many industrial manufacturing applications. The execution time and precision of the object positioning task have a significant impact on the overall performance and throughput, especially in cost-sensitive industries such as semiconductor manufacturing. In addition, the object positioning algorithm must adapt to changes in object size, features, and environmental conditions in real-time. While traditional sensors struggle to cope with dynamic conditions, vision-based perception is more adaptable and robust. Vision-based perception can capture and analyze visual information by using cameras and image processing algorithms, providing a robust way to locate objects in dynamic environments. However, classical perception algorithms based on vision cannot handle objects with different characteristics, and modern object detectors that rely on deep neural networks struggle to adapt to image sizes, resulting in unnecessary computations. To address these challenges, this paper proposes an approach for designing a branched multi-input deep neural network (DNN) that considers variations in input image sizes to adapt the input branches. In essence, the proposed DNN reduces the computation time for images with lower dimensions. To validate the proposed approach, an IC dataset is created that represents the variations in object sizes as seen in semiconductor manufacturing machines. Depending on the choice of input branches, the average inference time is reduced by over 30% with a slight gain in detection accuracy. Vibhor Jain, Sajid Mohamed, Dip Goswami, Sander Stuijk |
DSD | 4 |
| 2023 | Dependability of Future Edge-AI Processors: Pandora's BoxabstractThis paper addresses one of the directions of the HORIZON EU CONVOLVE project being dependability of smart edge processors based on computation-in-memory and emerging memristor devices such as RRAM. It discusses how how this alternative computing paradigm will change the way we used to do manufacturing test. In addition, it describes how these emerging devices inherently suffering from many non-idealities are calling for new solutions in order to ensure accurate and reliable edge computing. Moreover, the paper also covers the security aspects for future edge processors and shows the challenges and the future directions. Manil Dev Gomony, Anteneh Gebregiorgis, Moritz Fieback, Marc Geilen, Sander Stuijk, Jan Richter-Brockmann, Rajendra Bishnoi, Sven Argo, Lara Arche Andradas, Tim Güneysu, Mottaqiallah Taouil, Henk Corporaal, Said Hamdioui |
ETS | 5 |
| 2023 | QMTS: Fixed-point Quantization for Multiple-timescale Spiking Neural Networks
Sherif Eissa, Federico Corradi, Floran de Putter, Sander Stuijk, Henk Corporaal |
ICANN (1) | 4 |
| 2023 | Dissecting Tensor Cores via Microbenchmarks: Latency, Throughput and Numeric BehaviorsabstractTensor Cores have been an important unit to accelerate Fused Matrix Multiplication Accumulation (MMA) in all NVIDIA GPUs since Volta Architecture. To program Tensor Cores, users have to use either legacy wmma APIs or current mma APIs. Legacy wmma APIs are more easy-to-use but can only exploit limited features and power of Tensor Cores. Specifically, wmma APIs support fewer operand shapes and can not leverage the new sparse matrix multiplication feature of the newest Ampere Tensor Cores. However, the performance of current programming interface has not been well explored. Furthermore, the computation numeric behaviors of low-precision floating points (TF32, BF16, and FP16) supported by the newest Ampere Tensor Cores are also mysterious. In this paper, we explore the throughput and latency of current programming APIs. We also intuitively study the numeric behaviors of Tensor Cores MMA and profile the intermediate operations including multiplication, addition of inner product, and accumulation. All codes used in this work can be found inhttps://github.com/sunlex0717/DissectingTensorCores. Ang Li 0006, Tong Geng, Sander Stuijk, Henk Corporaal |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | SySCIM: SystemC-AMS Simulation of Memristive Computation In-MemoryabstractComputation-in-memory (CIM) is one of the most appealing computing paradigms, especially for implementing artificial neural networks. Non-volatile memories like ReRAMs, PCMs, etc., have proven to be promising candidates for the realization of CIM processors. However, these devices and their driving circuits are subject to non-idealities. This paper presents a comprehensive platform, named SysCIM, for simulating memristor-based CIM systems. SySCIM considers the impact of the non-idealities of the CIM components, including memristor device, memristor crossbar (interconnects), analog-to-digital converter, and transimpedance amplifier, on the vector-matrix multiplication performed by the CIM unit. The CIM modules are described in SystemC and SystemC-AMS to reach a higher simulation speed while maintaining high simulation accuracy. Experiments under different crossbar sizes show SySCIM performs simulations up to 117 x faster than HSPICE with less than 4% accuracy loss. The modular design of SySCIM provides researchers with an easy design-space exploration tool to investigate the effects of various non-idealities. Seyed Hossein Hashemi Shadmehri, Ali BanaGozar, Mehdi Kamal, Sander Stuijk, Ali Afzali-Kusha, Massoud Pedram, Henk Corporaal |
DATE | 4 |
| 2022 | DNAsim: Evaluation Framework for Digital Neuromorphic ArchitecturesabstractNeuromorphic architectures implement low-power machine learning applications using spike-based biological neuron models trained with bio-inspired or machine learning algorithms. Prior work on simulating Spiking Neural Networks (SNNs) focused on simulating emerging compute in-memory (CIM) architectures, while prior work on mapping SNNs focused mainly on minimizing inter-core communication or resource utilization and targeted either emerging CIM architectures or specific target platforms. SNN mapping choices on a neuromoprhic multi-processor platform can impact performance and energy consumption. In this paper, we introduce a simulation framework that evaluates application mapping on a user-defined NoC-based multi-core digital neuromorphic architecture. Our simulator evaluates latency and energy based on mapping and abstract spike activity traces which indicate the firing of neurons at specific discrete timesteps defined by the application. We create two hardware models based on reported work in literature and show the evaluation of different mapping scenarios for a state-of-the-art SNN benchmark. Sherif Eissa, Sander Stuijk, Henk Corporaal |
DSD | 2 |
| 2022 | Partial Evaluation in Junction TreesabstractOne prominent method to perform inference on probabilistic graphical models is the probability propagation in trees of clusters (PPTC) algorithm. In this paper, we demonstrate the use of partial evaluation, an established technique from the compiler domain, to improve the performance of online Bayesian inference using the PPTC algorithm in the context of observed evidence. We present a metaprogramming-based method to transform a base program into an optimized version by precomputing the static input at compile time while guaranteeing behavioral equivalence. We achieve an inference time reduction of 21% on average for the Promedas benchmark. Martin Roa Villescas, Patrick W. A. Wijnings, Sander Stuijk, Henk Corporaal |
DSD | 3 |
| 2022 | LEAPER: Fast and Accurate FPGA-based System Performance Prediction via Transfer LearningabstractMachine learning has recently gained traction as a way to overcome the slow accelerator generation and implementation process on an FPGA. It can be used to build performance and resource usage models that enable fast early-stage design space exploration. However, these models suffer from three main limitations. First, training requires large amounts of data (features extracted from design synthesis and implementation tools), which is cost-inefficient because of the time-consuming accelerator design and implementation process. Second, a model trained for a specific environment cannot predict performance or resource usage for a new, unknown environment. In a cloud system, renting a platform for data collection to build an ML model can significantly increase the total-cost-ownership (TCO) of a system. Third, ML-based models trained using a limited number of samples are prone to overfitting. To overcome these limitations, we propose LEAPER, a transfer learning-based approach for prediction of performance and resource usage in FPGA-based systems. The key idea of LEAPER is to transfer an ML-based performance and resource usage model trained for a low-end edge environment to a new, high-end cloud environment to provide fast and accurate predictions for accelerator implementation. Experimental results show that LEAPER (1) provides, on average across six workloads and five FPGAs, 85% accuracy when we use our transferred model for prediction in a cloud environment with 5-shot learning and (2) reduces design-space exploration time for accelerator implementation on an FPGA by 10×, from days to only a few hours. Gagandeep Singh 0002, Dionysios Diamantopoulos, Juan Gómez-Luna, Sander Stuijk, Henk Corporaal, Onur Mutlu |
ICCD | 4 |
| 2022 | Sibyl: adaptive and extensible data placement in hybrid storage systems using online reinforcement learningabstractHybrid storage systems (HSS) use multiple different storage devices to provide high and scalable storage capacity at high performance. Data placement across different devices is critical to maximize the benefits of such a hybrid system. Recent research proposes various techniques that aim to accurately identify performance-critical data to place it in a "best-fit" storage device. Unfortunately, most of these techniques are rigid, which (1) limits their adaptivity to perform well for a wide range of workloads and storage device configurations, and (2) makes it difficult for designers to extend these techniques to different storage system configurations (e.g., with a different number or different types of storage devices) than the configuration they are designed for. Our goal is to design a new data placement technique for hybrid storage systems that overcomes these issues and provides: (1) adaptivity, by continuously learning from and adapting to the workload and the storage device characteristics, and (2) easy extensibility to a wide range of workloads and HSS configurations. Gagandeep Singh 0002, Rakesh Nadig, Jisung Park 0001, Rahul Bera, Nastaran Hajinazar, David Novo, Juan Gómez-Luna, Sander Stuijk, Henk Corporaal, Onur Mutlu |
ISCA | 8 |
| 2022 | Accelerating Weather Prediction Using Near-Memory Reconfigurable FabricabstractOngoing climate change calls for fast and accurate weather and climate modeling. However, when solving large-scale weather prediction simulations, state-of-the-art CPU and GPU implementations suffer from limited performance and high energy consumption. These implementations are dominated by complex irregular memory access patterns and low arithmetic intensity that pose fundamental challenges to acceleration. To overcome these challenges, we propose and evaluate the use of near-memory acceleration using a reconfigurable fabric with high-bandwidth memory (HBM). We focus on compound stencils that are fundamental kernels in weather prediction models. By using high-level synthesis techniques, we develop NERO, an field-programmable gate array+HBM-based accelerator connected through Open Coherent Accelerator Processor Interface to an IBM POWER9 host system. Our experimental results show that NERO outperforms a 16-core POWER9 system by \( 5.3\times \) and \( 12.7\times \) when running two different compound stencil kernels. NERO reduces the energy consumption by \( 12\times \) and \( 35\times \) for the same two kernels over the POWER9 system with an energy efficiency of 1.61 GFLOPS/W and 21.01 GFLOPS/W. We conclude that employing near-memory acceleration solutions for weather prediction modeling is promising as a means to achieve both high performance and high energy efficiency. Gagandeep Singh 0002, Dionysios Diamantopoulos, Juan Gómez-Luna, Christoph Hagleitner, Sander Stuijk, Henk Corporaal, Onur Mutlu |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2021 | Efficient Tensor Cores support in TVM for Low-Latency Deep learningabstractDeep learning algorithms are gaining popularity in autonomous systems. These systems typically have stringent latency constraints that are challenging to meet given the high computational demands of these algorithms. Nvidia introduced Tensor Cores (TCs) to speed up some of the most commonly used operations in deep learning algorithms. Compilers (e.g., TVM) and libraries (e.g., cuDNN) focus on the efficient usage of TCs when performing batch processing. Latency sensitive applications can however not exploit large batch processing. This paper presents an extension to the TVM compiler that generates low latency TCs implementations, particularly for batch size 1. Experimental results show that our solution reduces the latency on average by 14% compared to the cuDNN library on a Desktop RTX2070 GPU, and by 49% on an Embedded Jetson Xavier GPU. Savvas Sioutas, Sander Stuijk, Andrew Nelson 0001, Henk Corporaal |
DATE | 3 |
| 2021 | Modeling FPGA-Based Systems via Few-Shot LearningabstractMachine-learning-based models have recently gained traction as a way to overcome the slow downstream implementation process of FPGAs by building models that provide fast and accurate performance predictions. However, these models suffer from two main limitations: (1) a model trained for a specific environment cannot predict for a new, unknown environment; (2) training requires large amounts of data (features extracted from FPGA synthesis and implementation reports), which is cost-inefficient because of the time-consuming FPGA design cycle. In various systems (e.g., cloud systems), where getting access to platforms is typically costly, error-prone, and sometimes infeasible, collecting enough data is even more difficult. Our research aims to answer the following question: for an FPGA-based system, can we leverage and transfer our ML-based performance models trained on a low-end local system to a new, unknown, high-end FPGA-based system, thereby avoiding the aforementioned two main limitations of traditional ML-based approaches? To this end, we propose a transfer-learning-based approach for FPGA-based systems that adapts an existing ML-based model to a new, unknown environment to provide fast and accurate performance and resource utilization predictions. Gagandeep Singh 0002, Dionysios Diamantopoulos, Juan Gómez-Luna, Sander Stuijk, Onur Mutlu, Henk Corporaal |
FPGA | 4 |
| 2021 | Characterization of Mems Microphone Sensitivity and Phase Distributions with Applications in Array ProcessingabstractAn array with MEMS microphones can distinguish individual noise sources in an environment through spatial filtering. Its effectiveness depends on the variations in microphone sensitivity and phase. Quantification of these variations is valuable, because it enables assessment and optimization of array performance. This is particularly important if the measurements are to be used for enforcement of noise regulations.Nominal microphone sensitivity and phase are manufacturer-specified, but the distribution (histogram) around these values is not. Hence, this work demonstrates a free-field comparison method for measuring these variations in a batch of arrays. We also provide the histograms at 1 kHz for a sample population of 8384 Knowles SPH0641LM4H-1 MEMS microphones (131 arrays of 64 microphones). The histograms follow t-distributions, resulting in 95% confidence intervals of ±0.39dB for sensitivity and ±0.82° for Finally, we phase. illustrate that delay-and-sum beamforming with these microphones results in a Gumbel-distributed gain with −0.13/+0.10dB 95% confidence interval. Patrick W. A. Wijnings, Sander Stuijk, Rick Scholte, Henk Corporaal |
ICASSP | 2 |
| 2021 | DominoSearch: Find layer-wise fine-grained N: M sparse schemes from dense neural networksabstractNeural pruning is a widely-used compression technique for Deep Neural Networks (DNNs). Recent innovations in Hardware Architectures (e.g. Nvidia Ampere Sparse Tensor Core) and N:M fine-grained Sparse Neural Network algorithms (i.e. every M-weights contains N non-zero values) reveal a promising research line of neural pruning. However, the existing N:M algorithms only address the challenge of how to train N:M sparse neural networks in a uniform fashion (i.e. every layer has the same N:M sparsity) and suffer from a significant accuracy drop for high sparsity (i.e. when sparsity > 80\%). To tackle this problem, we present a novel technique -- \textbf{\textit{DominoSearch}} to find mixed N:M sparsity schemes from pre-trained dense deep neural networks to achieve higher accuracy than the uniform-sparsity scheme with equivalent complexity constraints (e.g. model size or FLOPs). For instance, for the same model size with 2.1M parameters (87.5\% sparsity), our layer-wise N:M sparse ResNet18 outperforms its uniform counterpart by 2.1\% top-1 accuracy, on the large-scale ImageNet dataset. For the same computational complexity of 227M FLOPs, our layer-wise sparse ResNet18 outperforms the uniform one by 1.3\% top-1 accuracy. Furthermore, our layer-wise fine-grained N:M sparse ResNet50 achieves 76.7\% top-1 accuracy with 5.0M parameters. {This is competitive to the results achieved by layer-wise unstructured sparsity} that is believed to be the upper-bound of Neural Network pruning with respect to the accuracy-sparsity trade-off. We believe that our work can build a strong baseline for further sparse DNN research and encourage future hardware-algorithm co-design work. Our code and models are publicly available at \url{https://github.com/NM-sparsity/DominoSearch}. Aojun Zhou, Sander Stuijk, Rob G. J. Wijnhoven, Andrew Nelson 0001, Hongsheng Li 0001, Henk Corporaal |
NeurIPS | 3 |
| 2021 | Taming the State-space Explosion in the Makespan Optimization of Flexible Manufacturing SystemsabstractThis article presents a modular automaton-based framework to specify flexible manufacturing systems and to optimize the makespan of product batches. The Batch Makespan Optimization (BMO) problem is NP-Hard and optimization can therefore take prohibitively long, depending on the size of the state-space induced by the specification. To tame the state-space explosion problem, we develop an algebra based on automata equivalence and inclusion relations that consider both behavior and structure. The algebra allows us to systematically relate the languages induced by the automata, their state-space sizes, and their solutions to the BMO problem. Further, we introduce a novel constraint-based approach to systematically prune the state-space based on the the notions of nonpermutation-repulsiveness and permutation-attractiveness. We prove that constraining a nonpermutation-repulsing automaton with a permutation-attracting constraint always reduces the state-space. This approach allows us to (i) compute optimal solutions of the BMO problem when the (additional) constraints are taken into account and (ii) compute bounds for the (original) BMO problem (without using the constraints). We demonstrate the effectiveness of our approach by optimizing an industrial wafer handling controller. João Bastos, Jeroen Voeten, Sander Stuijk, Ramon R. H. Schiffelers, Henk Corporaal |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2021 | Camera-Based Vital Signs Monitoring During Sleep - A Proof of Concept StudyabstractPolysomnography (PSG) is the current gold standard for the diagnosis of sleep disorders. However, this multi-parametric sleep monitoring tool also has some drawbacks, e.g. it limits the patient's mobility during the night and it requires the patient to come to a specialized sleep clinic or hospital to attach the sensors. Unobtrusive techniques for the detection of sleep disorders such as sleep apnea are therefore gaining increasing interest. Remote photoplethysmography using video is a technique which enables contactless detection of hemodynamic information. Promising results in near-infrared have been reported for the monitoring of sleep-relevant physiological parameters pulse rate, respiration and blood oxygen saturation. In this study we validate a contactless monitoring system on eight patients with a high likelihood of relevant obstructive sleep apnea, which are enrolled for a sleep study at a specialized sleep center. The dataset includes 46.5 hours of video recordings, full polysomnography and metadata. The camera can detect pulse and respiratory rate within 2 beats/breaths per minute accuracy 92% and 91% of the time, respectively. Estimated blood oxygen values are within 4 percentage points of the finger-oximeter 89% of the time. These results demonstrate the potential of a camera as a convenient diagnostic tool for sleep apnea, and sleep disorders in general. Mark van Gastel, Sander Stuijk, Sebastiaan Overeem, Johannes P. van Dijk, Merel van Gilst, Gerard de Haan |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | NERO: A Near High-Bandwidth Memory Stencil Accelerator for Weather Prediction ModelingabstractOngoing climate change calls for fast and accurate weather and climate modeling. However, when solving large-scale weather prediction simulations, state-of-the-art CPU and GPU implementations suffer from limited performance and high energy consumption. These implementations are dominated by complex irregular memory access patterns and low arithmetic intensity that pose fundamental challenges to acceleration. To overcome these challenges, we propose and evaluate the use of near-memory acceleration using a reconfigurable fabric with high-bandwidth memory (HBM). We focus on compound stencils that are fundamental kernels in weather prediction models. By using high-level synthesis techniques, we develop NERO, an FPGA+HBM-based accelerator connected through IBM CAPI2 (Coherent Accelerator Processor Interface) to an IBM POWER9 host system. Our experimental results show that NERO outperforms a 16-core POWER9 system by 4.2x and 8.3x when running two different compound stencil kernels. NERO reduces the energy consumption by 22x and 29x for the same two kernels over the POWER9 system with an energy efficiency of 1.5 GFLOPS/Watt and 17.3 GFLOPS/Watt. We conclude that employing near-memory acceleration solutions for weather prediction modeling is promising as a means to achieve both high performance and high energy efficiency. Gagandeep Singh 0002, Dionysios Diamantopoulos, Christoph Hagleitner, Juan Gómez-Luna, Sander Stuijk, Onur Mutlu, Henk Corporaal |
FPL | 5 |
| 2020 | Approximate Inference by Kullback-Leibler Tensor Belief PropagationabstractProbabilistic programming provides a structured approach to signal processing algorithm design. The design task is formulated as a generative model, and the algorithm is derived through automatic inference. Efficient inference is a major challenge; e.g., the Shafer-Shenoy algorithm (SS) performs badly on models with large treewidth, which arise from various real-world problems. We focus on reducing the size of discrete models with large treewidth by storing intermediate factors in compressed form, thereby decoupling the variables through conditioning on introduced weights. This work proposes pruning of these weights using Kullback-Leibler divergence. We adapt a strategy from the Gaussian mixture reduction literature, leading to Kullback-Leibler Tensor Belief Propagation (KL-TBP), in which we use agglomerative hierarchical clustering to subsequently merge pairs of weights. Experiments using benchmark problems show KL-TBP consistently achieves lower approximation error than existing methods with competitive runtime. Patrick W. A. Wijnings, Sander Stuijk, Bert de Vries, Henk Corporaal |
ICASSP | 2 |
| 2020 | Programming tensor cores from an image processing DSLabstractTensor Cores (TCUs) are specialized units first introduced by NVIDIA in the Volta microarchitecture in order to accelerate matrix multiplications for deep learning and linear algebra workloads. While these units have proved to be capable of providing significant speedups for specific applications, their programmability remains difficult for the average user. In this paper, we extend the Halide DSL and compiler with the ability to utilize these units when generating code for a CUDA based NVIDIA GPGPU. To this end, we introduce a new scheduling directive along with custom lowering passes that automatically transform a Halide AST in order to be able to generate code for the TCUs. We evaluate the generated code and show that it can achieve over 5X speedup compared to Halide manual schedules without TCU support, while it remains within 20% of the NVIDIA cuBLAS implementations for mixed precision GEMM and within 10% of manual CUDA implementations with WMMA intrinsics. Savvas Sioutas, Sander Stuijk, Twan Basten, Lou J. Somers, Henk Corporaal |
SCOPES | 2 |
| 2020 | Reviewing inference performance of state-of-the-art deep learning frameworksabstractDeep learning models have replaced conventional methods for machine learning tasks. Efficient inference on edge devices with limited resources is key for broader deployment. In this work, we focus on the tool selection challenge for inference deployment. We present an extensive evaluation of the inference performance of deep learning software tools using state-of-the-art CNN architectures for multiple hardware platforms. We benchmark these hardware-software pairs for a broad range of network architectures, inference batch sizes, and floating-point precision, focusing on latency and throughput. Our results reveal interesting combinations for optimal tool selection, resulting in different optima when considering minimum latency and maximum throughput. Berk Ulker, Sander Stuijk, Henk Corporaal, Rob G. J. Wijnhoven |
SCOPES | 2 |
| 2020 | Real-time audio processing for hearing aids using a model-based bayesian inference frameworkabstractDevelopment of hearing aid (HA) signal processing algorithms entails an iterative process between two design steps, namely algorithm development and the embedded implementation. Algorithm designers favor high-level programming languages for several reasons including higher productivity, code readability and, perhaps most importantly, availability of state-of-the-art signal processing frameworks that open new research directions. Embedded software, on the other hand, is preferably implemented using a low-level programming language to allow finer control of the hardware, an essential trait in real-time processing applications. In this paper we present a technique that allows deploying DSP algorithms written in Julia, a modern high-level programming language, on a real-time HA processing platform known as openMHA. We demonstrate this technique by using a model-based Bayesian inference framework to perform real-time audio processing. Martin Roa Villescas, Bert de Vries, Sander Stuijk, Henk Corporaal |
SCOPES | 3 |
| 2020 | Schedule Synthesis for Halide Pipelines on GPUsabstractThe Halide DSL and compiler have enabled high-performance code generation for image processing pipelines targeting heterogeneous architectures through the separation of algorithmic description and optimization schedule. However, automatic schedule generation is currently only possible for multi-core CPU architectures. As a result, expert knowledge is still required when optimizing for platforms with GPU capabilities. In this work, we extend the current Halide Autoscheduler with novel optimization passes to efficiently generate schedules for CUDA-based GPU architectures. We evaluate our proposed method across a variety of applications and show that it can achieve performance competitive with that of manually tuned Halide schedules, or in many cases even better performance. Experimental results show that our schedules are on average 10% faster than manual schedules and over 2× faster than previous autoscheduling attempts. Savvas Sioutas, Sander Stuijk, Twan Basten, Henk Corporaal, Lou J. Somers |
ACM Trans. Archit. Code Optim. | 2 |
| 2019 | NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble LearningabstractThe cost of moving data between the memory/storage units and the compute units is a major contributor to the execution time and energy consumption of modern workloads in computing systems. A promising paradigm to alleviate this data movement bottleneck is near-memory computing (NMC), which consists of placing compute units close to the memory/storage units. There is substantial research effort that proposes NMC architectures and identifies workloads that can benefit from NMC. System architects typically use simulation techniques to evaluate the performance and energy consumption of their designs. However, simulation is extremely slow, imposing long times for design space exploration. In order to enable fast early-stage design space exploration of NMC architectures, we need high-level performance and energy models. Gagandeep Singh 0002, Juan Gómez-Luna, Giovanni Mariani, Geraldo F. Oliveira, Stefano Corda, Sander Stuijk, Onur Mutlu, Henk Corporaal |
DAC | 6 |
| 2019 | Implementation-aware design of image-based control with on-line measurable variable-delayabstractImage-based control uses image-processing algorithms to acquire sensing information. The sensing delay associated with the image-processing algorithm is typically platform-dependent and time-varying. Modern embedded platforms allow to characterize the sensing delay at design-time obtaining a delay histogram, and at run-time measuring its precise value. We exploit this knowledge to design variable-delay controllers. This design also takes into account the resource configuration of the image processing algorithm: sequential (with one processing resource) or pipelined (with multiprocessing capabilities). Since the control performance strongly depends on the model quality, we present a simulation benchmark that uses the model uncertainty and the delay histogram to obtain bounds on control performance. Our benchmark is used to select a variable-delay controller and a resource configuration that outperform a constant worst-case delay controller. Róbinson Medina Sánchez, Sander Stuijk, Dip Goswami, Twan Basten |
DATE | 2 |
| 2019 | NARMADA: Near-Memory Horizontal Diffusion Accelerator for Scalable Stencil ComputationsabstractReal-world weather forecasting applications consist of compound stencil kernels that do not perform well on conventional architectures. This behavior is due to their complex data access patterns, limited data reusability, and low arithmetic intensity. To overcome these issues, we harness the potential of near-memory computing by offloading a horizontal diffusion kernel, which is a compound stencil kernel, from the COSMO weather prediction application to a reconfigurable fabric. We use a heterogeneous system that comprises a CPU and an FPGA with on-chip SRAM memory and on-board DRAM memory. By introducing a memory hierarchy tailored to the targeted application and using a coherent memory model, we move the computation close to the memory, which improves memory efficiency. Our hardware design on the FPGA uses high-level synthesis techniques and results in an accelerator with IBM CAPI 2.0 (Coherent Accelerator Processor Interface) technology. We evaluate it against a tuned software implementation running on an IBM POWER9 host system. The experimental results show that these kernels on an FPGA can outperform a complete 16-core POWER9 node (configured with 64 threads) by 3.3x. Moreover, our solution provides an 18x improvement in the active energy consumption. Gagandeep Singh 0002, Dionysios Diamantopoulos, Christoph Hagleitner, Sander Stuijk, Henk Corporaal |
FPL | 4 |
| 2019 | Robust Bayesian Beamforming for Sources at Different Distances with Applications in Urban MonitoringabstractAcoustic smart sensor networks can provide valuable actionable intelligence to authorities for managing safety in the urban environment. A spatial filter (beamformer) for localization and separation of acoustic sources is a key component of such a network. However, classical methods such as delay-and-sum beamforming fail, because sources are located at varying distances from the sensor array. This causes a regularization problem where either far-away sources are wrongly attenuated, or noise is wrongly amplified. We solve this by considering source strength and location as random variables. The posterior distributions are approximated using Gibbs sampling. Each marginal is computed by combining importance sampling and inverse transform sampling using Chebyshev polynomial approximation. This leads to an iterative algorithm with similarities to deconvolution beamforming. Our method is robust against deviations in manifold model, can deal with sources at different distances and power levels, and does not require an a priori known number of sources. Patrick W. A. Wijnings, Sander Stuijk, Bert de Vries, Henk Corporaal |
ICASSP | 2 |
| 2019 | CIM-SIM: Computation In Memory SIMuIatorabstractComputation-in-memory reverses the trend in von-Neumann processors by bringing the computation closer to the data, to even within the memory array, as opposed to introducing new memory hierarchies to keep (frequently used) data closer to a central processing unit (CPU). In recent years, new non-volatile memory (NVM) technologies, e.g., memristor, PCM, etc., have proven that they can function as memories and perform computations on the stored data as well. In particular, when they are combined with a modest set of (digital) peripheral modules, a wider range of operations can be supported, e.g., vector matrix multiply and Boolean logic. In this paper, we are introducing the CIM-SIM, an open source simulator written in SystemC, which is capable of simulating the functional behaviour of such architectures. The architecture includes the definition of a set of technology-agnostic nano-instructions. Ali BanaGozar, Kanishkan Vadivel, Sander Stuijk, Henk Corporaal, Stephan Wong, Muath Abu Lebdeh, Said Hamdioui |
SCOPES | 3 |
| 2019 | Towards Efficient Code Generation for Exposed Datapath ArchitecturesabstractCoarse-grained reconfigurable architectures and other exposed datapath architectures such as transport-triggered architectures come with a high energy efficiency promise for accelerating data oriented workloads. Their main drawback results from the push of complexity from the architecture to the programmer; compiler techniques that allow starting from a higher-level programming language and generate code efficiently to such architectures robustly is still an open research area. In this article we survey the known main sources of challenges and outline a generic processor architecture template that covers the most common architecture variations along with a proposal for a common code generation framework for such challenging architectures. Kanishkan Vadivel, Roel Jordans, Sander Stuijk, Henk Corporaal, Pekka Jääskeläinen, Heikki Kultala |
SCOPES | 3 |
| 2019 | Schedule Synthesis for Halide Pipelines through Reuse AnalysisabstractEfficient code generation for image processing applications continues to pose a challenge in a domain where high performance is often necessary to meet real-time constraints. The inherently complex structure found in most image-processing pipelines, the plethora of transformations that can be applied to optimize the performance of an implementation, as well as the interaction of these optimizations with locality, redundant computation and parallelism, can be indentified as the key reasons behind this issue. Recent domain-specific languages (DSL) such as the Halide DSL and compiler attempt to encourage high-level design-space exploration to facilitate the optimization process. We propose a novel optimization strategy that aims to maximize producer-consumer locality by exploiting reuse in image-processing pipelines. We implement our analysis as a tool that can be used alongside the Halide DSL to automatically generate schedules for pipelines implemented in Halide and test it on a variety of benchmarks. Experimental results on three different multi-core architectures show an average performance improvement of 40% over the Halide Auto-Scheduler and 75% over a state-of-the art approach that targets the PolyMage DSL. Savvas Sioutas, Sander Stuijk, Luc Waeijen, Twan Basten, Henk Corporaal, Lou J. Somers |
ACM Trans. Archit. Code Optim. | 2 |
| 2019 | Designing a Controller with Image-based Pipelined Sensing and Additive UncertaintiesabstractPipelined image-based control uses parallel instances of its image-processing algorithm in a pipelined fashion to improve the quality of control. A performance-oriented control design improves the controller settling time with each additional processing resource, which creates a resources-performance trade-off. In real-life applications, it is common to have a continuous-time model with additive uncertainties in one or more parameters that may affect the controller performance and the aforementioned trade-off. We present a robustness analysis framework for performance-oriented pipelined controllers with additive model uncertainties. We present a technique to obtain discrete-time uncertainties based on the continuous-time uncertainties for given uncertainty bounds. To benchmark such uncertainty bounds for a real system, we consider uncertainties in one element of the system, potentially caused by multiple uncertain parameters in the model. Robustness and its impact in the trade-off analysis are studied. We also provide a robustness-oriented pipelined controller design that takes into account the benchmarked uncertainties. Our results show that in performance-oriented designs, the tolerable uncertainties for a pipelined controller decrease when increasing the number of pipes. In robustness-oriented designs, the controller robustness is enhanced with each newly added pipe. We show the feasibility of our technique by implementing a realistic example in a Hardware-in-the-Loop simulation. Róbinson Medina Sánchez, Juan Valencia, Sander Stuijk, Dip Goswami, Twan Basten |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2018 | Loop transformations leveraging hardware prefetchingabstractMemory-bound applications heavily depend on the bandwidth of the system in order to achieve high performance. Improving temporal and/or spatial locality through loop transformations is a common way of mitigating this dependency. However, choosing the right combination of optimizations is not a trivial task, due to the fact that most of them alter the memory access pattern of the application and as a result interfere with the efficiency of the hardware prefetching mechanisms present in modern architectures. We propose an optimization algorithm that analytically classifies an algorithmic description of a loop nest in order to decide whether it should be optimized stressing its temporal or spatial locality, while also taking hardware prefetching into account. We implement our technique as a tool to be used with the Halide compiler and test it on a variety of benchmarks. We find an average performance improvement of over 40% compared to previous analytical models targeting the Halide language and compiler. Savvas Sioutas, Sander Stuijk, Henk Corporaal, Twan Basten, Lou J. Somers |
CGO | 2 |
| 2018 | Fault-Tolerant Deployment of Dataflow Applications Using Virtual ProcessorsabstractMulti-processors are suited to host a dynamic mix of real-time dataflow applications, but are increasingly subject to faults because of the decreasing feature size. Applications can start and stop as needed if they execute on a private set of Virtual Processors (VPs) that are deployed on the physical processors. This allows online software updates, but makes it impossible to predict the deployment. If a fault renders a processor unusable, the free resources on other processors may be too fragmented to allow its VPs to be re-deployed. We show that mapping an application to more VPs reduces the maximum VP size. This increases the probability of successfully dealing with faults, at the cost of an increase of the total size. Such a mapping can either be run from the start, or we can split the VPs only when a fault occurs. Experiments confirm the feasibility of our approach, and show a trade-off between improved fault-tolerance and resource usage for both strategies. J. Reinier van Kampenhout, Sander Stuijk, Kees Goossens |
DSD | 2 |
| 2018 | A Review of Near-Memory Computing Architectures: Opportunities and ChallengesabstractThe conventional approach of moving stored data to the CPU for computation has become a major performance bottleneck for emerging scale-out data-intensive applications due to their limited data reuse. At the same time, the advancement in integration technologies have made the decade-old concept of coupling compute units close to the memory (called Near-Memory Computing) more viable. Processing right at the "home" of data can completely diminish the data movement problem of data-intensive applications. This paper focuses on analyzing and organizing the extensive body of literature on near-memory computing across various dimensions: starting from the memory level where this paradigm is applied, to the granularity of the application that could be executed on the near-memory units. We highlight the challenges as well as the critical need of evaluation methodologies that can be employed in designing these special architectures. Using a case study, we present our methodology and also identify topics for future research to unlock the full potential of near-memory computing. Gagandeep Singh 0002, Lorenzo Chelini, Stefano Corda, Ahsan Javed Awan, Sander Stuijk, Roel Jordans, Henk Corporaal, Albert-Jan Boonstra |
DSD | 5 |
| 2018 | A Unified Programming Model for Time- and Data-Driven Embedded ApplicationsabstractModern embedded systems encompass a fast increasing range of applications, spanning from automotive to multimedia, and industrial automation. To tackle the increasing design complexity, the model-based design paradigm promotes the use of Models of Computation (MoCs) to capture the essential application properties. Existing MoCs are split between the event/time-triggered paradigm and the data-driven paradigm. However, time and data are two inter-related dimensions that are essential for defining the correct application behavior. In this paper we advocate a unified MoC that integrates the notions of time and data while accounting for imperfect clocks. We present the formal properties of our model and show how the Synchronous Data Flow (SDF) MoC can be used to analyze the time performance guarantees. Gabriela Breaban, Sander Stuijk, Kees Goossens |
PDP | 2 |
| 2018 | Exploiting Specification Modularity to Prune the Optimization-Space of Manufacturing SystemsabstractIn this paper we address the makespan optimization of industrial-sized manufacturing systems. We introduce a framework which specifies functional system requirements in a compositional way and automatically computes makespan optimal solutions respecting these requirements. We show the optimization problem to be NP-Hard. To scale towards systems of industrial complexity, we propose a novel approach based on a subclass of compositional requirements which we call constraints. We prove that these constraints always prune the worst-case optimization-space thereby increasing the odds of finding an optimal solution (with respect to the additional constraints). We demonstrate the applicability of the framework on an industrial-sized manufacturing system. João Bastos, Sander Stuijk, Jeroen Voeten, Ramon R. H. Schiffelers, Henk Corporaal |
SCOPES | 2 |
| 2017 | Efficient synchronization methods for LET-based applications on a Multi-Processor System on ChipabstractDistributed control applications cover a wide range of areas such as automotive, avionics, and automation. The Logical Execution Time (LET) Model of Computation (MoC) was proposed as a formal method to describe the functional and timing behavior of such applications. However, modern Multi-Processor Systems on Chip (MPSOC) do not have a shared notion of time between processors, due to their use of Globally Asynchronous Locally Synchronous (GALS) architecture. In this paper we propose two methods (based on FIFO channels and barriers) to implement time and data synchronization on a MPSOC. While a barrier synchronizes the execution flows of tasks at predefined points in their executions, a FIFO is an asynchronous data communication method between two tasks. First, they are used to implement LET applications. Next, we show how dataflow applications and mixed LET-dataflow applications are supported too. We implemented both methods on a MPSOC prototyped on a FPGA, and show that the data synchronization outperforms the related work by 67% in terms of software overhead. Gabriela Breaban, Sander Stuijk, Kees Goossens |
DATE | 2 |
| 2017 | Programming and analysing scenario-aware dataflow on a multi-processor platformabstractThe FSM-SADF model of computation is especially suitable for analysing real-time applications with input-dependent behaviour such as different modes, variable execution times and scalable parallelism. Although FSM-SADF specifies which scenario transitions are possible, it does not specify how and when they are decided at runtime. Multiple actors of a scenario, e.g. video stream header parsing, may have to fire before it is known which scenario the application is in. We solve this causality dilemma with a concept for executing a sequence of scenarios, and demonstrate an implementation on multiple processors with rolling static-order scheduling. We furthermore present a platform-aware analysis model that covers concept and implementation, and integrate the contributions in a toolflow. A proof-of-concept confirms the low overhead of the implementation and the exact timing analysis of our model. J. Reinier van Kampenhout, Sander Stuijk, Kees Goossens |
DATE | 2 |
| 2017 | Identifying bottlenecks in manufacturing systems using stochastic criticality analysisabstractSystem design is a difficult process with many design-choices for which the impact may be difficult to foresee. Manufacturing system design is no exception to this. Increased use of flexible manufacturing systems which are able to perform different operations/use-cases further raises the design complexity. One important criterion to consider is the overall makespan and associated critical path for the different use-cases of the system. Stochastic critical path analysis plays a fundamental role in providing useful feedback for system designers to evaluate alternative specifications, which traditional fixed-time analysis cannot. In this paper, we extend our formal model-based framework, for the specification and design of manufacturing systems, with stochastic analysis abilities by associating a criticality index to each action performed by the system. This index can then be visualized and used within the framework such that a system designer can make better informed decisions. We propose a Monte-Carlo method as an estimation algorithm and we explicitly define and use confidence intervals to achieve an acceptable estimation error. We further demonstrate the use of the extended framework and stochastic analysis with an example manufacturing system. João Bastos, Bram van der Sanden, Olaf Donk, Jeroen Voeten, Sander Stuijk, Ramon R. H. Schiffelers, Henk Corporaal |
FDL | 5 |
| 2017 | Color-Distortion Filtering for Remote PhotoplethysmographyabstractThis paper introduces a powerful filtering method that exploits the physiological and optical properties of skin reflections to improve the performance of remote photoplethysmography (rPPG). Based on the fact that the pulsatile and nonpulsatile (e.g., intensity and specular changes) components have different reflection-spectra in a multi-wavelength camera, we propose to use their different characteristic color changes as a soft criterion to filter the RGB-signals in the frequency domain, such that the AC-components containing clear color distortions can be suppressed before the actual pulse extraction. This leads to a novel “Color-Distortion Filter” (CDF) that can be used as a common pre-processing step for arbitrary rPPG algorithms to increase their robustness. The benchmark in challenging fitness recordings shows that CDF brings significant and consistent improvements to all benchmarked rPPG algorithms, and drives all multi-channel approaches to a similar high quality-level. Wenjin Wang 0002, Albertus C. den Brinker, Sander Stuijk, Gerard de Haan |
FG | 3 |
| 2017 | Mapping of synchronous dataflow graphs on MPSoCs based on parallelism enhancement
Qi Tang 0002, Twan Basten, Marc Geilen, Sander Stuijk, Jibo Wei |
J. Parallel Distributed Comput. | 4 |
| 2017 | Task-FIFO Co-Scheduling of Streaming Applications on MPSoCs with Predictable Memory HierarchyabstractThis article studies the scheduling of real-time streaming applications on multiprocessor systems-on-chips with predictable memory hierarchy. An iteration-based task-FIFO co-scheduling framework is proposed for this problem. We obtain FIFO size distributions using Pareto space searching, based on which the task-to-processor mapping is obtained with the potential FIFO allocation being taken into account; then, the FIFO-to-memory allocation is optimized to minimize the total memory access cost; finally, a self-timed throughput analysis method that considers memory and direct memory access controller contention is utilized to analyze the throughput. Our methods are validated by a set of synthesized and practical applications on different platforms. Qi Tang 0002, Twan Basten, Marc Geilen, Sander Stuijk, Jibo Wei |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2017 | Reducing the Complexity of Dataflow Graphs Using Slack-Based MergingabstractThere exist many dataflow applications with timing constraints that require real-time guarantees on safe execution without violating their deadlines. Extraction of timing parameters (offsets, deadlines, periods) from these applications enables the use of real-time scheduling and analysis techniques, and provides guarantees on satisfying timing constraints. However, existing extraction techniques require the transformation of the dataflow application from highly expressive dataflow computational models, for example, Synchronous Dataflow (SDF) and Cyclo-Static Dataflow (CSDF) to Homogeneous Synchronous Dataflow (HSDF). This transformation can lead to an exponential increase in the size of the application graph that significantly increases the runtime of the analysis. In this article, we address this problem by proposing an offline heuristic algorithm called slack-based merging . The algorithm is a novel graph reduction technique that helps in speeding up the process of timing parameter extraction and finding a feasible real-time schedule, thereby reducing the overall design time of the real-time system. It uses two main concepts: (a) the difference between the worst-case execution time of the SDF graph’s firings and its timing constraints (slack) to merge firings together and generate a reduced-size HSDF graph, and (b) the novel concept of merging called safe merge , which is a merge operation that we formally prove cannot cause a live HSDF graph to deadlock. The results show that the reduced graph (1) respects the throughput and latency constraints of the original application graph and (2) typically speeds up the process of extracting timing parameters and finding a feasible real-time schedule for real-time dataflow applications. They also show that when the throughput constraint is relaxed with respect to the maximal throughput of the graph, the merging algorithm is able to achieve a larger reduction in graph size, which in turn results in a larger speedup of the real-time scheduling algorithms. Hazem Ismail Abdel Aziz Ali, Sander Stuijk, Benny Akesson, Luís Miguel Pinho |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | A Fast Estimator of Performance with Respect to the Design Parameters of Self Re-Entrant FlowshopsabstractSelf re-entrant flowshops consist of machines which process jobs several times. They are found in applications like TFT-LCD assembly, LED manufacturing and industrial printing. The structure of a self re-entrant flowshop influences its performance. To get better performance while reducing costs a fast performance estimation method can be used to explore the trade-offs between the structure and the performance during the design process. We present a novel performance estimator that uses the information in the jobs being processed to analyse the trade-offs. We study the impact of the design parameters of an industrial printer using the performance estimator with an average estimation time of 1.1 milliseconds per job and with an average accuracy of not less than 96%. Umar Waqas, Marc Geilen, Sander Stuijk, Joost van Pinxten, Twan Basten, Lou J. Somers, Henk Corporaal |
DSD | 3 |
| 2016 | Robust online face tracking-by-detectionabstractThe problem of online face tracking from unconstrained videos is still unresolved. Challenges range from coping with severe online appearance variations to coping with occlusion. We propose RFTD (Robust Face Tracking-by-Detection), a system which combines tracking and detection into a single framework to robustly track a face from unconstrained videos. RFTD is based on the idea that adaptive and stable algorithmic components can complement each other in the task of online tracking. An online Structured Output SVM (SO-SVM) is combined with an offline trained face detector to break the self-learning loop typical in tracking. In turn, the face detector is supervised by a Deformable Part Model (DPM) landmark detector to asses the reliability of the face detection output. Extensive evaluation shows that RFTD delivers consistently good tracking performances across different scenarios, i.e., high mean success rate and lowest standard deviation across benchmark videos. Francesco Comaschi, Sander Stuijk, Twan Basten, Henk Corporaal |
ICME | 2 |
| 2016 | Multiconstraint Static Scheduling of Synchronous Dataflow Graphs Via Retiming and UnfoldingabstractSynchronous dataflow graphs (SDFGs) are widely used to represent digital signal processing algorithms and streaming media applications. This paper presents several methods for binding and scheduling SDFGs on a multiprocessor platform. Exploring the state space generated by a self-timed execution (STE) of an SDFG, we present an exact method for static rate-optimal scheduling of SDFGs via implicit retiming and unfolding. By modeling a constraint as an extra enabling condition for the STE, we get a constrained STE which implies a schedule under the constraint. We present a general framework for scheduling SDFGs under constraints on the number of processors, buffer sizes, auto-concurrency, or combinations of them. Exploring the state space generated by the constrained STE, we can check whether a retiming, which leads to a rate-optimal schedule under the processor (or memory) constraint, exists. Combining this with a binary search strategy, we present heuristic methods to find a proper retiming and a static scheduling that schedules the retimed SDFG with optimal rate and with as few processors (or as little storage space) as possible. None of the methods explicitly converts an SDFG to its equivalent homogenous SDFG, the size of which may be tremendously larger than the original SDFG. We perform experiments on several models of real applications and hundreds of synthetic SDFGs. The results show that the exact method outperforms existing methods significantly; our heuristics reduce the resources used and are computationally efficient. Xue-Yang Zhu, Marc Geilen, Twan Basten, Sander Stuijk |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | Online multi-face detection and tracking using detector confidence and structured SVMsabstractOnline detection and tracking of a variable number of faces in video is a crucial component in many real-world applications ranging from video-surveillance to online gaming. In this paper we propose FAST-DT, a fully automated system capable of detecting and tracking a variable number of faces online without relying on any scene-specific cues. FAST-DT integrates a generic face detector with an adaptive structured output SVM tracker and uses the detector's continuous confidence to solve the target creation and removal problem. We improve in recall and precision over a state-of-the-art method on a video dataset of more than two hours while providing in addition an increase in throughput. Francesco Comaschi, Sander Stuijk, Twan Basten, Henk Corporaal |
AVSS | 2 |
| 2015 | A re-entrant flowshop heuristic for online scheduling of the paper path in a large scale printer
Umar Waqas, Marc Geilen, Jack Kandelaars, Lou J. Somers, Twan Basten, Sander Stuijk, Patrick Vestjens, Henk Corporaal |
DATE | 6 |
| 2015 | A Scenario-Aware Dataflow Programming ModelabstractThe FSM-SADF model of computation allows to find a tight bound on the throughput of firm real-time applications by capturing dynamic variations in scenarios. We explore an FSM-SADF programming model, and propose three different alternatives for scenario switching. The best candidate for our CompSOC platform was implemented, and experiments confirm that the tight throughput bound results in a reduced resource budget. This comes at the cost of a predictable overhead at run-time as well as increased communication and memory budgets. We show that design choices offer interesting trade-offs between run-time cost and resource budgets. J. Reinier van Kampenhout, Sander Stuijk, Kees Goossens |
DSD | 2 |
| 2015 | Modeling resource sharing using FSM-SADFabstractThis paper proposes a modeling approach to capture the mapping of an application on a platform. The approach is based on Scenario-Aware Dataflow (SADF) models. In contrast to the related work, we express the complete design-space in a single formal SADF model. This allows us to have a compact and explorable state-space linked with an executable model capable of symbolically analyzing different mappings for their timing behavior. We can model different bindings for application tasks, different static-orders schedules for tasks bound in shared resources, as well as naturally capturing resource claiming/unclaiming using SADF semantics. Moreover, by using the inherent properties of dataflow graphs and the dynamic behavior of a Finite-State Machine, we can model different levels of pipelining, such as full application pipelining and interleaved pipelining of consecutive executions of the application. The size of the model is independent of the number of executions of the application. Since we are able to capture all this behavior in a single SADF model we can use available dataflow analysis, such as worst-case and best-case throughput and deadlock-freedom checking. Furthermore, since the model captures the design-space independently of the analysis technique, one can use different exploration approaches to analyze different sets of requirements. João Bastos, Sander Stuijk, Jeroen Voeten, Ramon R. H. Schiffelers, Johan Jacobs, Henk Corporaal |
MEMOCODE | 2 |
| 2015 | Maximizing the Number of Good Dies for Streaming Applications in NoC-Based MPSoCs Under Process VariationabstractScaling CMOS technology into nanometer feature-size nodes has made it practically impossible to precisely control the manufacturing process. This results in variation in the speed and power consumption of a circuit. As a solution to process-induced variations, circuits are conventionally implemented with conservative design margins to guarantee the target frequency of each hardware component in manufactured multiprocessor chips. This approach, referred to as worst-case design, results in a considerable circuit upsizing, in turn reducing the number of dies on a wafer. This work deals with the design of real-time systems for streaming applications (e.g., video decoders) constrained by a throughput requirement (e.g., frames per second) with reduced design margins, referred to as better-than-worst-case design . To this end, the first contribution of this work is a complete modeling framework that captures a streaming application mapped to an NoC-based multiprocessor system with voltage-frequency islands under process-induced die-to-die and within-die frequency variations . The framework is used to analyze the impact of variations in the frequency of hardware components on application throughput at the system level . The second contribution of this work is a methodology to use the proposed framework and estimate the impact of reducing circuit design margins on the number of good dies that satisfy the throughput requirement of a real-time streaming application . We show on both synthetic and real applications that the proposed better-than-worst-case design approach can increase the number of good dies by up to 9.6% and 18.8% for designs with and without fixed SRAM and IO blocks, respectively. Davit Mirzoyan, Benny Akesson, Sander Stuijk, Kees Goossens |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2014 | Timing analysis of First-Come First-Served scheduled interval-timed Directed Acyclic GraphsabstractAnalyzing worst-case application timing for systems with shared resources is difficult, especially when non-monotonic arbitration policies like First-Come-First-Served (FCFS) scheduling are used in combination with varying task execution times. Analysis methods that conservatively analyze these systems are often based on state-space exploration, which is not scalable due to its inherent susceptibility to combinatorial explosion. We propose a scalable timing analysis method on periodically restarted Directed Acyclic Task Graphs, that can provide conservative bounds on task timing properties when shared resources with FCFS scheduling are used. By expressing task enabling and completion times in intervals, denoting best-case and worst-case timing properties, contention on the shared resources can be estimated using conservative approximations. With an industrial case study we show that our approach can easily analyze models with thousands of tasks in less than 10 seconds, and the worst-case bounds obtained show an average improvement of 46% compared to bounds obtained by static worst-case analysis. Raymond Frijns, Shreya Adyanthaya, Sander Stuijk, Jeroen Voeten, Marc Geilen, Ramon R. H. Schiffelers, Henk Corporaal |
DATE | 3 |
| 2014 | Memory-constrained static rate-optimal scheduling of synchronous dataflow graphs via retimingabstractSynchronous dataflow graphs (SDFGs) are widely used to model digital signal processing (DSP) and streaming media applications. In this paper, we use retiming to optimize SDFGs to achieve a high throughput with low storage requirement. Using a memory constraint as an additional enabling condition, we define a memory constrained self-timed execution of an SDFG. Exploring the state-space generated by the execution, we can check whether a retiming exists that leads to a rate-optimal schedule under the memory constraint. Combining this with a binary search strategy, we present a heuristic method to find a proper retiming and a static scheduling which schedules the retimed SDFG with optimal rate (i.e., maximal throughput) and with as little storage space as possible. Our experiments are carried out on hundreds of synthetic SDFGs and several models of real applications. Differential synthetic graph results and real application results show that, in 79% of the tested models, our method leads to a retimed SDFG whose rate-optimal schedule requires less storage space than the proven minimal storage requirement of the original graph, and in 20% of the cases, the returned storage requirements equal the minimal ones. The average improvement is about 7.3%. The results also show that our method is computationally efficient. Xue-Yang Zhu, Marc Geilen, Twan Basten, Sander Stuijk |
DATE | 4 |
| 2014 | ContoExam: an ontology on context-aware examinationsabstractPatient observations in health care, subjective surveys in social research or dyke sensor data in water management are all examples of measurements. Several ontologies already exist to express measurements, W3C's SSN ontology being a prominent example. However, these ontologies address quantities and properties as being equal, and ignore the foundation required to establish comparability between sensor data. Moreover, a measure of an observation in itself is almost always inconclusive without the context in which the measure was obtained. ContoExam addresses these aspects, providing for a unifying capability for context-aware expressions of observations about quantities and properties alike, by aligning them to ontological foundations, and by binding observations inextricably with their context. Paul Brandt, Twan Basten, Sander Stuijk |
FOIS | 3 |
| 2014 | A tool for fast ground truth generation for object detection and tracking from videoabstractObject detection and tracking is one of the most important components in computer vision applications. To carefully evaluate the performance of detection and tracking algorithms, it is important to develop benchmark data sets. One of the most tedious and error-prone aspects when developing benchmarks, is the generation of the ground truth. This paper presents FAST-GT (FAst Semi-automatic Tool for Ground Truth generation), a new generic framework for the semiautomatic generation of ground truths. FAST-GT reduces the need for manual intervention thus speeding-up the ground-truthing process. Francesco Comaschi, Sander Stuijk, Twan Basten, Henk Corporaal |
ICIP | 2 |
| 2013 | Dataflow-Based Multi-ASIP Platform Approach for Digital Control ApplicationsabstractTo provide a good balance between the performance and flexibility of future digital control platforms, we propose an FPGA-based heterogeneous multiprocessor approach, in which the platform is composed of processing elements from a set of parameterizable heterogeneous Application-Specific Instruction-set Processors (ASIPs), connected with an hierarchical interconnect. With a case-study treating two different industrial-scale controllers, we show that a platform generated from our template using only a small library of instantiable ASIP types outperforms an optimized 8-core general-purpose implementation by a factor 4.9 on sampling frequency and reduces IO-delay with 37.5%. Raymond Frijns, A. L. J. Kamp, Sander Stuijk, Jeroen Voeten, M. Bontekoe, K. J. A. Gemei, Henk Corporaal |
DSD | 3 |
| 2013 | MAMPSX: A demonstration of rapid, predictable HMPSOC synthesisabstractHeterogeneous Multiprocessor systems-on-chip (HMPSoC) are becoming popular as a means of meeting energy efficiency requirements of modern embedded systems. However, as these HMPSoCs run multimedia applications as well, they also need to meet realtime requirements. Designing HMPSoCs with predictable timing behavior is a key challenge, as the current design methods for these platforms are semi-automated, non-predictable, or support limited heterogeneity. In this demonstration, we present a design framework to rapidly generate and implement predictable HMPSoC designs. It takes the application specifications and the architecture model as input and generates the entire HMPSoC, for FPGA prototyping, that meets the throughput constraints of the application. We also present results of a case study that computes the performance-power tradeoffs of an industrial vision application. A tool-chain targeting the Xilinx Zynq FPGA is also presented. Shakith Fernando, Mark Wijtvliet, Firew Siyoum, Yifan He 0002, Sander Stuijk, Akash Kumar 0001, Henk Corporaal |
FPL | 5 |
| 2013 | Throughput-constrained DVFS for scenario-aware dataflow graphsabstractDynamic behavior of streaming applications can be effectively modeled by scenario-aware dataflow graphs (SADFs). Many streaming applications must provide timing guarantees (e.g., throughput) to assure their quality-of-service. For instance, a video decoder which is running on a mobile device is expected to deliver a video stream with a specific frame rate. Moreover, the energy consumption of such applications on handheld devices should be as low as possible. This paper proposes a technique to select a suitable multiprocessor DVFS point for each mode (scenario) of a dynamic application described by an SADF. The technique assures strict timing guarantees while minimizing energy consumption. The technique is evaluated by applying it to several streaming applications. It solves the problem faster than the state of the art technique for dataflow graphs. Moreover, the DVFS controller devised using the proposed technique is more compact and reduces energy consumption compared to the controller devised using the counterpart technique. Morteza Damavandpeyma, Sander Stuijk, Twan Basten, Marc Geilen, Henk Corporaal |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2013 | Efficient communication support in predictable heterogeneous MPSoC designs for streaming applications
Yifan He 0002, Dongrui She, Sander Stuijk, Henk Corporaal |
J. Syst. Archit. | 3 |
| 2013 | Schedule-Extended Synchronous Dataflow GraphsabstractSynchronous dataflow graphs (SDFGs) are used extensively to model streaming applications. An SDFG can be extended with scheduling decisions, allowing SDFG analysis to obtain properties, such as throughput or buffer sizes for the scheduled graphs. Analysis times depend strongly on the size of the SDFG. SDFGs can be statically scheduled using static-order schedules. The only generally applicable technique to model a static-order schedule in an SDFG is to convert it to a homogeneous SDFG (HSDFG). This may lead to an exponential increase in the size of the graph and to suboptimal analysis results (e.g., for buffer sizes in multiprocessors). We present techniques to model two types of static-order schedules, i.e., periodic schedules and periodic single appearance schedules, directly in an SDFG. Experiments show that both techniques produce more compact graphs compared to the technique that relies on a conversion to an HSDFG. This results in reduced analysis times for performance properties and tighter resource requirements. Morteza Damavandpeyma, Sander Stuijk, Twan Basten, Marc Geilen, Henk Corporaal |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Modeling static-order schedules in synchronous dataflow graphsabstractSynchronous dataflow graphs (SDFGs) are used extensively to model streaming applications. An SDFG can be extended with scheduling decisions, allowing SDFG analysis to obtain properties like throughput or buffer sizes for the scheduled graphs. Analysis times depend strongly on the size of the SDFG. SDFGs can be statically scheduled using static-order schedules. The only generally applicable technique to model a static-order schedule in an SDFG is to convert it to a homogeneous SDFG (HSDFG). This conversion may lead to an exponential increase in the size of the graph and to sub-optimal analysis results (e.g., for buffer sizes in multi-processors). We present a technique to model periodic static-order schedules directly in an SDFG. Experiments show that our technique produces more compact graphs compared to the technique that relies on a conversion to an HSDFG. This results in reduced analysis times for performance properties and tighter resource requirements. Morteza Damavandpeyma, Sander Stuijk, Twan Basten, Marc Geilen, Henk Corporaal |
DATE | 2 |
| 2012 | Playing games with scenario- and resource-aware SDF graphs through policy iterationabstractThe two-player mean-payoff game is a well-known game theoretic model that is widely used, for instance in economics and control theory. For controller synthesis, a controller is modeled as a player while the environment, or plant, is modeled as the opponent player (adversary). Synthesizing an optimal controller that satisfies a given criterion corresponds to finding a winning strategy for the controller player. Emerging streaming applications (audio, video, communication, etc.) for embedded systems exhibit both input sensitive and controller sensitive runtime behavior, where the controller's role is runtime management or scheduling. Embedded controllers need to be optimized for dynamic inputs, while guaranteeing throughput constraints. In this paper, we consider this design task for scenario- and resource-aware dataflow graphs that model streaming applications. Scenarios in these models capture classes of dynamic environment behavior. We demonstrate how to model and solve the controller synthesis problem by constructing a winning strategy in a two-player mean payoff throughput game. Marc Geilen, Twan Basten, Sander Stuijk, Henk Corporaal |
DATE | 4 |
| 2012 | Parametric throughput analysis of scenario-aware dataflow graphsabstractScenario-aware dataflow graphs (SADFs) efficiently model dynamic applications. The throughput of an application is an important metric to determine the performance of the system. For example, the number of frames per second output by a video decoder should always stay above a threshold that determines the quality of the system. During design-space exploration (DSE) or run-time management (RTM), numerous throughput calculations have to be performed. Throughput calculations have to be performed as fast as possible. For synchronous dataflow graphs (SDFs), a technique exists that extracts throughput expressions from a parameterized SDF in which the execution time of the tasks (actors) is a function of some parameters. Evaluation of these expressions can be done in a negligible amount of time and provides the throughput for a specific set of parameter values. This technique is not applicable to SADFs. In this paper, we present a technique, based on Max-Plus automata, that finds throughput expressions for a parameterized SADF. Experimental evaluation shows that our technique can be applied to realistic applications. These results also show that our technique is better scalable and faster compared to the available parametric throughput analysis technique for SDFs. Morteza Damavandpeyma, Sander Stuijk, Marc Geilen, Twan Basten, Henk Corporaal |
ICCD | 2 |
| 2012 | Static Rate-Optimal Scheduling of Multirate DSP Algorithms via Retiming and UnfoldingabstractThis paper presents an exact method and a heuristic method for static rate-optimal multiprocessor scheduling of real-time multi rate DSP algorithms represented by synchronous data flow graphs (SDFGs). Through exploring the state-space generated by a self-timed execution (STE) of an SDFG, a static rate-optimal schedule via explicit retiming and implicit unfolding can be found by our exact method. By constraining the number of concurrent firings of actors of an STE, the number of processors used in a schedule can be limited. Using this, we present a heuristic method for processor-constrained rate-optimal scheduling of SDFGs. Both methods do not explicitly convert an SDFG to its equivalent homogenous SDFG. Our experimental results show that the exact method gives a significant improvement compared to the existing methods, our heuristic method further reduces the number of processors used. Xue-Yang Zhu, Marc Geilen, Twan Basten, Sander Stuijk |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2012 | Efficient Retiming of Multirate DSP AlgorithmsabstractMultirate digital signal processing (DSP) algorithms are often modeled with synchronous dataflow graphs (SDFGs). A lower iteration period implies a faster execution of a DSP algorithm. Retiming is a simple but efficient graph transformation technique for performance optimization, which can decrease the iteration period without affecting functionality. In this paper, we deal with two problems: feasible retiming-retiming a SDFG to meet a given iteration period constraint, and optimal retiming-retiming a SDFG to achieve the smallest iteration period. We present a novel algorithm for feasible retiming and based on that one, a new algorithm for optimal retiming, and prove their correctness. Both methods work directly on SDFGs, without explicitly converting them to their equivalent homogeneous SDFGs. Experimental results show that our methods give a significant improvement compared to the earlier methods. Xue-Yang Zhu, Twan Basten, Marc Geilen, Sander Stuijk |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2011 | Hybrid Code-Data Prefetch-Aware Multiprocessor Task Graph SchedulingabstractThe ever increasing performance gap between processors and memories is one of the biggest performance bottlenecks for computer systems. In this paper, we propose a task scheduling technique that schedules an application, modeled with a task graph, on a multiprocessor system-on chip (MPSoC) that contains a limited on-chip memory. The proposed scheduling technique explores the trade-off between executing tasks in a code-driven (i.e. executing parallel tasks) or data-driven (i.e. executing pipelined tasks) manner to minimize the run-time of the application. Our static scheduler identifies those task sequences in which it is useful to use a code-driven execution and those task sequences that benefit from a data-driven execution. We extend the proposed technique to consider prefetching when choosing a suitable task order. The technique is implemented using an integer linear programming framework. To evaluate the effectiveness of the technique, we use an application from the multimedia domain and a synthetic task graph that is used in related work. Our experimental results show that our scheduler is able to reduce the run-time of an MP3 decoder application by 8% compared to a commonly used heuristic scheduler. Morteza Damavandpeyma, Sander Stuijk, Twan Basten, Marc Geilen, Henk Corporaal |
DSD | 2 |
| 2011 | Exploiting Inter and Intra Application Dynamism to Save EnergyabstractThe dynamism inside applications can be exploited to save energy. A proactive scheduler that exploits this dynamism through Dynamic Frequency and Voltage Scaling (DVFS) has been presented in [1][2]. So far, the claimed energy savings of this scheduler have never been demonstrated on a real hardware platform. In this paper, we show for the first time that the proactive scheduler from [1][2] is able to realize the claimed energy savings. Our experimental results show that this scheduler reduces the energy consumption of a MP3 decoder running on a TI Omap3530 board by 18%. The proactive scheduler from [1][2] can only be used on a system that is running a single application. In this paper, we extend this scheduler such that it can deal with multiple applications that are running concurrently. Our scheduler exploits both inter and intra application dynamism to save energy while providing timing guarantees to all applications. Experimental results show that our scheduler is able to achieve the same energy savings, 38%, as an optimized version of the Linux on demand scheduler when running two H.263 decoders concurrently. However, our scheduler achieves this result without any deadline misses, the on demand scheduler fails 10% of its deadlines leading to a substantial quality loss. Martijn Koedam, Sander Stuijk, Henk Corporaal |
DSD | 2 |
| 2011 | Power Minimisation for Real-Time Dataflow ApplicationsabstractEnergy efficient execution of applications is important for many reasons, e.g. time between battery charges, device temperature. Voltage and Frequency Scaling (VFS) enables applications to be run at lower frequencies on hardware resources thereby consuming less power. Real-time applications have deadlines that must be met otherwise their output is devalued. Dataflow modelling of real-time applications enables off-line verification of the application's temporal requirements. In this paper we describe a method to reduce the combined static and dynamic energy consumption using a Dynamic VFS (DVFS) technique for dataflow modelled real-time applications that may be mapped onto multiple hardware resources. We achieve this by using an application's static slack in order to perform DVFS while still satisfying the application's temporal requirements. We show that by formulating a dataflow modelled application and its mapping as a convex optimisation problem, with energy consumption as the objective function, the problem can be solved with a generic convex optimisation solver, producing an energy optimal constant frequency per application task. Our method allows task frequencies to be constrained such that, e.g. one frequency per application or per processor may be achieved. Andrew Nelson 0001, Orlando Moreira, Anca Mariana Molnos, Sander Stuijk, Ba Thang Nguyen, Kees Goossens |
DSD | 4 |
| 2011 | Iteration-Based Trade-Off Analysis of Resource-Aware SDFabstractSynchronous dataflow graphs (SDFGs) are widely used to model streaming applications such as signal processing and multimedia applications in embedded systems. Trade-off analysis between performance and resource usage of SDFGs allows designers to explore implementation alternatives of a system while meeting its performance requirements and resource constraints. This type of analysis is computationally very challenging, particularly when resources may be shared among computations. With resource sharing, system scheduling decisions lead to a combinatorial explosion in the number of scheduling alternatives to be explored. We present a new approach to explore the trade-offs in a such systems. It breaks analysis down in iterations of dataflow graph execution and uses a max-plus algebra semantics. The experimental results on a set of realistic benchmark models show that the new iteration-based approach and the traditional time-based analysis approach complement each other. None of the two approaches dominates the other in terms of quality of the analysis results and analysis time. The two approaches combined give the highest quality result. Marc Geilen, Twan Basten, Sander Stuijk, Henk Corporaal |
DSD | 4 |
| 2011 | Resource-Efficient Real-Time Scheduling Using Credit-Controlled Static-Priority ArbitrationabstractA present-day System-on-Chip (SoC) runs a wide range of applications with diverse real-time requirements. Resources, such as processors, interconnects and memories, are shared between these applications to reduce cost. Resource sharing causes temporal interference, which must be bounded by a suitable resource arbiter. System-level analysis techniques use the service guarantee of the arbiter to ensure that real-time requirements of these applications are satisfied. A service guarantee that underestimates the minimum service provided by an arbiter results in more allocation of resources than needed to satisfy latency and throughput requirements. For instance, a linear service guarantee cannot accurately capture burst service provision by many priority-based schedulers, such as Credit-Controlled Static Priority (CCSP) and Priority-Budget Scheduling (PBS). As a result, the timing analysis of these arbiters becomes too pessimistic. This leads to unnecessary cost penalties since some SoC resources, such as SDRAM bandwidth, are scarce and expensive. This paper addresses this problem for the CCSP arbiter. The two main contributions are: (1) a piecewise linear service guarantee that accurately captures bursty service provisioning, and (2) an equivalent dataflow model of the new service guarantee, which is an essential component to integrate the arbiter with dataflow-based system-level design techniques that analyze the worst-case latency and throughput of real-time applications. The new service guarantee enables efficient resource utilization under CCSP arbitration. Experimental results of an H.263 video decoder application show that memory bandwidth savings from 26% up to 67% can be achieved by using the new service guarantee as compared to the existing linear service guarantee. Firew Siyoum, Benny Akesson, Sander Stuijk, Kees Goossens, Henk Corporaal |
RTCSA (1) | 3 |
| 2010 | MNEMEE: a framework for memory management and optimization of static and dynamic data in MPSoCsabstractAs embedded systems are becoming the center of our digital life, system design becomes progressively harder. The integration of multiple features on devices with limited resources requires careful and exhaustive exploration of the design search space in order to efficiently map modern applications to an embedded multi-processor platform. The MNEMEE project [1] addresses this challenge by offering a unique integrated tool flow that performs source-to-source transformations to automatically optimize the original source code and map it on the target platform. The optimizations aim at reducing the number of memory accesses and the required memory storage of both dynamically and statically allocated data. Furthermore, the MNEMEE tool flow performs optimal assignment of all data on the memory hierarchy of the target platform. Overall, the MNEMEE techniques embedded in it will lead to more cost efficient systems that offer a better performance and lower energy consumption. This tutorial gives an overview of the MNEMEE tool flow. The objective of the tutorial is to familiarize the audience with the tool framework and the optimizations used in the individual tools. The tutorial also features a demonstration of the tool flow. This demonstration shows that the tools developed in the MNEMEE project provide a user-friendly and efficient framework for MPSoC programming and memory management. Arindam Mallik, Peter Marwedel, Dimitrios Soudris, Sander Stuijk |
CASES | 4 |
| 2010 | Automated bottleneck-driven design-space exploration of media processing systemsabstractMedia processing systems often have limited resources and strict performance requirements. An implementation must meet those design constraints while minimizing resource usage and energy consumption. Design-space exploration techniques help system designers to pinpoint bottlenecks in a system for a given configuration. The trade-offs between performance and resources in the design space can guide designers to tailor and tune the system. Many applications in those systems are computationally intensive and can be modeled by a synchronous dataflow graph. We present a bottleneck-analysis-driven technique to explore the design space of those systems automatically and incrementally. The feasibility and efficiency of the technique is demonstrated with experiments on a set of realistic application models ranging from multimedia to digital printing. Marc Geilen, Twan Basten, Sander Stuijk, Henk Corporaal |
DATE | 4 |
| 2010 | A Predictable Multiprocessor Design Flow for Streaming Applications with Dynamic BehaviourabstractThe design of new embedded systems is getting more and more complex as more functionality is integrated into these systems. To deal with the design complexity, a predictable design flow is needed. The result should be a system that guarantees that an application can perform its own tasks within strict timing deadlines, independent of other applications running on the system. Synchronous Dataflow Graphs (SDFGs) provide predictability and are often used to model time-constrained streaming applications that are mapped onto a multiprocessor platform. However, the model abstracts from the dynamic application behaviour which may lead to a large overestimation of its resource requirements. We present a design flow that takes the dynamic behaviour of applications into account when mapping them onto a multiprocessor platform. The design flow provides throughput guarantees for each application independent of the other applications while taking into account the available processing capacity, memory and communication bandwidth. The design flow generates a set of mappings that provide a trade-off in their resource usage. This trade-off can be used by a run-time mechanism to adapt the mapping in different use-cases to the available resource. The experimental results show that our design flow reduces the resource requirements of an MPEG-4 decoder by 66% compared to a state-of-the-art design flow based on SDFGs. Sander Stuijk, Marc Geilen, Twan Basten |
DSD | 1 |
| 2010 | Thermal-aware scratchpad memory design and allocationabstractScratchpad memories (SPMs) have become a promising on-chip storage solution for embedded systems from an energy, performance and predictability perspective. The thermal behavior of these types of memories has not been considered in detail. This thermal behavior plays an important role in the reliability of silicon devices and in their static (leakage) power consumption. In this paper, we propose two different techniques to improve the thermal behavior of SPMs. First, we propose a hardware-based, thermal-aware address translation technique that physically distributes memory accesses to consecutive addresses evenly over the whole memory area. Second, we propose a software-based, thermal-aware address generation technique. This technique tries to distribute the variables that are allocated to the SPM in such a way that an even thermal distribution is achieved. The first technique works particularly well for applications with a regular access pattern, whereas the second technique can also improve the behavior of applications with irregular access patterns. The two techniques thus complement each other and work well together. Using the first technique we show that the peak temperature of an SPM in 65nm technology, when running a typical streaming application, is decreased by up-to 10.0°C. Temperature cycling is reduced from up-to 14.8°C to almost zero in comparison with a non-thermal-aware solution. For our benchmark applications with an irregular access pattern, the second technique is able to reduce the peak temperature by up-to 3.5°C. These savings for both techniques are obtained without any performance degradation or extra silicon area. Morteza Damavandpeyma, Sander Stuijk, Twan Basten, Marc Geilen, Henk Corporaal |
ICCD | 2 |
| 2010 | CA-MPSoC: An automated design flow for predictable multi-processor architectures for multiple applications
Ahsan Shabbir, Akash Kumar 0001, Sander Stuijk, Bart Mesman, Henk Corporaal |
J. Syst. Archit. | 3 |
| 2010 | Buffer Sizing for Rate-Optimal Single-Rate Data-Flow Scheduling RevisitedabstractSingle-Rate Data-Flow (SRDF) graphs, also known as Homogeneous Synchronous Data-Flow (HSDF) graphs or Marked Graphs, are often used to model the implementation and do temporal analysis of concurrent DSP and multimedia applications. An important problem in implementing applications expressed as SRDF graphs is the computation of the minimal amount of buffering needed to implement a static periodic schedule (SPS) that is optimal in terms of execution rate, or throughput. Ning and Gao [1] propose a linear-programming-based polynomial algorithm to compute this minimal storage amount, claiming optimality. We show via a counterexample that the proposed algorithm is not optimal. We prove that the problem is, in fact, NP-complete. We give an exact solution, and experimentally evaluate the degree of inaccuracy of the algorithm of Ning and Gao. Orlando Moreira, Twan Basten, Marc Geilen, Sander Stuijk |
IEEE Trans. Computers | 4 |
| 2009 | A parameterized compositional multi-dimensional multiple-choice knapsack heuristic for CMP run-time managementabstractModern embedded systems typically contain chip-multiprocessors (CMPs) and support a variety of applications. Applications may run concurrently and can be started and stopped over time. Each application may typically have multiple feasible configurations, trading off quality aspects (energy consumption, audio-visual quality) with resource usage for various types of resources. Overall system quality needs to be guaranteed and optimized at all times. This leads to the need for a run-time management solution that selects an appropriate system configuration from all the application configurations of active applications. This run-time management problem can be phrased as a multi-dimensional multiple-choice knapsack (MMKP) problem. We present a compositional heuristic to solve MMKP, that due to the compositionality is better suited to CMP run-time management than existing heuristics that are all not compositional. Our heuristic outperforms the best-known heuristic to date. The heuristic is parameterized, leading to the additional advantage that it allows to trade off execution time vs. solution quality, and to bound the time needed to compute a solution. The latter makes it particularly well-suited for resource-constrained embedded platforms. Hamid Shojaei, Amir Hossein Ghamarian, Twan Basten, Marc Geilen, Sander Stuijk, Rob Hoes |
DAC | 5 |
| 2008 | Parametric Throughput Analysis of Synchronous Data Flow GraphsabstractSynchronous data flow graphs (SDFGs) have proved to be a very successful tool for modeling, analysis and synthesis of multimedia applications targeted at both single- and multiprocessor platforms. One of the most prominent performance constraints of concurrent real-time applications is throughput. For given actor execution times, throughput can be verified by analyzing the SDFG models of such applications, for instance using maximum cycle mean analysis or state space analysis. In various contexts, such as design space exploration or run-time reconfiguration, many fast throughput computations are required for varying actor execution times. We present methods to compute throughput of an SDFG where actor execution times can be parameters. The throughput of these graphs is obtained in the form of a function of these parameters. Recalculation of throughput is then merely an evaluation of this function for specific parameter values, which is much faster than the standard throughput analysis. We propose three different algorithms for parametric throughput analysis and evaluate these algorithms experimentally, showing the feasibility of the approach and showing that a divide and conquer algorithm performs best. Amir Hossein Ghamarian, Marc Geilen, Twan Basten, Sander Stuijk |
DATE | 4 |
| 2008 | Analyzing concurrency in streaming applications
Sander Stuijk, Twan Basten |
J. Syst. Archit. | 1 |
| 2008 | Resource-efficient routing and scheduling of time-constrained streaming communication on networks-on-chip
Sander Stuijk, Twan Basten, Marc Geilen, Amir Hossein Ghamarian, Bart D. Theelen |
J. Syst. Archit. | 1 |
| 2008 | Throughput-Buffering Trade-Off Exploration for Cyclo-Static and Synchronous Dataflow GraphsabstractMultimedia applications usually have throughput constraints. An implementation must meet these constraints, while it minimizes resource usage and energy consumption. The compute intensive kernels of these applications are often specified as cyclo-static or synchronous dataflow graphs. Communication between nodes in these graphs requires storage space which influences throughput. We present an exact technique to chart the Pareto space of throughput and storage trade-offs, which can be used to determine the minimal buffer space needed to execute a graph under a given throughput constraint. The feasibility of the exact technique is demonstrated with experiments on a set of realistic DSP and multimedia applications. To increase scalability of the approach, a fast approximation technique is developed that guarantees both throughput and a, tight, bound on the maximal overestimation of buffer requirements. The approximation technique allows to trade off worst-case overestimation versus run-time. Sander Stuijk, Marc Geilen, Twan Basten |
IEEE Trans. Computers | 1 |
| 2007 | Multiprocessor Resource Allocation for Throughput-Constrained Synchronous Dataflow GraphsabstractEmbedded multimedia systems often run multiple time-constrained applications simultaneously. These systems use multiprocessor systems-on-chip of which it must be guaranteed that enough resources are available for each application to meet its throughput constraints. This requires a task binding and scheduling mechanism that provides timing guarantees for each application independent of other applications while taking into account the available processor space, memory and communication bandwidth. Sander Stuijk, Twan Basten, Marc Geilen, Henk Corporaal |
DAC | 1 |
| 2007 | Latency Minimization for Synchronous Data Flow GraphsabstractSynchronous data flow graphs (SDFGs) are a very useful means for modeling and analyzing streaming applications. Some performance indicators, such as throughput, have been studied before. Although throughput is a very useful performance indicator for concurrent real-time applications, another important metric is latency. Especially for applications such as video conferencing, telephony and games, latency beyond a certain limit cannot be tolerated. This paper proposes an algorithm to determine the minimal achievable latency, providing an execution scheme for executing an SDFG with this latency. In addition, a heuristic is proposed for optimizing latency under a throughput constraint. Experimental results show that latency computations are efficient despite the theoretical complexity of the problem. Substantial latency improvements are obtained, of 24-54% on average for a synthetic benchmark of 900 models, and up to 37% for a benchmark of six real DSP and multimedia models. The heuristic for minimizing latency under a throughput constraint gives optimal latency and throughput results under a constraint of maximal throughput for all DSP and multimedia models, and for over 95% of the synthetic models. Amir Hossein Ghamarian, Sander Stuijk, Twan Basten, Marc Geilen, Bart D. Theelen |
DSD | 2 |
| 2006 | Exploring trade-offs in buffer requirements and throughput constraints for synchronous dataflow graphsabstractMultimedia applications usually have throughput constraints. An implementation must meet these constraints, while it minimizes resource usage and energy consumption. The compute intensive kernels of these applications are often specified as Synchronous Dataflow Graphs. Communication between nodes in these graphs requires storage space which influences throughput. We present exact techniques to chart the Pareto space of throughput and storage trade-offs, which can be used to determine the minimal storage space needed to execute a graph under a given throughput constraint. The feasibility of the approach is demonstrated with a number of examples. Sander Stuijk, Marc Geilen, Twan Basten |
DAC | 1 |
| 2006 | Resource-Efficient Routing and Scheduling of Time-Constrained Network-on-Chip CommunicationabstractNetwork-on-chip-based multiprocessor systems-on-chip are considered as future embedded systems platforms. One of the steps in mapping an application onto such a parallel platform involves scheduling the communication on the network-on-chip. This paper presents different scheduling strategies that minimize resource usage by exploiting all scheduling freedom offered by networks-on-chip. Our experiments show that resource-utilization is improved when compared to existing techniques Sander Stuijk, Twan Basten, Marc Geilen, Amir Hossein Ghamarian, Bart D. Theelen |
DSD | 1 |
| 2006 | Liveness and Boundedness of Synchronous Data Flow GraphsabstractSynchronous data flow graphs (SDFGs) have proven to be suitable for specifying and analyzing streaming applications that run on single- or multi-processor platforms. Streaming applications essentially continue their execution indefinitely. Therefore, one of the key properties of an SDFG is liveness, i.e., whether all parts of the SDFG can run infinitely often. Another elementary requirement is whether an implementation of an SDFG is feasible using a limited amount of memory. In this paper, we study two interpretations of this property, called boundedness and strict boundedness, that were either already introduced in the SDFG literature or studied for other models. A third and new definition is introduced, namely self-timed boundedness, which is very important to SDFGs, because self-timed execution results in the maximal throughput of an SDFG. Necessary and sufficient conditions for liveness in combination with all variants of boundedness are given, as well as algorithms for checking those conditions. As a by-product, we obtain an algorithm to compute the maximal achievable throughput of an SDFG that relaxes the requirement of strong connectedness in earlier work on throughput analysis Amir Hossein Ghamarian, Marc Geilen, Twan Basten, Bart D. Theelen, Mohammad Reza Mousavi 0001, Sander Stuijk |
FMCAD | 6 |
| 2006 | A scenario-aware data flow model for combined long-run average and worst-case performance analysisabstractData flow models are used for specifying and analysing signal processing and streaming applications. However, traditional data flow models are either not capable of expressing the dynamic aspects of modern streaming applications or they do not support relevant analysis techniques. The dynamism in modern streaming applications often originates from different modes of operation (scenarios) in which data production and consumption rates and/or execution times may differ. This paper introduces a scenario-aware generalisation of the synchronous data flow model, which uses a stochastic approach to model the order in which scenarios occur. The formally defined operational semantics of a scenario-aware data flow model implies a Markov chain, which can be analysed for both long-run average and worst-case performance metrics using existing exhaustive or simulation-based techniques. The potential of using scenario-aware data flow models for performance analysis of modern streaming applications is illustrated with an MPEG-4 decoder example Bart D. Theelen, Marc Geilen, Twan Basten, Jeroen Voeten, Stefan Valentin Gheorghita, Sander Stuijk |
MEMOCODE | 6 |
| 2005 | Minimising buffer requirements of synchronous dataflow graphs with model checkingabstractSignal processing and multimedia applications are often implemented on resource constrained embedded systems. It is therefore important to find implementations that use as little resources as possible. These applications are frequently specified as synchronous dataflow graphs. Communication between actors of these graphs requires storage capacity. In this paper, we present an exact method to determine the minimum storage capacity required to execute the graph using model-checking techniques. This can be done for different measures of storage capacity. The problem is known to be NP-complete and because of this, existing buffer minimisation techniques are heuristics and hence not exact. Modern model-checking tools are quite efficient and they have been successfully applied to scheduling-related problems. We study the feasibility of this approach with examples. Marc Geilen, Twan Basten, Sander Stuijk |
DAC | 3 |
| 2005 | Automatic scenario detection for improved WCET estimationabstractModern embedded applications usually have real-time constraints and they are implemented using heterogeneous multiprocessor systems-on-chip. Dimensioning a system requires accurate estimations of the worst-case execution time (WCET). Overestimation leads to over-dimensioning. This paper introduces a method for automatic discovery of scenarios that incorporate correlations between different parts of applications. It is based on the application parameters with a large impact on the execution time. We show on a benchmark that, using scenarios, the estimated WCET may be reduced with 16%. Stefan Valentin Gheorghita, Sander Stuijk, Twan Basten, Henk Corporaal |
DAC | 2 |
| 2005 | Predictable Embedding of Large Data Structures in Multiprocessor Networks-on-ChipabstractThis extended abstract presents models to derive timing and resource usage numbers for an application when distant, shared memories are used in an important class of future embedded platforms, namely network-on-chip-based multiprocessors. Sander Stuijk, Twan Basten, Bart Mesman, Marc Geilen |
DATE | 1 |
| 2005 | Predictable embedding of large data structures in multiprocessor networks-on-chipabstractPredictable, tile-based multiprocessor networks-on-chip are considered as future embedded systems platforms. Each tile contains one or a few processors and local memories. These memories are typically too small to store large data structures (e.g. a video frame). A solution to this is to embed tiles with large memories in the architecture. However, fetching data from these memories is slow because of the large network delays. The delay can be hidden by using prefetching. Our main contributions are models that allow timing analysis to provide guaranteed quality and performance when using remote memories and prefetching. We use two realistic video applications to show that our models can be used in practice to derive a predictable system using large memory tiles and prefetching, and to provide guaranteed real-time performance. Sander Stuijk, Twan Basten, Bart Mesman, Marc Geilen |
DSD | 1 |
| 2003 | Analyzing Concurrency in Computational NetworksabstractWe present a concurrency model that allows reasoning about concurrency in executable specifications. The model mainly focuses on data-flow and streaming applications and at task-level concurrency. The aim of the model is to provide insight in concurrency bottlenecks in an application and to provide support for performing implementation independent concurrency optimization. Sander Stuijk, Twan Basten |
MEMOCODE | 1 |