EDBT 2026 Demo / reviewers in the wild / expert
Andrew Nelson 0001
dblp:46/6957-1
· DBLP profile ↗
15ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-4071-3502ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Time-Aware Shaper Scheduling for In-Vehicle Networks via Deep Reinforcement LearningabstractModern vehicles increasingly rely on distributed computing platforms that exchange large volumes of sensor and control data with strict timing requirements. Ensuring that this traffic meets its deadlines over Ethernet-based in-vehicle networks requires Time-Sensitive Networking (TSN) and, in particular, effective configuration of the Time-Aware Shaper (TAS). However, generating and updating TAS schedules that remain valid as traffic patterns evolve is an NP-hard problem that traditional optimization or heuristic methods address only partially. This paper introduces a Deep Reinforcement Learning (DRL) scheduler that learns to configure TAS schedules directly from network state while preserving standard compliance through analytical validation. The proposed DRL scheduler encodes the scenario (network topology and workload) of the in-vehicle network using a Graph Neural Network (GNN) and learns scheduling policies that balance deadline satisfaction, latency, and resource utilization. Evaluation on a comprehensive benchmark shows that the proposed approach consistently outperforms state-of-the-art heuristics and a topology-specific DRL baseline, achieving higher success rate and lower delay while maintaining efficient bandwidth use. Once trained, it can adapt to new traffic scenarios within milliseconds, demonstrating the potential of the DRL-based scheduler as a foundation for adaptive and reliable communication in next-generation software-defined vehicles. Mohammad Parsa Karimi, Majid Nabi, Andrew Nelson 0001, Kees Goossens, Twan Basten |
IEEE Internet Things J. | 3 |
| 2025 | Deep-Reinforcement-Learning-Based Scheduler for Time-Aware Shaper in In-Vehicle NetworksabstractAs vehicles develop into software-defined platforms with powerful automated driving capabilities and driver support systems, their in-vehicle networks become significantly more complicated. A key technique for ensuring deterministic, low-latency connectivity for crucial data traffic in such settings is Time-Sensitive Networking (TSN), and specifically the Time-Aware Shaper (TAS). However, current TAS scheduling techniques have difficulty adjusting schedules to dynamically shifting traffic patterns and changing operating conditions. This paper presents an adaptive scheduler using Deep Reinforcement Learning (DRL), which aims to meet strict deadlines, reducing latency and providing near-ideal resource usage. Experimental results for different vehicle scenarios show that our DRL-based scheduler performs better in terms of success rate, low latency, and overall network performance than state-of-the-art heuristic algorithms such as earliest deadline first (EDF) scheduling. Mohammadparsa Karimi, Majid Nabi, Andrew Nelson 0001, Kees Goossens, Twan Basten |
VTC2025-Spring | 3 |
| 2025 | INSIM: A Modular Simulation Platform for TSN-based In-Vehicle NetworksabstractIn-vehicle networks (IVNs) are rapidly evolving to support increasingly complex automotive applications, demanding higher bandwidth and deterministic timing bounds. Time-Sensitive Networking (TSN) has emerged as a promising Ethernet-based technology that addresses these stringent requirements. However, evaluating TSN-based IVN strategies remains a challenge due to the lack of standardized benchmarks and simulation tools. This paper introduces INSIM, a modular simulation platform specifically designed for TSN-based IVNs, providing an intuitive graphical interface, an extensible plug-in architecture, and integrated benchmarking features. INSIM integrates analytical performance models and discrete-event simulations (as plug-ins), enhancing the workflow for engineers by refining topology design, adjusting parameters, conducting simulations, and assessing performance, while providing researchers with a flexible platform to plug in, analyze, and compare custom network resource managers or analytical performance models. Mohammadparsa Karimi, Majid Nabi, Andrew Nelson 0001, Kees Goossens, Twan Basten |
VTC2025-Fall | 3 |
| 2022 | An Evaluation Framework for Vision-in-the-Loop Motion Control SystemsabstractIndustrial applications and processes such as quality inspections, pick and place operations, and semiconductor manufacturing require accurate positioning control for achieving the high throughput of the assembly machines. Vision-based sensing is considered to be a potential means to achieve robust positioning control which is referred to as a vision-in-the-loop (VIL) system. In such motion systems, the point-of-control and the point-of-interest are often different due to several physical factors. In this case, validation of a system is done only when a machine prototype is available. A physical prototype is often expensive and infeasible in real-life. This paper proposes an evaluation framework for VIL systems targeting a predictable multi-core embedded platform. The presented framework offers model-in-the-loop (MIL), software-in-the-loop (SIL), and processor-in-the-loop (PIL) simulation features for evaluating the closed-loop performance of industrial motion control systems. As a deployment platform, we consider a predictable embedded platform CompSOC. The predictable nature of the CompSOC platform guarantees periodic and deterministic execution of the control applications and allows verification of the timing properties and performance of the VIL system. Additionally, the framework offers automatic code generation feature targeting the CompSOC platform. Closed-loop simulation setup models the system dynamics and camera position in the CoppeliaSim physics simulation engine and simulates the system software in C and MATLAB. CoppeliaSim runs as a server and MATLAB as a client in synchronous mode. We show the effectiveness of our framework using a vision-based motion control example. Chaitanya Jugade, Daniel Hartgers, Phan Dúc Anh, Sajid Mohamed, Mojtaba Haghi, Dip Goswami, Andrew Nelson 0001, Gijs van der Veen, Kees Goossens |
ETFA | 7 |
| 2021 | Modeling, implementation, and analysis of XRCE-DDS applications in distributed multi-processor real-time embedded systemsabstractThe Publish-Subscribe paradigm is a design pattern for transparent communication in many recent distributed applications. Data Distribution Service (DDS) is a machine-to-machine communication standard that aims to provide reliable, highperformance, inter-operable, and real-time data exchange based on publish-subscribe paradigm. However, the high resource requirement of DDS limits its usage in low-cost embedded systems. XRCE-DDS is a Client-Agent based standard to enable resource-constrained small embedded systems to connect to the DDS global data space. Current XRCE-DDS implementations suffer from dependencies with host operating systems, target only single processing units, and lack performance analysis methods. In this paper, we present a bare-metal implementation of XRCE-DDS standard on the CompSOC platform as an instance of Multi-Processor System on Chip (MPSoC). The proposed framework includes a hard real-time side hosting the XRCE-DDS Client, and a soft real-time side hosting the XRCE-DDS Agent. A Scenario Aware Data Flow (SADF) model is proposed to capture the dynamism of the system behavior in terms of different execution scenarios. We analyze the long-term expected value for throughput by capturing the probabilistic scenario switching using a proposed Markov model which is experimentally validated. Saeid Dehnavi, Dip Goswami, Martijn Koedam, Andrew Nelson 0001, Kees Goossens |
DATE | 4 |
| 2021 | Efficient Tensor Cores support in TVM for Low-Latency Deep learningabstractDeep learning algorithms are gaining popularity in autonomous systems. These systems typically have stringent latency constraints that are challenging to meet given the high computational demands of these algorithms. Nvidia introduced Tensor Cores (TCs) to speed up some of the most commonly used operations in deep learning algorithms. Compilers (e.g., TVM) and libraries (e.g., cuDNN) focus on the efficient usage of TCs when performing batch processing. Latency sensitive applications can however not exploit large batch processing. This paper presents an extension to the TVM compiler that generates low latency TCs implementations, particularly for batch size 1. Experimental results show that our solution reduces the latency on average by 14% compared to the cuDNN library on a Desktop RTX2070 GPU, and by 49% on an Embedded Jetson Xavier GPU. Savvas Sioutas, Sander Stuijk, Andrew Nelson 0001, Henk Corporaal |
DATE | 4 |
| 2021 | A Deployment Framework for Quality-Sensitive Applications in Resource-Constrained Dynamic EnvironmentsabstractTraditional embedded systems and recent platforms used in emerging computing paradigms (e.g., fog computing) have resource limits and require their applications and services to be dynamically added (i.e., deployed) and removed at run-time. These applications often have non-functional (quality) requirements (e.g., end-to-end latency) which are only satisfied when sufficient resources are allocated to them. Hence, a run-time decision-maker is needed to optimize the deployments, in terms of resource budgets that are allocated to applications. Additionally, computing platforms have become heterogeneous in terms of their resources and the applications they execute. However, the existing deployment solutions are limited to specific resources and services. In this paper, we propose a run-time deployment framework that is more flexible in defining constraints and optimization goals and works with more heterogeneous resources and resource models than existing solutions. The framework is implemented on an embedded platform as a proof of concept. Shayan Tabatabaei Nikkhah, Marc Geilen, Dip Goswami, Martijn Koedam, Andrew Nelson 0001, Kees Goossens |
DSD | 5 |
| 2021 | CompROS: A composable ROS2 based architecture for real-time embedded robotic developmentabstractRobot Operating System (ROS) is a de-facto standard robot middleware in many academic and industrial use cases. However, utilizing ROS/ROS2 in safety-critical embedded applications with real-time requirement is challenging because of C1) Non-real-time underlying hardware, C2) No control on the host OS scheduler, C3) Unpredictable dynamic memory allocation, C4) High resource requirement, and C5) Unpredictable execution model for ROS nodes. In this paper, we address these limiting factors by proposing a hardwaresoftware architecture -CompROS- for ROS2 based robotic development in a Multi-Processor System on Chip (MPSoC) platform. The proposed hardware architecture consists of a Hard Real-Time (HRT) RISC-V based subsystem implemented in the Programmable Logic (PL) part of the MPSoC platform, a Soft Real-Time (SRT) ARM-based subsystem in the Processing System (PS) part of the MPSoC platform, and a Non-Real-Time (NRT) PC. While the proposed hardware architecture along with a partitioning layer overcomes the first two limiting factors, the rest are managed by the proposed multi-layer software architecture. We make a bare-metal implementation of XRCE-DDS standard for PL-PS communication, while peer-to-peer PL-PL communication is done through a proposed real-time publish-subscribe approach. The reliable communication for PS-PL communication is done through utilizing C-HEAP protocol. Further, we integrate ROS2 software layers on top of the proposed hardware and software layers. Finally, with respect to C5, we present a real-time execution model of ROS2 nodes by a mapping of ROS2 entities to CompROS entities, which is validated through experimental results. We run ROS2 middleware with an executable size of less than 200 KB on an MPSoC platform. Saeid Dehnavi, Martijn Koedam, Andrew Nelson 0001, Dip Goswami, Kees Goossens |
IROS | 3 |
| 2021 | DominoSearch: Find layer-wise fine-grained N: M sparse schemes from dense neural networksabstractNeural pruning is a widely-used compression technique for Deep Neural Networks (DNNs). Recent innovations in Hardware Architectures (e.g. Nvidia Ampere Sparse Tensor Core) and N:M fine-grained Sparse Neural Network algorithms (i.e. every M-weights contains N non-zero values) reveal a promising research line of neural pruning. However, the existing N:M algorithms only address the challenge of how to train N:M sparse neural networks in a uniform fashion (i.e. every layer has the same N:M sparsity) and suffer from a significant accuracy drop for high sparsity (i.e. when sparsity > 80\%). To tackle this problem, we present a novel technique -- \textbf{\textit{DominoSearch}} to find mixed N:M sparsity schemes from pre-trained dense deep neural networks to achieve higher accuracy than the uniform-sparsity scheme with equivalent complexity constraints (e.g. model size or FLOPs). For instance, for the same model size with 2.1M parameters (87.5\% sparsity), our layer-wise N:M sparse ResNet18 outperforms its uniform counterpart by 2.1\% top-1 accuracy, on the large-scale ImageNet dataset. For the same computational complexity of 227M FLOPs, our layer-wise sparse ResNet18 outperforms the uniform one by 1.3\% top-1 accuracy. Furthermore, our layer-wise fine-grained N:M sparse ResNet50 achieves 76.7\% top-1 accuracy with 5.0M parameters. {This is competitive to the results achieved by layer-wise unstructured sparsity} that is believed to be the upper-bound of Neural Network pruning with respect to the accuracy-sparsity trade-off. We believe that our work can build a strong baseline for further sparse DNN research and encourage future hardware-algorithm co-design work. Our code and models are publicly available at \url{https://github.com/NM-sparsity/DominoSearch}. Aojun Zhou, Sander Stuijk, Rob G. J. Wijnhoven, Andrew Nelson 0001, Hongsheng Li 0001, Henk Corporaal |
NeurIPS | 5 |
| 2015 | Distributed power management of real-time applications on a GALS multiprocessor SOCabstractIt is generally desirable to reduce the power consumption of embedded systems. Dynamic Voltage and Frequency Scaling (DVFS) is a commonly applied technique to achieve power reduction at the cost of computational performance. Multiprocessor System on Chips (MPSoCs) can have multiple voltage and frequency domains, e.g. per-core. When DVFS is applied to real-time applications, the effects must be accounted for in the associated formal timing model. In this work, we contribute our distributed multi-core run-time power-management technique for real-time dataflow applications that uses per-core lookup-tables to select low-power DVFS operating points that meet the application's timing requirement. We describe in detail how timing slack is observed locally at run-time on each core and is used to select a local DVFS operating point that meets the application's timing requirement. We further describe our static off-line formal analysis technique to generate these per-core lookup-tables that link timing slack to low-power DVFS operating points. We provide an experimental analysis of our proposed technique using an H.263 decoder application that is mapped onto an FPGA prototyped hardware platform. Andrew Nelson 0001, Kees Goossens |
EMSOFT | 1 |
| 2015 | An efficient configuration methodology for time-division multiplexed single resourcesabstractComplex contemporary systems contain multiple applications, some which have firm real-time requirements while others do not. These applications are deployed on multi-core platforms with shared resources, such as processors, interconnect, and memories. However, resource sharing causes contention between sharing applications that must be resolved by a resource arbiter. Time-Division Multiplexing (TDM) is a commonly used arbiter, but it is challenging to configure such that the bandwidth and latency requirements of the real-time resource clients are satisfied, while minimizing their total allocation to improve the performance of non-real-time clients. This work addresses this problem by presenting an efficient TDM configuration methodology. The five main contributions are: 1) An analysis to derive a bandwidth and latency guarantee for a TDM schedule with arbitrary slot assignment, 2) A formulation of the TDM configuration problem and a proof that it is NP-hard, 3) An integer-linear programming model that optimally solves the configuration problem by exhaustively evaluating all possible TDM schedule sizes, 4) A heuristic method to choose candidate schedule sizes that substantially reduces computation time with only a slight decrease in efficiency, 5) An experimental evaluation of the methodology that examines its scalability and quantifies the trade-off between computation time and total allocation for the optimal and the heuristic algorithms. The approach is also demonstrated on a case study of a HD video and graphics processing system, where a memory controller is shared by a number of processing elements. Benny Akesson, Anna Minaeva, Premysl Sucha, Andrew Nelson 0001, Zdenek Hanzálek |
RTAS | 4 |
| 2015 | Dataflow formalisation of real-time streaming applications on a Composable and Predictable Multi-Processor SOC
Andrew Nelson 0001, Kees Goossens, Benny Akesson |
J. Syst. Archit. | 1 |
| 2014 | CoMik: A predictable and cycle-accurately composable real-time microkernelabstractThe functionality of embedded systems is ever increasing. This has lead to mixed time-criticality systems, where applications with a variety of real-time requirements co-exist on the same platform and share resources. Due to inter-application interference, verifying the real-time requirements of such systems is generally non trivial. In this paper, we present the CoMik microkernel that provides temporally predictable and composable processor virtualisation. CoMik's virtual processors are cycle-accurately composable, i.e. their timing cannot affect the timing of co-existing virtual processors by even a single cycle. Real-time applications executing on dedicated virtual processors can therefore be verified and executed in isolation, simplifying the verification of mixed time-criticality systems. We demonstrate these properties through experimentation on an FPGA prototyped hardware platform. Andrew Nelson 0001, Ashkan Beyranvand Nejad, Anca Mariana Molnos, Martijn Koedam, Kees Goossens |
DATE | 1 |
| 2014 | Composable and Predictable Dynamic Loading for Time-Critical Partitioned SystemsabstractIn time-critical systems such as in avionics, for safety and timing guarantees, applications are isolated from each other. Resources are partitioned in time and space creating a partition per application. Such isolation allows fault containment and independent development, testing and verification of applications. Current partitioned systems do not allow dynamically adding applications. Applications are statically loaded in their respective partitions. However dynamic loading can be useful or even necessary for scenarios such as on-board software updates, dynamic reconfiguration or re-loading applications in case of a fault. In this paper we propose a software architecture to dynamically create and manage partitions and a method for compostable dynamic loading which ensures that loading applications do not affect the running applications and vice versa. Furthermore the loading time is also predictable i.e. the loading time can be bounded a priori. We achieve this by splitting the loading process into parts, wherein only a small part which reserves minimum required resources is executed in the system partition and the other parts are executed in the allocated application partition which ensures isolation from other applications. We implement the software architecture for a SoC prototype on an FPGA board and demonstrate its composability and predictability properties. Shubhendu Sinha, Martijn Koedam, Rob van Wijk, Andrew Nelson 0001, Ashkan Beyranvand Nejad, Marc Geilen, Kees Goossens |
DSD | 4 |
| 2011 | Power Minimisation for Real-Time Dataflow ApplicationsabstractEnergy efficient execution of applications is important for many reasons, e.g. time between battery charges, device temperature. Voltage and Frequency Scaling (VFS) enables applications to be run at lower frequencies on hardware resources thereby consuming less power. Real-time applications have deadlines that must be met otherwise their output is devalued. Dataflow modelling of real-time applications enables off-line verification of the application's temporal requirements. In this paper we describe a method to reduce the combined static and dynamic energy consumption using a Dynamic VFS (DVFS) technique for dataflow modelled real-time applications that may be mapped onto multiple hardware resources. We achieve this by using an application's static slack in order to perform DVFS while still satisfying the application's temporal requirements. We show that by formulating a dataflow modelled application and its mapping as a convex optimisation problem, with energy consumption as the objective function, the problem can be solved with a generic convex optimisation solver, producing an energy optimal constant frequency per application task. Our method allows task frequencies to be constrained such that, e.g. one frequency per application or per processor may be achieved. Andrew Nelson 0001, Orlando Moreira, Anca Mariana Molnos, Sander Stuijk, Ba Thang Nguyen, Kees Goossens |
DSD | 1 |