EDBT 2026 Demo / reviewers in the wild / expert
Tomasz Kloda
dblp:156/5559
· DBLP profile ↗
22ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-0822-4976ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 first-author · 12 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI Inference in the Heat: Thermal-Aware Strict Partitioning for Configurable Real-Time Gang Tasks
Binqi Sun, Jinyang Li 0004, Tomasz Kloda, Tarek F. Abdelzaher, Marco Caccamo |
RTAS | 3 |
| 2025 | Multi-Objective Memory Bandwidth Regulation and Cache Partitioning for Multicore Real-Time Systems
Binqi Sun, Zhihang Wei, Andrea Bastoni, Debayan Roy, Mirco Theile, Tomasz Kloda, Rodolfo Pellizzoni, Marco Caccamo |
ECRTS | 6 |
| 2025 | Memguard-RW: Improved Real-Time Memory Bandwidth Regulation Within a HypervisorabstractMemory bandwidth is a critical factor in the performance of DRAM-based computing architectures, particularly in memory-intensive computations. Modern multi-core processors share critical resources, such as main memory and cache, which impact the predictability of real-time systems due to resource contention. Techniques like memory access regulation, cache partitioning, and static hypervisors aim to mitigate this contention. This paper presents an improvement of a memory control mechanism based on MemGuard, named MemGuard-RW, designed and implemented within a hypervisor. MemGuard-RW uses two performance counters for measuring the memory accesses, one for memory readings and another one for writings, thus decreasing the pessimism on the original memory access budget from MemGuard. Additionally, we extend the MemGuard schedulability analysis considering the two-counter approach. To evaluate the effectiveness of the implementation and analysis, we deployed FreeRTOS as a guest alongside three stress-generating guests, measuring the interference experienced by the FreeRTOS instance and comparing the analysis with two counters with the original analysis with one counter, using modern benchmarks. The results demonstrate that the proposed mechanism successfully regulates memory accesses, showing its potential for enhancing the predictability and performance of real-time systems in multi-core environments. Our proposed analysis with two counters reduces the upper bound of around 20 % for tasks having medium and high memory usage. Everaldo P. Gomes, Adrien Jakubiak, Giovani Gracioli, Tomasz Kloda |
ISORC | 4 |
| 2025 | SAPar: A Surrogate-Assisted DNN Partitioner for Efficient Inferences on Edge TPU PipelinesabstractPipelining deep neural networks (DNNs) across multiple Edge Tensor Processing Units (TPUs) can enhance on-device performance by increasing the capacity for DNN parameters caching and enabling pipeline parallelism. Effective deployment on pipelined Edge TPUs requires a partitioning tool to divide the DNN into segments, each assigned to a different Edge TPU in the pipeline. Achieving balanced workload distribution across these segments is crucial for optimal timing performance. However, workload balancing across Edge TPUs is challenging, as DNN execution time is influenced by proprietary hardware architecture and compiler internals, forming a black-box function inaccessible to partitioning tools. To address this challenge, this article introduces SAPar , a new surrogate-assisted DNN partitioner that integrates a neighborhood search engine with a surrogate-assisted evaluator for effective and efficient DNN partitioning. The neighborhood search engine systematically explores the decision space, guided by knowledge obtained from empirical insights and neighborhood evaluation feedback provided by the surrogate-assisted evaluator. The evaluator cooperatively applies an accurate yet time-consuming latency profiler and an efficient graph transformer-based surrogate model , achieving both precision and scalability. Experiments on real Edge TPU hardware demonstrate that SAPar achieves significantly better pipeline performance than Google’s current profiling-based partitioner with an 8.82× to 110× speedup in partitioning time. Moreover, SAPar reduces the bottleneck latency by 8.93% to 44.15% across five classic DNN models compared with a state-of-the-art reinforcement learning-based partitioner. Binqi Sun, Bohua Zou, Yigong Hu, Tomasz Kloda, Ling Wang 0001, Tarek F. Abdelzaher, Marco Caccamo |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | Response Time Analysis and Optimal Priority Assignment for Global Non-Preemptive Fixed-Priority Rigid Gang SchedulingabstractNon-preemptive rigid gang scheduling combines the efficiency of parallel execution with the reduced overhead of non-preemptive scheduling. This approach is particularly advantageous for parallel hardware accelerators, such as Google's Edge Tensor Processing Unit (TPU), which is widely used for deep neural network (DNN) inference on embedded systems. This paper studies sporadic global non-preemptive fixed-priority (NP-FP) rigid gang scheduling, which is well-suited for DNN applications in Edge TPU pipelines. Each gang task spawns a fixed number of threads that must execute concurrently across distinct processing units. We introduce the first carry-in limitation technique specifically designed for gang task response time analysis, addressing the unique challenges posed by intra-task parallelism. This technique is formulated as a generalized knapsack problem, and we develop both a linear programming relaxation and a dynamic programming approach to solve it under different time complexities. Additionally, we propose the first optimal priority assignment policy for NP-FP gang schedulability tests. Our proposed schedulability analysis and optimal priority assignment policy are evaluated through extensive experiments, including both synthetic task sets and a case study using DNN benchmarks on commercial off-the-shelf Edge TPU accelerators. The results demonstrate that the proposed approaches effectively enhance the state-of-the-art global NP-FP gang schedulability tests, achieving improvements of up to 57.9% for synthetic task sets and 76.7% for Edge TPU benchmarks. Furthermore, we conduct an ablations study to examine the impact of different algorithmic components in the proposed technique, providing valuable insights for future research. Binqi Sun, Tomasz Kloda, Jiyang Chen, Cen Lu, Marco Caccamo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUabstractPipelining on Edge Tensor Processing Units (TPUs) optimizes the deep neural network (DNN) inference by breaking it down into multiple stages processed concurrently on multiple accelerators. Such DNN inference tasks can be modeled as sporadic non-preemptive gangs with execution times that vary with their parallelism levels. This paper proposes a strict partitioning strategy for deploying DNN inferences in real-time systems. The strategy determines tasks' parallelism levels and assigns tasks to disjoint processor partitions. Configuring the tasks in the same partition with a uniform parallelism level avoids scheduling anomalies and enables schedulability verification using well-understood uniprocessor analyses. Evaluation using real-world Edge TPU benchmarks demonstrated that the proposed method achieves a higher schedulability ratio than state-of-the-art gang scheduling techniques. Binqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco Caccamo |
DAC | 2 |
| 2024 | Response Time Analysis for Fixed-Priority Preemptive Uniform Multiprocessor SystemsabstractWe present a response time analysis for global fixed-priority preemptive scheduling of constrained-deadline tasks upon a uniform multiprocessor where each processor can be characterized by a different speed. A fixed-priority scheduler assigns the jobs with the highest priorities to the fastest processors. Since determining whether all tasks can meet their deadlines is generally intractable even with identical processors, we propose two sufficient schedulability tests that calculate upper bounds on the task’s worst-case response time within polynomial and pseudo-polynomial time. The proposed tests leverage the linear programming model to upper bound the interference of the higher-priority tasks. Furthermore, we identify specific conditions and platforms upon which the problem can be solved more efficiently within linear time. These formulations are used to iteratively evaluate and refine possible solutions until a safe upper bound on the task’s worst-case response time is found. Additionally, we demonstrate that, with specific minor modifications, the proposed tests are compatible with Audsley’s optimal priority assignment. Experimental evaluations performed on synthetic task sets show that the proposed approach outperforms the state-of-the-art methods. Binqi Sun, Tomasz Kloda, Marco Caccamo |
ECRTS | 2 |
| 2024 | Strict Partitioning for Sporadic Rigid Gang TasksabstractThe rigid gang task model is based on the idea of executing multiple threads simultaneously on a fixed number of processors to increase efficiency and performance. Although there is extensive literature on global rigid gang scheduling, partitioned approaches have several practical advantages (e.g., task isolation and reduced scheduling overheads). In this paper, we propose a new partitioned scheduling strategy for rigid gang tasks, named strict partitioning. The method creates disjoint partitions of tasks and processors to avoid inter-partition interference. Moreover, it tries to assign tasks with similar volumes (i.e., parallelisms) to the same partition so that the intra-partition interference can be reduced. Within each partition, the tasks can be scheduled using any type of scheduler, which allows the use of a less pessimistic schedulability test. Extensive synthetic experiments and a case study based on Edge TPU benchmarks show that strict partitioning achieves better schedulability performance than state-of-the-art global gang schedulability analyses for both preemptive and non-preemptive rigid gang task sets. Binqi Sun, Tomasz Kloda, Marco Caccamo |
RTAS | 2 |
| 2024 | Minimizing cache usage with fixed-priority and earliest deadline first schedulingabstractAbstract Cache partitioning is a technique to reduce interference among tasks running on the processors with shared caches. To make this technique effective, cache segments should be allocated to tasks that will benefit the most from having their data and instructions stored in the cache. The requests for cached data and instructions can be retrieved faster from the cache memory instead of fetching them from the main memory, thereby reducing overall execution time. The existing partitioning schemes for real-time systems divide the available cache among the tasks to guarantee their schedulability as the sole and primary optimization criterion. However, it is also preferable, particularly in systems with power constraints or mixed criticalities where low- and high-criticality workloads are executing alongside, to reduce the total cache usage for real-time tasks. Cache minimization as part of design space exploration can also help in achieving optimal system performance and resource utilization in embedded systems. In this paper, we develop optimization algorithms for cache partitioning that, besides ensuring schedulability, also minimize cache usage. We consider both preemptive and non-preemptive scheduling policies on single-processor systems with fixed- and dynamic-priority scheduling algorithms ( Rate Monotonic ( RM ) and Earliest Deadline First ( EDF ), respectively). For preemptive scheduling, we formulate the problem as an integer quadratically constrained program and propose an efficient heuristic achieving near-optimal solutions. For non-preemptive scheduling, we combine linear and binary search techniques with different fixed-priority schedulability tests and Quick Processor-demand Analysis (QPA) for EDF. Our experiments based on synthetic task sets with parameters from real-world embedded applications show that the proposed heuristic: (i) achieves an average optimality gap of 0.79% within 0.1× run time of a mathematical programming solver and (ii) reduces average cache usage by 39.15% compared to existing cache partitioning approaches. Besides, we find that for large task sets with high utilization, non-preemptive scheduling can use less cache than preemptive to guarantee schedulability. Binqi Sun, Tomasz Kloda, Sergio Arribas García, Giovani Gracioli, Marco Caccamo |
Real Time Syst. | 2 |
| 2023 | Schedulability Analysis of Non-preemptive Sporadic Gang Tasks on Hardware AcceleratorsabstractNon-preemptive rigid gang scheduling combines the performance benefits of parallel execution with the low overhead of non-preemptive scheduling and rigid task programming model. This approach appears particularly well-suited for parallel hardware accelerators where the context switch and migration overheads are critical and should be avoided. One of the most notable examples today is Google's Edge Tensor Processing Unit (TPU) used for neural network inference on embedded boards. The paper studies sporadic non-preemptive rigid gang scheduling applied to multi-TPU edge AI accelerators. Each gang task spawns a fixed number of threads that must execute simultaneously on distinct processing units. We consider non-preemptive fixed-priority gang (NP-FP-Gang) scheduling and propose the first carry-in limitation for gang task response time analysis. The gang task carry-in limitation differs from conventional sequential tasks due to the intra-task parallelism. We formulate it as a generalized knapsack problem and develop a linear programming relaxation and a dynamic programming approach to solve the problem under different time complexities. The performance of the proposed schedulability analysis is evaluated through randomly generated synthetic task sets and a case study using neural network benchmarks executed on commercial off-the-shelf multi-TPU edge AI accelerators. The evaluation results show that the proposed response time analysis effectively improves the state of-the-art NP-FP-Gang schedulability test even by 85.7% for the Edge TPU benchmarks in particular. Binqi Sun, Tomasz Kloda, Jiyang Chen, Cen Lu, Marco Caccamo |
RTAS | 2 |
| 2023 | Co-Optimizing Cache Partitioning and Multi-Core Task Scheduling: Exploit Cache Sensitivity or Not?abstractCache partitioning techniques have been successfully adopted to mitigate interference among concurrently executing real-time tasks on multi-core processors. Considering that the execution time of a cache-sensitive task strongly depends on the cache available for it to use, co-optimizing cache partitioning and task allocation improves the system's schedulability. In this paper, we propose a hybrid multi-layer design space exploration technique to solve this multi-resource management problem. We explore the interplay between cache partitioning and schedulability by systematically interleaving three optimization layers, viz., (i) in the outer layer, we perform a breadth-first search combined with proactive pruning for cache partitioning; (ii) in the middle layer, we exploit a first-fit heuristic for allocating tasks to cores; and (iii) in the inner layer, we use the well-known recurrence relation for the schedulability analysis of non-preemptive fixed-priority (NP-FP) tasks in a uniprocessor setting. Although our focus is on NP-FP scheduling, we evaluate the flexibility of our framework in supporting different scheduling policies (NP-EDF, P-EDF) by plugging in appropriate analysis methods in the inner layer. Experiments show that, compared to the state-of-the-art techniques, the proposed framework can improve the real-time schedulability of NP-FP task sets by an average of 15.2% with a maximum improvement of 233.6% (when tasks are highly cache-sensitive) and a minimum of 1.6% (when cache sensitivity is low). For such task sets, we found that clustering similar- period (or mutually compatible) tasks often leads to higher schedulability (on average 7.6 %) than clustering by cache sensitivity. In our evaluation, the framework also achieves good results for preemptive and dynamic-priority scheduling policies. Binqi Sun, Debayan Roy, Tomasz Kloda, Andrea Bastoni, Rodolfo Pellizzoni, Marco Caccamo |
RTSS | 3 |
| 2023 | SchedGuard++: Protecting against Schedule Leaks Using Linux Containers on Multi-Core ProcessorsabstractTiming correctness is crucial in a multi-criticality real-time system, such as an autonomous driving system. It has been recently shown that these systems can be vulnerable to timing inference attacks, mainly due to their predictable behavioral patterns. Existing solutions like schedule randomization cannot protect against such attacks, often limited by the system’s real-time nature. This article presents “ SchedGuard++ ”: a temporal protection framework for Linux-based real-time systems that protects against posterior schedule-based attacks by preventing untrusted tasks from executing during specific time intervals. SchedGuard++ supports multi-core platforms and is implemented using Linux containers and a customized Linux kernel real-time scheduler. We provide schedulability analysis assuming the Logical Execution Time (LET) paradigm, which enforces I/O predictability. The proposed response time analysis takes into account the interference from trusted and untrusted tasks and the impact of the protection mechanism. We demonstrate the effectiveness of our system using a realistic radio-controlled rover platform. Not only is “ SchedGuard++ ” able to protect against the posterior schedule-based attacks, but it also ensures that the real-time tasks/containers meet their temporal requirements. Jiyang Chen, Tomasz Kloda, Rohan Tabish, Ayoosh Bansal, Chien-Ying Chen, Bo Liu 0044, Sibin Mohan, Marco Caccamo, Lui Sha |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2023 | Lazy Load Scheduling for Mixed-criticality Applications in Heterogeneous MPSoCsabstractNewly emerging multiprocessor system-on-a-chip (MPSoC) platforms provide hard processing cores with programmable logic (PL) for high-performance computing applications. In this article, we take a deep look into these commercially available heterogeneous platforms and show how to design mixed-criticality applications such that different processing components can be isolated to avoid contention on the shared resources such as last-level cache and main memory. Our approach involves software/hardware co-design to achieve isolation between the different criticality domains. At the hardware level, we use a scratchpad memory (SPM) with dedicated interfaces inside the PL to avoid conflicts in the main memory. At the software level, we employ a hypervisor to support cache-coloring such that conflicts at the shared L2 cache can be avoided. In order to move the tasks in/out of the SPM memory, we rely on a DMA engine and propose a new CPU-DMA co-scheduling policy, called Lazy Load , for which we also derive the response time analysis. The results of a case study on image processing demonstrate that the contention on the shared memory subsystem can be avoided when running with our proposed architecture. Moreover, comprehensive schedulability evaluations show that the newly proposed Lazy Load policy outperforms the existing CPU-DMA scheduling approaches and is effective in mitigating the main memory interference in our proposed architecture. Tomasz Kloda, Giovani Gracioli, Rohan Tabish, Reza Mirosanlou, Renato Mancuso 0001, Rodolfo Pellizzoni, Marco Caccamo |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2022 | Latency analysis of self-suspending task chainsabstractMany cyber-physical systems are offloading computation-heavy programs to hardware accelerators (e.g., GPU and TPU) to reduce execution time. These applications will self-suspend between offloading data to the accelerators and obtaining the returned results. Previous efforts have shown that self-suspending tasks can cause scheduling anomalies, but none has examined inter-task communication. This paper aims to explore self-suspending tasks' data chain latency with periodic activation and asynchronous message passing. We first present the cause for suspension-induced delays and worst-case latency analysis. We then propose a rule for utilizing the hardware co-processors to reduce data chain latency and schedulability analysis. Simulation results show that the proposed strategy can improve overall latency while preserving system schedulability. Tomasz Kloda, Jiyang Chen, Antoine Bertout, Lui Sha, Marco Caccamo |
DATE | 1 |
| 2022 | Memory allocation for low-power real-time embedded microcontroller: a case studyabstractMemory allocation of instructions and data can affect the program execution speed. This paper tests various memory-intensive benchmarks under different memory allocations on a Cortex-M4-based microcontroller and solves the allocation problem using integer linear programming. Zhishen Zhang, Yuwen Shen, Binqi Sun, Tomasz Kloda, Marco Caccamo |
ETFA | 4 |
| 2021 | Flexible Cache Partitioning for Multi-Mode Real-Time SystemsabstractCache partitioning is a well-studied technique that mitigates the inter-processor cache interference in multiprocessor systems. The resulting optimization problem involves allocating portions of the cache to individual processors. In multi-mode applications (e.g., flight control system that runs in take-off, cruise, or landing mode), the cache memory requirement can change over time, making runtime cache repartitioning necessary. This paper presents a cache partition allocation framework enabling flexible cache partitioning for multi-mode real-time systems. The main objective is to guarantee timing predictability in the steady states and during mode changes. We evaluate the effectiveness of our approach for multiple embedded benchmarks with different ranges of cache size sensitivity. The results show increased schedulability compared to static partitioning approaches. Ohchul Kwon, Gero Schwäricke, Tomasz Kloda, Denis Hoornaert, Giovani Gracioli, Marco Caccamo |
DATE | 3 |
| 2021 | SchedGuard: Protecting against Schedule Leaks Using Linux ContainersabstractReal-time systems have recently been shown to be vulnerable to timing inference attacks, mainly due to their predictable behavioral patterns. Existing solutions such as schedule randomization lack the ability to protect against such attacks, often limited by the system's real-time nature. This paper presents “SchedGuard”: a temporal protection framework for Linux-based hard real-time systems that protects against posterior scheduler side-channel attacks by preventing untrusted tasks from executing during specific time segments. SchedGuard is integrated into the Linux kernel using cgroups, making it amenable to use with container frameworks. We demonstrate the effectiveness of our system using a realistic radio-controlled rover platform and synthetically generated workloads. Not only is SchedGuard able to protect against the attacks mentioned above, but it also ensures that the real-time tasks/containers meet their temporal requirements. Jiyang Chen, Tomasz Kloda, Ayoosh Bansal, Rohan Tabish, Chien-Ying Chen, Bo Liu 0044, Sibin Mohan, Marco Caccamo, Lui Sha |
RTAS | 2 |
| 2020 | Fixed-Priority Memory-Centric Scheduler for COTS-Based MultiprocessorsabstractMemory-centric scheduling attempts to guarantee temporal predictability on commercial-off-the-shelf (COTS) multiprocessor systems to exploit their high performance for real-time applications. Several solutions proposed in the real-time literature have hardware requirements that are not easily satisfied by modern COTS platforms, like hardware support for strict memory partitioning or the presence of scratchpads. However, even without said hardware support, it is possible to design an efficient memory-centric scheduler. In this article, we design, implement, and analyze a memory-centric scheduler for deterministic memory management on COTS multiprocessor platforms without any hardware support. Our approach uses fixed-priority scheduling and proposes a global "memory preemption" scheme to boost real-time schedulability. The proposed scheduling protocol is implemented in the Jailhouse hypervisor and Erika real-time kernel. Measurements of the scheduler overhead demonstrate the applicability of the proposed approach, and schedulability experiments show a 20% gain in terms of schedulability when compared to contention-based and static fair-share approaches. Gero Schwäricke, Tomasz Kloda, Giovani Gracioli, Marko Bertogna, Marco Caccamo |
ECRTS | 2 |
| 2020 | An Automatic Scenario Generator for Validation of Automated Valet Parking SystemsabstractA primary goal of self-driving car manufacturers is to create an autonomous car system that is clearly and demonstrably safer than an average human-controlled car.The real-world tests are expensive, time-consuming and potentially dangerous.The virtual simulation is therefore required.The autonomous driving valet parking is expected to be the first commercially available automated driving function without a human driver at the wheel (SAE Level 4).Although many simulation solutions for the automotive market already exist, none of them features the parking environments.In this paper, we propose a new software virtual scenario generator for the parking sites.The tool populates the synthetics parking maps with objects and actions related to these environments: the cars driving from the drop-off point towards the vacant slots and the randomly placed parked cars, each with a given probability of exiting its slot.The generated scenarios are in the OpenSCENARIO format and are fully simulated in the Virtual Test Drive simulator. Andrea Tagliavini, Donato Ferraro, Tomasz Kloda, Paolo Burgio |
VEHITS | 3 |
| 2020 | Latency upper bound for data chains of real-time periodic tasks
Tomasz Kloda, Antoine Bertout, Yves Sorel |
J. Syst. Archit. | 1 |
| 2019 | Deterministic Memory Hierarchy and Virtualization for Modern Multi-Core Embedded SystemsabstractOne of the main predictability bottlenecks of modern multi-core embedded systems is contention for access to shared memory resources. Partitioning and software-driven allocation of memory resources is an effective strategy to mitigate contention in the memory hierarchy. Unfortunately, however, many of the strategies adopted so far can have unforeseen side-effects when practically implemented latest-generation, high-performance embedded platforms. Predictability is further jeopardized by cache eviction policies based on random replacement, targeting average performance instead of timing determinism. In this paper, we present a framework of software-based techniques to restore memory access determinism in high-performance embedded systems. Our approach leverages OS-transparent and DMA-friendly cache coloring, in combination with an invalidation-driven allocation (IDA) technique. The proposed method allows protecting important cache blocks from (i) external eviction by tasks concurrently executing on different cores, and (ii) internal eviction by tasks running on the same core. A working implementation obtained by extending the Jailhouse partitioning hypervisor is presented and evaluated with a combination of synthetic and real benchmarks. Tomasz Kloda, Marco Solieri, Renato Mancuso 0001, Nicola Capodieci, Paolo Valente, Marko Bertogna |
RTAS | 1 |
| 2018 | Latency analysis for data chains of real-time periodic tasksabstractA data chain is a sequence of periodic realtime communicating tasks that are processing the data from sensors up to actuators. It determines an order in which the tasks propagate data but not in which they are executed: inter-task communication and scheduling are independent. In this paper, we focus on the latency computation, considered as the time elapsed from getting the data from an input and processing it to an output of a data chain. We propose a method for the worst-case latency calculation of periodic tasks' data chains executed by a partitioned fixed-priority preemptive scheduler upon a multiprocessor platform. As far as we know, there is no such formal approach based on closed-form expression for communicating real-time tasks. Tomasz Kloda, Antoine Bertout, Yves Sorel |
ETFA | 1 |