Syed Aftab Rashid

dblp:153/1818 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0002-1739-6895ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Performance Evaluation of Brokerless Messaging Libraries
abstract
Messaging systems are essential for efficiently transferring large volumes of data, ensuring rapid response times and high-throughput communication. The state-of-the-art on messaging systems mainly focuses on the performance evaluation of brokered messaging systems, which use an intermediate broker to guarantee reliability and quality of service. However, over the past decade, brokerless messaging systems have emerged, eliminating the single point of failure and trading off reliability guarantees for higher performance. Still, the state-of-the-art on evaluating the performance of brokerless systems is scarce. In this work, we solely focus on brokerless messaging systems. First, we perform a qualitative analysis of several possible candidates, to find the most promising ones. We then design and implement an extensive open-source benchmarking suite to systematically and fairly evaluate the performance of the chosen libraries, namely, ZeroMQ, NanoMsg, and NanoMsg-Next-Generation (NNG). We evaluate these libraries considering different metrics and workload conditions, and provide useful insights into their limitations. Our analysis enables practitioners to select the most suitable library for their requirements.
Lorenzo La Corte, Syed Aftab Rashid, Andrei-Marian Dan
SRDS2
2025 Evaluating Differential Firmware Updates for Embedded IoT Device Fleets
abstract
Remote firmware updates are critical for maintaining the performance and security of Internet-of-Things (IoT) device fleets. Remote updates are especially critical from a fleet management perspective, requiring thousands of devices to be kept up-to-date. When considering a fleet of devices, transferring large update files is inefficient. Therefore many modern systems consider differential updates. In this work, we describe and evaluate a secure differential update implementation for fleets of embedded IoT devices. We demonstrate how differential updates can be performed on a real hardware platform using a popular open-source embedded update framework, SWUpdate. We also show how differential updates can integrate with redundancy concepts such as A/B partitioning and on-device security mechanisms such as secure boot. Our experiments, performed using both real hardware and a device fleet simulator, identify crucial differences between differential and full-image updates: the differential updates can decrease the file size by 49%, the install time by 78%, and require up to 77% more temporary disk memory compared to full-image updates. These key insights are essential for developers and practitioners when selecting the IoT firmware update strategy.
Jayden Renee Sorensen, Syed Aftab Rashid, Hossam ElHussini, Andrei-Marian Dan
WFCS2
2024 Improved Memory Contention Analysis for the 3-Phase Task Model
abstract
In multiprocessor-based real-time systems, main memory is identified as a major bottleneck in the worst-case timing analysis of tasks. Phased execution models such as the 3-phase task model, i.e., that divides the execution of tasks into distinct computation and memory phases, have shown to be a good candidate to tackle the memory contention problem. The 3-phase execution model in particular has gained much attention from both academia and industry as it limits when tasks can access main memory to pre-defined phases. Information on when those phases may happen and their length can then be leveraged to build a fine-grained memory contention analysis. However, the existing work that focus on the memory contention analysis for 3-phase tasks may overestimate the memory contention caused by interfering write requests. This yields pessimistic bounds on the total memory contention suffered by tasks which in turn leads to pessimistic worst-case execution time (WCET) and worst-case response time (WCRT) bounds. In this work, we improve the state-of-the-art memory contention analysis for 3-phase tasks by (i) tightly bounding the memory contention that can be suffered due to write requests; and (ii) providing a new memory contention-aware WCET analysis.
Jatin Arora 0006, Syed Aftab Rashid, Geoffrey Nelissen, Cláudio Maia, Eduardo Tovar
RTCSA2
2023 Improved Bus Contention Analysis for 3-Phase Tasks
abstract
The 3-phase task execution model has shown to be a good candidate to tackle the memory bus contention problem. It divides the execution of tasks into computation and memory phases that enable a fine-grained memory bus contention analysis. However, existing works that focus on the bus contention analysis for 3-phase tasks, neglect the fact that memory bus contention strongly relates to the number of bus/memory requests generated by tasks, which, in turn, depends on the content of the cache memories during the execution of those tasks. These existing works assume that the worst-case number of bus/memory requests will be generated during all the memory phases of all tasks, irrespective of the already existing content in the cache memory. This overestimates the memory bus contention of tasks, leading to pessimistic worst-case response time (WCRT) bounds. This work proposes a holistic approach towards bus contention analysis for 3-phase tasks by (1) deriving an upper bound on the actual cache misses of tasks that lead to bus/memory requests; (2) improving State-of-the-Art (SOTA) bus contention analysis of two bus arbitration schemes that dominate all existing works on the bus contention analysis for 3-phase tasks; and (3) performing an extensive experimental evaluation under different settings to compare the proposed analysis against the SOTA. Results show that incorporating a tighter bound on the number of cache misses of tasks into the bus contention analysis can lead to a significant improvement in the task set schedulability.
Jatin Arora 0006, Syed Aftab Rashid, Geoffrey Nelissen, Cláudio Maia, Eduardo Tovar
RTCSA2
2022 Cache-aware Schedulability Analysis of PREM Compliant Tasks
abstract
The Predictable Execution Model (PREM) is useful for mitigating inter-core interference due to shared resources such as the main memory. However, it is cache-agnostic, which makes schedulabulity analysis pessimistic, via overestimation of prefetches and write-backs. In response, we present cache-aware schedulability analysis for PREM tasks on fixed-task-priority partitioned multicores, that bounds the number of cache prefetches and write-backs. Our approach identifies memory blocks loaded in the execution of a previous scheduling interval of each task, that remain in the cache until its next scheduling interval. Doing so, greatly reduces the estimated prefetches and write backs. In experimental evaluations, our analysis improves the schedulability of PREM tasks by up to 55 percentage points.
Syed Aftab Rashid, Muhammad Ali Awan, Pedro F. Souto, Konstantinos Bletsas 0001, Eduardo Tovar
DATE1
2022 Analyzing Fixed Task Priority Based Memory Centric Scheduler for the 3-Phase Task Model
abstract
The sharing of main memory among concurrently executing tasks on a multicore platform results in increasing the execution times of those tasks in a non-deterministic manner. The use of phased execution models that divide the execution of tasks into distinct execution and memory phase(s), e.g., the PRedictable Execution Model (PREM) and the 3-Phase task model, along with Memory Centric Scheduling (MCS) present a promising solution to reduce main memory interference among tasks.Existing works in the state-of-the-art that focus on MCS have considered (i) a TDMA-based memory scheduler, i.e., tasks’ memory requests are served under a static TDMA schedule, and (ii) Processor-Priority (PP) based memory scheduler, i.e., tasks’ memory requests are served depending on the priority of the processor/core on which the task is executing. This paper extends MCS by considering a Task-Priority (TP) based memory scheduler, i.e., tasks’ memory requests are served under a global priority order depending on the priority of the task that issues the requests. We present an analysis to bound the total memory interference that can be suffered by the tasks under the TPbased MCS. In contrast to the recent works on MCS that considers non-preemptive tasks, our analysis considers limited preemptive scheduling. Additionally, we investigate the impact of different preemption points on the memory interference of tasks. Experimental results show that our proposed TP-based MCS can significantly reduce the memory interference that can be suffered by the tasks in comparison to the PP-based MCS.
Jatin Arora 0006, Syed Aftab Rashid, Cláudio Maia, Eduardo Tovar
RTCSA2
2022 Work-in-Progress: A Holistic Approach to WCRT Analysis for Multicore Systems
abstract
Commercial-off-the-shelf (COTS) multicore processors have become a preferable choice for modern systems to meet the increasing functionalities and computational demand of modern applications. However, the adoption of multicore platforms in hard real-time systems, i.e., systems that run applications with stringent timing requirements, is still under scrutiny. The main challenge that hinders the use of COTS multicore platforms in hard real-time systems is their unpredictability, which originates from the sharing of different hardware resources. A task executing on one core of a multicore platform has to compete with other co-running tasks (running on other cores) to access hardware resources such as the last-level cache (LLC), the interconnect (e.g., memory bus), and the main memory. This competition leads to inter-core contention which can significantly impact the Worst-Case Execution Time (WCET) and Worst-Case Response Time (WCRT) of tasks.
Jatin Arora 0006, Syed Aftab Rashid, Cláudio Maia, Geoffrey Nelissen, Eduardo Tovar
RTSS2
2022 Bus-contention aware WCRT analysis for the 3-phase task model considering a work-conserving bus arbitration scheme
Jatin Arora 0006, Cláudio Maia, Syed Aftab Rashid, Geoffrey Nelissen, Eduardo Tovar
J. Syst. Archit.3
2022 Schedulability analysis for 3-phase tasks with partitioned fixed-priority scheduling
abstract
Multicore platforms are being increasingly adopted in Cyber-Physical Systems (CPS) due to their advantages over single-core processors, such as raw computing power and energy efficiency. Typically, multicore platforms use a shared memory bus that connects the cores to the off-chip main memory. This sharing of memory bus may cause tasks running on different cores to compete for access to the main memory whenever data/instructions are need to be read/written from/to the main memory. Such competition is problematic, as it may cause variations in the execution time of tasks in a non-deterministic way. To reduce the complexity of analyzing this problem, the 3-phase task model was proposed that divides tasks’ executions into distinct memory and execution phases. The distinctive memory phases are then scheduled to eliminate/minimize main memory contention between concurrently executing tasks. However, 3-phase tasks running on different cores may still compete to access the shared memory bus/main memory in order to execute memory phases. This paper presents a partitioned scheduling-based approach that allows one to derive memory bus contention-aware worst-case response time of tasks that follow the 3-phase task model. In particular, the bus-contention analysis is derived by considering two memory access models, i.e., (i) dedicated memory access model, where a core having allowed to access the main memory via memory bus is permitted to execute more than one memory phase, and (ii) fair memory access model, that restrict each core to execute only one memory phase in its allocated bus access. Both these models represent different system and application requirements, and the resulting bus contention of tasks may vary depending on the considered model. To evaluate the effectiveness of the proposed bus contention analysis, we compare its performance against an existing analysis in the state-of-the-art by performing (i) case-study experiments, using benchmarks from the Mälardalen Benchmark suite, and (ii) empirical evaluation using synthetic task sets. Results show that our proposed analysis can improve task set schedulability of 3-phase tasks by up to 88 percentage points.
Jatin Arora 0006, Cláudio Maia, Syed Aftab Rashid, Geoffrey Nelissen, Eduardo Tovar
J. Syst. Archit.3
2022 Tightening the CRPD bound for multilevel non-inclusive caches
abstract
Tasks running on microprocessors with cache memories are often subjected to cache related preemption delays (CRPDs). CRPDs may significantly increase task execution times, thereby, affecting their schedulability. Schedulability analysis accounting for the impact of CRPD has been extensively studied over the past two decades for systems with a single level of cache. Yet, the literature on CRPD for multilevel non-inclusive caches is relatively scarce. Two main challenges exist when analyzing multilevel caches: (1) characterization of the indirect effect of preemption, i.e., capturing the increase in cache interference at lower cache levels (e.g., level-two or L2 cache) due to the evictions of cache content from a higher cache level (e.g., level-one or L1 cache), and (2) upper bounding the maximum CRPD suffered by tasks at lower cache levels (e.g., L2 cache), i.e., determining the cache content of tasks that can be evicted from lower cache levels in case of preemptions. Existing analysis that focus on bounding CRPD for multilevel non-inclusive caches overestimate the values of (1) and (2) leading to pessimistic worst-case response time (WCRT) estimations. In this work, we reduce the excessive pessimism of the state-of-the-art CRPD analysis for multilevel non-inclusive caches by (i) introducing the notion of multi-level useful cache blocks, i.e., cache blocks that can cause CRPD at different cache levels, and use it to compute a tighter bound on the indirect effect of preemption of tasks; and (ii) deriving a new analysis to compute tighter bounds on the CRPD of tasks at lower cache levels (e.g., L2 cache). We performed a thorough experimental evaluation using benchmarks to compare the performance of our proposed CRPD analysis against the state-of-the-art CRPD analysis. Experimental results show that our proposed CRPD analysis dominates the existing analysis and improves task set schedulability by up to 20% percentage points.
Syed Aftab Rashid, Geoffrey Nelissen, Eduardo Tovar
J. Syst. Archit.1
2020 Cache Persistence-Aware Memory Bus Contention Analysis for Multicore Systems
abstract
Memory bus contention strongly relates to the number of main memory requests generated by tasks running on different cores of a multicore platform, which, in turn, depends on the content of the cache memories during the execution of those tasks. Recent works have shown that due to cache persistence the memory access demand of multiple jobs of a task may not always be equal to its worst-case memory access demand in isolation. Analysis of the variable memory access demand of tasks due to cache persistence leads to significantly tighter worst-case response time (WCRT) of tasks.In this work, we show how the notion of cache persistence can be extended from single-core to multicore systems. In particular, we focus on analyzing the impact of cache persistence on the memory bus contention suffered by tasks executing on a multi-core platform considering both work conserving and non-work conserving bus arbitration policies. Experimental evaluation shows that cache persistence-aware analyses of bus arbitration policies increase the number of task sets deemed schedulable by up to 70 percentage points in comparison to their respective counterparts that do not account for cache persistence.
Syed Aftab Rashid, Geoffrey Nelissen, Eduardo Tovar
DATE1
2020 Bounding Cache Persistence Reload Overheads for Set-Associative Caches
abstract
Cache memories have a strong impact on the response time of tasks executed on modern computing platforms. For tasks scheduled under fixed-priority preemptive scheduling (FPPS), the worst-case response time (WCRT) analyses that account for cache persistence between jobs along with cache related preemption delays (CRPDs) have been shown to dominate analyses that only consider CRPDs. Yet, the existing approaches that analyze cache persistence in the context of WCRT analysis can only support direct-mapped caches. In this work, we analyze cache persistence in the context of WCRT analysis for set-associative caches. The main contributions of this work are: (i) to propose a solution to find persistent cache blocks (PCBs) of tasks considering set-associative caches, (ii) to present three different approaches to calculate cache persistence reload overheads (CPROs), i.e., the memory overhead due to eviction of PCBs of tasks, under set-associative caches, and (iii) an experimental evaluation showing that our proposed approaches result in up to 22 percentage points higher task set schedulability than the state-of-the-art approaches.
Syed Aftab Rashid, Geoffrey Nelissen, Eduardo Tovar
RTCSA1
2020 Work-In-Progress: WCRT Analysis for the 3-Phase Task Model in Partitioned Scheduling
abstract
Multicore platforms are being increasingly adopted in Cyber-Physical Systems (CPS) due to their advantages over single-core processors, such as raw computing power and energy efficiency. Typically, multicore platforms use a shared system bus that connects the cores to the memory hierarchy (including caches and main memory). However, such hierarchy causes tasks running on different cores to compete for access to the shared system bus whenever data reads or writes need to be made. Such competition is problematic as it may cause large variations in the execution time of tasks in a non-deterministic way. This paper presents an analysis that allows one to derive bus contention-aware worst-case response-time of tasks that follow the 3-phase task model executing under partitioned scheduling.
Jatin Arora 0006, Cláudio Maia, Syed Aftab Rashid, Geoffrey Nelissen, Eduardo Tovar
RTSS3
2017 Integrated Analysis of Cache Related Preemption Delays and Cache Persistence Reload Overheads
abstract
Schedulability analysis for tasks running on micro- processors with cache memory is incomplete without a treatment of Cache Related Preemption Delays (CRPD) and Cache Persistence Reload Overheads (CPRO). State-of-the-art analyses compute CRPD and CPRO independently, which might result in counting the same overhead more than once. In this paper, we analyze the pessimism associated with the independent calculation of CRPD and CPRO in comparison to an integrated approach. We answer two main questions: (1) Is it benecial to integrate the calculation of CRPD and CPRO? (2) When and to what extent can we gain in terms of schedulability by integrating the calculation of CRPD and CPRO? To achieve this, we (i) identify situations where considering CRPD and CPRO separately might result in overestimating the total memory overhead suffered by tasks, (ii) derive new analyses that integrate the calculation of CRPD and CPRO; and (iii) perform a thorough experimental evaluation using benchmarks to compare the performance of the integrated analysis against the separate calculation of CRPD and CPRO.
Syed Aftab Rashid, Geoffrey Nelissen, Sebastian Altmeyer, Robert I. Davis 0001, Eduardo Tovar
RTSS1
2016 Cache-Persistence-Aware Response-Time Analysis for Fixed-Priority Preemptive Systems
abstract
A task can be preempted by several jobs of higherpriority tasks during its response time. Assuming the worst-casememory demand for each of these jobs leads to pessimistic worst-case response time (WCRT) estimations. Indeed, there is a bigchance that a large portion of the instructions and data associatedwith the preempting task Tj are still available in the cache when Tj releases its next jobs. Accounting for this observation allowsthe pessimism of WCRT analysis to be significantly reduced, which is not considered by existing work. The four main contributions of this paper are: 1) The conceptof persistent cache blocks is introduced in the context of WCRTanalysis, which allows re-use of cache blocks to be captured,2) A cache-persistence-aware WCRT analysis for fixed-prioritypreemptive systems exploiting the PCBs to reduce the WCRTbound, 3) A multi-set extension of the analysis that furtherimproves the WCRT bound and 4) An evaluation showing thatour cache-persistence-aware WCRT analysis results in up to 10%higher schedulability than state-of-the-art approaches.
Syed Aftab Rashid, Geoffrey Nelissen, Damien Hardy, Benny Akesson, Isabelle Puaut, Eduardo Tovar
ECRTS1
2016 Poster Abstract: Cache Persistence Aware Response Time Analysis for Fixed Priority Preemptive Systems
abstract
Summary form only given. The existing gap between the processor and main memory operating speeds necessitates the use of intermediate cache memories to accelerate the average case access time to instructions and data that must be executed or treated on the processor. However, the introduction of cache memories in modern computing platforms is the cause of big variations in the execution time of each instruction depending on whether the instruction and the data it treats are already loaded in the cache or not. During the worst-case response time (WCRT) analysis, the existing works assume that each job released by the preempting tasks will ask for their worst-case memory demand. This is however pessimistic since there is a high chance that a big portion of the instructions and data associated with the preempting task τj, are still available in the cache when τjreleases its next jobs. We call this content persistent cache blocks (PCBs). In this work, we propose a method to accurately bound the memory overhead incurred by a low priority task due to high priority tasks executing during its response time. For this purpose, we first identify the existence of persistent and nonpersistent cache blocks (i.e., PCBs and nPCBs) associated with each task. We then show with an example that due to the existence of PCBs, the memory demand of a task can significantly vary over time. Therefore, accounting for PCBs in the memory demand of the preempting task allows to reduce the pessimism on the total memory demand considered by the WCRT analysis. Finally, we propose a refined WCRT analysis for fixed priority preemptive systems considering (i) the effect of PCBs on the memory demand of the preempting task, and (ii) accounting for the number of PCBs that can be evicted by the preempted tasks between two successive job releases of the preempting tasks.
Syed Aftab Rashid, Geoffrey Nelissen, Eduardo Tovar
RTAS1
2016 Integrating the calculation of preemption and persistence related cache overhead
abstract
In this work, we highlight the pessimism of independently calculating cache-related preemption delays (CRPDs) and cache persistence reload overheads (CPROs). We propose a first solution to reduce that pessimism by integrating the calculation of CRPDs and CPROs. However, the proposed result is limited to the useful memory blocks (UCB)-union and CPRO-union approaches. Two methods that are known to be simple but pessimistic.
Syed Aftab Rashid, Geoffrey Nelissen, Eduardo Tovar
RTSS1