VLDB 2026 Research / reviewers in the wild / expert
Mohamed Hassan 0002
dblp:65/4298-2
· DBLP profile ↗
36ranked-venue papers
14as first author
19since 2021 · last 2026
0000-0001-5926-5861ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 11 first-author · 15 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InterStellar 2.0: Fine-grained stream-guided HW/SW co-design for multi-channel DRAM performance steeringabstractThe gap between processor speed and memory latency limits system scalability, especially in data-intensive and artificial intelligence workloads where memory-level parallelism and bandwidth efficiency are critical. Prior work ( InterStellar ) showed that hardware/software (HW/SW) co-design can expose program-level access streams to the memory system, enabling more informed memory-controller (MC) scheduling. This work extends that approach to high-bandwidth, multi-channel platforms. We present InterStellar 2.0 , a scalable HW/SW co-design that: (1) supports multi-channel dynamic random-access memory (DRAM) by partitioning stream batches across channels, allowing each channel to operate independently without cross-channel coordination; and (2) introduces fine-grained stream descriptors so software can distinguish distinct access patterns, even within the same data structure. These capabilities improve DRAM locality management and allow the MC to issue future requests efficiently. We evaluate InterStellar 2.0 on an 8-core RISC-V platform across DRAM configurations from 1 to 32 channels. At 32 channels, InterStellar 2.0 improves performance by up to 2.92 × and increases memory bandwidth by up to 2.83 × over a commercial off-the-shelf (COTS) controller. Fine-grained stream tracking alone improves performance by up to 1 . 35 × . Overall, InterStellar 2.0 shows that stream-aware HW/SW co-design is practical, compatible, and scalable for multi-channel memory systems without ISA changes or inter-channel communication. Abdelrhman Mohamed Abotaleb, Maziar Goudarzi, Tomasz S. Czajkowski, Mohamed Hassan 0002 |
J. Syst. Archit. | 5 |
| 2025 | Criticality and Requirement Aware Heterogeneous Coherence for Mixed Criticality SystemsabstractWe propose$\mathsf{CoHoRT}$, as the first heterogeneous cache coherent solution for mixed criticality systems (MCS) equipped with several features that targets the characteristics and requirements of such systems.$\mathsf{CoHoRT}$is requirement-aware. It provides an optimization engine to optimally configure the architecture based on system requirements.$\mathsf{CoHoRT}$is also criticality-aware. It introduces a low-cost novel architecture to enable cores to heterogeneously run different coherence protocols (time-based and MSI-based protocols). Moreover, it enables a run-time switch between these protocols to provide hardware support for mode operation switch, which is a common chal-lenge in MCS. Our evaluation shows that$\mathsf{CoHoRT}$outperforms existing solutions both from worst-case memory latency as well as overall average performance. It also illustrates that$\mathsf{CoHoRT}$is able to meet timing requirements in various MCS setups and showcases$\mathsf{CoHoRT}$'s ability to adapt to mode switches. Safin Bayes, Mohamed Hassan 0002 |
DATE | 2 |
| 2025 | FPGA-Based MPSoCs for High-Performance Sensor Fusion: Accelerating Covariance Intersection
Hazem M. Sharf, Mohamed Hassan 0002 |
FPL | 2 |
| 2025 | The Case for HW/SW Harmony in Real-Time Systems: Tightening Memory Latency of Streaming ApplicationsabstractModern critical cyber-physical systems such as autonomous vehicles, drones, and real-time medical monitoring, demand not only intensive data processing but also stringent adherence to real-time performance constraints. These applications often involve continuous or sequential data streams (e.g., images, videos, and sensor readings), which require frequent memory accesses. Despite advancements in processing power, huge variable interference delay is incurred within the Dynamic Random Access Memory (DRAM) accesses. However, achieving a tight bound of memory latency remains a significant challenge, yet it is essential for ensuring safe and predictable execution of these critical tasks. To address this bottleneck, we propose InterStellarRT , a novel hardware/software harmony methodology that provides data-aware optimizations across the entire memory hierarchy. Leveraging a software layer that communicates data access patterns to the memory controller, InterStellarRT achieves significant reductions in memory access times, ensuring tightly bounded and predictable times. We perform the theoretical analysis of the memory latency bound. Then, we prove that InterStellarRT provides remarkable tighter memory latency bound for in-isolation and interference latencies compared to the state-of-the-art real-time systems based on the Commercial-Off-The-Shelf (COTS) Double Data Rate 4 (DDR4) memory devices and is also applicable to DDR5. We evaluate InterStellarRT on a RISC-V-based quad-core system on GEM5 and DDR4 in Ramulator. Analyzing benchmark results from Polybench, LAPACK, Phoenix, and HPCG Suites, InterStellarRT achieves a 3.8× tighter average bound for in-isolation memory latency and 13.5× for interference latency under affine workloads, while for mixed-affinity workloads, the bounds are 2.15× and 4×, respectively. Moreover, InterStellarRT achieves average 1.72× end-to-end speedup, and 1.9× bandwidth improvement, and 14% DRAM energy reduction against the baseline. Abdelrhman Mohamed Abotaleb, Mohamed Hassan 0002 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2025 | OpenDRAM: A Modular, High-performance Soft Memory Controller for DDR4 DRAMabstractWe propose OpenDRAM , a synthesizable high-performance DDR4 DRAM soft Memory Controller (MC) for FPGAs. Since DRAMs usually operate at a higher frequency compared to MCs (usually \(4\times\) ), to fully utilize DRAM’s bandwidth, the hardened DDR4 physical interface expects the controller to issue four DRAM commands in a single clock cycle. OpenDRAM is a modular, extensible MC, implementing high-performance bank-parallel schedulers. We detail the design of OpenDRAM ’s logic blocks in RTL and their integration with existing AMD’s Memory Interface Generator (MIG) modules for initialization, maintenance, and interfacing. The integrated project was comprehensively validated on an AMD Virtex UltraScale+ FPGA. We evaluate and compare the performance of OpenDRAM with AMD’s MIG controller and another open source controller, OPRECOMP, using synthetic and accelerator kernels. Results show that OpenDRAM surpasses both commercial and open source counterparts, offering performance improvements of up to 157% over AMD’s MIG and 267% over OPRECOMP, primarily owing to its reordering and scheduling mechanisms. To demonstrate its research use case, we prototype five distinct command schedulers, exploring tradeoffs between scheduling aggressiveness and maximum frequency, and show how FPGA-aware design can enhance timing closure. Finally, we release OpenDRAM as the first high-performance, extensible, open source MC for researchers to utilize, extend, and build upon. Danesh Germchi, Amin Katani, Mohamed Hassan 0002, Rodolfo Pellizzoni |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2024 | Shared Data Kills Real-Time Cache Analysis. How to Resurrect It?abstractWhile data sharing is becoming a necessity in modern multi-core real-time systems, it complicates system analyzability and leads to significantly pessimistic latency bounds. This work is a step towards facilitating high-performance and coherent data sharing in real-time systems by tackling two main problems. The first is a well-acknowledged one: shared caches render cache analysis techniques useless and all cache accesses have to be assumed misses. The second is a new one, where we show that coherence interference voids classical cache analysis techniques. We contribute a solution that tackles both problems by leveraging time-based cache coherence and a novel methodology to integrate its effect into cache analysis. Thanks to this solution, we enable the usage of shared memory hierarchy with coherent shared data, while we prove that we are able to restore cache analysis; and hence, provide much tighter memory latency bounds. Safin Bayes, Mohamed Hossam, Mohamed Hassan 0002 |
DATE | 3 |
| 2024 | Event Monitor Validation in High-Integrity SystemsabstractPlatforms for modern embedded systems equip an increasing number of high-performance features to provide the required levels of performance. Timing analysis solutions handle the complexity of these platforms by relying on hardware event monitors (HEMs) that provide insightful information about resource utilization and, hence, contention among tasks. As a result, HEMs have become a key element to warrant a safe timing behavior of a system, for which reason they must be validated. While some initial works target HEMs validation, they consider one HEM at a time and focus on those HEMs for which an expert can establish an expected value for relatively small code snippets. In this paper, we propose a methodology for the validation of those HEMs for which a specific expected value cannot be established a priori even for simple cases and, instead, needs to be validated in conjunction with other HEMs. Our method also deals with the natural variability of the HEMs' values in high-performance platforms when collected in different experiments. We illustrate the effectiveness of our proposed technique for validating HEMs related to cache coherence in a relevant platform in the avionics domain. Roger Pujol, Sergi Vilardell, Enrico Mezzetti, Mohamed Hassan 0002, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2024 | HW/SW Collaborative Techniques for Accelerating TinyML Inference Time at No CostabstractWith the unprecedented boom in TinyML development, optimizing Artificial Intelligence (AI) inference on resource-constrained microcontrollers (M CU s) is of paramount importance. Most of the existing works focus on peak memory or computation reduction. The tasks are partitioned in the patch-based or device-based during the execution. However, it comes with a price of the latency and communication overhead. In this paper, we propose several techniques to accelerate the Convolutional Neural Networks (CNN s) inference process. These techniques are both architecture- and application-aware. From the application perspective, 1) we maximize computation reuse through instruction reordering, 2) fuse several linear layers together to improve computation patterns, and 3) enable memory reuse of intermediate buffers for improving memory behavior. From the architecture perspective, we propose techniques that take into account knowledge about underlying architecture of the MCU including 1) cache-aware and 2) multi-core parallelism-aware techniques. Those solutions only require the general MCUs features thus demonstrating board generalization across various networks and devices. These techniques come at no additional cost. It improve the inference latency without any compromise of the model accuracy or the model size. Our evaluation on a use-case from the health-care domain with real-data set for four CNNs - LeNet, AlexNet, ResNet20, and SqueezeNet - show that we achieve up to 71 % reduction in inference latency. Bailian Sun, Mohamed Hassan 0002 |
DSD | 2 |
| 2024 | A Framework for Explainable, Comprehensive, and Customizable Memory-Centric WorkloadsabstractExplainable workloads with analyzable memory traffic patterns are key for accurate performance estimates at early design exploration phase for novel memory solutions. This paper proposes RAMify: a tunable framework for generating explainable memory-centric workloads. By being memory-aware: RAMify offers several tuning knobs enabling the generation of an extensive set of different workloads, each of them is low-level tuned to produce a particular DRAM access pattern. RAMify enables a systematic way to explore and evaluate novel memory subsystem proposals at early design phases, validate their performance, stress their behaviour, and qualitatively compare them against other policies under various memory-aware scenarios to facilitate data-driven design choices. We evaluated with extensive experiments across three different cycle-accurate memory simulators and a full-system multi-core simulator. Results show that using RAMify, we were able to 1) make interesting observations about the comparative behavior of two of the state-of-the-art memory technologies (DDR4 and HBM) that were not possible to make in non memory-centric benchmarks, and 2) We managed to reveal discrepancies in state-of-the-art memory simulator policies and scheduling techniques. Mohamed Abuelala, Mohamed Hassan 0002 |
ICCAD | 2 |
| 2023 | A Tight Holistic Memory Latency Bound Through Coordinated Management of Memory Resources
Shorouk Abdelhalim, Danesh Germchi, Mohamed Hossam, Rodolfo Pellizzoni, Mohamed Hassan 0002 |
ECRTS | 5 |
| 2023 | Improving Timing-Related Guarantees for Main Memory in Multicore Critical Embedded SystemsabstractMain memory is one of the most complex resources to analyze in multicore-based embedded real-time systems, with contention in the memory controller and the timing constraints of the main memory device as the main contributors to that complexity. One of the main challenges in multicore real-time systems is producing the required evidence on the management of contention delay for the certification. This stems from the fact that current MPSoCs barely provide any event monitors on how tasks interact and delay each other in memory. Besides, even if hardware and software mechanisms are in place to mitigate contention in the memory system, it is hard - if at all possible - to provide evidence about their correctness. In this work, we cover this gap by proposing a lightweight hardware mechanism that tightly tracks inter-core contention in memory. The proposed hardware mechanism, which we evaluate in detail, improves the quality of timing-related evidence that must be provided on how contention in main memory of multicore real-time systems is handled in adherence to applicable safety standards. Asier Fernández de Lecea, Mohamed Hassan 0002, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 2 |
| 2023 | DISCO: Time-Compositional Cache Coherence for Multi-Core Real-Time Embedded SystemsabstractTasks in modern embedded systems share data and communicate among each other. Nonetheless, the majority of research in real-time systems either assumes that tasks do not share data or prohibits data sharing by design. Only recently, some works investigated solutions to address this limitation and enable data sharing. However, we find these works to suffer from severe limitations. In particular, proposed predictable cache coherence protocols increase the worst-case memory latency (WCL) quadratically due to coherence interference and breaks compositionality by coupling the design and timing analysis of the coherence with the underlying bus arbitration policy. In this paper, we argue that a protocol that distinguishes between non-modifying (read) and modifying (write) memory accesses is key towards reducing the effects of coherence interference on WCL. Accordingly, we propose DISCO, a discriminative coherence solution that capitalizes on this observation to 1) balance average-case performance and WCL, and 2) more importantly, achieves compositionality in the existing of coherence by enabling the decomposition of the effects from coherence and arbitration components. DISCO achieves 7.2× lower latency bounds compared to the state-of-the-art predictable coherence protocol. DISCO also achieves up to 11.4× (5.3× on average) better performance than private cache bypassing for the SPLASH-3 benchmarks. Mohamed Hassan 0002 |
IEEE Trans. Computers | 1 |
| 2023 | PISCOT: A Pipelined Split-Transaction COTS-Coherent Bus for Multi-Core Real-Time SystemsabstractTasks in modern embedded systems such as automotive and avionics communicate among each other using shared data towards achieving the desired functionality of the whole system. In commodity platforms, cores communicate data through the shared memory hierarchy and correctness is maintained by a cache coherence protocol. Recent works investigated the deployment of coherence protocols in real-time systems and showed significant performance improvements. Nonetheless, we find these works to require modifications to commodity coherence protocols, assume simple in-order pipelines, and most importantly suffer from significant latency delays due to coherence interference along with average performance degradation. In this work, we propose PISCOT : a predictable and coherent bus architecture that (i) provides a considerably tighter bound compared to the state-of-the-art predictable coherent solutions (4× tighter bounds in a quad-core system). (ii) It does so with a negligible performance loss compared to conventional high-performance architecture coherence delays (less than 4% for SPLASH-3 benchmarks). This improves average performance by up to 5× (2.8× on average) compared to its predictable coherence counterpart. Finally, (iii) it achieves that without requiring any modifications to conventional coherence protocols. We show this by integrating PISCOT on top of two protocols with a detailed implementation with complete transient states: MSI and MESI. Salah Hessien, Mohamed Hassan 0002 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | Predictably and Efficiently Integrating COTS Cache Coherence in Real-Time SystemsabstractThe adoption of multi-core platforms in embedded real-time systems mandates predictable system components. Such components must guarantee the satisfaction of the timing constraints of various applications running on the system. One of the components that can break the system predictability is cache coherence, which ensures the correctness of shared data. This paper proposes a solution towards the enablement of predictable cache coherent real-time systems. The solution uses existing COTS coherence protocols and proposes a methodology to integrate them with legacy real-time arbiters without imposing any required modification to either of them. Doing so, the paper also works as an exploratory study of the integration of various coherence protocols with various predictable arbitration schemes leading to a total of 12 different architecture configurations. Evaluation against four state-of-the-art predictable coherence solutions as well as COTS-based solutions show that the proposed approach achieves the tightest existing latency bounds among predictable solutions with minimal performance degradation over the COTS ones. Mohamed Hossam, Mohamed Hassan 0002 |
ECRTS | 2 |
| 2022 | Parallelism-Aware High-Performance Cache Coherence with Tight Latency Bounds
Reza Mirosanlou, Mohamed Hassan 0002, Rodolfo Pellizzoni |
ECRTS | 2 |
| 2021 | Duetto: Latency Guarantees at Minimal Performance CostabstractThe management of shared hardware resources in multi-core platforms has been characterized by a fundamental trade-off: high-performance arbiters typically employed in COTS systems offer no worst-case guarantees, while dedicated real-time controllers provide timing guarantees at the cost of significantly degrading system performance. In this paper, we overcome this trade-off by introducing Duetto, a novel hardware resource management paradigm. Duetto pairs a real-time arbiter with a high-performance arbiter and a latency estimator module. Based on the observation that the resource is rarely overloaded, Duetto executes the high-performance arbiter most of the time, switching to the real-time arbiter only in the rare cases when the latency estimator deems that timing guarantees risk being violated. We demonstrate our approach on the case study of a multi-bank memory. Our evaluation based on cycle-accurate simulations shows that Duetto can provide the same latency guarantees as the real-time arbiter with limited loss of performance compared to the high-performance arbiter. Reza Mirosanlou, Mohamed Hassan 0002, Rodolfo Pellizzoni |
DATE | 2 |
| 2021 | Empirical Evidence for MPSoCs in Critical Systems: The Case of NXP's T2080 Cache CoherenceabstractThe adoption of complex MPSoCs in critical realtime embedded systems mandates a detailed analysis of their architecture to facilitate certification. This analysis is hindered by the lack of a thorough understanding of the MPSoC system due to the unobvious and/or insufficiently documented behavior of some key hardware features. Confidence in those features can only be regained by building specific tests to both, assess whether their behavior matches specifications and unveil their behavior when it is not fully known a priori. In this line, in this work we develop a thorough understanding of the cache coherence protocol in the avionics-relevant NXP T2080 architecture. Roger Pujol, Hamid Tabani, Jaume Abella 0001, Mohamed Hassan 0002, Francisco J. Cazorla |
DATE | 4 |
| 2021 | Demystifying the Characteristics of High Bandwidth Memory for Real-Time SystemsabstractThe number of functionalities controlled by software on every critical real-time product is on the rise in domains like automotive, avionics and space. To implement these advanced functionalities, software applications increasingly adopt artificial intelligence algorithms that manage massive amounts of data transmitted from various sensors. This translates into unprecedented memory performance requirements in critical systems that the commonly used DRAM memories struggle to provide. High-Bandwidth Memory (HBM) can satisfy these requirements offering high bandwidth, low power and high-integration capacity features. However, it remains unclear whether the predictability and isolation properties of HBM are compatible with the requirements of critical embedded systems. In this work, we perform to our knowledge the first timing analysis of HBM. We show the unique structural and timing characteristics of HBM with respect to DRAM memories and how they can be exploited for better time predictability, with emphasis on increased isolation among tasks and reduced worst-case memory latency. Kazi Asifuzzaman, Mohamed Abuelala, Mohamed Hassan 0002, Francisco J. Cazorla |
ICCAD | 3 |
| 2021 | Designing Predictable Cache Coherence Protocols for Multi-Core Real-Time SystemsabstractThis article addresses the challenge of allowing simultaneous and predictable accesses to shared data on multi-core systems. We propose a collection of predictable cache coherence protocols, which mandate the use of certain design invariants to ensure predictability. In particular, we enforce these invariants by augmenting the classic modify-share-invalid (MSI) protocol and modify-exclusive-share-invalid (MESI) protocol. This allows us to derive worst-case latency bounds on the resulting predictable MSI (PMSI) and predictable MESI (PMESI) protocols. Our analysis shows that while the arbitration latency scales linearly, the coherence latency scales quadratically with the number of cores, which emphasizes the importance of accounting for cache coherence effects on latency bounds. We implement PMSI and PMESI in a detailed micro-architectural simulator, and execute SPLASH-2 and synthetic workloads. Results show that our approach is always within the analytical worst-case latency bounds, and that PMSI and PMESI improve average-case performance by up to 4× over cache bypassing mechanisms that disallow caching of shared data in the cores’ private caches. PMSI and PMESI have average slowdowns of 1.45× and 1.46× compared to conventional MSI and MESI protocols, respectively. Anirudh M. Kaushik, Mohamed Hassan 0002, Hiren D. Patel |
IEEE Trans. Computers | 2 |
| 2020 | Discriminative Coherence: Balancing Performance and Latency Bounds in Data-Sharing Multi-Core Real-Time Systems
Mohamed Hassan 0002 |
ECRTS | 1 |
| 2020 | Analysis of Memory-Contention in Heterogeneous COTS MPSoCsabstractMultiple-Processors Systems-on-Chip (MPSoCs) provide an appealing platform to execute Mixed Criticality Systems (MCS) with both time-sensitive critical tasks and performance-oriented non-critical tasks. Their heterogeneity with a variety of processing elements can address the conflicting requirements of those tasks. Nonetheless, the complex (and hence hard-to-analyze) architecture of Commercial-Off-The-Shelf (COTS) MPSoCs presents a challenge encumbering their adoption for MCS. In this paper, we propose a framework to analyze the memory contention in COTS MPSoCs and provide safe and tight bounds to the delays suffered by any critical task due to this contention. Unlike existing analyses, our solution is based on two main novel approaches. 1) It conducts a hybrid analysis that blends both request-level and task-level analyses into the same framework. 2) It leverages available knowledge about the types of memory requests of the task under analysis as well as contending tasks; specifically, we consider information that is already obtainable by applying existing static analysis tools to each task in isolation. Thanks to these novel techniques, our comparisons with the state-of-the art approaches show that the proposed analysis provides the tightest bounds across all evaluated access scenarios. Mohamed Hassan 0002, Rodolfo Pellizzoni |
ECRTS | 1 |
| 2020 | DRAMbulism: Balancing Performance and Predictability through Dynamic PipeliningabstractWorst-case execution bounds for real-time programs are profoundly impacted by the latency of accessing hardware shared resources, such as off-chip DRAM. While many different memory controller designs have been proposed in the literature, there is a trade-off between average-case performance and predictable worst-case bounds, as techniques targeted at improving the former can harm the latter and vice-versa. We find that taking advantage of pipelining between different commands can improve both, but incorporating pipelining effects in worst-case analysis is challenging. In this work, we introduce a novel DRAM controller that successfully balances performance and predictability by employing a dynamic pipelining scheme. We show that the schedule of DRAM commands is akin to a two-stage two-mode pipeline, and hence, design an easily-implementable admission rule that allows us to dynamically add requests to the pipeline without hurting worst-case bounds. Reza Mirosanlou, Mohamed Hassan 0002, Rodolfo Pellizzoni |
RTAS | 2 |
| 2020 | The Best of All Worlds: Improving Predictability at the Performance of Conventional Coherence with No Protocol ModificationsabstractTasks in modern embedded systems such as automotive and avionics communicate among each other using shared data towards achieving the desired functionality of the whole system. In commodity platforms, cores communicate data through the shared memory hierarchy and correctness is maintained by a cache coherence protocol. Recent works investigated the deployment of coherence protocols in real-time systems and showed significant performance improvements. Nonetheless, we find these works to suffer from two main drawbacks. 1) They suffer from significant latency delays due to coherence interference. 2) They require amendments to existing coherence protocols. This represents a significant obstruction hindering the industry adoption of these proposals since it requires to re-verify the coherence protocol. Coherence verification is considered one of the most complex challenges in computer architecture, which makes it inconceivable for chip manufacturers to adopt modifications to their already verified protocols that they have stable for decades.In this work, we propose PISCOT: a predictable and coherent bus architecture that (i) provides a considerably tighter bound compared to the state-of-the-art predictable coherent solutions (4× tighter bounds in a quad-core system). (ii) It does so with a negligible performance loss compared to conventional high-performance architecture coherence delays (less than 4% for SPLASH-3 benchmarks). This improves average performance by up to 5× (2.8× on average) compared to its predictable coherence counterpart. Finally, (iii) it achieves that without requiring any modifications to conventional coherence protocols. Salah Hessien, Mohamed Hassan 0002 |
RTSS | 2 |
| 2020 | Reduced latency DRAM for multi-core safety-critical real-time systems
Mohamed Hassan 0002 |
Real Time Syst. | 1 |
| 2019 | Enabling Predictable, Simultaneous and Coherent Data Sharing in Mixed Criticality SystemsabstractEmerging embedded systems deployed in the automotive and avionics domains execute applications with different criticalities, comprising what is known as Mixed Criticality Systems (MCS). Applications in MCS often share data between tasks (coming from sensors for instance). Data sharing is challenging because it can lead to increased response times or even unpredictable behaviors if not carefully addressed. Therefore, several prior works in MCS either assumed that tasks do not share data or disallowed it by design. Recent solutions attempt to mitigate the effects of data sharing, albeit by introducing new restrictions on the system either by prohibiting applications from caching shared data or prohibiting the operating system from running tasks with shared data in parallel. We find these solutions also to have limited applicability as they deteriorate system schedulability and prohibit simultaneous access to shared data. In this paper, we propose PENDULUM: a time-based cache coherence protocol to enable simultaneous and predictable access to shared data in MCS. Our evaluation shows that PENDULUM, achieves flexibility and better performance compared to existing solutions, while maintaining system predictability. Nivedita Sritharan, Anirudh M. Kaushik, Mohamed Hassan 0002, Hiren D. Patel |
RTSS | 3 |
| 2018 | On the Off-Chip Memory Latency of Real-Time Systems: Is DDR DRAM Really the Best Option?abstractPredictable execution time upon accessing shared memories in multi-core real-time systems is a stringent requirement. A plethora of existing works focus on the analysis of Double Data Rate Dynamic Random Access Memories (DDR DRAMs), or redesigning its memory to provide predictable memory behavior. In this paper, we show that DDR DRAMs by construction suffer inherent limitations associated with achieving such predictability. These limitations lead to 1) highly variable access latencies that fluctuate based on various factors such as access patterns and memory state from previous accesses, and 2) overly pessimistic latency bounds. As a result, DDR DRAMs can be ill-suited for some real-time systems that mandate a strict predictable performance with tight timing constraints. Targeting these systems, we promote an alternative off-chip memory solution that is based on the emerging Reduced Latency DRAM (RLDRAM) protocol, and propose a predictable memory controller (RLDC) managing accesses to this memory. Comparing with the state-of-the-art predictable DDR controllers, the proposed solution provides up to 11× less timing variability and 6.4× reduction in the worst case memory latency. Mohamed Hassan 0002 |
RTSS | 1 |
| 2018 | MCXplore: Automating the Validation Process of DRAM Memory Controller DesignsabstractWe present an automated framework for the validation of memory controllers (MCs) called MCXplore. In developing this framework, we construct formal models for memory requests and command interactions. MCXplore enables validation engineers to define their test plans precisely using temporal logic specifications. We use the NuSMV model-checker to generate counterexamples that serve as test templates. MCXplore uses these test templates to generate memory tests to validate the correctness properties of the MC. We show the effectiveness of MCXplore by validating various state-of-the-art MC features as well as hard-to-detect timing violations. We also provide a set of predefined test plans, and regression test suites that validate essential properties of modern MCs. MCXplore is an open-source framework to allow validation engineers and researchers to extend and use. Mohamed Hassan 0002, Hiren D. Patel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Bounding DRAM Interference in COTS Heterogeneous MPSoCs for Mixed Criticality SystemsabstractCommercial off-the-shelf (COTS) heterogeneous multiple processors systems-on-chip (MPSoCs) are appealing platforms for emerging mixed criticality systems (MCSs). To satisfy MCS requirements, the platform must guarantee predictable timing bounds for critical applications, without degrading average performance for noncritical applications. In particular, this paper studies the main memory subsystem, which in modern MPSoCs is typically based on double data rate synchronous dynamic access memory. While there exists previous work on worst-case DRAM latency analysis, such work only covers a small subset of possible COTS configurations, which are not targeted at MCS. Therefore, we derive a generalized interference delay analysis for DRAM main memory that accounts for a breadth of features deployed in COTS platforms. We then explore the design space by studying the effects of each feature on both the worst-case delay for critical applications, and the bandwidth for noncritical applications. Mohamed Hassan 0002, Rodolfo Pellizzoni |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | A Comparative Study of Predictable DRAM ControllersabstractRecently, the research community has introduced several predictable dynamic random-access memory (DRAM) controller designs that provide improved worst-case timing guarantees for real-time embedded systems. The proposed controllers significantly differ in terms of arbitration, configuration, and simulation environment, making it difficult to assess the contribution of each approach. To bridge this gap, this article provides the first comprehensive evaluation of state-of-the-art predictable DRAM controllers. We propose a categorization of available controllers, and introduce an analytical performance model based on worst-case latency. We then conduct an extensive evaluation for all state-of-the-art controllers based on a common simulation platform, and discuss findings and recommendations. Danlu Guo, Mohamed Hassan 0002, Rodolfo Pellizzoni, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Exposing Implementation Details of Embedded DRAM Memory Controllers through Latency-based AnalysisabstractWe explore techniques to reverse-engineer DRAM embedded memory controllers (MCs), including page policies, address mapping, and command arbitration. There are several benefits to knowing this information: They allow tightening worst-case bounds of embedded systems and platform-aware optimizations at the operating system, source-code, and compiler levels. We develop a latency-based analysis, which we use to devise algorithms and C programs to extract MC properties. We show the effectiveness of the proposed approach by reverse-engineering the MC details in the XUPV5-LX110T Xilinx platform. Furthermore, to cover a breadth of policies, we use a simulation framework and document our findings. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2017 | Predictable Cache Coherence for Multi-core Real-Time SystemsabstractThis work addresses the challenge of allowing simultaneous and predictable accesses to shared data on multicore systems. We propose a predictable cache coherence protocol, which mandates the use of certain invariants to ensure predictability. In particular, we enforce these invariants by augmenting the classic modify-share-invalid (MSI) protocol with transient coherence states, and minimal architectural changes. This allows us to derive worst-case latency bounds on predictable MSI (PMSI) protocol. Our analysis shows that while the arbitration latency scales linearly, the coherence latency scales quadratically with the number of cores, which emphasizes that importance of accounting for cache coherence effects on latency bounds. We implement PMSI in gem5, and execute SPLASH-2 and synthetic workloads. Results show that our approach is always within the analytical worst-case latency bounds, and that PMSI improves averagecase performance by up to 4 over the next best predictable alternative. PMSI has average slowdowns of 1.45 and 1.46 compared to MSI and MESI protocols, respectively. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 1 |
| 2017 | PMC: A Requirement-Aware DRAM Controller for Multicore Mixed Criticality SystemsabstractWe propose a novel approach to schedule memory requests in Mixed Criticality Systems (MCS). This approach supports an arbitrary number of criticality levels by enabling the MCS designer to specify memory requirements per task. It retains locality within large-size requests to satisfy memory requirements of all tasks. To achieve this target, we introduce a compact time-division-multiplexing scheduler, and a framework that constructs optimal schedules to manage requests to off-chip memory. We also present a static analysis that guarantees meeting requirements of all tasks. We compare the proposed controller against state-of-the-art memory controllers using both a case study and synthetic experiments. Mohamed Hassan 0002, Hiren D. Patel, Rodolfo Pellizzoni |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | MCXplore: An automated framework for validating memory controller designs
Mohamed Hassan 0002, Hiren D. Patel |
DATE | 1 |
| 2016 | Criticality- and Requirement-Aware Bus Arbitration for Multi-Core Mixed Criticality SystemsabstractThis work presents CArb, an arbiter for controlling accesses to the shared memory bus in multi-core mixed criticality systems. CArb is a requirement-aware arbiter that optimally allocates service to tasks based on their requirements. It is also criticality-aware since it incorporates criticality as a first-class principle in arbitration decisions. CArb supports any number of criticality levels and does not impose any restrictions on mapping tasks to processors. Hence, it operates in tandem with existing processor scheduling policies. In addition, CArb is able to dynamically adapt memory bus arbitration at run time to respond to increases in the monitored execution times of tasks. Utilizing this adaptation, CArb is able to offset these increases; hence, postpones the system need to switch to a degraded mode. We prototype CArb, and evaluate it with an avionics case-study from Honeywell as well as synthetic experiments. Mohamed Hassan 0002, Hiren D. Patel |
RTAS | 1 |
| 2015 | Reverse-engineering embedded memory controllers through latency-based analysisabstractWe explore techniques to reverse-engineer properties of DRAM memory controllers (MCs). This includes page policies, address mapping schemes and command arbitration schemes. There are several benefits to knowing this information: they allow analysis techniques to effectively compute worst-case bounds, and they allow customizations to be made in software for predictability. We develop a latency-based analysis, and use this analysis to devise algorithms for micro-benchmarks to extract properties of MCs. In order to cover a breadth of page policies, address mappings and command arbitration schemes, we explore our technique using a micro-architecture simulation framework and document our findings. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 1 |
| 2015 | A framework for scheduling DRAM memory accesses for multi-core mixed-time critical systemsabstractMixed-time critical systems are real-time systems that accommodate both hard real-time (HRT) and soft realtime (SRT) tasks. HRT tasks mandate a gurantee on the worstcase latency, while SRT tasks have average-case bandwidth (BW) demands. Memory requests in mixed-time critical systems usually have different transaction sizes based on whether the issuer task is HRT or SRT. For example, HRT tasks often issue requests with a cache line size. On the other side, SRT tasks may issue requests with a size of KBs. Requests from multimedia cores, cores controlling network interfaces and direct memory accesses (DMAs) are obvious examples of these large-size requests. Based on these observations, we promote in this work a new approach to schedule memory requests. This approach retains locality within large-size requests to minimize the worst-case latency, while maintaining the average-case BW as high as required. To achieve this target, we introduce a novel and compact time-division-multiplexing scheduler that is adequate for mixed-time critical systems. We also present a novel framework that constructs optimal offchip DRAM memory controller schedules for multi-core mixedtime critical systems. These schedules are loaded to the memory controller during boot-time. Based on the proposed schedule, we provide a detailed static analysis that guarantees predictability. We compare the proposed controller against state-of-the-art realtime memory controllers using synthetic experiments as well as a practical use-case from multimedia systems. Mohamed Hassan 0002, Hiren D. Patel, Rodolfo Pellizzoni |
RTAS | 1 |