VLDB 2026 Research / reviewers in the wild / expert
Hiren D. Patel
dblp:65/5699
· DBLP profile ↗
78ranked-venue papers
10as first author
24since 2021 · last 2025
0000-0003-2750-4471ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 64 · 8 first-author · 23 since 2021Software engineering, systems software and programming languages · 19 · 4 first-author · 3 since 2021Theory of computation · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
Ruimin Shi, Gabin Schieffer, Maya B. Gokhale, Pei-Hung Lin, Hiren D. Patel, Ivy Bo Peng |
Euro-Par (2) | 5 |
| 2025 | Consistency-Aware and Predictable Memory Processing for Safety-Critical Out-of-Order MulticoresabstractWe introduce an approach that facilitates predictable processing of multiple outstanding memory requests in safety-critical out-of-order multicores. A primary challenge addressed by this work is ensuring that multiple outstanding memory requests maintain a memory consistent model while minimizing a low worst-case latency. Adhering to a memory consistency model is crucial for ensuring the correctness of programs executed on such multicores. Our approach, termed predictable processing of multiple outstanding requests$(\mathsf{PPP}$), Teverages micro-architectural enhancements to tackle this challenge. Experimental results show that$\mathsf{PPP}$delivers speedups of$2.07 \times, 2.79 \times$, and$3.38 \times$over serialization for 2,4, and 8 cores, while maintaining the worst-case latency. Zhuanhao Wu, Hiren D. Patel |
RTAS | 2 |
| 2025 | Optimal Split Point Placement for Predictable GPU Wavefront SplittingabstractPredictable wavefront splitting (PWS) is an optimization technique for graphics processing units (GPUs) to address the performance and worst-case execution time (WCET) impacts of branch divergence. PWS relies on manual annotation by the GPU programmer; these choices affect the resulting WCET. This work automates this process with two key approaches. First, we formulate the optimal annotation as an integer quadratic programming (IQP) problem such that the solution guarantees the lowest WCET. Second, we show that the problem can be solved with an optimal polynomial-time dynamic programming algorithm that achieves the same solutions as the IQP. We implement our algorithm in a compiler flow for an AMD GPU, and we deploy the annotated executable on a gem5 micro-architectural implementation of the AMD GCN3 GPU. We evaluate our implementation on a benchmark suite provided by AMD and supplement it with an extensive set of synthetic benchmarks. Our evaluation shows that these two approaches are able to reduce the WCET by between 13% and 31% compared to five baseline algorithms. Artem Klashtorny, Mahesh Tripunitara, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | A Compiler Phase to Optimally Split GPU Wavefronts for Safety-Critical SystemsabstractWe present a compiler phase for GPUs enabled with predictable wavefront splitting (PWSG) that implements an optimal algorithm to split diverging GPU wavefronts into separate scheduleable entities. This algorithm selects branches in the GPU kernel that guarantee the lowest worst-case execution time (WCET) for the kernel. We implement our algorithm in a compiler flow for an AMD GPU, and we deploy the resulting binary on a gem5 micro-architectural implementation of the AMD GCN3 GPU. We evaluate our implementation on an extensive set of synthetic benchmarks. Our experiments show that by automatically selecting points in the GPU kernel to split wavefronts, we are able to reduce the WCET ranging from 34% to 52% reduction compared to five alternative approaches. Artem Klashtorny, Mahesh Tripunitara, Hiren D. Patel |
DATE | 3 |
| 2024 | Exclusive Hierarchies for Predictable Sharing in Last-Level CacheabstractThis work presents an approach to use a last-level cache (LLC) in a memory hierarchy for cache-coherent real-time multicores that delivers a low worst-case latency (WCL) and higher performance than all of its counterparts. Our approach relies on the key observation that an exclusive memory hierarchy, by definition, eliminates back invalidations, which are one of the largest contributors to the WCL when using inclusive memory hierarchies. However, to the best of our knowledge, there are no prior efforts that ensure the predictability of exclusive hierarchies for cache-coherent multicores. Consequently, in this work, we propose PECC, a predictable exclusive cache coherence mechanism, that achieves a lower average data access latency while providing a low WCL bound that scales linearly in the number of cores. Our evaluation shows that PECC reduces the bound by 6% and improves the average performance by 2.33× over the predictable solution with an inclusive LLC. Zhuanhao Wu, Rodolfo Pellizzoni, Hiren D. Patel |
RTAS | 4 |
| 2024 | High Performance and Predictable Shared Last-level Cache for Safety-Critical SystemsabstractWe propose ZeroCost-LLC (ZCLLC), a novel shared inclusive last-level cache (LLC) design for timing predictable multi-core platforms that offers lower worst-case latency (WCL) when compared with a traditional shared inclusive LLC design. ZCLLC achieves low WCL by eliminating certain memory operations in the form of cache line invalidations across the cache hierarchy that are a consequence of a core’s memory request that misses in the cache hierarchy and when there is no vacant entry in the LLC to accommodate the fetched data for this request. In addition to low WCL, ZCLLC offers performance benefits in the form of additional caching capacity and unlike state-of-the-art approaches, ZCLLC does not impose any constraints on its usage across multiple cores. In this work, we describe the impact of LLC cache line invalidations on the WCL and systematically build solutions to eliminate these invalidations resulting in ZCLLC. We also present ZCLLC-OPT, an optimized variant of ZCLLC that offers lower WCL and improved average-case performance over ZCLLC. We apply optimizations to the shared bus arbitration mechanism and extend the micro-architecture of ZCLLC to allow for overlapping memory requests to the main memory. Our analysis reveals that the analytical WCL of a memory request under ZCLLC-OPT is 87.0%, 93.8%, and 97.1% lower than that under state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively. ZCLLC-OPT shows average-case performance speedups of 1.89×, 3.36×, and 6.24× compared with the state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively. When compared with the original ZCLLC that does not have any optimizations, ZCLLC-OPT shows lower analytical WCLs that are 76.5%, 82.6%, and 86.2% lower compared with ZCLLC-NORMAL for 2, 4, and 8 cores, respectively. Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Enabling Kubernetes Orchestration of Mixed-Criticality Software for Autonomous Mobile RobotsabstractContainerization and orchestration have become two key requirements in software development best practices. Containerization allows for better resource utilization, platform-independent development, and secure deployment of software. Orchestration automates the deployment, networking, scaling, and availability of containerized workloads and services. While containerization is increasingly being adopted in the robotic community, the use of task orchestration platforms (e.g., Kubernetes) is still an open challenge. The biggest limitation is due to the fact that state-of-the-art orchestrators do not support real-time containers, while advanced robotic software often consists of a mix of heterogeneous tasks (i.e., ROS nodes) with different levels of temporal constraints (i.e., mixed-criticality systems). This work addresses this challenge by presenting RT-Kube, a platform that extends the de-facto reference standard for container orchestration, Kubernetes, to schedule tasks with mixed-criticality requirements. It implements monitoring of tasks and detects missed deadlines for those with real-time constraints. It selects low-priority tasks to be migrated at runtime to different units of the computing cluster to free resources and recover from temporal violations. We present quantitative experimental results on the software implementing the mission of a Robotnik RB-Kairos mobile robot to demonstrate the effectiveness of the proposed approach. The source code is publicly available on GitHub. Francesco Lumpp, Franco Fummi, Hiren D. Patel, Nicola Bombieri |
IEEE Trans. Robotics | 3 |
| 2023 | Ditty: Directory-based Cache Coherence for Multicore Safety-critical SystemsabstractDitty is a predictable directory-based cache co-herence mechanism for multicore safety-critical systems that guarantees a worst-case latency (WCL) on data accesses. Prior approaches for predictable cache coherence use a shared snooping bus to interconnect cores. This restricts the number of cores in the multicore to typically four or eight due to scalability concerns. Ditty takes a first step towards a scalable cache coherence mechanism that is predictable and one that can support a larger number of cores. In designing Ditty, we propose a coherence protocol and micro-architecture additions to deliver a WCL bound that is lower than a naive approach. Our WCL analysis reveals that the resulting bounds are comparable to state-of-the-art bus-based predictable coherence approaches. We prototype Ditty in hardware and empirically evaluate it on an FPGA. Our evaluation shows the observed WCL is within computed WCL bound for both the synthetic and SPLASH-3 benchmarks. We release our implementation to the public domain. Zhuanhao Wu, Marat Bekmyrza, Nachiket Kapre, Hiren D. Patel |
DATE | 4 |
| 2023 | SCCL: An open-source SystemC to RTL translatorabstractWe present SCCL, an open-source tool that translates SystemC designs into synthesizable register-transfer level (RTL). SCCL supports a subset of Accellera's SystemC synthesis standard based on the 2011 revision of C++. We use LLVM's Clang front-end to parse SystemC designs, and a suite of analysis passes to construct a SystemC-specific intermediate abstract syntax tree representation called Hcode. Hcode simplifies translation to other intermediate forms such as FIRRTL as well as direct transcription to SystemVerilog or VHDL. Currently, SCCL provides a translation phase to generate synthesizable SystemVerilog. Distinguishing aspects of SCCL include support for complex templated class descriptions that facilitate concise, parameterized hardware specification; introduction and full support for a new type of synthesizable channel called sc_stream that maps directly to standards such as AXI Stream, and a complete reference implementation targeting the Xilinx Vivado toolchain. We demonstrate SCCL's capabilities with a series of case studies including a highly templated SystemC implementation of the ZFP [1] floating-point codec. All case studies are deployed and executed on a Xilinx Zynq UltraScale+ FPGA platform. Zhuanhao Wu, Maya B. Gokhale, Hiren D. Patel |
FCCM | 4 |
| 2023 | ZeroCost-LLC: Shared LLCs at No Cost to WCLabstractZeroCost-LLC (ZCLLC) is a shared inclusive lastlevel cache (LLC) architecture for predictable multicore platforms that does not incur additional cost to the worst-case latency (WCL) of memory requests when compared to the memory hierarchy without an LLC. Thus, the WCL remains the same as without an LLC in the memory hierarchy, but with the performance benefits of having an LLC, in the form of additional caching capacity. ZCLLC achieves this by eliminating all cache line invalidations, and proactively updating the main memory with cache lines to preserve an important vacancy invariant. Furthermore, ZCLLC does not impose any constraints on the way the LLC is used unlike other approaches such as LLC partitioning. Our analysis reveals that the WCL is 55.6%, 68.0%, and 80.2% lower, and the performance is 2.4%, 7.2%, and 25.6% better than the state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively. Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 3 |
| 2023 | Predictable GPU Wavefront Splitting for Safety-Critical SystemsabstractWe present a predictable wavefront splitting (PWS) technique for graphics processing units (GPUs). PWS improves the performance of GPU applications by reducing the impact of branch divergence while ensuring that worst-case execution time (WCET) estimates can be computed. This makes PWS an appropriate technique to use in safety-critical applications, such as autonomous driving systems, avionics, and space, that require strict temporal guarantees. In developing PWS on an AMD-based GPU, we propose microarchitectural enhancements to the GPU, and a compiler pass that eliminates branch serializations to reduce the WCET of a wavefront. Our analysis of PWS exhibits a performance improvement of 11% over existing architectures with a lower WCET than prior works in wavefront splitting. Artem Klashtorny, Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | Enhancing Strong PUF Security With Nonmonotonic Response QuantizationabstractStrong physical unclonable functions (PUFs) provide a low-cost authentication primitive for resource-constrained devices. However, most strong PUF architectures can be modeled through learning algorithms with a limited number of CRPs. In this article, we introduce the concept of nonmonotonic response quantization for strong PUFs. Responses depend not only on which path is faster but also on the distance between the arriving signals. Our experiments show that the resulting PUF has increased security against learning attacks. To demonstrate, we designed and implemented a nonmonotonically quantized ring oscillator-based PUF in 65-nm technology. Measurement results show nearly ideal uniformity and uniqueness with a bit error rate of 13.4% over the temperature range from 0 °C to 50 °C. Kleber Stangherlin, Zhuanhao Wu, Hiren D. Patel, Manoj Sachdev |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Predictable sharing of last-level cache partitions for multi-core safety-critical systemsabstractLast-level cache (LLC) partitioning is a technique to provide temporal isolation and low worst-case latency (WCL) bounds when cores access the shared LLC in multicore safety-critical systems. A typical approach to cache partitioning involves allocating a separate partition to a distinct core. A central criticism of this approach is its poor utilization of cache storage. Today's trend of integrating a larger number of cores exacerbates this issue such that we are forced to consider shared LLC partitions for effective deployments. This work presents an approach to share LLC partitions among multiple cores while being able to provide low WCL bounds. Zhuanhao Wu, Hiren D. Patel |
DAC | 2 |
| 2022 | Managing HBM Bandwidth on Multi-Die FPGAs with FPGA Overlay NoCsabstractWe can improve HBM bandwidth distribution and utilization on a multi-die FPGA like Xilinx Alveo U280 by using Overlay Network-on-Chips (NoCs). The HBM in Xilinx Alveo U280 offers 8 GB of memory capacity with a theoretical maximum bandwidth of 460 GBps, but exposed all the HBM ports to the FPGA fabric in only one die. As a result, computing elements assigned to other dies must use the scarce Super Long Lines (SLLs) to access HBM bandwidth. Furthermore, HBM is fractured internally into thirty-two smaller memories called pseudo channels, connected together by a hardened and performance-limited crossbar. The crossbar enables global accesses from any of the HBM ports, but introduces several throughput bottlenecks. An Overlay Hybrid NoC combining Hoplite NoC with Butterfly Fat Trees (BFT) NoCs offers a high-performance solution for distributing HBM bandwidth across all three dies. The routing capability of the NoC can be modified to supplant the internal crossbar of Xilinx HBM for global accesses. We demonstrate this in Xilinx Alveo U280 with BFT, Hoplite, and Hybrid NoC, using synthetic benchmarks and two application-based benchmarks, Dense matrix-matrix multiplication (DMM) and Sparse Matrix-Vector multiplication (SPMV). Our experiments show that Overlay NoCs can improve the throughput by 1.26× for synthetic benchmarks and up to 1.4× for SpMV workloads. Srinirdheeshwar Kuttuva Prakash, Hiren D. Patel, Nachiket Kapre |
FCCM | 2 |
| 2022 | ZHW: A Numerical CODEC for Big Data Scientific ComputationabstractDistributed big data in scientific computing presents a major I/O performance bottleneck when exploiting data paral-lelism. Consumer and producer compute nodes are often throttled by saturated data channels when processing large numerical data. We describe ZHW, a hardware implementation of the ZFP numerical CODEC that can greatly reduce I/O pressure caused by large scientific datasets. Our ZHW design overcomes barriers that have prevented prior ZFP-like hardware accelerators from obtaining maximum compression in their implementations. The SystemC ZHW hardware library is available in an open source public repository. We demonstrate the practicality of ZHW by synthesizing our CODEC on an Ultrascale+ FPGA and analyzing performance. Michael Barrow, Zhuanhao Wu, Maya B. Gokhale, Hiren D. Patel, Peter Lindstrom 0001 |
FPT | 5 |
| 2022 | Containerization and Orchestration of Software for Autonomous Mobile Robots: a Case Study of Mixed-Criticality Tasks across Edge-Cloud Computing PlatformsabstractContainerization promises to strengthen platform-independent development, better resource utilization, and secure deployment of software. As these benefits come with negligible overhead in CPU and memory utilization, containerization is increasingly being adopted in mobile robotic applications. An open challenge is supporting software tasks that have mixed-criticality requirements. Even more challenging is the combination of real-time containers with orchestration, which is an emerging paradigm to automate the deployment, networking, scaling, and availability of containerized workloads and services. This paper addresses this challenge by presenting a framework that extends the de-facto reference standard for container orchestration, Kubernetes, to schedule tasks with mixed-criticality requirements. Quantitative experimental results on the software implementing the mission of a Robotnik RB-Kairos mobile robot demonstrate the effectiveness of the proposed approach. The source code is publicly available on GitHub. Francesco Lumpp, Franco Fummi, Hiren D. Patel, Nicola Bombieri |
IROS | 3 |
| 2022 | Automatic Construction of Predictable and High-Performance Cache Coherence Protocols for Multicore Real-Time SystemsabstractPredictable hardware cache coherence is a viable shared data communication mechanism between cores for multicore real-time platforms. Prior works have established that predictable hardware cache coherence protocols offer significant performance advantages over alternative predictable data communication mechanisms while ensuring predictability. Unlike alternative predictable data communication mechanisms, designing predictable cache coherence protocols is nontrivial as it requires detailed understanding of the impact of different memory activity patterns to shared data for predictable and coherent data communication. Furthermore, designing predictable cache coherence protocols that deliver high average-case performance is even more challenging as it entails identifying opportunities such that a core’s access to a data is not stalled in the presence of interleaving memory activity from other cores to the same data. To this end, we present SYNTHIA, an open and automated tool for synthesizing predictable and high-performance snooping bus-based cache coherence protocols for multicore platforms deployed in real-time systems. SYNTHIA automates the complex analysis associated with designing predictable and high-performance cache coherence protocols, and constructs the complete protocol implementation (coherence states and transitions) that achieve predictability and performance. We use SYNTHIA to construct complete protocol implementations from simple specifications of common protocols (modified-shared-invalid (MSI), MESI, and MOESI protocols) and a predictable variant of the MESIF cache coherence protocol, which was recently found to be deployed in an existing multicore platform designed for real-time platforms. We validated the correctness, predictability, and performance guarantees of the generated protocol implementations from SYNTHIA using manually implemented versions, and a micro-architectural simulator. Anirudh M. Kaushik, Hiren D. Patel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | A Framework for Optimizing CPU-iGPU Communication on Embedded PlatformsabstractMany modern programmable embedded devices contain CPUs and a GPU that share the same system memory on a single die. Such a unified memory architecture allows the explicit data copying between CPU and integrated GPU (iGPU) to be eliminated with the benefit of significantly improving performance and energy savings. However, to enable such a “zero-copy” communication model, many devices either implement intricate cache coherence protocols or they may disable the last level caches. This often leads to strong performance degradation of cache-dependent applications, for which CPU-iGPU data transfer based on standard copy remains the best solution. This paper presents a framework based on a performance model, a set of micro-benchmarks, and a novel zero-copy communication pattern to accurately estimate the potential speedup a CPU-iGPU application may have by considering different communication models (i.e., standard copy, unified memory, or pinned “zerocopy”). It shows how the framework can be combined with standard profiler information to efficiently drive the application tuning for a given programmable embedded device. Francesco Lumpp, Hiren D. Patel, Nicola Bombieri |
DAC | 2 |
| 2021 | Automated Synthesis of Predictable and High-Performance Cache Coherence ProtocolsabstractWe present SYNTHIA, an open and automated tool for synthesizing predictable and high-performance snooping bus-based cache coherence protocols for multi-core processors in multi-processor system-on-chips (MPSoCs) deployed in real-time systems. SYNTHIA automates the complex analysis associated with designing predictable and high-performance cache coherence protocols, and constructs new states (transient states) and corresponding transitions that achieve predictability and performance. We use SYNTHIA to construct complete protocol implementations from simple specifications of common protocols (MSI, MESI, and MOESI protocols). We validated the correctness, predictability, and performance guarantees of the generated protocol implementations from SYNTHIA using manually implemented versions, and a micro-architectural simulator. Anirudh M. Kaushik, Hiren D. Patel |
DATE | 2 |
| 2021 | A Systematic Approach to Achieving Tight Worst-Case Latency and High-Performance Under Predictable Cache CoherenceabstractPredictable hardware cache coherence is an attractive data communication mechanism between safety-critical tasks deployed on real-time multi-core platforms due to its predictability and high-performance benefits. However, from a worst-case analysis standpoint, alternative data communication mechanisms appear in favorable light for adoption in real-time multi-core platforms. This is because alternative data communication mechanisms such as cache bypassing offer tighter worstcase latency (WCL) bounds for memory requests compared to predictable hardware cache coherence mechanisms. We present a systematic approach towards designing predictable cache coherence mechanisms that offer tight WCL and high-performance. Our approach consists of a formal framework that concisely captures the key reasons behind the high WCL in existing predictable cache coherence mechanisms. Guided by this formal framework, we describe one technique that employs micro-architectural extensions and protocol changes to achieve tight WCL and high-performance. We apply this technique to two existing cache coherence mechanisms. Our evaluation shows that the new cache coherence mechanisms resulting from our technique have the same tight WCL as alternative mechanisms, and still maintain a significant average-case performance advantage (up to 5× speedup) over the alternative mechanisms. Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 2 |
| 2021 | A Hardware Platform for Exploring Predictable Cache Coherence Protocols for Real-time MulticoresabstractThis work presents MapleBoard: a set of opensource hardware tools to implement predictable cache coherence protocols in hardware. MapleBoard consists of the following: (1) a novel domain-specific language (DSL) for specifying coherence protocols and synthesizing the corresponding hardware; and, (2) a real-time multicore hardware platform that seamlessly integrates the coherence protocols synthesized from the DSL. This platform has a memory hierarchy, a real-time bus interconnect between cores, and various predictable arbiters. As a demonstration of MapleBoard's efficacy, we explore hardware implementations of data bus organizations, and their impact on the worstcase communication latency (WCL). An important discovery we make is that a dedicated data bus (DDB) organization that allows bidirectional data communication offers lower analytical WCL bounds than any state-of-the-art predictable cache coherence protocols. The analytical WCL bounds are improved by 84%, 90% and 94% for 2-core, 4-core and 8-core systems respectively compared to prior works. We synthesize MapleBoard on the Xilinx Virtex Ultrascale+ VCU1525 board, and validate using both synthetic workloads and SPLASH-2 benchmark suite. Zhuanhao Wu, Anirudh M. Kaushik, Paulos Tegegn, Hiren D. Patel |
RTAS | 4 |
| 2021 | Gretch: A Hardware Prefetcher for Graph AnalyticsabstractData-dependent memory accesses (DDAs) pose an important challenge for high-performance graph analytics (GA). This is because such memory accesses do not exhibit enough temporal and spatial locality resulting in low cache performance. Prior efforts that focused on improving the performance of DDAs for GA are not applicable across various GA frameworks. This is because (1) they only focus on one particular graph representation, and (2) they require workload changes to communicate specific information to the hardware for their effective operation. In this work, we propose a hardware-only solution to improving the performance of DDAs for GA across multiple GA frameworks. We present a hardware prefetcher for GA called Gretch, that addresses the above limitations. An important observation we make is that identifying certain DDAs without hardware-software communication is sensitive to the instruction scheduling. A key contribution of this work is a hardware mechanism that activates Gretch to identify DDAs when using either in-order or out-of-order instruction scheduling. Our evaluation shows that Gretch provides an average speedup of 38% over no prefetching, 25% over conventional stride prefetcher, and outperforms prior DDAs prefetchers by 22% with only 1% increase in power consumption when executed on different GA workloads and frameworks. Anirudh M. Kaushik, Gennady Pekhimenko, Hiren D. Patel |
ACM Trans. Archit. Code Optim. | 3 |
| 2021 | Designing Predictable Cache Coherence Protocols for Multi-Core Real-Time SystemsabstractThis article addresses the challenge of allowing simultaneous and predictable accesses to shared data on multi-core systems. We propose a collection of predictable cache coherence protocols, which mandate the use of certain design invariants to ensure predictability. In particular, we enforce these invariants by augmenting the classic modify-share-invalid (MSI) protocol and modify-exclusive-share-invalid (MESI) protocol. This allows us to derive worst-case latency bounds on the resulting predictable MSI (PMSI) and predictable MESI (PMESI) protocols. Our analysis shows that while the arbitration latency scales linearly, the coherence latency scales quadratically with the number of cores, which emphasizes the importance of accounting for cache coherence effects on latency bounds. We implement PMSI and PMESI in a detailed micro-architectural simulator, and execute SPLASH-2 and synthetic workloads. Results show that our approach is always within the analytical worst-case latency bounds, and that PMSI and PMESI improve average-case performance by up to 4× over cache bypassing mechanisms that disallow caching of shared data in the cores’ private caches. PMSI and PMESI have average slowdowns of 1.45× and 1.46× compared to conventional MSI and MESI protocols, respectively. Anirudh M. Kaushik, Mohamed Hassan 0002, Hiren D. Patel |
IEEE Trans. Computers | 3 |
| 2021 | Task Mapping and Scheduling for OpenVX Applications on Heterogeneous Multi/Many-Core ArchitecturesabstractComputer vision applications have stringent performance constraints that must be satisfied when they are run at the edge on programmable low-power embedded devices. OpenVX has emerged as the de-facto reference standard to develop such applications. OpenVX uses a primitive-based programming model that results in a directed-acyclic graph (DAG) representation of the application, which can then be used for automatic system-level optimizations and synthesis to heterogeneous multi- and many-core platforms. Although OpenVX has been standardized, its state-of-the-art algorithm for task mapping and scheduling does not deliver the performance necessary for such applications to be deployed on heterogeneous multi-/many-core platforms. This article focuses on addressing this challenge with three main contributions: First, we implemented a static task scheduling and mapping approach for OpenVX using the heterogeneous earliest finish time (HEFT) heuristic. We show that HEFT allows us to improve the system performance up to 70 percent on one of the most widespread smart systems for applying computer vision and intelligent video analytics in general at the edge (i.e., NVIDIA VisionWorks on NVIDIA Jetson TX2). Second, we show that HEFT, in the context of a vision application for edge computing where some primitives may have multiple implementations (e.g., for CPU and GPU), can lead to load imbalance amongst heterogeneous computing elements (CEs), thus suffering from degraded performance. Third, we present an algorithm called exclusive earliest finish time (XEFT) that introduces the notion of exclusive overlap between single implementation primitives to improve the load balancing. We show that XEFT can further improve the system performance up to 33 percent over HEFT, and 82 percent over the native OpenVX scheduler. We present the results on a large set of benchmarks, including a real-world localization and mapping application (ORB-SLAM) combined with an NVIDIA inference application based on convolutional neural networks (CNNs) for object detection. Francesco Lumpp, Stefano Aldegheri, Hiren D. Patel, Nicola Bombieri |
IEEE Trans. Computers | 3 |
| 2020 | On the Task Mapping and Scheduling for DAG-based Embedded Vision Applications on Heterogeneous Multi/Many-core ArchitecturesabstractIn this work, we show that applying the heterogeneous earliest finish time (HEFT) heuristic for the task scheduling of embedded vision applications can improve the system performance up to 70% w.r.t. the scheduling solutions at the state of the art. We propose an algorithm called exclusive earliest finish time (XEFT) that introduces the notion of exclusive overlap between application primitives to improve the load balancing. We show that XEFT can improve the system performance up to 33% over HEFT, and 82% over the state of the art approaches. We present the results on different benchmarks, including a real-world localization and mapping application (ORB-SLAM) combined with the NVIDIA object detection application based on deep-learning. Stefano Aldegheri, Nicola Bombieri, Hiren D. Patel |
DATE | 3 |
| 2019 | Strengthening PUFs using CompositionabstractWe explore the idea of composing PUFs with the intent that the resultant PUF is stronger than the constituent PUFs. Prior work has proposed a construction, which subsequent work has shown to be weak. We revisit this prior construction and observe that it is actually weaker than previously thought when the constituent PUFs are arbiter PUFs. This weakness is demonstrated via our adaptation of the previously proposed Logistic Regression (LR) attack. We then propose new constructions called PUFs-composed-with-PUFs (PoP). In particular, we retain a two-layer construction, but allow the same input to the composite PUF to be input to more than one constituent PUF at the first layer. We explore this family of constructions, with arbiter PUFs serving as the constituent PUFs. In particular, we identify several axes which we can vary, and empirically study the resilience of our constructions compared to the prior construction and one another from the standpoint of LR attacks. As insight into why our family of constructions is stronger, we prove, under some idealized conditions, that the lower-bound on an attacker is indeed higher under our constructions than the upper-bound on an attacker for the prior construction. As such, our work suggests that composition can be a promising approach to strengthening PUFs, contrary to what prior work suggests. Zhuanhao Wu, Hiren D. Patel, Manoj Sachdev, Mahesh Tripunitara |
ICCAD | 2 |
| 2019 | CARP: A Data Communication Mechanism for Multi-core Mixed-Criticality SystemsabstractWe present CARP, a predictable and high-performance data communication mechanism for multi-core mixed-criticality systems (MCS). CARP is realized as a hardware cache coherence protocol that enables communication between critical and non-critical tasks while ensuring that non-critical tasks do not interfere with the safety requirements of critical tasks. The key novelty of CARP is that it is criticality-aware, and hence, handles communication patterns between critical and non-critical tasks appropriately. We derive the analytical worst-case latency bounds for requests using CARP and note that the observed per-request latencies are within the analytical worst-case latency bounds. We compare CARP against prior data communication mechanisms using synthetic and SPLASH-2 benchmarks. Our evaluation shows that CARP improves the average-case performance of MCS compared to prior data communication mechanisms, while maintaining the safety requirements of critical tasks. Anirudh M. Kaushik, Paulos Tegegn, Zhuanhao Wu, Hiren D. Patel |
RTSS | 4 |
| 2019 | Enabling Predictable, Simultaneous and Coherent Data Sharing in Mixed Criticality SystemsabstractEmerging embedded systems deployed in the automotive and avionics domains execute applications with different criticalities, comprising what is known as Mixed Criticality Systems (MCS). Applications in MCS often share data between tasks (coming from sensors for instance). Data sharing is challenging because it can lead to increased response times or even unpredictable behaviors if not carefully addressed. Therefore, several prior works in MCS either assumed that tasks do not share data or disallowed it by design. Recent solutions attempt to mitigate the effects of data sharing, albeit by introducing new restrictions on the system either by prohibiting applications from caching shared data or prohibiting the operating system from running tasks with shared data in parallel. We find these solutions also to have limited applicability as they deteriorate system schedulability and prohibit simultaneous access to shared data. In this paper, we propose PENDULUM: a time-based cache coherence protocol to enable simultaneous and predictable access to shared data in MCS. Our evaluation shows that PENDULUM, achieves flexibility and better performance compared to existing solutions, while maintaining system predictability. Nivedita Sritharan, Anirudh M. Kaushik, Mohamed Hassan 0002, Hiren D. Patel |
RTSS | 4 |
| 2018 | MCXplore: Automating the Validation Process of DRAM Memory Controller DesignsabstractWe present an automated framework for the validation of memory controllers (MCs) called MCXplore. In developing this framework, we construct formal models for memory requests and command interactions. MCXplore enables validation engineers to define their test plans precisely using temporal logic specifications. We use the NuSMV model-checker to generate counterexamples that serve as test templates. MCXplore uses these test templates to generate memory tests to validate the correctness properties of the MC. We show the effectiveness of MCXplore by validating various state-of-the-art MC features as well as hard-to-detect timing violations. We also provide a set of predefined test plans, and regression test suites that validate essential properties of modern MCs. MCXplore is an open-source framework to allow validation engineers and researchers to extend and use. Mohamed Hassan 0002, Hiren D. Patel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | A Comparative Study of Predictable DRAM ControllersabstractRecently, the research community has introduced several predictable dynamic random-access memory (DRAM) controller designs that provide improved worst-case timing guarantees for real-time embedded systems. The proposed controllers significantly differ in terms of arbitration, configuration, and simulation environment, making it difficult to assess the contribution of each approach. To bridge this gap, this article provides the first comprehensive evaluation of state-of-the-art predictable DRAM controllers. We propose a categorization of available controllers, and introduce an analytical performance model based on worst-case latency. We then conduct an extensive evaluation for all state-of-the-art controllers based on a common simulation platform, and discuss findings and recommendations. Danlu Guo, Mohamed Hassan 0002, Rodolfo Pellizzoni, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | Exposing Implementation Details of Embedded DRAM Memory Controllers through Latency-based AnalysisabstractWe explore techniques to reverse-engineer DRAM embedded memory controllers (MCs), including page policies, address mapping, and command arbitration. There are several benefits to knowing this information: They allow tightening worst-case bounds of embedded systems and platform-aware optimizations at the operating system, source-code, and compiler levels. We develop a latency-based analysis, which we use to devise algorithms and C programs to extract MC properties. We show the effectiveness of the proposed approach by reverse-engineering the MC details in the XUPV5-LX110T Xilinx platform. Furthermore, to cover a breadth of policies, we use a simulation framework and document our findings. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | Predictable Cache Coherence for Multi-core Real-Time SystemsabstractThis work addresses the challenge of allowing simultaneous and predictable accesses to shared data on multicore systems. We propose a predictable cache coherence protocol, which mandates the use of certain invariants to ensure predictability. In particular, we enforce these invariants by augmenting the classic modify-share-invalid (MSI) protocol with transient coherence states, and minimal architectural changes. This allows us to derive worst-case latency bounds on predictable MSI (PMSI) protocol. Our analysis shows that while the arbitration latency scales linearly, the coherence latency scales quadratically with the number of cores, which emphasizes that importance of accounting for cache coherence effects on latency bounds. We implement PMSI in gem5, and execute SPLASH-2 and synthetic workloads. Results show that our approach is always within the analytical worst-case latency bounds, and that PMSI improves averagecase performance by up to 4 over the next best predictable alternative. PMSI has average slowdowns of 1.45 and 1.46 compared to MSI and MESI protocols, respectively. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 3 |
| 2017 | PMC: A Requirement-Aware DRAM Controller for Multicore Mixed Criticality SystemsabstractWe propose a novel approach to schedule memory requests in Mixed Criticality Systems (MCS). This approach supports an arbitrary number of criticality levels by enabling the MCS designer to specify memory requirements per task. It retains locality within large-size requests to satisfy memory requirements of all tasks. To achieve this target, we introduce a compact time-division-multiplexing scheduler, and a framework that constructs optimal schedules to manage requests to off-chip memory. We also present a static analysis that guarantees meeting requirements of all tasks. We compare the proposed controller against state-of-the-art memory controllers using both a case study and synthetic experiments. Mohamed Hassan 0002, Hiren D. Patel, Rodolfo Pellizzoni |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | MCXplore: An automated framework for validating memory controller designs
Mohamed Hassan 0002, Hiren D. Patel |
DATE | 2 |
| 2016 | Criticality- and Requirement-Aware Bus Arbitration for Multi-Core Mixed Criticality SystemsabstractThis work presents CArb, an arbiter for controlling accesses to the shared memory bus in multi-core mixed criticality systems. CArb is a requirement-aware arbiter that optimally allocates service to tasks based on their requirements. It is also criticality-aware since it incorporates criticality as a first-class principle in arbitration decisions. CArb supports any number of criticality levels and does not impose any restrictions on mapping tasks to processors. Hence, it operates in tandem with existing processor scheduling policies. In addition, CArb is able to dynamically adapt memory bus arbitration at run time to respond to increases in the monitored execution times of tasks. Utilizing this adaptation, CArb is able to offset these increases; hence, postpones the system need to switch to a degraded mode. We prototype CArb, and evaluate it with an avionics case-study from Honeywell as well as synthetic experiments. Mohamed Hassan 0002, Hiren D. Patel |
RTAS | 2 |
| 2016 | Buffer Space Allocation for Real-Time Priority-Aware NetworksabstractIn this work, we address the challenge of incorporating buffer space constraints in worst-case latency analysis for priority-aware networks. A priority-aware network is a wormhole-switched network-on-chip with distinct virtual channels per priority. Prior worst-case latency analyses assume that the routers have infinite buffer space allocated to the virtual channels. This assumption renders these analyses impractical when considering actual deployments. This is because an implementation of the priority-aware network imposes buffer constraints on the application. These constraints can result in back pressure on the communication, which the analyses must incorporate. Consequently, we extend a worst- case latency analysis for priority-aware networks to include buffer space constraints. We provide the theory for these extensions and prove their correctness. We experiment on a large set of synthetic benchmarks, and show that we can deploy applications on priority-aware networks with virtual channels of sizes as small as two flits. In addition, we propose a polynomial time buffer space allocation algorithm. This algorithm minimizes the buffer space required at the virtual channels while scheduling the application sets on the target priority-aware network. Our empirical evaluation shows that the proposed algorithm reduces buffer space requirements in the virtual channels by approximately 85% on average. Hany Kashif, Hiren D. Patel |
RTAS | 2 |
| 2016 | Path Selection for Real-Time Communication on Priority-Aware NoCsabstractThis work investigates selecting paths for communication flows when deploying a hard real-time application on a chip-multiprocessor system. This chip-multiprocessor system uses a priority-aware real-time network-on-chip interconnect between the processors. Given a mapping of the computation tasks onto the chip-multiprocessor, the problem we address in this work is to discover paths the communication flows take such that hard real-time deadlines of flows are met. Furthermore, we must ensure that deadlines are met even in the presence of direct and indirect interference from other flows sharing network links on the path. To achieve this, our algorithm utilizes a stage-level analysis for real-time communication to determine the impact of a network link being used by a flow, and its effect on other flows sharing the link. The path selection algorithm uses heuristics such as selecting links with least interference, and considering lower-priority flows when dedicating links to paths of higher-priority flows since an optimal one is intractable. The algorithm also considers constraints on the number of virtual channels at each router port in the network. The statistically significant experimental results show an improvement in schedulability by 5% and 12% over existing path selection algorithms such as Minimum Interference Routing and Widest Shortest Path algorithms, respectively. We also present a set-top box case study to further illustrate the benefits of using the proposed algorithm. Hany Kashif, Hiren D. Patel, Sebastian Fischmeister |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2015 | Static slack-based instrumentation of programsabstractReal-time embedded programs are time sensitive and, to trace such programs, the instrumentation mechanism must honor the programs' timing constraints. We present a time-aware instrumentation technique that injects program code with slack-based conditional instrumentation. The central idea is to execute instrumentation code only when its execution does not increase the worst-case execution time beyond a program's deadline. This occurs at run-time. Unlike previous efforts, this work allows instrumenting on the path that results in the worst-case execution time of the program. We propose a software, and a hardware method of allowing for slack-based conditional instrumentation. We evaluate and compare these two alternatives using a common benchmark suite for real-time systems. Our results show that, on average, the two proposed methods achieve 57% and 80% instrumentation coverage, respectively, compared to only a 3% coverage by previous work. Hany Kashif, Johnson J. Thomas, Hiren D. Patel, Sebastian Fischmeister |
ETFA | 3 |
| 2015 | Reverse-engineering embedded memory controllers through latency-based analysisabstractWe explore techniques to reverse-engineer properties of DRAM memory controllers (MCs). This includes page policies, address mapping schemes and command arbitration schemes. There are several benefits to knowing this information: they allow analysis techniques to effectively compute worst-case bounds, and they allow customizations to be made in software for predictability. We develop a latency-based analysis, and use this analysis to devise algorithms for micro-benchmarks to extract properties of MCs. In order to cover a breadth of page policies, address mappings and command arbitration schemes, we explore our technique using a micro-architecture simulation framework and document our findings. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 3 |
| 2015 | A framework for scheduling DRAM memory accesses for multi-core mixed-time critical systemsabstractMixed-time critical systems are real-time systems that accommodate both hard real-time (HRT) and soft realtime (SRT) tasks. HRT tasks mandate a gurantee on the worstcase latency, while SRT tasks have average-case bandwidth (BW) demands. Memory requests in mixed-time critical systems usually have different transaction sizes based on whether the issuer task is HRT or SRT. For example, HRT tasks often issue requests with a cache line size. On the other side, SRT tasks may issue requests with a size of KBs. Requests from multimedia cores, cores controlling network interfaces and direct memory accesses (DMAs) are obvious examples of these large-size requests. Based on these observations, we promote in this work a new approach to schedule memory requests. This approach retains locality within large-size requests to minimize the worst-case latency, while maintaining the average-case BW as high as required. To achieve this target, we introduce a novel and compact time-division-multiplexing scheduler that is adequate for mixed-time critical systems. We also present a novel framework that constructs optimal offchip DRAM memory controller schedules for multi-core mixedtime critical systems. These schedules are loaded to the memory controller during boot-time. Based on the proposed schedule, we provide a detailed static analysis that guarantees predictability. We compare the proposed controller against state-of-the-art realtime memory controllers using synthetic experiments as well as a practical use-case from multimedia systems. Mohamed Hassan 0002, Hiren D. Patel, Rodolfo Pellizzoni |
RTAS | 2 |
| 2015 | SLA: A Stage-Level Latency Analysisfor Real-Time Communicationin a Pipelined Resource ModelabstractWe present a communication analysis for hard real-time systems interconnects. The objective is to provide tight estimates on the worst-case communication latency between communicating processing elements that use a priority-aware communication medium for data transmission. The communication model consists of communication tasks transmitting data across a series of pipelined resources. The analysis incorporates interferences caused by multiple communication tasks requesting the pipelined resources, and it captures parallel transmission of data between multiple pipeline stages. We call this analysis a stage-level analysis. We evaluate the proposed analysis through simulation of synthetic benchmarks, and we apply the analysis to an instantiation of a platform proposed by Shi and Burns. Our experiments confirm that stage-level analysis provides tight upper-bounds when compared to previous work and improves schedulability by 34 percent. Hany Kashif, Sina Gholamian, Hiren D. Patel |
IEEE Trans. Computers | 3 |
| 2014 | Bounding buffer space requirements for real-time priority-aware networksabstractOne implementation alternative for network interconnects in modern chip-multiprocessor systems is priority-aware arbitration networks. To enable the deployment of real-time applications to priority-aware networks, recent research proposes worst-case latency (WCL) analyses for such networks. Buffer space requirements in priority-aware networks, however, are seldom addressed. In this work, we bound the buffer space required for valid WCL analyses and consequently optimize router design for application specifications by computing the required buffer space at each virtual channel in priority-aware routers. In addition to the obvious advantage of bounding buffer space while providing valid WCL bounds, buffer space reduction decreases chip area and saves energy in priority-aware networks. Our experiments show that the proposed buffer space computation reduces the number of unfeasible implementations by 42% compared to an existing buffer space analysis technique. It also reduces the required buffer space in priority-aware routers by up to 79%. Hany Kashif, Hiren D. Patel |
ASP-DAC | 2 |
| 2014 | MEMOCODE 2014 software design contest: Space Invaders emulatorabstractThe MEMOCODE design contest for 2014 was centered around the emulation of the 1978 Taito video game Space Invaders. The challenge is to improve the speed of a cycle-accurate software emulator for the game. Contestants had a month toope improve the provided code, which already ran fairly well on the ARM-based Raspberry Pi platform. Entries were judged on how much faster their code ran and its quality. The winning groups used a variety of optimization techniques ranging from dynamic binary translation, data-structure restructuring, and improving instruction and data caching. Stephen A. Edwards, Hiren D. Patel |
MEMOCODE | 2 |
| 2013 | Low cost permanent fault detection using ultra-reduced instruction set co-processorsabstractIn this paper, we propose a new, low hardware overhead solution for permanent fault detection at the micro-architecture/instruction level. The proposed technique is based on an ultra-reduced instruction set co-processor (URISC) that, in its simplest form, executes only one Turing complete instruction — the subleq instruction. Thus, any instruction on the main core can be redundantly executed on the URISC using a sequence of subleq instructions, and the results can be compared, also on the URISC, to detect faults. A number of novel software and hardware techniques are proposed to decrease the performance overhead of online fault detection while keeping the error detection latency bounded including: (i) URISC routines and hardware support to check both control and data flow instructions; (ii) checking only a subset of instructions in the code based on a novel check window criterion; and (iii) URISC instruction set extensions. Our experimental results, based on FPGA synthesis and RTL simulations, illustrate the benefits of the proposed techniques. Sundaram Ananthanarayanan, Siddharth Garg, Hiren D. Patel |
DATE | 3 |
| 2013 | On the use of GP-GPUs for accelerating compute-intensive EDA applicationsabstractGeneral purpose graphics processing units (GP-GPUs) have recently been explored as a new computing paradigm for accelerating compute-intensive EDA applications. Such massively parallel architectures have been applied in accelerating the simulation of digital designs during several phases of their development - corresponding to different abstraction levels, specifically: (i) gate-level netlist descriptions, (ii) register-transfer level and (iii) transaction-level descriptions. This embedded tutorial presents a comprehensive analysis of the best results obtained by adopting GP-GPUs in all these EDA applications. Valeria Bertacco, Debapriya Chatterjee, Nicola Bombieri, Franco Fummi, Sara Vinco, Anirudh M. Kaushik, Hiren D. Patel |
DATE | 7 |
| 2013 | Systemc-clang: An open-source framework for analyzing mixed-abstraction SystemC models
Anirudh M. Kaushik, Hiren D. Patel |
FDL | 2 |
| 2013 | ORTAP: An Offset-based response time analysis for a pipelined communication resource modelabstractThis work addresses the challenge of computing worst-case response times of hard real-time applications deployed on multiprocessor systems. In particular, the worst-case response time analysis (WCRTA) focuses on the communication between distributed tasks of hard real-time applications. The proposed WCRTA models the communication as a pipelined communication resource model. This model incorporates the effect of pipelining, and the parallel transmission of data. Applications of such a model include multiprocessor systems that use complex interconnects such as network-on-chips (NoC)s with priorities. In this paper, we present an exponential analysis, and a polynomial analysis, and prove its correctness. As an application, we apply the pipelined communication resource model to priority-aware NoCs, and we compare the proposed analyses against prior analysis techniques. Our experimental evaluation on two instances of 4 × 4 and 8 × 8 NoCs with 512,000 synthetic benchmarks shows 48.3% and 66.7% improvement in schedulability for the two NoC sizes over prior work. Hany Kashif, Sina Gholamian, Rodolfo Pellizzoni, Hiren D. Patel, Sebastian Fischmeister |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2013 | An Instruction Scratchpad Memory Allocation for the Precision Timed ArchitectureabstractThis paper presents a static instruction scratchpad memory allocation scheme for the precision timed architecture (PRET). Since PRET provides timing instructions to control the temporal execution of programs, the objective of the allocation scheme is to ensure that the explicitly specified temporal requirements are met. Furthermore, this allocation incorporates the timing requirements from the multiple hardware threads of the PRET architecture. We formulate the allocation problem as an integer-linear programming problem, and we implement a tool that takes compiled ARMv4 binaries, constructs a timing-requirements-aware control-flow graph, performs a WCET analysis and SPM allocation, and rewrites the binaries with the allocation. We evaluate our approach using a modified version of the Malardalen benchmarks to show the benefits of the proposed approach. We also present a UAV benchmark derived from the PapaBench benchmark. Aayush Prakash, Hiren D. Patel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Using link-level latency analysis for path selection for real-time communication on NoCsabstractWe present a path selection algorithm that is used when deploying hard real-time traffic flows onto a chip-multiprocessor system. This chip-multiprocessor system uses a priority-based real-time network-on-chip interconnect between the multiple processors. The problem we address is the following: given a mapping of the tasks onto a chip-multiprocessor system, we need to determine the paths that the traffic flows take such that the flows meet there deadlines. Furthermore, we must ensure that the deadline is met even in the presence of direct and indirect interference from other flows sharing network links on the path. To achieve this, our algorithm utilizes a link-level analysis to determine the impact of a link being used by a flow, and its affect on other flows sharing the link. Our experimental results show that we can improve schedulability by about 8% and 15% over Minimum Interference Routing and Widest Shortest Path algorithms, respectively. Hany Kashif, Hiren D. Patel, Sebastian Fischmeister |
ASP-DAC | 2 |
| 2012 | Parallel simulation of mixed-abstraction SystemC models on GPUs and multicore CPUsabstractThis work presents a methodology that parallelizes the simulation of mixed-abstraction level SystemC models across multicore CPUs, and graphics processing units (GPUs) for improved simulation performance. Given a SystemC model, we partition it into processes suitable for GPU execution and CPU execution. We convert the processes identified for GPU execution into GPU kernels with additional SystemC wrapper processes that invoke these kernels. The wrappers enable seamless communication of events in all directions between the GPUs and CPUs. We alter the OSCI SystemC simulation kernel to allow parallel execution of processes. Hence, we co-simulate in parallel, the SystemC processes on multiple CPUs, and the GPU kernels on the GPUs; exploit both the CPUs, and GPUs for faster simulation. We experiment with synthetic benchmarks and a set-top box case study. Rohit Sinha 0001, Aayush Prakash, Hiren D. Patel |
ASP-DAC | 3 |
| 2012 | Reliable computing with ultra-reduced instruction set co-processorsabstractThis work presents a method to reliably perform computations in the presence of hard faults arising from aggressive technology scaling, and design defects from human error. Our method is based on an observation that a single Turing-complete instruction can mirror the semantics of any other instruction. One such instruction is the subleq instruction, which has been used for instructional purposes in the past. We find that the scope for using such a Turing-complete instruction is far greater, and in this paper, we present its applicability to fault tolerance. In particular, we extend a MIPS processor with a co-processor (called ultra-reduced instruction set co-processor -- URISC) that implements the subleq instruction. We use the URISC to execute sequences of subleq that are semantically equivalent to the faulty instructions. We formally prove this, and implement the translations in the back-end of the LLVM compiler. We generate binaries for our hardware prototype called MIPS-URISC, which we synthesize and execute on an Altera FPGA. Our experiments indicate the performance and area overheads, and the efficacy of the proposed approach. Aravindkumar Rajendiran, Sundaram Ananthanarayanan, Hiren D. Patel, Mahesh Tripunitara, Siddharth Garg |
DAC | 3 |
| 2012 | An instruction scratchpad memory allocation for the precision timed architectureabstractThis work presents a static instruction allocation scheme for the precision timed architecture's (PRET) scratchpad memory. Since PRET provides timing instructions to control the temporal execution of programs, the objective of the allocation scheme is to ensure that the explicitly specified temporal requirements are met. Furthermore, this allocation incorporates instructions from multiple hardware threads of the PRET architecture. We formulate the allocation as an integer-linear programming problem, and we implement a tool that takes binaries, constructs a control-flow graph, performs the allocation, rewrites the binary with the new allocation, and generates an output binary for the PRET architecture. We carry out experiments on a subset of a modified version of the Malardalen benchmarks to show the benefits of performing the allocation across multiple threads. Aayush Prakash, Hiren D. Patel |
DATE | 2 |
| 2012 | synASM: A High-Level Synthesis Framework With Support for Parallel and Timed ConstructsabstractThis paper presents a high-level synthesis framework called synASM that synthesizes abstract state machines (ASMs) to VHDL for field-programmable gate arrays (FPGAs). In particular, this paper focuses on the specification, scheduling, and synthesis of parallel and timed constructs. ASMs possess well-defined formal semantics for sequential and parallel computation, and their composition. We extend ASMs to support the specification of timing requirements, which we call timed constructs. We also describe the composition of timed constructs with sequential and parallel computation. A key contribution of this paper is the extension of the force-directed scheduling algorithm to support both parallel and timed constructs. We implement the synthesis back-end in synASM that targets FPGAs. Our experiments show improvements of up to 52% in lookup table usage and 34% in total area for certain examples. Rohit Sinha 0001, Hiren D. Patel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | Temporal isolation on multiprocessing architecturesabstractMultiprocessing architectures provide hardware for executing multiple tasks simultaneously via techniques such as simultaneous multithreading and symmetric multiprocessing. The problem addressed by this paper is that even when tasks that are executing concurrently do not communicate, they may interfere by affecting each others' timing. For cyber-physical system applications, such interference can nullify many of the advantages offered by parallel hardware and can enormously complicate synthesis of software from models. This paper examines what changes need to be made at lower levels of abstraction to support temporal isolation for effective software synthesis. We discuss techniques at the microarchitecture level, in the memory hierarchy, in on-chip communication, and in the instruction-set architecture that can facilitate temporal isolation. Dai N. Bui, Edward A. Lee, Isaac Liu, Hiren D. Patel, Jan Reineke 0001 |
DAC | 4 |
| 2011 | Abstract state machines as an intermediate representation for high-level synthesisabstractThis work presents a high-level synthesis methodology that uses the abstract state machines (ASMs) formalism as an intermediate representation (IR). We perform scheduling and allocation on this IR, and generate synthesizable VHDL. We have the following advantages when using ASMs as an IR: 1) it allows the specification of both sequential and parallel computation, 2) it supports an extension of a clean timing model based on an interpretation of the sequential semantics, and 3) it has well-defined formal semantics, which allows the integration of formal methods into the methodology. While we specify our designs using ASMs, we do not mandate this. Instead, one can create translators that convert the algorithmic specifications from C-like languages into their equivalent ASM specifications. This makes the hardware synthesis transparent to the designer. We experiment our methodology with examples of a FIR, microprocessor, and an edge detector. We synthesize these designs and validate our designs on an FPGA. Rohit Sinha 0001, Hiren D. Patel |
DATE | 2 |
| 2011 | Extending Force-Directed Scheduling with Explicit Parallel and Timed Constructs for High-Level SynthesisabstractThis work extends force-directed scheduling (FDS) to support specification constructs that express parallelism and timing behaviours. We select the FDS algorithm because it maximizes the amount of resource sharing, and it naturally supports constructs for parallelism. However, timed constructs are not supported. As a result, we propose timed FDS (TFDS) that optimizes over parallel, timed and untimed constructs. In doing so, we make the following four contributions: 1) we extend the definition of control data flow graphs (CDFGs) to define timed CDFGs (TCDFGs), 2) we define a scheduling algorithm for timed constructs called TIME, 3) we extend the definition of mobility used in FDS, and 4) we present optimizations for a composition of parallel, timed and untimed constructs to better aid FDS. We implement our extensions in a high-level synthesis framework based on the abstract state machine formalism, and we generate synthesizable VHDL. We experiment with several examples such as FIR, edge detector, and a differential equation solver, and target them onto an Altera DE2 FPGA. Some of these experiments show improvements of up to 52% in circuit area when compared to their unoptimized counterparts. Rohit Sinha 0001, Hiren D. Patel |
FCCM | 2 |
| 2011 | An authorization scheme for version control systemsabstractWe present gitolite, an authorization scheme for Version Control Systems (VCSes). We have implemented it for the Git VCS. A VCS enables versioning, distributed collaboration and several other features, and is an important context for authorization and access control. Our main consideration behind the design of gitolite is the balance between expressive power, correctness and usability in realistic settings. We discuss our design of gitolite, and in particular the four user-classes in its delegation model, and the administrative actions a user at each class performs. We discuss also our ongoing work on expressing gitolite precisely in first-order logic, to thereby give it a precise semantics and establish correctness properties. gitolite has been adopted in open-source software development, university and industry settings. We discuss our experience with these deployments, and present some performance results related to access enforcement from a real deployment. Sitaram Chamarty, Hiren D. Patel, Mahesh Tripunitara |
SACMAT | 2 |
| 2010 | SCGPSim: a fast SystemC simulator on GPUsabstractThe main objective of this paper is to speed up the simulation performance of SystemC designs at the RTL abstraction level by exploiting the high degree of parallelism afforded by today's general purpose graphics processors (GPGPUs). Our approach parallelizes SystemC's discrete-event simulation (DES) on GPGPUs by transforming the model of computation of DES into a model of concurrent threads that synchronize as and when necessary. Unlike the cooperative threading model employed in the SystemC reference implementation, our threading model is capable of executing in parallel on the large number of simple processing units available on GPUs. Our simulation infrastructure is called SCGPSim and it includes a source-to-source (S2S) translator to transform synthesizable SystemC models into parallelly executable programs targeting an NVIDIA GPU. The translator retains the simulation semantics of the original designs by applying semantics preserving transformations. The resulting transformed models mapped onto the massively parallel architecture of GPUs improve simulation efficiency quite substantially. Preliminary experiments with varying-sized examples such as AES, ALU, and FIR have shown simulation speed-ups ranging from 30× to 100×. Considering that our transformations are not yet optimized, we believe that optimizing them will improve the simulation performance even further. Mahesh Nanjundappa, Hiren D. Patel, Bijoy Antony Jose, Sandeep K. Shukla |
ASP-DAC | 2 |
| 2010 | Deploying Hard Real-Time Control Software on Chip-MultiprocessorsabstractDeploying real-time control systems software on multiprocessors requires distributing tasks on multiple processing nodes and coordinating their executions using a protocol. One such protocol is the discrete-event (DE) model of computation. In this paper, we investigate distributed discrete-event (DE) with null-message protocol (NMP) on a multicore system for real-time control software. We illustrate analytically and experimentally that even with the null-message deadlock avoidance scheme in the protocol, the system can deadlock due to inter-core message dependencies. We identify two central reasons for such deadlocks: 1) the lack of an upper-bound on packet transmission rates and processing capability, and 2) an unknown upper-bound on the communication network delay. To address these, we propose using architectural features such as timing control and real-time network-on-chips to prevent such message-dependent deadlocks. We employ these architectural techniques in conjunction with a distributed DE strategy called PTIDES for an illustrative car wash station example and later follow it with a more realistic tunnelling ball device application. Dai N. Bui, Hiren D. Patel, Edward A. Lee |
RTCSA | 2 |
| 2009 | A disruptive computer design idea: Architectures with repeatable timingabstractThis paper argues that repeatable timing is more important and more achievable than predictable timing. It describes microarchitecture approaches to pipelining and memory hierarchy that deliver repeatable timing and promise comparable or better performance compared to established techniques. Specifically, threads are interleaved in a pipeline to eliminate pipeline hazards, and a hierarchical memory architecture is outlined that hides memory latencies. Stephen A. Edwards, Edward A. Lee, Isaac Liu, Hiren D. Patel, Martin Schoeberl |
ICCD | 5 |
| 2008 | Exploring power management in multi-core systemsabstractPower dissipation has become a critical design metric in microprocessor-based system design. In a multi-core system, running multiple applications, power and performance can be dynamically traded off using an integrated power management (PM) unit. This PM unit monitors the performance and power of each core and dynamically adjusts the individual voltages and frequencies in order to maximize system performance under a given power budget (usually set by the operating system). This paper presents a performance and power analysis methodology, featuring a simulation model for multi-core systems that can be easily reconfigured for different scenarios and a PM infrastructure for the exploration and analysis of PM algorithms. Two algorithms have been implemented: one for discrete and one for continuous power modes based on non-linear programming. Extensive experiments are reported, illustrating the effect of power management both at the core and the chip level. Reinaldo A. Bergamaschi, Guoling Han, Alper Buyuktosunoglu, Hiren D. Patel, Indira Nair, Gero Dittmann, Geert Janssen, Nagu R. Dhanwada, Pradip Bose, John A. Darringer |
ASP-DAC | 4 |
| 2008 | Predictable programming on a precision timed architectureabstractIn a hard real-time embedded system, the time at which a result is computed is as important as the result itself. Modern processors go to extreme lengths to ensure their function is predictable, but have abandoned predictable timing in favor of average-case performance. Real-time operating systems provide timing-aware scheduling policies, but without precise worst-case execution time bounds they cannot provide guarantees. Ben Lickly, Isaac Liu, Hiren D. Patel, Stephen A. Edwards, Edward A. Lee |
CASES | 4 |
| 2008 | An Automated Mapping of Timed Functional Specification to a Precision Timed ArchitectureabstractMost common real-time embedded programming languages provide a means to specify functionality; however, they have few constructs to specify precise timing constraints. LabVIEW is one example of a graphical programming language that supports timing specifications in the form of timed-loops. In this work, we present a plug-in for LabVIEW Embedded that maps the LabVIEW G graphical programming language and its timing specifications to the PREcision Timed machine (PRET), an architecture that exposes timing instructions in its instruction set architecture. We demonstrate the use of the plug-in with a simple producer/consumer example that uses timing to enforce synchronization. Shanna-Shaye Forbes, Hiren D. Patel, Edward A. Lee, Hugo A. Andrade |
DS-RT | 2 |
| 2008 | On the Deterministic Multi-threaded Software Synthesis from Polychronous SpecificationsabstractIn order to exploit the emerging multi-core processors, creating multi-threaded applications is going to be a necessity. However, resolving concurrency, synchronization, and coordination issues, and tackling the non-determinism germane in multi-threaded software is extremely difficult. Ensuring deterministic behavior and correctness with respect to the specification is necessary for safe execution of such code. It is desirable to synthesize multi-threaded code from formal specifications using a provably 'correct-by- construction' approach. In the past, reasonable success has been achieved in the 'correct-by-construction' sequential software synthesis for embedded reactive systems from synchronous programming models. Here we target deterministic multi-threaded software synthesis from deterministic specifications, such that the behavior of the code is semantically equivalent to that of the specification. We choose the polychronous model of computation for specification because (i) such specifications are multi-rate, reactive, concurrent and can be made deterministic through constraints on the environment, and (ii) formal verification methodologies and tools exist for such specifications. In this paper, we analyze under what condition a polychronous specification can be synthesized into multi-threaded C-code preserving its semantics. We also discuss how the synchronous data flow graph structure for a polychronous specification can be used to infer the threading structure of the resulting C-code. Bijoy Antony Jose, Sandeep K. Shukla, Hiren D. Patel, Jean-Pierre Talpin |
MEMOCODE | 3 |
| 2008 | On Cosimulating Multiple Abstraction-Level System-Level ModelsabstractSystemC's growing community for system-level design exploration is a result of SystemC's capability of modeling at register transfer level (RTL) and above RTL abstraction levels. However, a synthesis path from SystemC at abstraction layers above RTL is still in its infancy. A recent extension of SystemC, which is called Bluespec-SystemC electronic system level (BS-ESL), counters this difficulty with itsmodelofcomputationemploying atomic rule-based specifications and synthesis to Verilog. In order to simulate a model consisting of one part designed in SystemC and another using BS-ESL, we require an interoperability semantics and implementation of such a semantics. To illustrate the problem, we formalize the simulation semantics of BS-ESL and discrete-event simulation of RTL SystemC, and provide a solution based on this formalization. Hiren D. Patel, Sandeep K. Shukla |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Model-driven Validation of SystemC DesignsabstractFunctional test generation for dynamic validation of current system level designs is a challenging task. Manual test writing or automated random test generation techniques are often used for such validation practices. However, directing tests to particular reachable states of a SystemC model is often difficult, especially when these models are large and complex. In this work, we present a model-driven methodology for generating directed tests that take the SystemC model under validation to specific reachable states. This allows the validation to uncover very specific scenarios which lead to different corner cases. Our formal modeling is done entirely within the Microsoft SpecExplorer tool to describe the specification of the system under validation in the notation of AsmL. We also exploit SpecExplorer's abilities for state space exploration for our test generations, and its APIs for connecting the model to implementation programs to drive the validation of SystemC models with the generated test cases. Hiren D. Patel, Sandeep K. Shukla |
DAC | 1 |
| 2007 | Tackling an abstraction gap: co-simulating SystemC DE with bluespec ESLabstractThe growing SystemC community for system level design exploration is a result of SystemC's capability of modeling at RTL and above RTL abstraction levels. However, managing shared state concurrency using multi-threading in large SystemC models is error prone. A recent extension of SystemC called Bluespec-SystemC (BS-ESL) counters this difficulty with its model of computation employing atomic rule-based specifications. However, for simulating a model that is partly designed in SystemC and partly using BS-ESL, an interoperability semantics and implementation of such a semantic is required. This paper views the interoperability problem as an abstraction gap closure problem. To illustrate the problem, the simulation semantics of BS-ESL and discrete-event simulation of RTL SystemC were formalized and provide a solution based on this formalization Hiren D. Patel, Sandeep K. Shukla |
DATE | 1 |
| 2007 | Heterogeneous Behavioral Hierarchy Extensions for SystemCabstractSystem level design methodology and language support for high-level modeling enhances productivity for designing complex embedded systems. For an effective methodology, efficiency of simulation and a sound refinement-based implementation path are also necessary. Although some of the recent system level design languages (SLDLs) such as SystemC, SystemVerilog, or SpecC have features for system level abstractions, several essential ingredients are missing from these. We consider: 1) explicit support for multiple models of computation (MoCs) or heterogeneity so that distributed reactive embedded systems with hardware and software components can be easily modeled; 2) the ability to build complex behaviors by hierarchically composing simpler behaviors and the ability to distinguish between structural and heterogeneous behavioral hierarchy; and 3) hierarchical composition of behaviors that belong to distinct MoCs, as essential for successful SLDLs. One important requirement for such an SLDL should be that the simulation semantics are compositional, and hence no flattening of hierarchically composed behaviors are needed for simulation. In this paper, we show how we designed SystemC extensions to facilitates for heterogeneous behavioral hierarchy, compositional simulation semantics, and a simulation kernel that shows up to 40% more efficient than standard SystemC simulation Hiren D. Patel, Sandeep K. Shukla, Reinaldo A. Bergamaschi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | EWD: A metamodeling driven customizable multi-MoC system modeling frameworkabstractWe present the EWD design environment and methodology, a modeling and simulation framework suited for complex and heterogeneous embedded systems with varying degrees of expressibility and modeling fidelity. This environment promotes the use of multiple models of computation (MoCs) to support heterogeneity and metamodeling for conformance tests of syntactic and static semantics during the process of modeling. Therefore, EWD is a multiple MoC modeling and simulation framework that ensures conformance of the MoC formalisms during model construction using a metamodeling approach. In addition, EWD provides a suite of translation tools that generate executable models for two simulation frameworks to demonstrate its language-independent modeling framework. The EWD methodology uses the Generic Modeling Environment for customization of the MoC-specific modeling syntax into a visual representation. To embed the execution semantics of the MoCs into the models, we have built parsing and translation tools that leverage an XML-based interoperability language. This interoperability language is then translated into executable Standard ML or Haskell models that can also be analyzed by existing simulation frameworks such as SML-Sys or ForSyDe. In summary, EWD is a metamodeling driven multitarget design environment with multi-MoC modeling capability. Deepak Mathaikutty, Hiren D. Patel, Sandeep K. Shukla, Axel Jantsch |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2006 | Heterogeneous behavioral hierarchy for system level designsabstractEnhancing productivity for designing complex embedded systems requires system level design methodology and language support for capturing complex design in high level models. For an effective methodology, efficiency of simulation and a sound refinement based implementation path are also necessary. Although some of the recent system level design languages for system level abstractions, several essential ingredients are missing from these. We consider (i) explicit support for multiple models of computation (MoCs) or heterogeneity; (ii) the ability to build complex behaviors by hierarchically composing simpler behaviors; and (iii) hierarchical composition of behaviors that belong to distinct models of computation, as essential for successful SLDLs. These render an SLDL with modeling fidelity that exploits both heterogeneity and hierarchy and allows for simpler modeling and efficient simulation. One important requirement for such an SLDL should be that the simulation semantics be also compositional, and hence no flattening of hierarchically composed behaviors be needed for simulation. In this paper we show how we designed SystemC extensions to provide facilities for heterogeneous behavioral hierarchy, compositional simulation semantics, and implemented a simulation kernel which we show experimentally as up to 50% more efficient than standard SystemC simulation Hiren D. Patel, Sandeep K. Shukla, Reinaldo A. Bergamaschi |
DATE | 1 |
| 2006 | A rule-based model of computation for SystemC: integrating SystemC and Bluespec for co-designabstractBluespec's rule-based model of computation (MoC) for hardware concurrency has gained attention for several reasons. From its basis in term rewriting systems, rules have the property of atomicity, which improves correctness by construction, particularly in large-scale concurrency with finegrained, dynamic resource sharing (typical in complex hardware). Rule-based interface methods extend atomicity across module boundaries, have a natural transactional reading, and precisely and formally characterize resource-sharing constraints. All this can be synthesized to hardware with competitive quality. SystemC expresses concurrency with threading and events, just like RTL, where it is difficult to deal with fine-grain concurrency and resource sharing. Further, there is no systematic methodology for module composition. Thus, while SystemC is suitable for very coarse modeling and for embedded software development, its limitations make it difficult to model correct by construction hardware systems accurately. In this paper, we show how to integrate Bluespec's rule-based MoC into SystemC. We augment SystemC modules with rules and rule-based interface methods, and augment the SystemC simulation kernel with a rule execution kernel. The integration is augmentative in that a model can contain both rule-based modules (where hardware accuracy is desired) as well as core SystemC or TLM modules (for embedded software, instruction-set simulators, existing SystemC IP, or pure behavioral models), thus providing the advantages of each MoC where appropriate Hiren D. Patel, Sandeep K. Shukla, Elliot Mednick, Rishiyur S. Nikhil |
MEMOCODE | 1 |
| 2006 | CARH: service-oriented architecture for validating system-level designsabstractExisting system-level design languages (SLDLs) and frameworks mainly provide a modeling and a simulation framework. However, there is an increasing demand for supporting tools to aid designers in quick and faster design space and architectural exploration. As a result, numerous tools such as integrated development environments (IDEs) and others that help in debugging, visualization, validation, and verification are commonly employed by designers. As with most tools, they are targeted for a specific purpose, making it difficult for designers to possess all desired features from one particular tool. Only public-domain tools can be easily extended or interfaced with other existing tools, which a lot of the existing commercial tools do not promote. Having an extendable framework allows designers to implement their own desirable features and incorporate them into their framework. However, for technology reuse and transfer, it is important to have a tidy infrastructure for interfacing the extension with the framework, such that the added solution is not highly coupled with the environment, making distribution and deployment to other frameworks difficult, if not impossible. This requires a plug-and-play framework where features can be easily integrated. These issues of extendibility, deployment, and the inadequacies in SLDLs and frameworks are tackled by presenting a service-oriented architecture for validating SLDs for SystemC, called CARH, We code name our software systems after famous computer scientists. CARH which uses a variety of open-source technologies such as Doxygen, Apache's Xerces extensible markup language parsers, SystemC, and the adaptive communication environment (ACE) object request broker. Hiren D. Patel, Deepak Mathaikutty, David Berner, Sandeep K. Shukla |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | SystemCXML: An Exstensible SystemC Front end Using XML
David Berner, Jean-Pierre Talpin, Hiren D. Patel, Deepak Mathaikutty, Sandeep K. Shukla |
FDL | 3 |
| 2005 | Modelling Environment for Heterogeneous Systems based on MoCs
Deepak Mathaikutty, Hiren D. Patel, Sandeep K. Shukla, Axel Jantsch |
FDL | 2 |
| 2005 | Towards Behavioural Hierarchy Extensions for SystemC
Hiren D. Patel, Sandeep K. Shukla |
FDL | 1 |
| 2005 | Towards a heterogeneous simulation kernel for system-level models: a SystemC kernel for synchronous data flow modelsabstractAs SystemC gains popularity as a modeling language of choice for system-on-chip (SoC) designs, heterogeneous modeling in SystemC and efficient simulation become increasingly important. However, in the current reference implementation, all SystemC models are simulated through a nondeterministic discrete-event (DE) simulation kernel that schedules events at run time mimicking other models of computation (MoCs) using DE, which may get cumbersome. This sometimes results in too many delta cycles hindering the simulation performance of the model. SystemC also uses this simulation kernel as the target simulation engine. This makes it difficult to express different MoCs naturally in SystemC. In an SoC model, different components may need to be naturally expressible in different MoCs. These components may be amenable to static scheduling-based simulation or other presimulation optimization techniques. The goal is to create a simulation framework for heterogeneous SystemC models and to gain efficiency and ease of use within the framework of SystemC reference implementation. In this paper, a synchronous data flow (SDF) kernel extension for SystemC is introduced. Experimental results showing improvement in simulation time are also presented. Hiren D. Patel, Sandeep K. Shukla |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | A Functional Programming Framework of Heterogeneous Model of Computation for System Design
Deepak Mathaikutty, Hiren D. Patel, Sandeep K. Shukla |
FDL | 2 |
| 2004 | Towards a heterogeneous simulation kernel for system level models: a SystemC kernel for synchronous data flow modelsabstractAs SystemC gains popularity as a modeling language of choice for system-on-chip (SOC) designs, heterogeneous modeling in SystemC and efficient simulation become increasingly important. However, in the current reference implementation, all SystemC models are simulated through a non-deterministic Discrete-Event simulation kernel, which schedules events at run-time. This sometimes results in too many delta cycles hindering the simulation performance of the model. The SystemC language also seems to target this simulation kernel as the target simulation engine. This makes it difficult to express different Models Of Computation naturally in SystemC. In an SOC model, different components may need to be naturally expressible in different Models Of Computations. Some of these components may be amenable to static scheduling based simulation or other pre-simulation optimization techniques. Our goal is to create a simulation framework for heterogeneous SystemC models, to gain efficiency and ease of use within the framework of SystemC reference implementation. In this paper, we focus on Synchronous Data Flow (SDF) models, where the rates of data produced and consumed by a data flow node/block are known a priori. In digital signal processing (DSP) applications where relative sample rates are specified for each DSP component, such models are quite common. Compile time knowledge of these rates allow the use of static scheduling resulting in significant improvement in simulation efficiency. We describe an alternate SystemC kernel that exploits such static scheduling of SDF models. Our experiments show improvement in simulation time over the original models and over the latest efficiency results from [20]. Hiren D. Patel, Sandeep K. Shukla |
ACM Great Lakes Symposium on VLSI | 1 |