EDBT 2026 Demo / reviewers in the wild / expert
Anirudh M. Kaushik
dblp:137/4253 · also Anirudh Mohan Kaushik
· DBLP profile ↗
16ranked-venue papers
7as first author
9since 2021 · last 2024
0000-0002-8347-0109ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | High Performance and Predictable Shared Last-level Cache for Safety-Critical SystemsabstractWe propose ZeroCost-LLC (ZCLLC), a novel shared inclusive last-level cache (LLC) design for timing predictable multi-core platforms that offers lower worst-case latency (WCL) when compared with a traditional shared inclusive LLC design. ZCLLC achieves low WCL by eliminating certain memory operations in the form of cache line invalidations across the cache hierarchy that are a consequence of a core’s memory request that misses in the cache hierarchy and when there is no vacant entry in the LLC to accommodate the fetched data for this request. In addition to low WCL, ZCLLC offers performance benefits in the form of additional caching capacity and unlike state-of-the-art approaches, ZCLLC does not impose any constraints on its usage across multiple cores. In this work, we describe the impact of LLC cache line invalidations on the WCL and systematically build solutions to eliminate these invalidations resulting in ZCLLC. We also present ZCLLC-OPT, an optimized variant of ZCLLC that offers lower WCL and improved average-case performance over ZCLLC. We apply optimizations to the shared bus arbitration mechanism and extend the micro-architecture of ZCLLC to allow for overlapping memory requests to the main memory. Our analysis reveals that the analytical WCL of a memory request under ZCLLC-OPT is 87.0%, 93.8%, and 97.1% lower than that under state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively. ZCLLC-OPT shows average-case performance speedups of 1.89×, 3.36×, and 6.24× compared with the state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively. When compared with the original ZCLLC that does not have any optimizations, ZCLLC-OPT shows lower analytical WCLs that are 76.5%, 82.6%, and 86.2% lower compared with ZCLLC-NORMAL for 2, 4, and 8 cores, respectively. Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | ZeroCost-LLC: Shared LLCs at No Cost to WCLabstractZeroCost-LLC (ZCLLC) is a shared inclusive lastlevel cache (LLC) architecture for predictable multicore platforms that does not incur additional cost to the worst-case latency (WCL) of memory requests when compared to the memory hierarchy without an LLC. Thus, the WCL remains the same as without an LLC in the memory hierarchy, but with the performance benefits of having an LLC, in the form of additional caching capacity. ZCLLC achieves this by eliminating all cache line invalidations, and proactively updating the main memory with cache lines to preserve an important vacancy invariant. Furthermore, ZCLLC does not impose any constraints on the way the LLC is used unlike other approaches such as LLC partitioning. Our analysis reveals that the WCL is 55.6%, 68.0%, and 80.2% lower, and the performance is 2.4%, 7.2%, and 25.6% better than the state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively. Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 2 |
| 2023 | Predictable GPU Wavefront Splitting for Safety-Critical SystemsabstractWe present a predictable wavefront splitting (PWS) technique for graphics processing units (GPUs). PWS improves the performance of GPU applications by reducing the impact of branch divergence while ensuring that worst-case execution time (WCET) estimates can be computed. This makes PWS an appropriate technique to use in safety-critical applications, such as autonomous driving systems, avionics, and space, that require strict temporal guarantees. In developing PWS on an AMD-based GPU, we propose microarchitectural enhancements to the GPU, and a compiler pass that eliminates branch serializations to reduce the WCET of a wavefront. Our analysis of PWS exhibits a performance improvement of 11% over existing architectures with a lower WCET than prior works in wavefront splitting. Artem Klashtorny, Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | Automatic Construction of Predictable and High-Performance Cache Coherence Protocols for Multicore Real-Time SystemsabstractPredictable hardware cache coherence is a viable shared data communication mechanism between cores for multicore real-time platforms. Prior works have established that predictable hardware cache coherence protocols offer significant performance advantages over alternative predictable data communication mechanisms while ensuring predictability. Unlike alternative predictable data communication mechanisms, designing predictable cache coherence protocols is nontrivial as it requires detailed understanding of the impact of different memory activity patterns to shared data for predictable and coherent data communication. Furthermore, designing predictable cache coherence protocols that deliver high average-case performance is even more challenging as it entails identifying opportunities such that a core’s access to a data is not stalled in the presence of interleaving memory activity from other cores to the same data. To this end, we present SYNTHIA, an open and automated tool for synthesizing predictable and high-performance snooping bus-based cache coherence protocols for multicore platforms deployed in real-time systems. SYNTHIA automates the complex analysis associated with designing predictable and high-performance cache coherence protocols, and constructs the complete protocol implementation (coherence states and transitions) that achieve predictability and performance. We use SYNTHIA to construct complete protocol implementations from simple specifications of common protocols (modified-shared-invalid (MSI), MESI, and MOESI protocols) and a predictable variant of the MESIF cache coherence protocol, which was recently found to be deployed in an existing multicore platform designed for real-time platforms. We validated the correctness, predictability, and performance guarantees of the generated protocol implementations from SYNTHIA using manually implemented versions, and a micro-architectural simulator. Anirudh M. Kaushik, Hiren D. Patel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Automated Synthesis of Predictable and High-Performance Cache Coherence ProtocolsabstractWe present SYNTHIA, an open and automated tool for synthesizing predictable and high-performance snooping bus-based cache coherence protocols for multi-core processors in multi-processor system-on-chips (MPSoCs) deployed in real-time systems. SYNTHIA automates the complex analysis associated with designing predictable and high-performance cache coherence protocols, and constructs new states (transient states) and corresponding transitions that achieve predictability and performance. We use SYNTHIA to construct complete protocol implementations from simple specifications of common protocols (MSI, MESI, and MOESI protocols). We validated the correctness, predictability, and performance guarantees of the generated protocol implementations from SYNTHIA using manually implemented versions, and a micro-architectural simulator. Anirudh M. Kaushik, Hiren D. Patel |
DATE | 1 |
| 2021 | A Systematic Approach to Achieving Tight Worst-Case Latency and High-Performance Under Predictable Cache CoherenceabstractPredictable hardware cache coherence is an attractive data communication mechanism between safety-critical tasks deployed on real-time multi-core platforms due to its predictability and high-performance benefits. However, from a worst-case analysis standpoint, alternative data communication mechanisms appear in favorable light for adoption in real-time multi-core platforms. This is because alternative data communication mechanisms such as cache bypassing offer tighter worstcase latency (WCL) bounds for memory requests compared to predictable hardware cache coherence mechanisms. We present a systematic approach towards designing predictable cache coherence mechanisms that offer tight WCL and high-performance. Our approach consists of a formal framework that concisely captures the key reasons behind the high WCL in existing predictable cache coherence mechanisms. Guided by this formal framework, we describe one technique that employs micro-architectural extensions and protocol changes to achieve tight WCL and high-performance. We apply this technique to two existing cache coherence mechanisms. Our evaluation shows that the new cache coherence mechanisms resulting from our technique have the same tight WCL as alternative mechanisms, and still maintain a significant average-case performance advantage (up to 5× speedup) over the alternative mechanisms. Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 1 |
| 2021 | A Hardware Platform for Exploring Predictable Cache Coherence Protocols for Real-time MulticoresabstractThis work presents MapleBoard: a set of opensource hardware tools to implement predictable cache coherence protocols in hardware. MapleBoard consists of the following: (1) a novel domain-specific language (DSL) for specifying coherence protocols and synthesizing the corresponding hardware; and, (2) a real-time multicore hardware platform that seamlessly integrates the coherence protocols synthesized from the DSL. This platform has a memory hierarchy, a real-time bus interconnect between cores, and various predictable arbiters. As a demonstration of MapleBoard's efficacy, we explore hardware implementations of data bus organizations, and their impact on the worstcase communication latency (WCL). An important discovery we make is that a dedicated data bus (DDB) organization that allows bidirectional data communication offers lower analytical WCL bounds than any state-of-the-art predictable cache coherence protocols. The analytical WCL bounds are improved by 84%, 90% and 94% for 2-core, 4-core and 8-core systems respectively compared to prior works. We synthesize MapleBoard on the Xilinx Virtex Ultrascale+ VCU1525 board, and validate using both synthetic workloads and SPLASH-2 benchmark suite. Zhuanhao Wu, Anirudh M. Kaushik, Paulos Tegegn, Hiren D. Patel |
RTAS | 2 |
| 2021 | Gretch: A Hardware Prefetcher for Graph AnalyticsabstractData-dependent memory accesses (DDAs) pose an important challenge for high-performance graph analytics (GA). This is because such memory accesses do not exhibit enough temporal and spatial locality resulting in low cache performance. Prior efforts that focused on improving the performance of DDAs for GA are not applicable across various GA frameworks. This is because (1) they only focus on one particular graph representation, and (2) they require workload changes to communicate specific information to the hardware for their effective operation. In this work, we propose a hardware-only solution to improving the performance of DDAs for GA across multiple GA frameworks. We present a hardware prefetcher for GA called Gretch, that addresses the above limitations. An important observation we make is that identifying certain DDAs without hardware-software communication is sensitive to the instruction scheduling. A key contribution of this work is a hardware mechanism that activates Gretch to identify DDAs when using either in-order or out-of-order instruction scheduling. Our evaluation shows that Gretch provides an average speedup of 38% over no prefetching, 25% over conventional stride prefetcher, and outperforms prior DDAs prefetchers by 22% with only 1% increase in power consumption when executed on different GA workloads and frameworks. Anirudh M. Kaushik, Gennady Pekhimenko, Hiren D. Patel |
ACM Trans. Archit. Code Optim. | 1 |
| 2021 | Designing Predictable Cache Coherence Protocols for Multi-Core Real-Time SystemsabstractThis article addresses the challenge of allowing simultaneous and predictable accesses to shared data on multi-core systems. We propose a collection of predictable cache coherence protocols, which mandate the use of certain design invariants to ensure predictability. In particular, we enforce these invariants by augmenting the classic modify-share-invalid (MSI) protocol and modify-exclusive-share-invalid (MESI) protocol. This allows us to derive worst-case latency bounds on the resulting predictable MSI (PMSI) and predictable MESI (PMESI) protocols. Our analysis shows that while the arbitration latency scales linearly, the coherence latency scales quadratically with the number of cores, which emphasizes the importance of accounting for cache coherence effects on latency bounds. We implement PMSI and PMESI in a detailed micro-architectural simulator, and execute SPLASH-2 and synthetic workloads. Results show that our approach is always within the analytical worst-case latency bounds, and that PMSI and PMESI improve average-case performance by up to 4× over cache bypassing mechanisms that disallow caching of shared data in the cores’ private caches. PMSI and PMESI have average slowdowns of 1.45× and 1.46× compared to conventional MSI and MESI protocols, respectively. Anirudh M. Kaushik, Mohamed Hassan 0002, Hiren D. Patel |
IEEE Trans. Computers | 1 |
| 2019 | CARP: A Data Communication Mechanism for Multi-core Mixed-Criticality SystemsabstractWe present CARP, a predictable and high-performance data communication mechanism for multi-core mixed-criticality systems (MCS). CARP is realized as a hardware cache coherence protocol that enables communication between critical and non-critical tasks while ensuring that non-critical tasks do not interfere with the safety requirements of critical tasks. The key novelty of CARP is that it is criticality-aware, and hence, handles communication patterns between critical and non-critical tasks appropriately. We derive the analytical worst-case latency bounds for requests using CARP and note that the observed per-request latencies are within the analytical worst-case latency bounds. We compare CARP against prior data communication mechanisms using synthetic and SPLASH-2 benchmarks. Our evaluation shows that CARP improves the average-case performance of MCS compared to prior data communication mechanisms, while maintaining the safety requirements of critical tasks. Anirudh M. Kaushik, Paulos Tegegn, Zhuanhao Wu, Hiren D. Patel |
RTSS | 1 |
| 2019 | Enabling Predictable, Simultaneous and Coherent Data Sharing in Mixed Criticality SystemsabstractEmerging embedded systems deployed in the automotive and avionics domains execute applications with different criticalities, comprising what is known as Mixed Criticality Systems (MCS). Applications in MCS often share data between tasks (coming from sensors for instance). Data sharing is challenging because it can lead to increased response times or even unpredictable behaviors if not carefully addressed. Therefore, several prior works in MCS either assumed that tasks do not share data or disallowed it by design. Recent solutions attempt to mitigate the effects of data sharing, albeit by introducing new restrictions on the system either by prohibiting applications from caching shared data or prohibiting the operating system from running tasks with shared data in parallel. We find these solutions also to have limited applicability as they deteriorate system schedulability and prohibit simultaneous access to shared data. In this paper, we propose PENDULUM: a time-based cache coherence protocol to enable simultaneous and predictable access to shared data in MCS. Our evaluation shows that PENDULUM, achieves flexibility and better performance compared to existing solutions, while maintaining system predictability. Nivedita Sritharan, Anirudh M. Kaushik, Mohamed Hassan 0002, Hiren D. Patel |
RTSS | 2 |
| 2018 | Exposing Implementation Details of Embedded DRAM Memory Controllers through Latency-based AnalysisabstractWe explore techniques to reverse-engineer DRAM embedded memory controllers (MCs), including page policies, address mapping, and command arbitration. There are several benefits to knowing this information: They allow tightening worst-case bounds of embedded systems and platform-aware optimizations at the operating system, source-code, and compiler levels. We develop a latency-based analysis, which we use to devise algorithms and C programs to extract MC properties. We show the effectiveness of the proposed approach by reverse-engineering the MC details in the XUPV5-LX110T Xilinx platform. Furthermore, to cover a breadth of policies, we use a simulation framework and document our findings. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | Predictable Cache Coherence for Multi-core Real-Time SystemsabstractThis work addresses the challenge of allowing simultaneous and predictable accesses to shared data on multicore systems. We propose a predictable cache coherence protocol, which mandates the use of certain invariants to ensure predictability. In particular, we enforce these invariants by augmenting the classic modify-share-invalid (MSI) protocol with transient coherence states, and minimal architectural changes. This allows us to derive worst-case latency bounds on predictable MSI (PMSI) protocol. Our analysis shows that while the arbitration latency scales linearly, the coherence latency scales quadratically with the number of cores, which emphasizes that importance of accounting for cache coherence effects on latency bounds. We implement PMSI in gem5, and execute SPLASH-2 and synthetic workloads. Results show that our approach is always within the analytical worst-case latency bounds, and that PMSI improves averagecase performance by up to 4 over the next best predictable alternative. PMSI has average slowdowns of 1.45 and 1.46 compared to MSI and MESI protocols, respectively. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 2 |
| 2015 | Reverse-engineering embedded memory controllers through latency-based analysisabstractWe explore techniques to reverse-engineer properties of DRAM memory controllers (MCs). This includes page policies, address mapping schemes and command arbitration schemes. There are several benefits to knowing this information: they allow analysis techniques to effectively compute worst-case bounds, and they allow customizations to be made in software for predictability. We develop a latency-based analysis, and use this analysis to devise algorithms for micro-benchmarks to extract properties of MCs. In order to cover a breadth of page policies, address mappings and command arbitration schemes, we explore our technique using a micro-architecture simulation framework and document our findings. Mohamed Hassan 0002, Anirudh M. Kaushik, Hiren D. Patel |
RTAS | 2 |
| 2013 | On the use of GP-GPUs for accelerating compute-intensive EDA applicationsabstractGeneral purpose graphics processing units (GP-GPUs) have recently been explored as a new computing paradigm for accelerating compute-intensive EDA applications. Such massively parallel architectures have been applied in accelerating the simulation of digital designs during several phases of their development - corresponding to different abstraction levels, specifically: (i) gate-level netlist descriptions, (ii) register-transfer level and (iii) transaction-level descriptions. This embedded tutorial presents a comprehensive analysis of the best results obtained by adopting GP-GPUs in all these EDA applications. Valeria Bertacco, Debapriya Chatterjee, Nicola Bombieri, Franco Fummi, Sara Vinco, Anirudh M. Kaushik, Hiren D. Patel |
DATE | 6 |
| 2013 | Systemc-clang: An open-source framework for analyzing mixed-abstraction SystemC models
Anirudh M. Kaushik, Hiren D. Patel |
FDL | 1 |