Zhuanhao Wu

dblp:258/6306 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0003-3272-062XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 8 first-author · 11 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Consistency-Aware and Predictable Memory Processing for Safety-Critical Out-of-Order Multicores
abstract
We introduce an approach that facilitates predictable processing of multiple outstanding memory requests in safety-critical out-of-order multicores. A primary challenge addressed by this work is ensuring that multiple outstanding memory requests maintain a memory consistent model while minimizing a low worst-case latency. Adhering to a memory consistency model is crucial for ensuring the correctness of programs executed on such multicores. Our approach, termed predictable processing of multiple outstanding requests$(\mathsf{PPP}$), Teverages micro-architectural enhancements to tackle this challenge. Experimental results show that$\mathsf{PPP}$delivers speedups of$2.07 \times, 2.79 \times$, and$3.38 \times$over serialization for 2,4, and 8 cores, while maintaining the worst-case latency.
Zhuanhao Wu, Hiren D. Patel
RTAS1
2024 Exclusive Hierarchies for Predictable Sharing in Last-Level Cache
abstract
This work presents an approach to use a last-level cache (LLC) in a memory hierarchy for cache-coherent real-time multicores that delivers a low worst-case latency (WCL) and higher performance than all of its counterparts. Our approach relies on the key observation that an exclusive memory hierarchy, by definition, eliminates back invalidations, which are one of the largest contributors to the WCL when using inclusive memory hierarchies. However, to the best of our knowledge, there are no prior efforts that ensure the predictability of exclusive hierarchies for cache-coherent multicores. Consequently, in this work, we propose PECC, a predictable exclusive cache coherence mechanism, that achieves a lower average data access latency while providing a low WCL bound that scales linearly in the number of cores. Our evaluation shows that PECC reduces the bound by 6% and improves the average performance by 2.33× over the predictable solution with an inclusive LLC.
Zhuanhao Wu, Rodolfo Pellizzoni, Hiren D. Patel
RTAS2
2024 High Performance and Predictable Shared Last-level Cache for Safety-Critical Systems
abstract
We propose ZeroCost-LLC (ZCLLC), a novel shared inclusive last-level cache (LLC) design for timing predictable multi-core platforms that offers lower worst-case latency (WCL) when compared with a traditional shared inclusive LLC design. ZCLLC achieves low WCL by eliminating certain memory operations in the form of cache line invalidations across the cache hierarchy that are a consequence of a core’s memory request that misses in the cache hierarchy and when there is no vacant entry in the LLC to accommodate the fetched data for this request. In addition to low WCL, ZCLLC offers performance benefits in the form of additional caching capacity and unlike state-of-the-art approaches, ZCLLC does not impose any constraints on its usage across multiple cores. In this work, we describe the impact of LLC cache line invalidations on the WCL and systematically build solutions to eliminate these invalidations resulting in ZCLLC. We also present ZCLLC-OPT, an optimized variant of ZCLLC that offers lower WCL and improved average-case performance over ZCLLC. We apply optimizations to the shared bus arbitration mechanism and extend the micro-architecture of ZCLLC to allow for overlapping memory requests to the main memory. Our analysis reveals that the analytical WCL of a memory request under ZCLLC-OPT is 87.0%, 93.8%, and 97.1% lower than that under state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively. ZCLLC-OPT shows average-case performance speedups of 1.89×, 3.36×, and 6.24× compared with the state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively. When compared with the original ZCLLC that does not have any optimizations, ZCLLC-OPT shows lower analytical WCLs that are 76.5%, 82.6%, and 86.2% lower compared with ZCLLC-NORMAL for 2, 4, and 8 cores, respectively.
Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel
ACM Trans. Embed. Comput. Syst.1
2023 Ditty: Directory-based Cache Coherence for Multicore Safety-critical Systems
abstract
Ditty is a predictable directory-based cache co-herence mechanism for multicore safety-critical systems that guarantees a worst-case latency (WCL) on data accesses. Prior approaches for predictable cache coherence use a shared snooping bus to interconnect cores. This restricts the number of cores in the multicore to typically four or eight due to scalability concerns. Ditty takes a first step towards a scalable cache coherence mechanism that is predictable and one that can support a larger number of cores. In designing Ditty, we propose a coherence protocol and micro-architecture additions to deliver a WCL bound that is lower than a naive approach. Our WCL analysis reveals that the resulting bounds are comparable to state-of-the-art bus-based predictable coherence approaches. We prototype Ditty in hardware and empirically evaluate it on an FPGA. Our evaluation shows the observed WCL is within computed WCL bound for both the synthetic and SPLASH-3 benchmarks. We release our implementation to the public domain.
Zhuanhao Wu, Marat Bekmyrza, Nachiket Kapre, Hiren D. Patel
DATE1
2023 SCCL: An open-source SystemC to RTL translator
abstract
We present SCCL, an open-source tool that translates SystemC designs into synthesizable register-transfer level (RTL). SCCL supports a subset of Accellera's SystemC synthesis standard based on the 2011 revision of C++. We use LLVM's Clang front-end to parse SystemC designs, and a suite of analysis passes to construct a SystemC-specific intermediate abstract syntax tree representation called Hcode. Hcode simplifies translation to other intermediate forms such as FIRRTL as well as direct transcription to SystemVerilog or VHDL. Currently, SCCL provides a translation phase to generate synthesizable SystemVerilog. Distinguishing aspects of SCCL include support for complex templated class descriptions that facilitate concise, parameterized hardware specification; introduction and full support for a new type of synthesizable channel called sc_stream that maps directly to standards such as AXI Stream, and a complete reference implementation targeting the Xilinx Vivado toolchain. We demonstrate SCCL's capabilities with a series of case studies including a highly templated SystemC implementation of the ZFP [1] floating-point codec. All case studies are deployed and executed on a Xilinx Zynq UltraScale+ FPGA platform.
Zhuanhao Wu, Maya B. Gokhale, Hiren D. Patel
FCCM1
2023 ZeroCost-LLC: Shared LLCs at No Cost to WCL
abstract
ZeroCost-LLC (ZCLLC) is a shared inclusive lastlevel cache (LLC) architecture for predictable multicore platforms that does not incur additional cost to the worst-case latency (WCL) of memory requests when compared to the memory hierarchy without an LLC. Thus, the WCL remains the same as without an LLC in the memory hierarchy, but with the performance benefits of having an LLC, in the form of additional caching capacity. ZCLLC achieves this by eliminating all cache line invalidations, and proactively updating the main memory with cache lines to preserve an important vacancy invariant. Furthermore, ZCLLC does not impose any constraints on the way the LLC is used unlike other approaches such as LLC partitioning. Our analysis reveals that the WCL is 55.6%, 68.0%, and 80.2% lower, and the performance is 2.4%, 7.2%, and 25.6% better than the state-of-the-art LLC partition sharing techniques for 2, 4, and 8 cores, respectively.
Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel
RTAS1
2023 Predictable GPU Wavefront Splitting for Safety-Critical Systems
abstract
We present a predictable wavefront splitting (PWS) technique for graphics processing units (GPUs). PWS improves the performance of GPU applications by reducing the impact of branch divergence while ensuring that worst-case execution time (WCET) estimates can be computed. This makes PWS an appropriate technique to use in safety-critical applications, such as autonomous driving systems, avionics, and space, that require strict temporal guarantees. In developing PWS on an AMD-based GPU, we propose microarchitectural enhancements to the GPU, and a compiler pass that eliminates branch serializations to reduce the WCET of a wavefront. Our analysis of PWS exhibits a performance improvement of 11% over existing architectures with a lower WCET than prior works in wavefront splitting.
Artem Klashtorny, Zhuanhao Wu, Anirudh M. Kaushik, Hiren D. Patel
ACM Trans. Embed. Comput. Syst.2
2023 Enhancing Strong PUF Security With Nonmonotonic Response Quantization
abstract
Strong physical unclonable functions (PUFs) provide a low-cost authentication primitive for resource-constrained devices. However, most strong PUF architectures can be modeled through learning algorithms with a limited number of CRPs. In this article, we introduce the concept of nonmonotonic response quantization for strong PUFs. Responses depend not only on which path is faster but also on the distance between the arriving signals. Our experiments show that the resulting PUF has increased security against learning attacks. To demonstrate, we designed and implemented a nonmonotonically quantized ring oscillator-based PUF in 65-nm technology. Measurement results show nearly ideal uniformity and uniqueness with a bit error rate of 13.4% over the temperature range from 0 °C to 50 °C.
Kleber Stangherlin, Zhuanhao Wu, Hiren D. Patel, Manoj Sachdev
IEEE Trans. Very Large Scale Integr. Syst.2
2022 Predictable sharing of last-level cache partitions for multi-core safety-critical systems
abstract
Last-level cache (LLC) partitioning is a technique to provide temporal isolation and low worst-case latency (WCL) bounds when cores access the shared LLC in multicore safety-critical systems. A typical approach to cache partitioning involves allocating a separate partition to a distinct core. A central criticism of this approach is its poor utilization of cache storage. Today's trend of integrating a larger number of cores exacerbates this issue such that we are forced to consider shared LLC partitions for effective deployments. This work presents an approach to share LLC partitions among multiple cores while being able to provide low WCL bounds.
Zhuanhao Wu, Hiren D. Patel
DAC1
2022 ZHW: A Numerical CODEC for Big Data Scientific Computation
abstract
Distributed big data in scientific computing presents a major I/O performance bottleneck when exploiting data paral-lelism. Consumer and producer compute nodes are often throttled by saturated data channels when processing large numerical data. We describe ZHW, a hardware implementation of the ZFP numerical CODEC that can greatly reduce I/O pressure caused by large scientific datasets. Our ZHW design overcomes barriers that have prevented prior ZFP-like hardware accelerators from obtaining maximum compression in their implementations. The SystemC ZHW hardware library is available in an open source public repository. We demonstrate the practicality of ZHW by synthesizing our CODEC on an Ultrascale+ FPGA and analyzing performance.
Michael Barrow, Zhuanhao Wu, Maya B. Gokhale, Hiren D. Patel, Peter Lindstrom 0001
FPT2
2021 A Hardware Platform for Exploring Predictable Cache Coherence Protocols for Real-time Multicores
abstract
This work presents MapleBoard: a set of opensource hardware tools to implement predictable cache coherence protocols in hardware. MapleBoard consists of the following: (1) a novel domain-specific language (DSL) for specifying coherence protocols and synthesizing the corresponding hardware; and, (2) a real-time multicore hardware platform that seamlessly integrates the coherence protocols synthesized from the DSL. This platform has a memory hierarchy, a real-time bus interconnect between cores, and various predictable arbiters. As a demonstration of MapleBoard's efficacy, we explore hardware implementations of data bus organizations, and their impact on the worstcase communication latency (WCL). An important discovery we make is that a dedicated data bus (DDB) organization that allows bidirectional data communication offers lower analytical WCL bounds than any state-of-the-art predictable cache coherence protocols. The analytical WCL bounds are improved by 84%, 90% and 94% for 2-core, 4-core and 8-core systems respectively compared to prior works. We synthesize MapleBoard on the Xilinx Virtex Ultrascale+ VCU1525 board, and validate using both synthetic workloads and SPLASH-2 benchmark suite.
Zhuanhao Wu, Anirudh M. Kaushik, Paulos Tegegn, Hiren D. Patel
RTAS1
2019 Strengthening PUFs using Composition
abstract
We explore the idea of composing PUFs with the intent that the resultant PUF is stronger than the constituent PUFs. Prior work has proposed a construction, which subsequent work has shown to be weak. We revisit this prior construction and observe that it is actually weaker than previously thought when the constituent PUFs are arbiter PUFs. This weakness is demonstrated via our adaptation of the previously proposed Logistic Regression (LR) attack. We then propose new constructions called PUFs-composed-with-PUFs (PoP). In particular, we retain a two-layer construction, but allow the same input to the composite PUF to be input to more than one constituent PUF at the first layer. We explore this family of constructions, with arbiter PUFs serving as the constituent PUFs. In particular, we identify several axes which we can vary, and empirically study the resilience of our constructions compared to the prior construction and one another from the standpoint of LR attacks. As insight into why our family of constructions is stronger, we prove, under some idealized conditions, that the lower-bound on an attacker is indeed higher under our constructions than the upper-bound on an attacker for the prior construction. As such, our work suggests that composition can be a promising approach to strengthening PUFs, contrary to what prior work suggests.
Zhuanhao Wu, Hiren D. Patel, Manoj Sachdev, Mahesh Tripunitara
ICCAD1
2019 CARP: A Data Communication Mechanism for Multi-core Mixed-Criticality Systems
abstract
We present CARP, a predictable and high-performance data communication mechanism for multi-core mixed-criticality systems (MCS). CARP is realized as a hardware cache coherence protocol that enables communication between critical and non-critical tasks while ensuring that non-critical tasks do not interfere with the safety requirements of critical tasks. The key novelty of CARP is that it is criticality-aware, and hence, handles communication patterns between critical and non-critical tasks appropriately. We derive the analytical worst-case latency bounds for requests using CARP and note that the observed per-request latencies are within the analytical worst-case latency bounds. We compare CARP against prior data communication mechanisms using synthetic and SPLASH-2 benchmarks. Our evaluation shows that CARP improves the average-case performance of MCS compared to prior data communication mechanisms, while maintaining the safety requirements of critical tasks.
Anirudh M. Kaushik, Paulos Tegegn, Zhuanhao Wu, Hiren D. Patel
RTSS3