Sanjay Gandham

dblp:294/0918 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-5819-1773ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-author · 11 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 AutoSkewBMT: Autonomously Synthesizing Optimized Integrity Authentication Mechanism for DNN Accelerators
abstract
As domain-specific accelerators for deep neural network (DNN) inference gain popularity due to their performance and flexibility advantages over general-purpose systems, the security of accelerator data in memory has emerged as a significant concern. However, the overhead associated with standard memory security measures, such as encryption and integrity authentication, presents a major challenge for accelerators, particularly given the high throughput demands of typical DNN applications [1]. In this work, we present AutoSkewBMT, a security framework that autonomously generates optimized integrity system configurations to enhance the Bonsai Merkle Tree (BMT)-based integrity authentication workflow for DNN accelerators. The framework leverages a novel and efficient design space generation algorithm to optimally skew the BMT for specific workloads. Configurations generated by AutoSkewBMT outperform recent state-of-the-art solutions by up to 32% on general DNN workloads.
Rakin Muhammad Shadab, Sanjay Gandham, Mingjie Lin
DAC2
2025 FlexTEE: Dynamically Enhancing Metadata Locality Through Affine Address Transformation for Heterogeneous & Secure AI Platforms
abstract
The Ubiquitous adoption of domain-specific acceleration for deep neural networks (DNNs) has exposed them to security threats and memory vulnerabilities. Since performance is critical, DNN accelerators rarely employ high-overhead, authentication-based countermeasures (such as an integrity tree), making them vulnerable to integrity-based memory adversaries [1], [2]. Although recent accelerators incorporate specialized low-overhead integrity solutions, these rely on specific accelerator characteristics, making them incompatible to be used with a processor in a shared secure-memory framework.In this paper, we introduce FlexTEE, a flexible security framework that dynamically adapts to the runtime characteristics of both processor and DNN accelerator to significantly reduce Bonsai Merkle Tree (BMT)-based integrity overheads for the accelerator. The unique memory access patterns of DNN accelerators i.e., the strided access patterns caused by accelerator data tiling lead to frequent metadata accesses from the memory, causing excessive BMT utilization overhead. Leveraging this insight, we propose a novel, processor-transparent, metadata address mapping scheme that reorganizes metadata-data relationship for the accelerator, transforming disjoint metadata accesses into sequential accesses for the encryption engine. This reorganization significantly reduces BMT authentication overhead for DNN accelerators with the same security guarantees. For CPU workloads, FlexTEE achieves comparable performance to conventional processor-TEE implementations with minimal resource overhead. For accelerators, FlexTEE reduces BMT verification costs by up to 87% compared to regular BMT-based accelerators. Additionally, it enhances system throughput by up to 30% on popular DNN models, outperforming state-of-the-art secure accelerators.
Rakin Muhammad Shadab, Sanjay Gandham, Mingjie Lin
ICCAD2
2024 CircuitSeer: RTL Post-PnR Delay Prediction via Coupling Functional and Structural Representation
abstract
Register transfer level (RTL) optimization is a critical design phase that ensures timing closure and performance. Although machine learning (ML) has been utilized to quickly predict post-synthesis delay metrics, estimating post-place and route (PnR) delay remains a significant challenge. This is due to the distinct functionality-preserving characteristics of logic synthesis and the structure-dependent aspect of physical design. Furthermore, Logic Synthesis heavily restructures the netlist, resulting in substantial structural disparities that hinder capturing the post-synthesis netlist structure.
Sanjay Gandham, Joe Walston, Sourav Samanta, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin, Stelios Diamantidis
ICCAD1
2024 SCALE: A Structure-Centric Accelerator for Message Passing Graph Neural Networks
abstract
Message passing paradigm has been widely used in developing complex Graph Neural Network (GNN) models, allowing for concise representations of edge and vertex-wise operations. Despite its pivotal role in theoretical advancement, the respective expression of edge and vertex operations, along with evolving GNN variants and datasets, has inevitably led to enormous computational complexity due to heterogeneous computation kernels. In particular, such inconsistent computation characteristics present new challenges in leveraging intermediate data reuse, ensuring both edge and vertex-wise workload balance, and sustaining system scalability. In this paper, we propose a structurecentric accelerator, SCALE, that can support a variety of message passing GNN models with improved parallelism, data reuse, and scalability. The central idea is to find latent similarities among GNN primitives such as shared dataflow structure, rather than strictly adhering to heterogeneous model structure. This serves as a hinge to homogenize inconsistencies in various GNN computation kernels. To accomplish this concept, SCALE consists of three unique designs, a novel systolic array-like architecture, a degree and vertex-aware scheduling, and a coherent dataflow tailored for fused graph and neural operations. The proposed systolic array-like architecture can support varying dataflows such as all-reduce, of distinct GNN operations improving parallelism, data reuse, and throughput. The degree and vertex-aware scheduling can remedy the workload imbalance encountered in vertex and edge-wise operations. Moreover, the proposed dataflow can unify the data movement of both graph and neural operators without extra communication and storage overheads. Our simulation results show that SCALE achieves 1.82× speedup and 38.9% energy reduction on average over the state-of-the-art GNN accelerators [1]–[4].
Lingxiang Yin, Sanjay Gandham, Mingjie Lin, Hao Zheng 0005
MICRO2
2024 A Secure Computing System With Hardware-Efficient Lazy Bonsai Merkle Tree for FPGA-Attached Embedded Memory
abstract
With high-impact cyber-attacks on the rise, provisioning cybersecurity to the emerging Internet of Things (IoT) systems typically comprising of modern embedded computing platforms becomes significantly more challenging to achieve. Contemporary secure-memory computing stipulates both content encryption and integrity protection that can seriously impede the computing performance and consume excessive amount of hardware resources. In this paper, we focus on hardware-efficient verification of the memory integrity in the mission-critical computing tasks executing on an FPGA-based secure embedded system, effectively mitigating adversarial attacks such as memory buffer replay. We proposed an innovative partitioned parallel cache structure that leverages the unique reconfigurable capability of modern FPGA devices and successfully circumvents the hardware implementation challenges due to the recursiveness that inherently exists in Merkle tree updating schemes. We designed and implemented a new Bonsai Merkle tree (BMT) lazy update controller specifically designed for FPGA to efficiently exploit the parallelism offered by its reconfigurable fabric. Our experimental results for the new system show up to 95x and 149x latency overhead reduction respectively for write and read and up to 17% better throughput in standard benchmarks compared to software-based approach. Critical system performance is also improved with the lowering of average evictions by up to 8%.
Rakin Muhammad Shadab, Sanjay Gandham, Amro Awad, Mingjie Lin
IEEE Trans. Dependable Secur. Comput.3
2023 OCMGen: Extended Design Space Exploration with Efficient FPGA Memory Inference
abstract
Deep learning applications demand high memory storage and computational power to operate on millions of parameters. Field Programmable Gate Arrays (FPGAs), with high compute resources and the ability to store data on-chip in their distributed memory components such as Block RAM (BRAM) and Ultra RAM (URAM), are good candidates to deploy such memory-intensive applications [1]. However, without careful tailoring of the hardware design for a target device, current synthesis tools (e.g., Xilinx Vivado) can severely underutilize these RAM primitives reducing the usable on-chip memory (OCM). Consequently, this forces the accelerator to perform more frequent expensive off-chip accesses, limiting its performance.
Sanjay Gandham, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin
FCCM1
2023 OMT: A Demand-Adaptive, Hardware-Targeted Bonsai Merkle Tree Framework for Embedded Heterogeneous Memory Platform
abstract
Novel flash-based, crash-tolerant, non-volatile memory (NVM) such as Intel's Optane DC memory brings about new and exciting use-case scenarios for both traditional and embedded computing systems involving Field-Programmable Gate Arrays (FPGA). However, NVMs cannot be proper replacement for existing DDR memory modules due to low write endurance and are more well-suited for a hybrid NVM + Volatile memory system. They are also well-known to be vulnerable to different memory-based adversaries that demand the use of a robust authentication method such as Bonsai Merkle Tree. However, typical update process of a BMT (eager update) requires updating the entire update chain frequently, affecting run-time performance even for the data that is not persistence-critical. The latest intermittent BMT update techniques can help provide better real-time throughput, but they lack crash-consistency.
Rakin Muhammad Shadab, Sanjay Gandham, Mingjie Lin
FPGA3
2023 SAGA: Sparsity-Agnostic Graph Convolutional Network Acceleration with Near-Optimal Workload Balance
abstract
Graph Convolutional Networks (GCNs) have shown much promise in resolving sophisticated scientific problems with non-Euclidean data, such as traffic prediction, disease classification, and many others. However, the irregular sparsity of real-world graphs remains a major challenge toward efficient GCN acceleration. In this paper, we propose SAGA, a Sparsity-Agnostic Graph Convolutional Accelerator with near-optimal workload balance. Specifically, it consists of two unique features, an NZ-based scheduling, and a novel accelerator architecture. Unlike conventional GCN accelerators with uneven distribution of sparse matrix, the proposed NZ-based scheduling leverages the metadata encoded in the compression format to enable even distribution of sparse matrix at runtime, thus achieving near-optimal workload balancing. In addition, the proposed architecture, including a task scheduler, an accumulation table, and a partial row accumulation unit, can support the proposed NZ-based scheduling without data preprocessing and reformatting with low overheads. We prototyped the proposed design through FPGAs, and our evaluation results show that SAGA achieves up to$\mathbf{1.56}\times$speedup and$\mathbf{2.05}\times$energy savings on average as compared to the prior art [1].
Sanjay Gandham, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin
ICCAD1
2023 HMT: A Hardware-centric Hybrid Bonsai Merkle Tree Algorithm for High-performance Authentication
abstract
The Bonsai Merkle tree (BMT) is a widely used tree structure for authentication of metadata such as encryption counters in a secure computing system. Common BMT algorithms were designed for traditional Von Neumann architectures with a software-centric implementation in mind and as such, they are predominantly recursive and sequential in nature. However, the modern heterogeneous computing platforms employing Field-Programmable Gate Array (FPGA) devices require concurrency-focused algorithms to fully utilize the versatility and parallel nature of such systems. The recursive nature of traditional BMT algorithms makes them challenging to implement in such hardware-based setups. Our goal for this work is to introduce HMT, a hardware-friendly BMT algorithm that enables the verification and update processes to function independently and provides the benefits of relaxed update while being comparable to the eager update in terms of update complexity. The methodology of HMT contributes both novel algorithmic revisions and innovative hardware techniques to implementing BMT. We mathematically demonstrate the challenges of potentially unbounded recursions in relaxed BMT updates. To solve this problem, we use a partitioned BMT caching scheme that allocates a separate write-back cache for each BMT level—thus allowing for low and fixed upper bounds for dirty evictions compared to the traditional BMT caches. Then we introduce the aforementioned hybrid BMT algorithm that is hardware-targeted, parallel, and relaxes the update depending on BMT cache hit but makes the update conditions more flexible compared to lazy update to save additional write-backs. Deploying this new algorithm, we have designed a new BMT controller with a dataflow architecture including speculative buffers and parallel write-back engines to facilitate performance-enhancing mechanisms (like multiple concurrent authentication and independent updates) that were not possible with the conventional lazy algorithm. Our empirical performance measurements on a Xilinx U200 accelerator FPGA have demonstrated that HMT can achieve up to 7× improvement in bandwidth and 4.5× reduction in latency over lazy-update BMT baseline and up to 14% faster execution in standard benchmarks compared to a state-of-the-art, eager-update BMT solution.
Rakin Muhammad Shadab, Sanjay Gandham, Amro Awad, Mingjie Lin
ACM Trans. Embed. Comput. Syst.3
2022 HMT: A Hardware-Centric Hybrid Bonsai Merkle Tree Algorithm for High-Performance Authentication
abstract
Merkle tree is a widely used tree structure for authentication of data/metadata in a secure system. Even though recent state-of-the art systems use MAC based authentication to protect the actual data, they still use smaller-sized MT, namely Bonsai Merkle Tree (BMT) to protect the metadata such as encryption counters. Common BMT algorithms were designed for traditional von Neumann architecture with software-centric implementations in mind, hence they use a lot of recursions and are often sequential in nature. The predominantly recursive and sequential nature of these traditional BMT algorithms make them largely unsuitable for use and challenging to implement in the modern heterogeneous computing platforms employing Field-Programmable Gate Array (FPGA) devices. Our goal for this work is to introduce HMT, a hardware-friendly BMT algorithm that enables the verification and update processes to function independently and provides the benefits of relaxed update while being comparable to eager update in terms of update complexity. Deploying this new algorithm, we have designed a new BMT controller with a dataflow architecture and speculative buffers that allow multiple parallel authentication on-flight which was not possible with the conventional algorithms. This new MT subsystem enables up to 7x improvement in bandwidth while also exhibiting up to 4.5x reduction in latency over the baseline.
Rakin Muhammad Shadab, Sanjay Gandham, Amro Awad, Mingjie Lin
FPGA3
2022 ARES: Persistently Secure Non-Volatile Memory with Processor-transparent and Hardware-friendly Integrity Verification and Metadata Recovery
abstract
Emerging byte-addressable Non-Volatile Memory (NVM) technology, although promising superior memory density and ultra-low energy consumption, poses unique challenges to achieving persistent data privacy and computing security, both of which are critically important to the embedded and IoT applications. Specifically, to successfully restore NVMs to their working states after unexpected system crashes or power failure, maintaining and recovering all the necessary security-related metadata can severely increase memory traffic, degrade runtime performance, exacerbate write endurance problem, and demand costly hardware changes to off-the-shelf processors. In this article, we designed and implemented ARES, a new FPGA-assisted processor-transparent security mechanism that aims at efficiently and effectively achieving all three aspects of a security triad—confidentiality, integrity, and recoverability—in modern embedded computing. Given the growing prominence of CPU-FPGA heterogeneous computing architectures, ARES leverages FPGA’s hardware reconfigurability to offload performance-critical and security-related functions to the programmable hardware without microprocessors’ involvement. In particular, recognizing that the traditional Merkle tree caching scheme cannot fully exploit FPGA’s parallelism due to its sequential and recursive function calls, we (1) proposed a Merkle tree cache architecture that partitions a unified cache into multiple levels with parallel accesses and (2) further designed a novel Merkle tree scheme that flattened and reorganized the computation in the traditional Merkle tree verification and update processes to fully exploit the parallel cache ports and to fully pipeline time-consuming hashing operations. Beyond that, to accelerate the metadata recovery process, multiple parallel recovery units are instantiated to recover counter metadata and multiple Merkle sub-trees. Our hardware prototype of the ARES system on a Xilinx U200 platform shows that ARES achieved up to 1.4× lower latency and 2.6× higher throughput against the baseline implementation, while metadata recovery time was shortened by 1.8 times. When integrated with an embedded processor, neither hardware changes nor software changes are required. We also developed a theoretical framework to analytically model and explain experimental results.
Kazi Abu Zubair, Mazen Al-Wadi, Rakin Muhammad Shadab, Sanjay Gandham, Amro Awad, Mingjie Lin
ACM Trans. Embed. Comput. Syst.5
2021 ARC: Reconfigurable Cache Security Assurance with Application-Specific Randomized Mapping in FPGA-Based Heterogeneous Computing
abstract
Modem general purpose processors suffer from cache side-channel attacks (SCA) such as Prime+Probe [1] where the attacker can infer the victim's information. Last-Level caches(LLC) are particularly vulnerable as they are shared between different cores of the processor. Encryption-based randomized caches such as CEASER [2] have been successful in mitigating conflict-based SCA by stopping the attackers from creating eviction sets but they have a few drawbacks 1) Encryption and remapping is done at all times, even when not performing security-critical tasks and 2) These mitigation techniques provide no defense against flush- based cache attacks such as Flush+Reload. Moreover, such randomized caches employing least-recently used (LRU) replacement policy incur impractical overheads to provide defense against conflict-based SCA. On the other hand, randomized caches employing random replacement policy can mitigate theses attacks with relatively low overhead but suffer from lower hit rate due to inefficient replacement policy. In this paper we show that moving the shared LLC of the processor to the programmable fabric of heterogeneous devices such as FPGA+CPU system-on-chips provides high degree of flexibility in terms of security and performance. To this end, we propose two randomized cache modes 1) Fast: Generic cache using LRU policy while providing no security against SCA and 2) Secure: Randomized cache using random replacement policy that can mitigate SCA. When the LLC is implemented on the reprogrammable fabric of the FPGA, modern FPGA+CPU SoCs ability to reconfigure the FPGA fabric during run-time allows the cache to switch between these two modes. Additionally, we propose a novel randomized cache mechanism, ARC, that can mitigate not only conflict-based attacks but also flush- based cache attacks.
Sanjay Gandham, Rakin Muhammad Shadab, Mingjie Lin
FCCM1