Archit Gajjar

dblp:207/9131 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-3759-7766ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Exploiting Power Side-Channel Vulnerabilities in XGBoost Accelerator
abstract
XGBoost (eXtreme Gradient Boosting), a widelyused decision tree algorithm, plays a crucial role in applications such as ransomware and fraud detection. While its performance is well-established, its security against model extraction on hardware platforms like Field Programmable Gate Arrays (FPGAs) has not been fully explored. In this paper, we demonstrate a significant vulnerability where sensitive model data can be leaked from an XGBoost implementation through side-channel attacks (SCAs). By analyzing variations in power consumption, we show how an attacker can infer node features within the XGBoost model, leading to the extraction of critical data. We conduct an experiment using the XGBoost accelerator FAXID on the Sakura-X platform, demonstrating a method to deduce model decisions by monitoring power consumptions. The results show that on average 367k tests are sufficient to leak sensitive values. Our findings underscore the need for improved hardware and algorithmic protections to safeguard machine learning models from these types of attacks.
Yimeng Xiao, Archit Gajjar, Aydin Aysu, Paul D. Franzon
DAC2
2025 Analog In-Memory Computing Enhanced FPGA for High-Throughput and Energy-Efficient Acceleration
abstract
The ever-growing demand for AI computing, coupled with slowing performance gains in chip manufacturing, has heightened the role of FPGA-based accelerators. FPGAs enable the implementation of application-customized parallel dataflows due to their reconfigurability, achieving high energy efficiency. However, the bit-level routing fabric on FPGAs often results in high overheads because large amounts of data must be shuttled between compute blocks and memory blocks on the FPGA. We propose enhancing FPGAs with in-memory computing macros, specifically analog Dot Product Engines based on non-volatile RRAM devices. Using the Verilog to Routing (VTR) framework, we simulate a novel 40 nm, 26.2 mm × 26.2 mm architecture and employ a custom event-driven simulator to evaluate its performance. Our design achieves 25.5 ×103TOPS/W, an average ×31.4 throughput improvement and an average ×9,380 energy efficiency improvement when compared to state-of-the-art FPGA implementations of AI models.
Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Xia Sheng, Giacomo Pedretti, Aman Arora 0001, Paolo Faraboschi, Jim Ignowski, Luca Buonanno
FCCM1
2025 Enhancing FPGAs with Analog In-Memory Computing Macros
abstract
While the AI computing needs are ever-increasing and the innovation in models generates tens of new architectures yearly, the performance gain from improvements in chip manufacturing has slowed down. Within this context, FPGA-based accelerators play a fundamental role. FPGAs are the backbone of specialized architectures, their reconfigurability being the key differentiation that enables an effective design space exploration. At the same time, to overcome the limitations induced by the memory bottleneck, the computing architectures community has proposed the in-memory computing paradigm: storage and computations are both performed in non-volatile memory devices.
Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Rand Jean, Xia Sheng, Giacomo Pedretti, Paolo Faraboschi, Jim Ignowski, Luca Buonanno
FPGA1
2025 RACE-IT: A Reconfigurable Analog Computing Engine for In-Memory Transformer Acceleration
abstract
Transformer models represent the cutting edge of Deep Neural Networks (DNNs) and excel in a wide range of machine learning tasks. However, processing these models demands significant computational resources and results in a substantial memory footprint. While In-memory Computing (IMC) offers promise for accelerating Vector-Matrix Multiplications (VMMs) with high computational parallelism and minimal data movement, employing it for other crucial DNN operators remains a formidable task. This challenge is exacerbated by the extensive use of complex activation functions, Softmax, and data-dependent matrix multiplications (DMMuls) within Transformer models. To address this challenge, we introduce a Reconfigurable Analog Computing Engine (RACE) by enhancing Analog Content Addressable Memories (ACAMs) to support broader operations. Based on the RACE, we propose the RACE-IT accelerator (meaning RACE for In-memory Transformers) to enable efficient analog-domain execution of all core operations of Transformer models. Given the flexibility of our proposed RACE in supporting arbitrary computations, RACE-IT is well-suited for adapting to emerging and non-traditional DNN architectures without requiring hardware modifications. We compare RACE-IT with various accelerators. Results show that RACE-IT increases performance by 453× and 15×, and reduces energy by 354× and 122× over the state-of-the-art GPUs and existing Transformer-specific IMC accelerators, respectively.
Aishwarya Natarajan, Luca Buonanno, Archit Gajjar, Ron M. Roth, Sergey Serebryakov, John Moon, Omar Eldash, Jim Ignowski, Giacomo Pedretti
ICCD4
2024 RD-FAXID: Ransomware Detection with FPGA-Accelerated XGBoost
abstract
Over the last decade, there has been a rise in cyberattacks, particularly ransomware, causing significant disruption and financial repercussions across public and private sectors. Tremendous efforts have been spent on developing techniques to detect ransomware to, ideally, protect data or have as minimum data loss as possible. Ransomware attacks are becoming more frequent and sophisticated as there is a constant tussle between attackers and cybersecurity defenders. Machine Learning (ML) approaches have proven more effective in detecting ransomware than classical signature-based detection. In particular, tree-based algorithms such as Decision Trees (DT), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) spike up interest among cybersecurity researchers. However, due to the nature of the problem, traditional CPUs and GPUs fail to keep up with the desired performance, especially for large data workloads. Thus, the problem demands a customized solution to detect the ransomware. Here, we propose an FPGA accelerated tree-based ML model for multi-dataset ransomware detection. We show the capability of the proposed prototype to address the problem from more than one set of features, reducing false positive and negative rates to have robust predictions by looking at Hardware Performance Counters (HPCs), Operating System (OS) calls, and network traffic information simultaneously. With 1,000 samples per batch, the FPGA prototype has 65.8 \({\times}\) and 4.1 \({\times}\) lower latency over the CPU and GPU, respectively. Moreover, the FPGA design is up to 11.3 \({\times}\) cost-effective and 643 \({\times}\) energy-efficient compared to the CPU and 3 \({\times}\) cost-effective and 16.8 \({\times}\) energy-efficient over the GPU.
Archit Gajjar, Priyank Kashyap, Aydin Aysu, Paul D. Franzon, Chris Cheng, Giacomo Pedretti, Jim Ignowski
ACM Trans. Reconfigurable Technol. Syst.1
2022 FAXID: FPGA-Accelerated XGBoost Inference for Data Centers using HLS
abstract
Advanced ensemble trees have proven quite effective in providing real-time predictions against ransomware detection, medical diagnosis, recommendation engines, fraud detection, failure predictions, crime risk, to name a few. Especially, XGBoost, one of the most prominent and widely used decision trees, has gained popularity due to various optimizations on gradient boosting framework that provides increased accuracy for classification and regression problems. XGBoost’s ability to train relatively faster, handling missing values, flexibility and parallel processing make it a better candidate to handle data center workload. Today’s data centers with enormous Input/Output Operations per Second (IOPS) demand a real-time accelerated inference with low latency and high throughput because of significant data processing due to applications such as ransomware detection or fraud detection.This paper showcases an FPGA-based XGBoost accelerator designed with High-Level Synthesis (HLS) tools and design flow accelerating binary classification inference. We employ Alveo U50 and U200 to demonstrate the performance of the proposed design and compare it with existing state-of-the-art CPU (Intel Xeon E5-2686 v4) and GPU (Nvidia Tensor Core T4) implementations with relevant datasets. We show a latency speedup of our proposed design over state-of-art CPU and GPU implementations, including energy efficiency and cost-effectiveness. The proposed accelerator is up to 65.8x and 5.3x faster, in terms of latency than CPU and GPU, respectively. The Alveo U50 is a more cost-effective device, and the Alveo U200 stands out as more energy-efficient.
Archit Gajjar, Priyank Kashyap, Aydin Aysu, Paul D. Franzon, Sumon Dey, Chris Cheng
FCCM1