EDBT 2026 Demo / reviewers in the wild / expert
Dhruv Parikh
dblp:313/5747
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient and Accurate Graph Classification with Hyperdimensional Computing on FPGAabstractReal-time, energy-efficient inference on edge devices is essential for graph classification across a range of applications. Hyperdimensional Computing (HDC) is a brain-inspired computing paradigm that encodes input features into low-precision, high-dimensional vectors with simple element-wise operations, making it well-suited for resource-constrained edge platforms. Recent work enhances HDC accuracy for graph classification via Nyström kernel approximations. Edge acceleration of such methods faces several challenges: (i) redundancy among (landmark) samples selected via uniform sampling, (ii) storing the Nyström projection matrix under limited on-chip memory, (iii) expensive, contention-prone codebook lookups, and (iv) load imbalance due to irregular sparsity in SpMV. Jebacyril Arockiaraj, Dhruv Parikh, Viktor Prasanna 0001 |
CF | 2 |
| 2026 | ImageHD: Energy-Efficient On-Device Continual Learning of Visual Representations via Hyperdimensional ComputingabstractOn-device continual learning (CL) enables edge devices to adapt to non-stationary data streams without offline retraining, which is critical for real-time edge AI. However, most existing CL methods rely on backpropagation-based updates or exemplar-heavy classifiers, incurring high compute, memory, and latency overheads that hinder deployment on resource-constrained devices. Hyperdimensional Computing (HDC) offers an alternative by enabling fast, non-iterative online updates. When paired with a lightweight convolutional neural network (CNN) feature extractor, HDC supports efficient on-device adaptation with strong visual representations. Despite this progress, prior HDC-based continual learning systems typically employ multi-tier memory hierarchies and complex cluster management, which complicate deployment on resource-constrained hardware platforms.In this paper, we propose ImageHD, an FPGA accelerator for on-device continual learning of visual data based on HDC, designed to address these limitations. ImageHD targets streaming continual learning under strict latency and on-chip memory constraints, without expensive iterative optimization. At the algorithmic level, we introduce a hardware-optimized continual learning method that bounds total class exemplars through a unified exemplar memory and a hardware-efficient cluster merging strategy, while integrating a quantized CNN front-end to reduce edge deployment overhead without sacrificing accuracy. At the system level, ImageHD is realized as a streaming dataflow architecture on the AMD Zynq ZCU104 FPGA, integrating hyperdimensional encoding, similarity search, and bounded cluster management using word-packed binary hypervector representations to enable massively parallel bitwise operations within tight on-chip resource budgets. Experimental results on the CORe50 dataset demonstrate up to 40.4× (4.84×) speedups and 383× (105.1×) energy efficiency gains over optimized CPU (GPU) baselines, establishing HDC-enabled continual learning as a practical foundation for real-time, on-device lifelong learning systems. Jebacyril Arockiaraj, Dhruv Parikh, Viktor Prasanna 0001 |
FCCM | 2 |
| 2026 | GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGAabstractVision Graph Neural Networks (ViGs) model an image as a graph of patch tokens, enabling adaptive, feature-driven neighborhoods. Unlike CNNs with fixed grid biases or Vision Transformers with global token interactions, ViGs rely on dynamic graph convolution: at each layer, a feature-dependent graph is built via k-nearest neighbor (kNN) search on current patch features, followed by message passing. This per-layer graph construction is the main bottleneck—consuming 50–95% of graph convolution time on CPUs and GPUs—and scales by O(N2) where N is the number of patches, creating a sequential dependency between graph construction and feature updates.In this paper, we introduce GraphLeap, a simple yet novel reformulation that removes this dependency by decoupling graph construction from feature update across layers. GraphLeap performs the feature update at layer ℓ using a graph constructed from the previous layer’s features, while simultaneously using the current layer’s features to construct the graph for layer ℓ+1. This one-layer lookahead dynamic graph construction enables concurrent graph construction and GNN message passing, yielding a new class of ViGs, which we call GraphLeap. While using prior-layer features for graph construction can introduce minor accuracy degradation, we show that lightweight fine-tuning for a few epochs is sufficient to recover original accuracy. Building on GraphLeap, we present the first end-to-end FPGA accelerator for Vision GNNs. Our design is a streaming, layer-pipelined architecture that overlaps a kNN graph construction engine with a feature update engine, exploits node and channel-level parallelism, and enables efficient on-chip dataflow without explicit materialization of edge features. Evaluated on isotropic and pyramidal ViG models deployed on an Alveo U280 FPGA, our approach achieves up to 95.7× speedup over CPU and 8.5× speedup over GPU baselines, demonstrating the feasibility of real-time Vision GNN inference and motivating future hardware– algorithm co-design for graph-based vision models. Code is available at https://github.com/anvitha305/GraphLeap. Anvitha Ramachandran, Dhruv Parikh, Viktor Prasanna 0001 |
FCCM | 2 |
| 2026 | NysX: An Accurate and Energy-Efficient FPGA Accelerator for Hyperdimensional Graph Classification at the EdgeabstractReal-time, energy-efficient inference on edge devices is essential for graph classification across a range of applications. Hyperdimensional Computing (HDC) is a brain-inspired computing paradigm that encodes input features into low-precision, high-dimensional vectors with simple element-wise operations, making it well-suited for resource-constrained edge platforms. Recent work enhances HDC accuracy for graph classification via Nyström kernel approximations. Edge acceleration of such methods faces several challenges: (i) redundancy among (landmark) samples selected via uniform sampling, (ii) storing the Nyström projection matrix under limited on-chip memory, (iii) expensive, contention-prone codebook lookups, and (iv) load imbalance due to irregular sparsity in SpMV. Jebacyril Arockiaraj, Dhruv Parikh, Viktor Prasanna 0001 |
FPGA | 2 |
| 2025 | Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing UnitsabstractThe proliferation of large language models has driven demand for long-context inference on resourceconstrained edge platforms. However, deploying these models on Neural Processing Units (NPUs) presents significant challenges due to architectural mismatch: the quadratic complexity of standard attention conflicts with NPU memory and compute patterns. This paper presents a comprehensive performance analysis of causal inference operators on a modern NPU, benchmarking quadratic attention against sub-quadratic alternatives including structured state-space models and causal convolutions. Our analysis reveals a spectrum of critical bottlenecks: quadratic attention becomes severely memory-bound with catastrophic cache inefficiency, while sub-quadratic variants span from computebound on programmable vector cores to memory-bound by data movement. These findings provide essential insights for codesigning hardware-aware models and optimization strategies to enable efficient long-context inference on edge platforms. Neelesh Gupta, Rakshith Jayanth, Dhruv Parikh, Viktor Prasanna 0001 |
HiPC | 3 |
| 2025 | Vision Transformers for End-to-End Vision-Based Quadrotor Obstacle AvoidanceabstractWe demonstrate the capabilities of an attentionbased end-to-end approach for high-speed vision-based quadrotor obstacle avoidance in dense, cluttered environments, with comparison to various state-of-the-art learning architectures. Quadrotor unmanned aerial vehicles (UAVs) have tremendous maneuverability when flown fast; however, as flight speed increases, traditional model-based approaches to navigation via independent perception, mapping, planning, and control modules breaks down due to increased sensor noise, compounding errors, and increased processing latency. Thus, learning-based, end-to-end vision-to-control networks have shown to have great potential for online control of these fast robots through cluttered environments. We train and compare convolutional, U-Net, and recurrent architectures against vision transformer (ViT) models for depth image-to-control in high-fidelity simulation, observing that ViT models are more effective than others as quadrotor speeds increase and in generalization to unseen environments, while the addition of recurrence further improves performance while reducing quadrotor energy cost across all tested flight speeds. We assess performance at speeds of up to 7m/s in simulation and hardware. To the best of our knowledge, this is the first work to utilize vision transformers for end-to-end vision-based quadrotor control. Anish Bhattacharya, Nishanth Rao, Dhruv Parikh, Pratik Kunapuli, Yuwei Wu 0005, Yuezhan Tao, Nikolai Matni, Vijay Kumar 0001 |
ICRA | 3 |
| 2024 | Accelerating ViT Inference on FPGA through Static and Dynamic PruningabstractVision Transformers (ViTs) have achieved state-of-the-art accuracy on various computer vision tasks. However, their high computational complexity prevents them from being applied to many real-world applications. Weight and token pruning methods are well-known in reducing ViT model complexity. However, naively combining and integrating both the methods results in irregular computation patterns leading to accuracy drops and difficulties in hardware acceleration. This limits the net complexity reduction offered by integrating such pruning methods. To address the above challenges, we propose a comprehensive algorithm-hardware codesign for accelerating ViT on FPGA through simultaneous pruning - combining static weight pruning and dynamic token pruning. For algorithm design, we systematically combine a hardware-aware structured block-pruning method for pruning model parameters and a dynamic token pruning method for removing unimportant token vectors. Moreover, we design a novel training algorithm to reduce the accuracy drop due to such simultaneous pruning. For hardware design, we develop a novel hardware accelerator for executing the pruned model. The proposed hardware design employs multi-level parallelism with a load-balancing strategy to efficiently deal with the irregular computation pattern presented by the two pruning approaches. Moreover, we develop an efficient hardware mechanism for executing the on-the-fly token pruning. We apply our codesign approach to the widely used DeiT-Small model. We implement the proposed accelerator on a state-of-the-art FPGA. The evaluation results show that the proposed algorithm reduces computation complexity by up to 3.4× with ≈ 3% accuracy drop and a model compression ratio of up to 1.6×. Compared with state-of-the-art implementation on CPU, GPU, and FPGA, our codesign on FPGA achieves an average latency reduction of 12.8×, 3.2×, and 0.7 – 2.1×, respectively. Dhruv Parikh, Shouyi Li, Bingyi Zhang, Rajgopal Kannan, Carl E. Busart, Viktor Prasanna 0001 |
FCCM | 1 |