EDBT 2026 Demo / reviewers in the wild / expert
Palash Das 0001
dblp:156/2488-1
· DBLP profile ↗
11ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-1979-0298ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeSTAR: Hardware Trojans and its mitigation strategy in NoC routers
Josna Philomina, Rekha K. James, Shirshendu Das, Palash Das 0001, Daleesha M. Viswanathan |
Integr. | 4 |
| 2026 | Exploiting virtual channel allocation policies in STT-RAM buffers of NoC routers through hardware Trojan
Josna Philomina, Rekha K. James, Palash Das 0001, Shirshendu Das, Daleesha M. Viswanathan |
J. Syst. Archit. | 3 |
| 2026 | GateAttn-ViT: Entropy-gated, attention-guided token pruning for resource-efficient Vision Transformer acceleration on FPGAs
Akshansh Yadav, Palash Das 0001 |
J. Syst. Archit. | 2 |
| 2025 | Hybrid Token Selector based Accelerator for ViTsabstractVision Transformers (ViTs) have shown great success in computer vision but suffer from high computational complexity due to the quadratic growth in the number of tokens processed. Token selection/pruning has emerged as a promising solution; however, early methods introduce significant overhead and complexity. Applying a token selector in the early layers of a ViT can yield substantial computational savings (GFLOPs) compared to using it in later layers. However, this approach often leads to significant accuracy loss, particularly with the popular Attention-based Token Selection (ATS) technique. To address these issues, we propose a hybrid token selection (HTS) strategy that integrates our Keypoint-based Token Selection (KTS) with the existing ATS method. KTS dynamically selects important tokens based on image content in the early layers, while ATS refines token pruning in the later layers. This hybrid approach reduces computational costs while maintaining accuracy. Additionally, we design custom hardware modules to accelerate the execution of the proposed methods and the ViT backbone. The proposed HTS delivers a 35.85% reduction in execution time relative to the baseline without any token selection. Furthermore, our results demonstrate that HTS achieves up to a 0.39% increase in accuracy and offers up to 6.05% savings in GFLOPs compared to existing method. Akshansh Yadav, Anadi Goyal, Palash Das 0001 |
DATE | 3 |
| 2023 | ALAMNI: Adaptive LookAside Memory Based Near-Memory Inference Engine for Eliminating Multiplications in Real-TimeabstractThe CNN algorithm seeks high performance and energy efficiency in real-time inference. The costly off-chip memory accesses put additional burdens on CNN's execution. Towards avoiding off-chip accesses, we propose ALAMNI, a novel near-memory architecture that expedites the CNNs in the logic layer of the Hybrid Memory Cube. We exploit intra- and inter-vault parallelism to accelerate the highly parallel CNN operations. The proposed ALAMNI replaces costly multiplications of CNNs with lookaside memory (LAM) based searches. The proposed ALAMNI policy is effective on unseen data as it discards the data pre-profiling overhead by an adaptive LAM update policy. The ALAMNI controller keeps the most frequent triplets of weight (W), activation (A), and multiplication result (M),, in the LAM to eliminate redundant computations. As an optimization, we incorporate a bitmasking concept to raise the hit rate of LAMs and further amortize computations. We also present a study on the relation between the amount of bitmasking and the loss of classification accuracy of the popular ConvNets. We keep the bitmasking as a reconfigurable feature of the ALAMNI units to achieve desired classification accuracy. Experimental results show substantial improvement in the system's performance and energy efficiency compared to the baseline and state-of-the-art. Palash Das 0001, Shashank Sharma 0002, Hemangee K. Kapoor |
IEEE Trans. Computers | 1 |
| 2022 | Hydra: A near hybrid memory accelerator for CNN inferenceabstractConvolutional neural network (CNN) accelerators often suffer from limited off-chip memory bandwidth and on-chip capacity constraints. One solution to this problem is near-memory or in-memory processing. Non-volatile memory, such as phase-change memory (PCM), has emerged as a promising DRAM alternative. It is also used in combination with DRAM, forming a hybrid memory. Though near-memory processing (NMP) has been used to accelerate the CNN inference, the feasibility/efficacy of NMP remained unexplored for a hybrid main memory system. Additionally, PCMs are also known to have low write endurance, and therefore, the tremendous amount of writes generated by the accelerators can drastically hamper the longevity of the PCM memory. In this work, we propose Hydra, a near hybrid memory accelerator integrated close to the DRAM to execute inference. The PCM banks store the models that are only read by the memory controller during the inference. For entire forward propagation (inference), the intermediate writes from Hydra are entirely performed to the DRAM, eliminating PCM-writes to enhance PCM lifetime. Unlike the other in-DRAM processing-based works, Hydra does not eliminate any multiplication operations by using binary or ternary neural networks, making it more suitable for the requirement of high accuracy. We also exploit inter- and intra-chip (DRAM chip) parallelism to improve the system's performance. On average, Hydra achieves around 20x performance improvements over the in-DRAM processing-based state-of-the-art works while accelerating the CNN inference. Palash Das 0001, Ajay Joshi, Hemangee K. Kapoor |
DATE | 1 |
| 2022 | ZaLoBI: Zero avoiding Load Balanced Inference acceleratorabstractConvolutional neural networks are prevalent machine learning tools used in computer vision. Their ubiquitous use and high compute requirement have given rise to the design and development of accelerators for the same. Among several approaches to improve the performance of these accelerators, exploiting data sparsity has become very popular. Along similar lines, this paper proposes a design that skips the computation of zero-valued data operands and achieves better speedup. The savings in zero-valued computations also results in energy savings. The proposed accelerator exploits two levels of data parallelism to distribute work across multiple processing elements (PEs). The random distribution of zero values results in certain PEs getting idle due to the skipping of computations, thus creating load imbalance in the system. To address this issue, we extend our contribution in performing load balancing by dynamically scheduling tasks to the idle PEs. Our zero avoiding load-balanced accelerator (ZaLoBI) achieves around 76% and 5.57% speedup over the respective baselines and also outperforms the state-of-the-art works while saving energy. Imlijungla Longchar, Palash Das 0001, Hemangee K. Kapoor |
VLSI-SoC | 2 |
| 2021 | CLU: A Near-Memory Accelerator Exploiting the Parallelism in Convolutional Neural NetworksabstractConvolutional/Deep Neural Networks (CNNs/DNNs) are rapidly growing workloads for the emerging AI-based systems. The gap between the processing speed and the memory-access latency in multi-core systems affects the performance and energy efficiency of the CNN/DNN tasks. This article aims to alleviate this gap by providing a simple and yet efficient near-memory accelerator-based system that expedites the CNN inference. Towards this goal, we first design an efficient parallel algorithm to accelerate CNN/DNN tasks. The data is partitioned across the multiple memory channels (vaults) to assist in the execution of the parallel algorithm. Second, we design a hardware unit, namely the convolutional logic unit (CLU), which implements the parallel algorithm. To optimize the inference, the CLU is designed, and it works in three phases for layer-wise processing of data. Last, to harness the benefits of near-memory processing (NMP), we integrate homogeneous CLUs on the logic layer of the 3D memory, specifically the Hybrid Memory Cube (HMC). The combined effect of these results in a high-performing and energy-efficient system for CNNs/DNNs. The proposed system achieves a substantial gain in the performance and energy reduction compared to multi-core CPU- and GPU-based systems with a minimal area overhead of 2.37%. Palash Das 0001, Hemangee K. Kapoor |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2021 | nZESPA: A Near-3D-Memory Zero Skipping Parallel Accelerator for CNNsabstractConvolutional neural networks (CNNs) are one of the most popular machine learning tools for computer vision. The ubiquitous use in several applications with its high computation-cost has made it lucrative for optimization through accelerated architecture. State-of-the-art has either exploited the parallelism of CNNs, or eliminated computations through sparsity or used near-memory processing (NMP) to accelerate the CNNs. We introduce NMP-fully sparse architecture, which acquires all three capabilities. The proposed architecture is parallel and hence processes the independent CNN tasks concurrently. To exploit the sparsity, the proposed system employs a dataflow, namely, Near-3D-Memory Zero Skipping Parallel dataflow or nZESPA dataflow. This dataflow maintains the compressed-sparse encoding of data that skips all ineffectual zero-valued computations of CNNs. We design a custom accelerator which employs the nZESPA dataflow. The grids of nZESPA modules are integrated into the logic layer of the hybrid memory cube. This integration saves a significant amount of off-chip communications while implementing the concept of NMP. We compare the proposed architecture with three other architectures which either do not exploit sparsity (NMP-dense) or do not employ NMP (traditional-fully sparse) or do not include both (traditional-dense). The proposed system outperforms the baselines in terms of performance and energy consumption while executing CNN inference. Palash Das 0001, Hemangee K. Kapoor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Dimming Hybrid Caches to Assist in Temperature Control of Chip MultiProcessorsabstractThe continuous rise of on-chip components like cores and caches has brought enormous computing capabilities at the cost of high leakage power and temperature. A recent study has shown a substantial spatial temperature variance in modern large on-chip caches. This high temperature elevates the cooling cost and becomes responsible for the thermal breakdown of the chip. One solution to reduce the leakage is the use of non-volatile memory (NVM) like STT-RAM. Other includes incorporating the concept of dark silicon. In this paper, we amalgamate the idea of using STT-RAM in the last level cache (LLC) and the dark silicon approach to shut down certain cache ways to leverage the benefits from both. We address the downsides like higher access latencies of STT-RAM by the use of hybrid cache (SRAM + STT-RAM) and weak endurance of the STT-RAM by wear leveling. We propose a system to handle three different temperature thresholds (high, medium, and low) by appropriately selecting the type of cache ways to be powered off. The proposed system delivers up to 5.38 K reduction in temperature compared to the baseline, 93% reduction in leakage power with an EDP gain up to 92%. Chirag Joshi, Palash Das 0001, Ashwini A. Kulkarni, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Towards Near-Data Processing of Compare Operations in 3D-Stacked MemoryabstractThe gap between the processing speed and memory access speed of the modern multi-core systems has become a bottleneck for the emerging data-intensive workloads. In this scenario, it has become a smarter idea to move some amount of computation closer to the data, thus stimulating the concept of near-data processing (NDP). Compare or scanning, the core operations of many applications, typically in a database, can leverage the benefits of NDP. We propose near-data compare unit (NDCU), a less-invasive hardware, that can be integrated with the existing ecosystem of the hybrid memory cube (HMC). While integrating NDCU, we have designed two full-system architectures, one is lighter NDP with no parallelism (NNP) and the second is NDP with vault level parallelism (NVLP). While the first architecture is more power and area efficient, the second one is very fast with negligible overheads. With the motive of carrying out scan operation, we have specifically implemented 'compare-n-hit', 'compare-n-count' and 'compare-n-max' operations on both row-store and column-store databases and found significant improvements over conventional CPU-based system. We get around 2.3x and 37x performance improvement in NNP and NVLP architectures respectively. In both the designs, we reduce the energy consumption by around 8x on an average. Palash Das 0001, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 1 |