Anirban Nag

dblp:173/5588 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0003-4905-8038ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 CoGraf: Fully Accelerating Graph Applications with Fine-Grained PIM
abstract
Processing-in-Memory (PIM) delivers enormous performance by taking advantage of internal DRAM bandwidth and parallelism. However, graph applications are difficult to adapt to PIM due to their irregular access patterns.We present the first Fine-Grained PIM (FGPIM) design that fully accelerates vertex-centric push-based graph applications by accelerating both their update (computing vertex updates) and apply (summing up the updates) phases. For the update phase, we design a tuple-based LLC that can coalesce at different granularities to group graph updates together and propose multi-DRAM column processing FGPIM instructions to match the cache coalescing to the row-level parallelism of the FGPIM. With this acceleration, the apply phase becomes the bottleneck, and we propose bank-parallel FGPIM instructions with predicates to allow FGPIM to accelerate the conditional updates as well. We achieve an average speedup in the region of interest of 1.8x/3x compared to naive FGPIM and 4.4x/9.8x compared to state-of-the-art non-PIM baseline (HBM2/DDR4), and DRAM energy reduction of 67%/86% and 88%/94%.These results show the importance of providing a complete solution that accelerates both the update and apply phases.
Ali Semi Yenimol, Anirban Nag, Chang Hyun Park 0001, David Black-Schaffer
ASPLOS (2)2
2026 GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read Mapping
abstract
Genome sequencing has become a central focus in computational biology due to its critical role in applications such as personalized medicine, disease outbreak tracking, and evolutionary research. A genome study typically begins with sequencing, which produces millions to billions of short DNA fragments known as reads. Extracting meaningful biological insights from these reads requires a computationally intensive step called read mapping, where each read is aligned to a reference genome. Read mapping for short reads comes in two forms: single-end and paired-end, with the latter being more prevalent due to its higher accuracy and support for advanced analysis. Read mapping remains a major performance bottleneck in genome analysis as a result of the extensive use of computationally intensive dynamic programming. Prior efforts have attempted to mitigate this cost by employing filters to identify and potentially discard computationally expensive matches and leveraging hardware accelerators to speed up the computations. While partially effective, these approaches have limitations. In particular, existing filters are often ineffective for paired-end reads, as they evaluate each read independently and exhibit relatively low filtering ratios. In this work, we propose GenPairX, a hardware-algorithm codesigned accelerator that efficiently minimizes the computational load of paired-end read mapping while enhancing the throughput of memory-intensive operations. GenPairX introduces: (1) a novel filtering algorithm that jointly considers both reads in a pair to improve filtering effectiveness, and a lightweight alignment algorithm to replace most of the computationally expensive dynamic programming operations, and (2) two specialized hardware mechanisms to support the proposed algorithms. The proposed hardware addresses the high memory bandwidth demands of the read filtering process via orchestration of memory accesses over high-bandwidth memory channels, and accelerates the alignment of candidate reads via simple vectorized logical XOR operators. Our evaluations show that GenPairX delivers substantial performance improvements over state-of-the-art solutions, achieving$1575 \times$and$1.43 \times$higher throughput per watt compared to leading CPU-based and accelerator-based read mappers, respectively, all without compromising accuracy.
Julien Eudine, Renzo Andri, Can Firtina, Mohammad Sadrosadati, Nika Mansouri-Ghiasi, Konstantina Koliogeorgi, Anirban Nag, Arash Tavakkol, Haiyu Mao, Onur Mutlu, Shai Bergman, Ji Zhang 0035
HPCA9
2021 ONT-X: An FPGA Approach to Real-time Portable Genomic Analysis
abstract
Oxford Nanopore Technologies (ONT) MinION is a pocket-sized portable DNA sequencer for on-the-field DNA sequencing. The ONT data analysis pipeline requires substantial amounts of computational resources for the basecalling and read alignment step in the pipeline. We study the compute and memory requirement of the basecalling step and present the detailed microarchitecture for accelerating it on FPGA. We also make the case that an embedded FPGA-based SoC with partial reconfiguration (PR) capability is an ideal substrate for accelerating ONT data analysis pipeline, due to the diversity in these application kernels and the evolving nature of the algorithms. Focusing on two widely used ONT data analysis pipelines, we present two modes of PR operation to cater to either latency or throughput requirements of these two pipelines and demonstrate how to amortize the partial reconfiguration cost, overlap the kernels, minimize the data transfer cost and achieve 111× higher throughput and 6× reduction in latency using these two PR modes.
C. N. Ramachandra, Anirban Nag, Rajeev Balasubramonian, Gurpreet S. Kalsi, Kamlesh R. Pillai, Sreenivas Subramoney
FCCM2
2021 OrderLight: Lightweight Memory-Ordering Primitive for Efficient Fine-Grained PIM Computations
abstract
Modern workloads such as neural networks, genomic analysis, and data analytics exhibit significant data-intensive phases (low compute to byte ratio) and, as such, stand to gain considerably by using processing-in-memory (PIM) solutions along with more traditional accelerators. While PIM has been researched extensively, the granularity of computation offload to PIM and the granularity of memory access arbitration between host and PIM, as well as their implications, have received relatively little attention. In this work, we first introduce a taxonomy to study the design space whilst considering these two aspects. Based on this taxonomy, we observe that much of PIM research to date has largely relied on coarse-grained approaches which, we argue, have steep costs (incompatibility with mainstream memory interfaces, prohibition of concurrent host accesses, and more). To this end, we believe that better support for fine-grained approaches is warranted in accelerators coupled with PIM-enabled memories.
Anirban Nag, Rajeev Balasubramonian
MICRO1
2019 GenCache: Leveraging In-Cache Operators for Efficient Sequence Alignment
abstract
Precision Medicine will rely on frequent genomic analysis, especially for patients undergoing cancer treatments or suffering from rare diseases. Sequence alignment is invoked in multiple stages of the genomic analysis pipeline. Recent projects have introduced accelerators, GenAx and Darwin, for 2nd and 3rd generation sequencers respectively. In this work, we improve upon the GenAx design by increasing its parallelism and reducing its memory bandwidth demands. This is achieved with a combination of hardware and software innovations. We first integrate in-cache operators from prior work into the GenAx memory hierarchy; we then augment the in-cache peripheral circuit to support additional new operators. We then re-structure the sequence alignment algorithm to (i) leverage the many in-cache operators, (ii) exploit the common case in genomic datasets, (iii) use Bloom Filters to reduce futile accesses, and (iv) maximize data reuse within a re-organized memory hierarchy. While the baseline GenAx accelerator processes a batch of reads in 194 seconds while nearly saturating the 153.6 GB/s memory bandwidth, the proposed GenCache architecture processes the same batch of reads in 37 seconds at an improved energy efficiency of 8.6×, while demanding 20 GB/s average memory bandwidth. Our hardware and software techniques thus interact synergistically to target both memory and compute bottlenecks, while not affecting the outputs of the application. We show that the basic principles in GenCache can also be exploited by 3rd generation sequence aligners.
Anirban Nag, C. N. Ramachandra, Rajeev Balasubramonian, Ryan Stutsman, Edouard Giacomin, Hari Kambalasubramanyam, Pierre-Emmanuel Gaillardon
MICRO1
2016 ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars
abstract
A number of recent efforts have attempted to design accelerators for popular machine learning algorithms, such as those involving convolutional and deep neural networks (CNNs and DNNs). These algorithms typically involve a large number of multiply-accumulate (dot-product) operations. A recent project, DaDianNao, adopts a near data processing approach, where a specialized neural functional unit performs all the digital arithmetic operations and receives input weights from adjacent eDRAM banks. This work explores an in-situ processing approach, where memristor crossbar arrays not only store input weights, but are also used to perform dot-product operations in an analog manner. While the use of crossbar memory as an analog dot-product engine is well known, no prior work has designed or characterized a full-fledged accelerator based on crossbars. In particular, our work makes the following contributions: (i) We design a pipelined architecture, with some crossbars dedicated for each neural network layer, and eDRAM buffers that aggregate data between pipeline stages. (ii) We define new data encoding techniques that are amenable to analog computations and that can reduce the high overheads of analog-to-digital conversion (ADC). (iii) We define the many supporting digital components required in an analog CNN accelerator and carry out a design space exploration to identify the best balance of memristor storage/compute, ADCs, and eDRAM storage on a chip. On a suite of CNN and DNN workloads, the proposed ISAAC architecture yields improvements of 14.8×, 5.5×, and 7.5× in throughput, energy, and computational density (respectively), relative to the state-of-the-art DaDianNao architecture.
Ali Shafiee, Anirban Nag, Naveen Muralimanohar, Rajeev Balasubramonian, John Paul Strachan, Miao Hu 0002, R. Stanley Williams, Vivek Srikumar
ISCA2