EDBT 2026 Demo / reviewers in the wild / expert
Nanditha P. Rao
dblp:140/7207
· DBLP profile ↗
5ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0003-2369-0836ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | MCSim: A Multi-Core Cache Simulator Accelerated on a Resource-constrained FPGAabstractPerformance evaluation of caches is an important component of the design process. Software or analytical model-based simulation approaches, although used by architects, are abstract models and are, therefore, not completely accurate. RTL-based simulators can be automatically mapped to FPGAs and can also be faster than software simulators. We present an FPGA accelerated multi-core cache simulator MCSim supporting a parameterized two-level cache structure. It can be partially reconfigured to include prefetchers and can simulate several cache configurations in parallel. We run a set of SPEC 2017 benchmarks on MCSim and find that it can run nearly 2.61x and 5.33x faster (on an average) as compared to ChampSim and Sniper for a 4-core framework, and 7x and 11.5x faster (on an average) as compared to ChampSim and Sniper in a single-core framework to generate hit/miss rates. We also present the scalability of our simulator on FPGAs with larger number of resources. Shivani Shah, Nanditha P. Rao |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | An FPGA based Tiled Systolic Array Generator to Accelerate CNNsabstractThe main computation in any CNN is convolution operation. This computation shows significant potential for massively parallel implementations on an FPGA. Systolic arrays with their intrinsic pipelining have been explored for CNN inference. In this paper, we present a systolic array architecture suitably designed for a novel method of convolution operation. We implement an image-kernel convolution and test it with representative image inputs to several models like LeNet-5, AlexNet, VGG-16, and Resnet-34. We compare the proposed design with conventional convolution and HLS based designs. We limit our implementation to resource constrained FPGA: AMD-Xilinx Zynq 7020 platform. We observe that the proposed architecture outperforms the direct convolution method and HLS pipelined designs by 2× and 2.1×, respectively, on average. Since DSP blocks are scarce resources, we constrain our implementation to avoid DSP blocks and use the LUTs instead. Thus, our implementation uses nearly 9× more LUTs than baseline convolution but 8× fewer LUTs than the HLS pipelined implementation. We further accelerate the convolution throughput by 11×. We achieve this by implementing a tiled systolic architecture that completely utilises the parallel computing resources of the FPGA. Veerendra S. Devaraddi, Nanditha P. Rao |
DSD | 2 |
| 2022 | An Automated Approach to Compare Bit Serial and Bit Parallel In-Memory Computing for DNNsabstractThis paper presents an exhaustive comparison of two different techniques for In-Memory Computing in SRAM: bit-serial arithmetic (BSA), and bit-parallel arithmetic (BPA). We have modeled both BSA and BPA and integrated them with CACTI for the CMOS 28nm technology node. The results are analyzed for both approaches with ten different sub-array configurations ranging from 128x128 to 2048x2048. We performed the convolution operation on the ImageNet dataset for comparison. The key observation is that the BPA begins to yield at least 25% better (lower) Energy Delay Product (EDP) as compared to that in BSA for large (2048x2048) sub-array sizes. The BSA yields $4 \times$ lower delay on an average, while the BPA yields $\sim 6 \times$ lower dynamic energy. Hence, the choice of the IMC architecture needs to be made depending on the application need (low energy/high performance). We have also simulated 12 different multi-bank IMC arrangements and show that just modifying the memory array structure can improve the EDP by up to $8 \times$. Alok Parmar, Kailash Prasad, Nanditha P. Rao, Joycee Mekie |
ISCAS | 3 |
| 2021 | Cache-accel: FPGA Accelerated Cache Simulator with Partially Reconfigurable PrefetcherabstractComputer architects need to choose the design configurations which will work effectively across most commonly used workloads. Design space exploration of caches enables the architect to choose the right configuration based on metrics such as hit rates, power, area, and timing. Although the idea of a cache simulator is not new, the hardware/FPGA implementation of such simulators has not been well explored. We implement an FPGA accelerated parameterized two-level cache simulator called Cache-accel which can be partially reconfigured to include prefetching. The key motivation behind the idea is the speed with which the design space exploration can be carried out by exploiting the parallelism available in an FPGA. Cache-accel reports cache metrics such as hit/miss rates with and without prefetching for nine cache configurations in parallel. We run a set of SPEC 2017 benchmarks on Cache-accel and find that it can run nearly 7x and 11.5x faster (on an average) as compared to ChampSim and Snipersim to generate hit/miss rates for the nine parallel configurations. Shivani Shah, Vaibhavi Mathur, Sahithi Meenakshi Vutakuru, Kavya Borra, Nanditha P. Rao |
DSD | 5 |
| 2015 | A Detailed Characterization of Errors in Logic Circuits due to Single-Event TransientsabstractWhen a high energy particle strikes an integrated circuit, the electron-hole pairs generated in the substrate get collected at source/drain regions. The net effect on the circuit is a transient current (single event transient or SET) injected into circuit nodes. This SET can propagate and cause an error in a state register, which is called a single-event upset (SEU). We perform a detailed characterization of the impact of an SET on a logic circuit. We observe that the impact of an SET can be understood only as a two cycle phenomenon, with several possible SEU outcomes. To understand the relative probabilities of these outcomes, we perform a detailed characterization of the impact of an SET using a Monte Carlo sampling scheme which uses post-layout circuit simulations. These simulations were run in 180nm and 65nm technologies with circuits from ISCAS'85 and ITC'99 benchmarks as test cases. Based on these simulations, we observe that all the SEU outcomes are represented to a substantial extent. Further, there is a substantial fraction of outcomes in which the SET affects multiple state registers. This indicates that the traditional view of impact of an SET on a circuit as a single bit-flip and a single cycle phenomenon may not be fully justified. The impact of an SET is in reality, distributed across two cycles and across multiple state register bits. Nanditha P. Rao, Madhav P. Desai |
DSD | 1 |