EDBT 2026 Demo / reviewers in the wild / expert
Nanditha Rao
dblp:259/3772
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-2369-0836ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesabstractNear-data processing (NDP) mitigates the data movement bottleneck in modern computing systems by performing computation close to where the data resides. Solid-state drives (SSDs) are well suited for NDP because they: (1) store large application datasets that exceed main memory capacity, and (2) contain multiple heterogeneous computation resources, e.g., general-purpose embedded cores in the SSD controller, DRAM chips, and NAND flash chips, which enable three NDP paradigms: in-storage processing (ISP), processing using DRAM in the SSD (PuD-SSD), and in-flash processing (IFP). These resources offer massive internal parallelism and enable in-place computation, which reduces data movement across the memory hierarchy. A large body of prior SSD-based NDP techniques operate in isolation, mapping computations to only one or two NDP paradigms (i.e., ISP, PuD-SSD, or IFP) within the SSD. These techniques (1) are tailored to specific workloads or kernels, (2) do not offload computations across all three NDP paradigms in the SSD and thus fail to exploit the full computational potential of an SSD, and (3) lack programmer-transparency, often requiring significant manual effort to identify offloadable code regions and map them to the SSD computation resources, which limits their general applicability and ease of deployment. While several prior works propose techniques to partition computation between the host and near-memory accelerators, adapting these techniques to SSDs offers limited benefits because they (1) ignore the heterogeneity of the SSD computation resources, and (2) make offloading decisions based on limited factors such as bandwidth utilization, data movement cost, or memory intensity, while ignoring key factors such as resource utilization. We propose Conduit, a general-purpose, programmertransparent NDP framework for SSDs that accelerates a broad range of workloads by leveraging available SSD computation resources. Conduit operates in two stages. At compile time, Conduit executes a custom compiler (e.g., LLVM) pass that (i) vectorizes suitable application code segments into single-instruction multiple-data (SIMD) operations that align with the SSD's page layout, and (ii) embeds metadata (e.g., operation type, operand sizes) into the vectorized instructions to guide runtime offloading decisions. At runtime, within the SSD, Conduit performs instruction-granularity offloading by evaluating six key application and system features (e.g., operation type, computation resource utilization, data dependence delay), and uses a cost function to select the most suitable SSD computation resource to execute each vectorized instruction. We evaluate Conduit and two prior NDP offloading techniques using an in-house event-driven SSD simulator on six data-intensive applications (e.g., large language model inference and training, encryption). Conduit outperforms the best-performing prior offloading policy by$1.8 \times$and reduces energy consumption by 46 %, with small latency and storage overheads, and no additional hardware cost. Rakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta, Rahul Bera, Nika Mansouri-Ghiasi, Nanditha Rao, Qingcai Jiang, Andreas Kosmas Kakolyris, Yu Liang 0004, Mohammad Sadrosadati, Onur Mutlu |
HPCA | 7 |
| 2025 | FlexPE: Flexible Processing Elements for Workload Optimization and Acceleration on FPGAsabstractModern applications in the fields of machine learning, and scientific computing benefit from custom hardware configurations. In these cases, rather than designing or remapping the algorithms onto new accelerators, we propose flexible Processing Elements (PEs) that can adapt or reconfigure themselves according to the data type and compute type of the workloads. In this paper, we propose FlexPE, a framework that can generate reconfigurable and flexible (shrinkable or expandable) PEs according to the workload. It can generate an Field Programmable Gate Array (FPGA)-based custom accelerator in Register Transfer Language (RTL) with exactly the type of computations/data types required for the workload so that idle resources are not instantiated. FlexPE is evaluated on AMD-Xilinx ZedBoard and ZCU-104 FPGAs on PolyBench and MachSuite workloads. Through this approach, we achieve a remarkable reduction of resources by nearly 35x on the FPGAs. We achieve an impressive throughput increase of 81x and 43x on the two FPGAs on average, and 35x reduction in resources compared to related work. Lasya Punya Sree Gundumogula, Rachana Kaparthi, Himanshu Rai, Nanditha Rao |
ISCAS | 4 |
| 2024 | VPU-CIM: A 130nm, 33.98 TOPS/W RRAM based Compute-In-Memory Vector Co-ProcessorabstractDeep Learning inference on edge devices requires reduced memory load/store latency and low bit-precision computations. To address these challenges, we present VPU-CIM: a novel RRAM-based Compute-In-Memory (CIM) variable bit-precision vector co-processor. We introduce vector extensions to the RISC-V ISA and implement it as an in-memory compute unit with a unique data mapping strategy. The design is implemented using open-source Skywater 130nm PDK, with area estimates provided for TSMC 28nm and ASAP 7nm PDKs. Our design achieves an energy efficiency of 33.98 TOPS/W for a 4,4 (I, W) precision configuration. The results demonstrate the potential of RRAM-based vector computations in memory. J. Chithambara Moorthii, Vinay Rayapati, Nanditha Rao, Manan Suri |
ISCAS | 3 |
| 2024 | FPGA-based Hardware Software Co-design to Accelerate Brain Tumour SegmentationabstractBrain tumors are a major concern, being the leading cause of cancer-related deaths. Computer-aided diagnosis significantly reduces the workload on physicians and improves cancer diagnosis and treatment. Brain tumor segmentation is a computationally intensive image-processing task. In this paper, we propose an FPGA-based Hardware-Software Co-design to accelerate this task using Watershed and Otsu thresholding algorithms. The FPGA handles parallel components, while the CPU manages sequential tasks in the same System-on-Chip (SoC). Using PolarFire Icicle FPGA platform, we process 20 MRI brain scan images (128x128) from the Kaggle dataset. Implementing both algorithms in parallel on the FPGA results in a 1.97× acceleration compared to a CPU-only implementation, mainly achieved by a 1973× reduction in latency when moving the Otsu algorithm from the CPU to the FPGA. This optimization employs DSP/MATH blocks, loop unrolling, and pipelining techniques. Vinay Rayapati, Gogireddy Ravi Kiran Reddy, Gandi Ajay Kumar, Saketh Gajawada, Sanampudi Gopala Krishna Reddy, Nanditha Rao |
ISCAS | 6 |
| 2024 | Hybrid Multi-tile Vector Systolic Architecture for Accelerating Convolution on FPGAsabstractTo enhance the efficiency of image-kernel convolution operations in convolutional neural networks, we introduce a Vector Systolic Array Accelerator with adaptable lane-width. This architecture utilizes multiple vector lanes for concurrent data element processing to address computational bottlenecks. To further improve the throughput, we propose a novel hybrid tiled vector systolic design in which we partition the hardware resources to efficiently utilize them along with a unique data mapping strategy. In this approach, we choose some tiles to use LUTs and others to use DSPs. We observe that the throughput of the vector systolic accelerator is 7x and 1.26x higher for single-tile and multi-tile configurations than their non-vector counterparts respectively. The hybrid tile design significantly increases tile count, achieving a competitive peak throughput of 1165 GOPs and 1072 GOPs for optimal lane width of Vector-6 and 8, which is 3.8x better than related work. Jay Shah, Nanditha Rao |
ISCAS | 2 |
| 2019 | Analysis of the Effect of QoS on Video Conferencing QoEabstractNetwork service providers tend to focus on the quality of service (QoS) they provide to their customers. This entails analysis of various QoS metrics (such as bandwidth, packet loss and jitter) in order to be able to improve their services. This is a single-dimensional approach to a problem that needs to be analyzed not only from a business improvement perspective but also from a customer satisfaction perspective. QoS metrics do not directly translate to customer experience, which is more qualitative than quantitative. Thus, it is necessary to correlate qualitative metrics that customers relate to with quantitative metrics that can be analyzed and improved upon by service providers. This is a non-trivial problem that needs deeper exploration. In this paper, we attempt to correlate video conferencing QoE (Quality of Experience) with network QoS. In order to do this, we developed a novel Docker image called Lime, to be able to automate the experiments and emulate the network environment. We performed 144 separate video conferences under predefined network handicaps (scenarios). We discovered that bandwidth is directly proportional to the perceived quality of the video implying that higher bandwidth is preferred. On the other hand, frequently fluctuating bandwidth quickly reduced the user-opinion, and also resulted in slower subsequent climb in opinion after a period of high fluctuation. This indicated that steady bandwidth is preferred over irregularly increasing bandwidth. Jitter and packet loss were found to contribute to negative user-opinion as well as low bandwidth. Conversely, increasing jitter and packet loss was mostly forgiven if the bandwidth stayed stable and high. Lime is shown to be a novel tool to fulfill requirements related to video conferencing experiments under pre-defined network scenarios. Nanditha Rao, Amir Haghighati Maleki, Fu Chi Chen, Anwar Haque |
IWCMC | 1 |