EDBT 2026 Demo / reviewers in the wild / expert
Manoj Pravakar Saha
dblp:289/8489
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
0009-0003-1120-2337ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LATTICE: Efficient In-Memory DNN Model VersioningabstractDNN model versions are used for various tasks such as fine-tuning for downstream tasks, explainability, and debugging. Numerous checkpointing solutions exist that can be adapted to persist intermediate versions of a model, as it is being trained, at different storage locations. Additionally, version management tools allow us to log, visualize, compare, and query metadata related to ML, tracking changes made to previously built models. However, the version creation process of existing methods incurs high runtime and storage overheads. In this paper, we introduce LATTICE, a low-latency, direct persistence-based DNN versioning library for Non-Volatile Memory (NVM) expansion devices. LATTICE minimizes stalls during model versioning and reduces end-to-end versioning time by reorganizing the version creation workflow, streamlining memory allocation and deallocation for efficient snapshot creation, and leveraging multi-threaded parallelism. We also develop a user-friendly versioning API that transparently implements direct persistence. Our comprehensive evaluation with diverse DNN models shows that LATTICE can reduce persistence time by as much as 99.99%, decrease end-to-end versioning time by up to 72%, reduce versioning stalls by up to 35%, and increase versioning frequency by 0.2×-3.84× compared to state-of-the-art solutions. LATTICE also reduces space utilization for different workloads. The space savings are from 23.8% to 43.2% for workloads where model layers are progressively frozen and from 84.8% to 98.9% for fine-tuning workloads where only the last layers are tuned. Manoj Pravakar Saha, Ashikee Ghosh, Raju Rangaswami, Yanzhao Wu 0001, Janki Bhimani |
SYSTOR | 1 |
| 2023 | MoKE: Modular Key-value Emulator for Realistic Studies on Emerging Storage DevicesabstractKey-value stores are widely used as building blocks in today's IT infrastructure for managing and storing large amounts of data. Storage technologies are undergoing continuous innovations to accelerate KV workloads. However, designing high-performance KV or object storage devices is challenging and still needs more research to address the performance bottlenecks of the existing designs. There is a void for an inexpensive and extendable research platform that enables in-depth exploration of the index management components within the KV devices. To fill this void, we design Modular Key-value Emulator (MoKE). MoKE is a software emulator for fostering future full-stack software/hardware KV and object storage device research. MoKE is cheap (software-based emulator), usable with SNIA KV API (supports popular host-device interfaces), extendable (supports internal KV device research), and adaptable (QEMU-based). Manoj Pravakar Saha, Danlin Jia, Janki Bhimani, Ningfang Mi |
CLOUD | 1 |
| 2023 | Allocation Policies Matter for Hybrid Memory SystemsabstractExisting tiered memory systems all use DRAM-Preferred as their allocation policy, whereby pages get allocated from higher-performing DRAM until it is filled, after which all future allocations are made from lower-performing persistent memory (PM). The novel insight of this work is that the right page allocation policy for a workload can help to lower the access latencies for the newly allocated pages. We design, implement, and evaluate three page allocation policies within the real system deployment of the state-of-the-art dynamic tiering system. We observe that the right page allocation policy can improve the performance of a tiered memory system by as much as 17x for certain workloads. Adnan Maruf, Daniel Carlson, Ashikee Ghosh, Manoj Pravakar Saha, Janki Bhimani, Raju Rangaswami |
HPDC | 4 |
| 2023 | Leveraging Keys In Key-Value SSD for Production WorkloadsabstractKey-Value SSDs reduce host-side resource utilization for unstructured data management by streamlining the I/O stack. However, designing a robust Key-Value SSD with resource constrained flash controllers has always been a challenge. The key-to-page (K2P) mapping inside KV-SSD, which consolidates multiple layers of indirection in the traditional block I/O storage, has its own shortcomings. The sparsely populated NVMe KV namespace leads to very large index, which cannot be optimized similar to hybrid- or block-FTL in block-SSDs. In addition, the background index management tasks (e.g. compaction on LSM-tree index) also lead to performance degradation. Moreover, existing KV index design is not equipped to tackle fast changing workload patterns. These shortcomings have stalled the adoption of KV-SSDs in production environments. In this work, we take the position that these shortcomings can be addressed by leveraging the information embedded inside keys about application keyspaces and groups as prefixes. The prefixes can be used to partition the monolithic large index into smaller ones. We demonstrate a naive prefix-based index partitioning mechanism inside KV-SSD that can reduce on-flash index accesses for multiple production workloads and discuss the shortcomings of this approach. Lastly, we discuss our proposed design of a society of indices that initialize, interact and evolve based on workload characteristics over time. Manoj Pravakar Saha, Omkar Desai, Bryan S. Kim, Janki Bhimani |
HPDC | 1 |
| 2023 | RHIK: Re-configurable Hash-based Indexing for KVSSDabstractKey-Value Solid State Drive (KV-SSD), a key addressable SSD technology, promises to simplify storage management for unstructured data and improve system performance with minimal host-side intervention. However, we find that the current state-of-the-art KV-SSD exhibits indexing peculiarities that limit their widespread adoption. Through experiments, we observe that the performance degrades as more data are stored, and the KV-SSD can only store a limited number of key-value pairs even though the amount of data stored on the device is significantly lower than its capacity. We introduce RHIK, a reconfigurable hash-bashed indexing for KV-SSD, for high performance and high occupancy. We implement our proposed indexing scheme on the open-source KV-SSD emulator that is validated against a real KV-SSD, and demonstrate its effectiveness using real workload traces and synthetic microbenchmarks. Manoj Pravakar Saha, Bryan S. Kim, Haryadi S. Gunawi, Janki Bhimani |
HPDC | 1 |
| 2021 | KV-SSD: What Is It Good For?abstractAn increasing concern that curbs the widespread adoption of KV-SSD is whether or not offloading host-side operations to the storage device changes device behavior, negatively affecting various applications’ overall performance. In this paper, we systematically measure, quantify, and understand the performance of KV-SSD by studying the impact of its distinct components such as indexing, data packing, and key handling on I/O concurrency, garbage collection, and space utilization. Our experiments and analysis uncover that KV-SSD’s behavior differs from well-known idiosyncrasies of block-SSD. Proper understanding of its characteristics will enable us to achieve better performance for random, read-heavy, and highly concurrent workloads. Manoj Pravakar Saha, Adnan Maruf, Bryan S. Kim, Janki Bhimani |
DAC | 1 |
| 2021 | Fine-grained control of concurrency within KV-SSDsabstractThe development of KV-SSDs allows simplifying the I/O stack compared to the traditional block-based SSDs. We propose a novel Key-Value-based Storage infrastructure for Parallel Computing(KV-SiPC)-a framework for multi-thread OpenMP applications to use NVMe-based KV-SSDs. We design a new capability to execute workloads with multiple parallel data threads along with traditional parallel compute threads, that allow us to improve the overall throughput of applications, utilizing the maximum possible storage bandwidth. We implement our KV-SiPC infrastructure in a real system by extending various processing layers (e.g., program, OS, and device layers) and evaluate the performance of KV-SiPC by using block-based NVMe SSDs in the traditional I/O stack as a baseline for comparisons. The experimental results show that KV-SiPC can better utilize the available device bandwidth and significantly increases application I/O throughput. Janki Bhimani, Jingpei Yang, Ningfang Mi, Changho Choi, Manoj Pravakar Saha, Adnan Maruf |
SYSTOR | 5 |
| 2020 | Performance and Consistency Analysis for Distributed Deep Learning ApplicationsabstractAccelerating the training of Deep Neural Network (DNN) models is very important for successfully using deep learning techniques in fields like computer vision and speech recognition. Distributed frameworks help to speed up the training process for large DNN models and datasets. Plenty of works have been done to improve model accuracy and training efficiency, based on mathematical analysis of computations in the Con-volutional Neural Networks (CNN). However, to run distributed deep learning applications in the real world, users and developers need to consider the impacts of system resource distribution. In this work, we deploy a real distributed deep learning cluster with multiple virtual machines. We conduct an in-depth analysis to understand the impacts of system configurations, distribution typologies, and application parameters, on the latency and correctness of the distributed deep learning applications. We analyze the performance diversity under different model consistency and data parallelism by profiling run-time system utilization and tracking application activities. Based on our observations and analysis, we develop design guidelines for accelerating distributed deep-learning training on virtualized environments. Danlin Jia, Manoj Pravakar Saha, Janki Bhimani, Ningfang Mi |
IPCCC | 2 |