EDBT 2026 Demo / reviewers in the wild / expert
Janki Bhimani
dblp:171/1567
· DBLP profile ↗
39ranked-venue papers
11as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 4 first-author · 19 since 2021Computer networks · 11 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
Omkar Desai, Ziyang Jiao, Shuyi Pei, Janki Bhimani, Bryan S. Kim |
FAST | 4 |
| 2026 | KORAL: Knowledge Graph Guided LLM Reasoning for SSD Operational Analysis
Mayur Akewar, Sandeep Madireddy, Janki Bhimani |
IPDPS | 4 |
| 2026 | Holpaca: Holistic and Adaptable Cache Management for Shared EnvironmentsabstractModern data-intensive systems rely on in-memory caching to achieve high throughput and low latency. CacheLib, Meta's general-purpose caching engine, provides high performance and flexibility for building specialized caches for a variety of applications. However, despite its wide adoption in large-scale infrastructures, CacheLib's data management mechanisms exhibit inefficiencies in shared environments. Particularly, its static and uncoordinated memory allocation leads to fragmented resource usage, unfair memory distribution, and degraded performance across tenants and instances. José Pedro Peixoto, Alexis González, Janki Bhimani, Raju Rangaswami, Cláudia Brito, João Paulo 0001, Ricardo Macedo |
ICPE | 3 |
| 2025 | Heimdall: Optimizing Storage I/O Admission with Extensive Machine Learning PipelineabstractThis paper introduces Heimdall, a highly accurate and efficient machine learning-powered I/O admission policy for flash storage, designed to operate in a black-box manner. We make domain-specific innovations in various ML stages by introducing accurate period-based labeling, 3-stage noise filtering, in-depth feature engineering, and fine-grained tuning, which together improve the decision accuracy from 67% up to 93%. We perform various deployment optimizations to reach a sub-μs inference latency and a small, 28KB, memory overhead. With 500 unbiased random experiments derived from production traces, we show Heimdall delivers 15-35% lower average I/O latency compared to the state of the art and up to 2x faster to a baseline. Heimdall is ready for user-level, in-kernel, and distributed deployments. Daniar Heri Kurniawan, Rani Ayu Putri, Peiran Qin, Kahfi S. Zulkifli, Ray A. O. Sinurat, Janki Bhimani, Sandeep Madireddy, Achmad I. Kistijantoro, Haryadi S. Gunawi |
EuroSys | 6 |
| 2025 | Can LLMs Model the Environmental Impact on SSD?abstractEnvironmental stressors such as temperature, humidity, vibration, and radiation can severely impact the performance and reliability of SSDs, particularly in edge, automotive, aerospace, and datacenter deployments. Capturing sensor data in the field and conducting accelerated lab experiments are challenging, as they are time-consuming, resource-intensive, and often destructive to hardware. Specialized setups, such as thermal chambers or vibration rigs, are also required, which is why few studies explore this area, and current storage management techniques like RAID, tiering, and deduplication do not consider environmental factors. Models to capture these impacts would open new research opportunities across various fields. However, accurately modeling these effects remains challenging due to, (1) the limited availability of experimental data, (2) the complex, domino-like impact of historical exposure, (3) the interrelated nature of environmental factors, such as temperature and humidity, which exhibit correlation, (4) different response of each type of NAND flash memory TLC, MLC, and SLC to environmental factors, and (5) the difficulty that analytical and simple machine learning models face in generalizing across devices, environments, and unseen combinations of stressors. We believe that LLMs may offer a transformative alternative to this complex problem, with embedded domain knowledge and reasoning capabilities, to facilitate prompt-based natural language interaction. We propose a hybrid framework that combines Chain-of-Thought prompting and Retrieval-Augmented Generation to guide LLMs using physical principles and prior experiments. It enables interpretable "what-if" analysis of SSD behavior under environmental changes. Our results show that the LLM can effectively model the impact of temperature, humidity, and vibration on SSD performance, producing tail latency and bandwidth predictions with minimal error. The code and data are available on GitHub at https://github.com/Damrl-lab/SSD_LLM. Mayur Akewar, Gang Quan, Sandeep Madireddy, Janki Bhimani |
HotStorage | 4 |
| 2025 | Quantum Neural Networks Need CheckpointingabstractQuantum Neural Networks (QNNs) harness quantum superposition and entanglement, offering promising advantages for machine learning tasks. However, noise in quantum computers frequently disrupts QNN training, wasting computational resources and extending queue times. This paper introduces the first QNN checkpointing framework to address this challenge. Through experiments on various quantum devices, we demonstrate that QNN behavior is fundamentally hardware-dependent, with the same model performing differently across platforms. This key finding shows that quantum checkpoints require additional metadata about hardware specifics and shot counts unique to quantum systems. Our framework requires minimal storage (only 186.6KB for a 100-qubit QNN) and negligible overhead, enabling frequent checkpointing to enhance training resilience and reproducibility in the NISQ era. Christopher Kverne, Mayur Akewar, Yuqian Huo, Tirthak Patel, Janki Bhimani |
HotStorage | 5 |
| 2025 | Revisiting Noise-adaptive Transpilation in Quantum Computing: How Much Impact Does it Have?abstractTranspilation, particularly noise-aware optimization, is widely regarded as essential for maximizing the performance of quantum circuits on superconducting quantum computers. The common wisdom is that each circuit should be transpiled using up-to-date noise calibration data to optimize fidelity. In this work, we revisit the necessity of frequent noise-adaptive transpilation, conducting an in-depth empirical study across five IBM 127-qubit quantum computers and 16 diverse quantum algorithms. Our findings reveal novel and interesting insights: (1) noise-aware transpilation leads to a heavy concentration of workloads on a small subset of qubits, which increases output error variability; (2) using random mapping can mitigate this effect while maintaining comparable average fidelity; and (3) circuits compiled once with calibration data can be reliably reused across multiple calibration cycles and time periods without significant loss in fidelity. These results suggest that the classical overhead associated with daily, per-circuit noise-aware transpilation may not be justified. We propose lightweight alternatives that reduce this overhead without sacrificing fidelity – offering a path to more efficient and scalable quantum workflows. Yuqian Huo, Jinbiao Wei, Christopher Kverne, Mayur Akewar, Janki Bhimani, Tirthak Patel |
ICCAD | 5 |
| 2025 | LATTICE: Efficient In-Memory DNN Model VersioningabstractDNN model versions are used for various tasks such as fine-tuning for downstream tasks, explainability, and debugging. Numerous checkpointing solutions exist that can be adapted to persist intermediate versions of a model, as it is being trained, at different storage locations. Additionally, version management tools allow us to log, visualize, compare, and query metadata related to ML, tracking changes made to previously built models. However, the version creation process of existing methods incurs high runtime and storage overheads. In this paper, we introduce LATTICE, a low-latency, direct persistence-based DNN versioning library for Non-Volatile Memory (NVM) expansion devices. LATTICE minimizes stalls during model versioning and reduces end-to-end versioning time by reorganizing the version creation workflow, streamlining memory allocation and deallocation for efficient snapshot creation, and leveraging multi-threaded parallelism. We also develop a user-friendly versioning API that transparently implements direct persistence. Our comprehensive evaluation with diverse DNN models shows that LATTICE can reduce persistence time by as much as 99.99%, decrease end-to-end versioning time by up to 72%, reduce versioning stalls by up to 35%, and increase versioning frequency by 0.2×-3.84× compared to state-of-the-art solutions. LATTICE also reduces space utilization for different workloads. The space savings are from 23.8% to 43.2% for workloads where model layers are progressively frozen and from 84.8% to 98.9% for fine-tuning workloads where only the last layers are tuned. Manoj Pravakar Saha, Ashikee Ghosh, Raju Rangaswami, Yanzhao Wu 0001, Janki Bhimani |
SYSTOR | 5 |
| 2025 | Storage Abstractions for SSDs: The Past, Present, and FutureabstractThis article traces the evolution of SSD (solid-state drive) interfaces, examining the transition from the block storage paradigm inherited from hard disk drives to SSD-specific standards customized to flash memory. Early SSDs conformed to the block abstraction for compatibility with the existing software storage stack, but studies and deployments show that this limits the performance potential for SSDs. As a result, new SSD-specific interface standards emerged to not only capitalize on the low latency and abundant internal parallelism of SSDs, but also include new command sets that diverge from the longstanding block abstraction. We first describe flash memory technology in the context of the block storage abstraction and the components within an SSD that provide the block storage illusion. We then describe the genealogy and relationships among academic research and industry standardization efforts for SSDs, along with some of their rise and fall in popularity. We classify these works into four evolving branches: (1) extending block abstraction with host-SSD hints/directives; (2) enhancing host-level control over SSDs; (3) offloading host-level management to SSDs; and (4) making SSDs byte-addressable. By dissecting these trajectories, the article also sheds light on the emerging challenges and opportunities, providing a roadmap for future research and development in SSD technologies. Xiangqun Zhang 0002, Janki Bhimani, Shuyi Pei, Sungjin Lee 0001, Yoon Jae Seong, Eui Jin Kim, Changho Choi, Eyee Hyun Nam, Jongmoo Choi, Bryan S. Kim |
ACM Trans. Storage | 2 |
| 2024 | Learning-Based Dynamic Memory Allocation Schemes for Apache Spark Data ProcessingabstractApache Spark is an in-memory analytic framework that has been adopted in the industry and research fields. Two memory managers, Static and Unified, are available in Spark to allocate memory for caching Resilient Distributed Datasets (RDDs) and executing tasks. However, we find that the static memory manager (SMM) lacks flexibility, while the unified memory manager (UMM) puts heavy pressure on the garbage collection of the JVM on which Spark resides. To address these issues, we design a learning-based bidirectional usage-bounded memory allocation scheme to support dynamic memory allocation with the consideration of both memory demands and latency introduced by garbage collection. We first develop an auto-tuning memory manager (ATuMm) that adopts an intuitive feedback-based learning solution. However, ATuMm is a slow learner that can only alter the states of Java Virtual Memory (JVM) Heap in a limited range. That is, ATuMm decides to increase or decrease the boundary between the execution and storage memory pools by a fixed portion of JVM Heap size. To overcome this shortcoming, we further develop a new reinforcement learning-based memory manager (Q-ATuMm) that uses a Q-learning intelligent agent to dynamically learn and tune the partition of JVM Heap. We implement our new memory managers in Spark 2.4.0 and evaluate them by conducting experiments in a real Spark cluster. Our experimental results show that our memory manager can reduce the total garbage collection time and thus further improve Spark applications’ performance (i.e., reduced latency) compared to the existing Spark memory management solutions. By integrating our machine learning-driven memory manager into Spark, we can further obtain around 1.3x times reduction in the latency. Danlin Jia, Natalia Valencia, Janki Bhimani, Bo Sheng, Ningfang Mi |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | MoKE: Modular Key-value Emulator for Realistic Studies on Emerging Storage DevicesabstractKey-value stores are widely used as building blocks in today's IT infrastructure for managing and storing large amounts of data. Storage technologies are undergoing continuous innovations to accelerate KV workloads. However, designing high-performance KV or object storage devices is challenging and still needs more research to address the performance bottlenecks of the existing designs. There is a void for an inexpensive and extendable research platform that enables in-depth exploration of the index management components within the KV devices. To fill this void, we design Modular Key-value Emulator (MoKE). MoKE is a software emulator for fostering future full-stack software/hardware KV and object storage device research. MoKE is cheap (software-based emulator), usable with SNIA KV API (supports popular host-device interfaces), extendable (supports internal KV device research), and adaptable (QEMU-based). Manoj Pravakar Saha, Danlin Jia, Janki Bhimani, Ningfang Mi |
CLOUD | 3 |
| 2023 | Allocation Policies Matter for Hybrid Memory SystemsabstractExisting tiered memory systems all use DRAM-Preferred as their allocation policy, whereby pages get allocated from higher-performing DRAM until it is filled, after which all future allocations are made from lower-performing persistent memory (PM). The novel insight of this work is that the right page allocation policy for a workload can help to lower the access latencies for the newly allocated pages. We design, implement, and evaluate three page allocation policies within the real system deployment of the state-of-the-art dynamic tiering system. We observe that the right page allocation policy can improve the performance of a tiered memory system by as much as 17x for certain workloads. Adnan Maruf, Daniel Carlson, Ashikee Ghosh, Manoj Pravakar Saha, Janki Bhimani, Raju Rangaswami |
HPDC | 5 |
| 2023 | Leveraging Keys In Key-Value SSD for Production WorkloadsabstractKey-Value SSDs reduce host-side resource utilization for unstructured data management by streamlining the I/O stack. However, designing a robust Key-Value SSD with resource constrained flash controllers has always been a challenge. The key-to-page (K2P) mapping inside KV-SSD, which consolidates multiple layers of indirection in the traditional block I/O storage, has its own shortcomings. The sparsely populated NVMe KV namespace leads to very large index, which cannot be optimized similar to hybrid- or block-FTL in block-SSDs. In addition, the background index management tasks (e.g. compaction on LSM-tree index) also lead to performance degradation. Moreover, existing KV index design is not equipped to tackle fast changing workload patterns. These shortcomings have stalled the adoption of KV-SSDs in production environments. In this work, we take the position that these shortcomings can be addressed by leveraging the information embedded inside keys about application keyspaces and groups as prefixes. The prefixes can be used to partition the monolithic large index into smaller ones. We demonstrate a naive prefix-based index partitioning mechanism inside KV-SSD that can reduce on-flash index accesses for multiple production workloads and discuss the shortcomings of this approach. Lastly, we discuss our proposed design of a society of indices that initialize, interact and evolve based on workload characteristics over time. Manoj Pravakar Saha, Omkar Desai, Bryan S. Kim, Janki Bhimani |
HPDC | 4 |
| 2023 | RHIK: Re-configurable Hash-based Indexing for KVSSDabstractKey-Value Solid State Drive (KV-SSD), a key addressable SSD technology, promises to simplify storage management for unstructured data and improve system performance with minimal host-side intervention. However, we find that the current state-of-the-art KV-SSD exhibits indexing peculiarities that limit their widespread adoption. Through experiments, we observe that the performance degrades as more data are stored, and the KV-SSD can only store a limited number of key-value pairs even though the amount of data stored on the device is significantly lower than its capacity. We introduce RHIK, a reconfigurable hash-bashed indexing for KV-SSD, for high performance and high occupancy. We implement our proposed indexing scheme on the open-source KV-SSD emulator that is validated against a real KV-SSD, and demonstrate its effectiveness using real workload traces and synthetic microbenchmarks. Manoj Pravakar Saha, Bryan S. Kim, Haryadi S. Gunawi, Janki Bhimani |
HPDC | 4 |
| 2022 | Do Temperature and Humidity Exposures Hurt or Benefit Your SSDs?abstractSSDs are becoming mainstream data storage de-vices, replacing HDDs in most data centers, consumer goods, and IoT gadgets. In this work, we ask an uncharted research question: What is the environmental conditions' impact on SSD performance? To answer it, we systematically measure, quantify, and characterize the impact of various commonly changing envi-ronmental conditions such as temperature and humidity on the performance of SSDs. Our experiments and analysis uncover that exposure to changes in temperature and humidity can significantly affect SSD performance. Adnan Maruf, Sashri Brahmakshatriya, Baolin Li 0001, Devesh Tiwari, Gang Quan, Janki Bhimani |
DATE | 6 |
| 2022 | Wear leveling in SSDs considered harmfulabstractWe argue that wear leveling in SSDs does more harm than good under modern settings where the endurance limit is in the hundreds. To support this claim, we evaluate existing wear leveling techniques and show that they exhibit anomalous behaviors and produce a high write amplification. These findings are consistent with a recent large-scale field study on the operational characteristics of SSDs. We discuss the option of forgoing wear leveling and instead adopting capacity variance in SSDs, and show that the capacity variance extends the lifetime of the SSD by up to 2.94×. Ziyang Jiao, Janki Bhimani, Bryan S. Kim |
HotStorage | 2 |
| 2022 | MULTI-CLOCK: Dynamic Tiering for Hybrid Memory SystemsabstractThe rapid growth of i-memory computing powered by data-intensive applications has increased demand for DRAM in servers. However, a DRAM-based system can be limiting for modern workloads because of its capacity, cost, and power consumption characteristics. Hybrid memory systems, which consist of different types of memory, such as DRAM and persistent memory, can help address many of these limitations. One promising direction that has been explored in the recent literature involves introducing persistent memory devices as a second memory tier that is directly exposed to the CPU. The resulting tiered memory design must address the fundamental challenge of placing the right data in the right memory tier at the right time while minimizing overhead. We present MULTI -CLOCK, an efficient, low-overhead hybrid memory system that relies on a unique page selection technique for tier placement. MULTl-CLOCK’s page selection captures both page access recency and frequency, and enables moving pages to appropriate tiers at the right time within hybrid memory systems. We implemented a Linux-based, NUMA-aware version of MULTI-CLOCK that is entirely transparent and backward compatible with any existing application. Our evaluation with diverse real-world applications such as graph processing and key-value stores shows that MULTI -CLOCK can improve the average throughput by as much as 352% when compared with several state-of-the-art techniques for tiered memory. Adnan Maruf, Ashikee Ghosh, Janki Bhimani, Daniel Campello, Andy Rudoff, Raju Rangaswami |
HPCA | 3 |
| 2022 | I/O Workload Management for All-Flash Datacenter Storage Systems Based on Total Cost of OwnershipabstractRecently, the capital expenditure of flash-based Solid State Driver (SSDs) keeps declining and the storage capacity of SSDs keeps increasing. As a result, all-flash storage systems have started to become more economically viable for large shared storage installations in datacenters, where metrics like Total Cost of Ownership (TCO) are of paramount importance. On the other hand, flash devices suffer from write amplification, which, if unaccounted, can substantially increase the TCO of a storage system. In this paper, we first develop a TCO model for datacenter all-flash storage systems, and then plug a Write Amplification model (WAF) of NVMe SSDs we build based on empirical data into this TCO model. Our new WAF model accounts for workload characteristics like write rate and percentage of sequential writes. Furthermore, using both the TCO and WAF models as the optimization criterion, we design new flash resource management schemes (minTCO) to guide datacenter managers to make workload allocation decisions under the consideration of TCO for SSDs. Based on that, we also developminTCO-RAIDto support RAID SSDs andminTCO-Offlineto optimize the offline workload-disk deployment problem during the initialization phase. Experimental results show thatminTCOcan reduce the TCO and keep relatively high throughput and space utilization of the entire datacenter storage resources. Zhengyu Yang 0001, Manu Awasthi, Mrinmoy Ghosh, Janki Bhimani, Ningfang Mi |
IEEE Trans. Big Data | 4 |
| 2022 | Auto-Tuning Parameters for Emerging Multi-Stream Flash-Based Storage Drives Through New I/O Pattern GenerationsabstractIn the era of big data processing, more and more data centers in cloud storage are now replacing traditional HDDs with enterprise SSDs. Both developers and users of these SSDs require thorough benchmarking to evaluate and configure the variable parameters of emerging technologies.[2]and[3]are the recent development of the SSD industry, which assists in placing data on SSDs in a smart way to improve application performance and SSD endurance. The challenging part to use multi-stream SSDs is to assign stream IDs to incoming writes, such that each stream consists of data with a similar lifetime. The benefit of the stream management algorithms varies over different workloads. Thus, first, we propose a new framework, calledPatternI/Ogenerator (PatIO), to capture the enterprise storage behavior that is prevailing across various user workloads, virtualization setup, file systems, and volume managers for the database server applications on flash-based storage. Second, usingPatIO, we study what type of applications may be benefited by which stream assignment algorithm. Third, we design the framework to automatically tune the variable parameters of different stream identification algorithms of the multi-stream SSDs. Our evaluation shows 20 to 110 percent of the reward function increase, measuring the cumulative impact on application performance and SSD endurance. Janki Bhimani, Adnan Maruf, Ningfang Mi, Rajinikanth Pandurangan, Vijay Balakrishnan |
IEEE Trans. Computers | 1 |
| 2022 | Automatic Stream Identification to Improve Flash Endurance in Data CentersabstractThe demand for high performance I/O in Storage-as-a-Service (SaaS) is increasing day by day. To address this demand, NAND Flash-based Solid-state Drives (SSDs) are commonly used in data centers as cache- or top-tiers in the storage rack ascribe to their superior performance compared to traditional hard disk drives (HDDs). Meanwhile, with the capital expenditure of SSDs declining and the storage capacity of SSDs increasing, all-flash data centers are evolving to serve cloud services better than SSD-HDD hybrid data centers. During this transition, the biggest challenge is how to reduce the Write Amplification Factor (WAF) as well as to improve the endurance of SSD since this device has a limited program/erase cycles. A specified case is that storing data with different lifetimes (i.e., I/O streams with similar temporal fetching patterns such as reaccess frequency) in one single SSD can cause high WAF, reduce the endurance, and downgrade the performance of SSDs. Motivated by this, multi-stream SSDs have been developed to enable data with a different lifetime to be stored in different SSD regions. The logic behind this is to reduce the internal movement of data—when garbage collection is triggered, there are high chances of having data blocks with either all the pages being invalid or valid. However, the limitation of this technology is that the system needs to manually assign the same streamID to data with a similar lifetime. Unfortunately, when data arrives, it is not known how important this data is and how long this data will stay unmodified. Moreover, according to our observation, with different definitions of a lifetime (i.e., different calculation formulas based on selected features previously exhibited by data, such as sequentiality, and frequency), streamID identification may have varying impacts on the final WAF of multi-stream SSDs. Thus, in this article, we first develop a portable and adaptable framework to study the impacts of different workload features and their combinations on write amplification. We then propose a feature-based stream identification approach, which automatically co-relates the measurable workload attributes (such as I/O size, I/O rate, and so on.) with high-level workload features (such as frequency, sequentiality, and so on.) and determines a right combination of workload features for assigning streamIDs . Finally, we develop an adaptable stream assignment technique to assign streamID for changing workloads dynamically. Our evaluation results show that our automation approach of stream detection and separation can effectively reduce the WAF by using appropriate features for stream assignment with minimal implementation overhead. Janki Bhimani, Zhengyu Yang 0001, Jingpei Yang, Adnan Maruf, Ningfang Mi, Rajinikanth Pandurangan, Changho Choi, Vijay Balakrishnan |
ACM Trans. Storage | 1 |
| 2021 | Understanding Flash-Based Storage I/O Behavior of GamesabstractComputer games are an extremely popular but overlooked workload. Cloud-gaming has been one of the biggest buzzwords in the gaming industry throughout 2020. The rapid growth of the video gaming industry and the diverse set of popular video games available today raises increasing concern to properly understand its I/O characteristics to improve their performance and design better gaming servers and consoles. To the best of our knowledge, this is the first attempt to systematically measure, quantify, and characterize the organization of game data into files, back-end storage access patterns, and the performance of gaming workloads. We explore the I/O behavior of 14 recent and famous games, producing a series of observations coming from measurements done on a real setup. Adnan Maruf, Zhengyu Yang 0001, Bridget Davis, Jeffrey Wong, Matthew Durand, Janki Bhimani |
CLOUD | 7 |
| 2021 | KV-SSD: What Is It Good For?abstractAn increasing concern that curbs the widespread adoption of KV-SSD is whether or not offloading host-side operations to the storage device changes device behavior, negatively affecting various applications’ overall performance. In this paper, we systematically measure, quantify, and understand the performance of KV-SSD by studying the impact of its distinct components such as indexing, data packing, and key handling on I/O concurrency, garbage collection, and space utilization. Our experiments and analysis uncover that KV-SSD’s behavior differs from well-known idiosyncrasies of block-SSD. Proper understanding of its characteristics will enable us to achieve better performance for random, read-heavy, and highly concurrent workloads. Manoj Pravakar Saha, Adnan Maruf, Bryan S. Kim, Janki Bhimani |
DAC | 4 |
| 2021 | Fine-grained control of concurrency within KV-SSDsabstractThe development of KV-SSDs allows simplifying the I/O stack compared to the traditional block-based SSDs. We propose a novel Key-Value-based Storage infrastructure for Parallel Computing(KV-SiPC)-a framework for multi-thread OpenMP applications to use NVMe-based KV-SSDs. We design a new capability to execute workloads with multiple parallel data threads along with traditional parallel compute threads, that allow us to improve the overall throughput of applications, utilizing the maximum possible storage bandwidth. We implement our KV-SiPC infrastructure in a real system by extending various processing layers (e.g., program, OS, and device layers) and evaluate the performance of KV-SiPC by using block-based NVMe SSDs in the traditional I/O stack as a baseline for comparisons. The experimental results show that KV-SiPC can better utilize the available device bandwidth and significantly increases application I/O throughput. Janki Bhimani, Jingpei Yang, Ningfang Mi, Changho Choi, Manoj Pravakar Saha, Adnan Maruf |
SYSTOR | 1 |
| 2020 | Performance and Consistency Analysis for Distributed Deep Learning ApplicationsabstractAccelerating the training of Deep Neural Network (DNN) models is very important for successfully using deep learning techniques in fields like computer vision and speech recognition. Distributed frameworks help to speed up the training process for large DNN models and datasets. Plenty of works have been done to improve model accuracy and training efficiency, based on mathematical analysis of computations in the Con-volutional Neural Networks (CNN). However, to run distributed deep learning applications in the real world, users and developers need to consider the impacts of system resource distribution. In this work, we deploy a real distributed deep learning cluster with multiple virtual machines. We conduct an in-depth analysis to understand the impacts of system configurations, distribution typologies, and application parameters, on the latency and correctness of the distributed deep learning applications. We analyze the performance diversity under different model consistency and data parallelism by profiling run-time system utilization and tracking application activities. Based on our observations and analysis, we develop design guidelines for accelerating distributed deep-learning training on virtualized environments. Danlin Jia, Manoj Pravakar Saha, Janki Bhimani, Ningfang Mi |
IPCCC | 3 |
| 2019 | What does Vibration do to Your SSD?abstractVibration generated in modern computing environments such as autonomous vehicles, edge computing infrastructure, and data center systems is an increasing concern. In this paper, we systematically measure, quantify and characterize the impact of vibration on the performance of SSD devices. Our experiments and analysis uncover that exposure to both short-term and long-term vibration, even within the vendor-specified limits, can significantly affect SSD I/O performance and reliability. Janki Bhimani, Tirthak Patel, Ningfang Mi, Devesh Tiwari |
DAC | 1 |
| 2019 | Emulate Processing of Assorted Database Server Applications on Flash-Based Storage in Datacenter InfrastructuresabstractIn the era of big data processing, more and more datacenters in cloud storages are now replacing traditional HDDs with enterprise SSDs. Both developers and users of these SSDs require thorough benchmarking to evaluate their performance impacts. I/O performance with synthetic workload or classic benchmark varies drastically from real I/O activities in the datacenter. Thus, we propose a new framework, called Pattern I/O generator (PatIO), to collectively capture the enterprise storage behavior that is prevailing across assorted user workloads and system configurations for different database server applications on flash-based storage. PatIO is designed to emulate the processing of real-world I/O activities easily with less time and resource requirements. Our methodology comprises three main steps: (1) dissect the overall I/O activities of various real workloads and identify the prevailing attributes in distinct visual I/O patterns; (2) construct a pattern warehouse as the collection of unique I/O patterns that are generated through various combinations of multiple I/O jobs; and (3) finally integrate different combinations of these synthetically generated I/O patterns to reproduce the comprehensive characteristics of various real workloads and system setup for the database server applications. To provide an easy-to-use experience, we develop a graphical user interface (GUI). We evaluate our framework by comparing I/O characteristics and I/O performance of generated workloads with those of real-world workloads for multiple database applications such as MySQL, Cassandra, and ForestDB. Janki Bhimani, Rajinikanth Pandurangan, Ningfang Mi, Vijay Balakrishnan |
IPCCC | 1 |
| 2019 | ATuMm: Auto-tuning Memory Manager in Apache SparkabstractApache Spark is an in-memory analytic framework that has been adopted in the industry and research fields. Two memory managers, Static and Unified, are available in Spark to allocate memory for caching Resilient Distributed Datasets (RDDs) and executing tasks. However, we found that the static memory manager (SMM) lacks flexibility, while the unified memory manager (UMM) puts heavy pressure on the garbage collection of JVM on which Spark resides. To address these issues, we design an auto-tuning memory manager (ATuMm) to support dynamic memory allocation with the consideration of both memory demands and latency introduced by garbage collection. We implement our new memory manager in Spark 2.2.0 and evaluate it by conducting experiments in a real Spark cluster. Our experimental results show that our auto-tuning memory manager can reduce the total garbage collection time and thus further improve the performance (i.e., reduced latency) of Spark applications, compared to the existing Spark memory management solutions. Danlin Jia, Janki Bhimani, Son Nam Nguyen, Bo Sheng, Ningfang Mi |
IPCCC | 2 |
| 2018 | BloomStream: Data Temperature Identification for Flash Based Memory Storage Using Bloom FiltersabstractData temperature identification is an importance issue of many fields like data caching and storage tiering in modern flash-based storage systems. With the technological advancement of memory and storage, data temperature identification is no longer just a classification of hot and cold, but instead becomes a "multistreaming" data categorization problem to classify data into multiple categories according to their temperature. Therefore, we propose a novel data temperature identification scheme that adopts bloom filters to efficiently capture both frequency and recency of data blocks and accurately identify the exact data temperature for each data block. Moreover, in bloom filter data structure we replace the original OR operation with the XOR masking operation such that our scheme can delete or reset bits in bloom filters and thus avoid high false positives due to saturation. We further utilize twin bloom filters to alternatively keep unmasked clean copies of data and thus ensure low false negative rate. Our extensive evaluation results show that our new scheme can accurately identify the exact data temperature with low false identification rates across different synthetic and real I/O workloads. More importantly, our scheme consumes less memory space compared to other existing data temperature identification schemes. Janki Bhimani, Ningfang Mi, Bo Sheng |
IEEE CLOUD | 1 |
| 2018 | FIOS: Feature Based I/O Stream Identification for Improving Endurance of Multi-Stream SSDsabstractThe demand for high speed 'Storage-as-a-Service' (SaaS) is increasing day-by-day. SSDs are commonly used in higher tiers of storage rack in data centers. Also, all flash data centers are evolving to better serve cloud services. Although SSDs guaranty better performance when compared to HDDs, but SSDs endurance is still a matter of concern. Storing data with different lifetime in an SSD can cause high write amplification and reduce the endurance and performance of SSDs. Recently, multi-stream SSDs have been developed to enable data with different lifetime to be stored in different SSD regions and thus reduce write amplification. To efficiently use this new multi-streaming technology, it is important to choose appropriate workload features to assign the same streamID to data with similar lifetime. However, we found that streamID identification using different features may have varying impacts on the final write amplification of multi-stream SSDs. Therefore, in this paper we develop a portable and adoptable framework to study the impacts of different workload features and their combinations on write amplification. We also introduce a new feature, named "coherency", to capture the friendship among write operations with respect to their update time. Finally, we propose a feature-based stream identification approach, which co-relates the measurable workload attributes (such as I/O size, I/O rate, etc.) with high level workload features (such as frequency, sequentiality etc.) and determines a good combination of workload features for assigning streamIDs. Our evaluation results show that our proposed approach can always reduce the Write Amplification Factor (WAF) by using appropriate features for stream assignment. Janki Bhimani, Ningfang Mi, Zhengyu Yang 0001, Jingpei Yang, Rajinikanth Pandurangan, Changho Choi, Vijay Balakrishnan |
IEEE CLOUD | 1 |
| 2017 | FIM: Performance Prediction for Parallel Computation in Iterative Data Processing ApplicationsabstractPredicting performance of an application running on high performance computing (HPC) platforms in a cloud environment is increasingly becoming important because of its influence on development time and resource management. However, predicting the performance with respect to parallel processes is complex for iterative, multi-stage applications. This research proposes a performance approximation approach FiM to model the computing performance of iterative, multi-stage applications running on a master-compute framework. FiM consists of two key components that are coupled with each other: 1) Stochastic Markov Model to capture non-deterministic runtime that often depends on parallel resources, e.g., number of processes. 2) Machine Learning Model that extrapolates the parameters for calibrating our Markov model when we have changes in application parameters such as dataset. Our new modeling approach considers different design choices along multiple dimensions, namely (i) process level parallelism, (ii) distribution of cores on multi-core processors in cloud computing, (iii) application related parameters, and (iv) characteristics of datasets. The major contribution of our prediction approach is that FiM is able to provide an accurate prediction of parallel computation time for the datasets which have much larger size than that of the training datasets. Such calculation prediction provides data analysts a useful insight of optimal configuration of parallel resources (e.g., number of processes and number of cores) and also helps system designers to investigate the impact of changes in application parameters on system performance. Janki Bhimani, Ningfang Mi, Miriam Leeser, Zhengyu Yang 0001 |
CLOUD | 1 |
| 2017 | AutoPath: Harnessing Parallel Execution Paths for Efficient Resource Allocation in Multi-Stage Big Data FrameworksabstractDue to the flexibility of data operations and scalability of in- memory cache, Spark has revealed the potential to become the standard distributed framework to replace Hadoop for data-intensive processing in both industry and academia. However, we observe that the built-in scheduling algorithms in Spark (i.e., FIFO and FAIR) are not optimized for the applications with multiple parallel and independent branches in stages. Specifically, the child stage needs to wait and collect data from all its parent branches, but this wait has no guaranteed upper bound since it is tightly coupled with each branch's workload characteristic, stage order, and their corresponding allocated computing resource. To address this challenge, we investigate a superior solution which ensures all branches acquire suitable resources according to their workload demand in order to let the finish time of each branch be as close as possible. Based on this, we propose a novel scheduling policy, named AutoPath, which can effectively reduce the overall makespan of such kind of applications by detecting and leveraging the parallel path, and adaptively assigning computing resources based on the estimated workload demands during runtime. We implemented the new scheduling scheme in Spark v1.5.0 and evaluated it with selected representative workloads. The experiments demonstrate that our new scheduler effectively reduces the makespan and improves resource utilizations for these applications, compared to the current FIFO and FAIR schedulers. Han Gao 0013, Zhengyu Yang 0001, Janki Bhimani, Bo Sheng, Ningfang Mi |
ICCCN | 3 |
| 2017 | Enhancing SSDs with multi-stream: What? why? how?abstractThe adoption of SSDs has become very prominent, but they still suffer from challenges to control write amplification. Traditional SSDs have single active append point where new data writes can be stored. Data of different lifetime stored together causes high write amplification. Recently, multi-stream SSDs are developed that allows multiple active append points. These multiple active append points can be used to store data of different lifetime in different locations within SSD. Such a data placement according to the lifetime of data would considerably reduce internal write amplification of SSD. For using multi-stream SSDs it is required to attach stream-id to each new incoming data writes. According to these stream-ids, the flash transition layer (FTL) of a multi-stream SSD then appends data to different erase blocks. Thus, multi-stream SSDs will help to reduce write amplification. But, to efficiently use this new multi-stream SSDs, it is important to properly identify streamids of data with respect to its lifetime. The lifetime of data is expected using different features that data exhibits like frequency, sequentiality etc. Stream-id identification using different features may have different impact on the final write amplification of multi-stream SSDs, depending on workload. Thus, it is required to quantify the impact of different data features that are used for stream-id identification on the resultant write amplification. Additionally, the combination of these data features may be used for stream-id identification, so it is also important to be able to study the impact of such different combinations. In order to address above challenges towards efficiently using multi-stream SSDs, here we propose a portable and adoptable framework to study the impact of stream-id identification using different workload data features and their combinations on write amplification of multi-stream SSDs. Our evaluation results show that use of appropriate features according to workload can considerably reduce the Write Amplification Factor (WAF) when compared to the legacy SSDs. Janki Bhimani, Jingpei Yang, Zhengyu Yang 0001, Ningfang Mi, N. H. V. Krishna Giri, Rajinikanth Pandurangan, Changho Choi, Vijay Balakrishnan |
IPCCC | 1 |
| 2017 | Cyber-physical system enabled nearby traffic flow modelling for autonomous vehiclesabstractWe propose a nearby traffic flow modelling solution based on built-in Cyber-Physical System (CPS) sensors of autonomous vehicles. Our goal is to enhance the offline route planning and driving decision adjustment based on the first-hand traffic information, especially during poor Internet connection moments. Specifically, our model helps to select the optimal speed on a road, the optimal distance for timing to brake, and the safe distance from other vehicles to keep. Moreover, our model can also assist neighboring autonomous vehicles by communicating required information through Ad-Hoc network communications or through a centralized cloud. In detail, we first focus on the unique characteristic of traffic flow (such as traffic rule, avoid collision behaviours), and then build a comprehensive model to handle multiple scenarios. Technically, our model uses density functions of velocities, the differential equation of traffic flows, and the traffic viscosity with information collected from the traffic flow, the distances between vehicles, the amount and density of vehicle, the instant velocity, the speed limit, and the momentum to analysis the the driving scene. We evaluate our model with real traffic data collected by in-vehicle CPS sensors to the proposed nearby traffic flow model. Results show that our work can accurately conduct offline estimation on nearby traffic signal influence, and reveal the correlations among velocity, density and (spatial and temporal) location to adjust route during runtime. Zhengyu Yang 0001, Siyu Huang, Xianzhi Du, Janki Bhimani, Ningfang Mi |
IPCCC | 6 |
| 2017 | AutoTiering: Automatic data placement manager in multi-tier all-flash datacenterabstractIn the year of 2017, the capital expenditure of Flash-based Solid State Drivers (SSDs) keeps declining and the storage capacity of SSDs keeps increasing. As a result, the “selling point” of traditional spinning Hard Disk Drives (HDDs) as a backend storage — low cost and large capacity — is no longer unique, and eventually they will be replaced by low-end SSDs which have large capacity but perform orders of magnitude better than HDDs. Thus, it is widely believed that all-flash multi-tier storage systems will be adopted in the enterprise datacenters in the near future. However, existing caching or tiering solutions for SSD-HDD hybrid storage systems are not suitable for all-flash storage systems. This is because that all-flash storage systems do not have a large speed difference (e.g., 10x) among each tier. Instead, different specialties (such as high performance, high capacity, etc.) of each tier should be taken into consideration. Motivated by this, we develop an automatic data placement manager called “AutoTiering” to handle virtual machine disk files (VMDK) allocation and migration in an all-flash multitier datacenter to best utilize the storage resource, optimize the performance, and reduce the migration overhead. AutoTiering is based on an optimization framework, whose core technique is to predict VM's performance change on different tiers with different specialties without conducting real migration. As far as we know, AutoTiering is the first optimization solution designed for all-flash multi-tier datacenters. We implement AutoTiering on VMware ESXi [1], and experimental results show that it can significantly improve the I/O performance compared to existing solutions. Zhengyu Yang 0001, Morteza Hoseinzadeh, Allen Andrews, Clay Mayers, David Thomas Evans, Rory Thomas Bolt, Janki Bhimani, Ningfang Mi, Steven Swanson |
IPCCC | 7 |
| 2017 | H-NVMe: A hybrid framework of NVMe-based storage system in cloud computing environmentabstractIn the year of 2017, more and more datacenters have started to replace traditional SATA and SAS SSDs with NVMe SSDs due to NVMe's outstanding performance [1]. However, for historical reasons, current popular deployments of NVMe in VM-hypervisor-based platforms (such as VMware ESXi [2]) have numbers of intermediate queues along the I/O stack. As a result, performance is bottlenecked by synchronization locks in these queues, cross-VM interference induces I/O latency, and most importantly, up-to-64K-queue capability of NVMe SSDs cannot be fully utilized. In this paper, we developed a hybrid framework of NVMe-based storage system called “H-NVMe”, which provides two VM I/O stack deployment modes “Parallel Queue Mode” and “Direct Access Mode”. The first mode increases parallelism and enables lock-free operations by implementing local lightweight queues in the NVMe driver. The second mode further bypasses the entire I/O stack in the hypervisor layer and allows trusted user applications whose hosting VMDK (Virtual Machine Disk) files are attached with our customized vSphere IOFilters [3] to directly access NVMe SSDs to improve the performance isolation. This suits premium users who have higher priorities and the permission to attach IOFilter to their VMDKs. H-NVMe is implemented on VMware EXSi 6.0.0, and our evaluation results show that the proposed H-NVMe framework can significant improve throughputs and bandwidths compared to the original inbox NVMe solution. Zhengyu Yang 0001, Morteza Hoseinzadeh, Ping Wong, John Artoux, Clay Mayers, David Thomas Evans, Rory Thomas Bolt, Janki Bhimani, Ningfang Mi, Steven Swanson |
IPCCC | 8 |
| 2017 | Docker characterization on high performance SSDsabstractDocker containers are becoming the mainstay for deploying applications in cloud platforms, having many desirable features like ease of deployment, developer friendliness and lightweight virtualization. Meanwhile, solid state disks (SSDs) have witnessed tremendous performance boost through recent innovations in industry such as Non-Volatile Memory Express (NVMe) standards. However, the performance of containerized applications on these high speed contemporary SSDs has not yet been investigated. In this paper, we present a characterization of the performance impact among a wide variety of the available storage options for deploying Docker containers and provide the configuration options to best utilize the high performance SSDs. Qiumin Xu, Manu Awasthi, Krishna T. Malladi, Janki Bhimani, Jingpei Yang, Murali Annavaram |
ISPASS | 4 |
| 2016 | Performance prediction techniques for scalable large data processing in distributed MPI systemsabstractPredicting performance of an application running on parallel computing platforms is increasingly becoming important due to the long development time of an application and the high resource management cost of parallel computing platforms. However, predicting overall performance is complex and must take into account both parallel calculation time and communication time. Difficulty in accurate performance modeling is compounded by myriad design choices along multiple dimensions, namely (i) process level parallelism, (ii) distribution of cores on multi-processor platforms, (iii) application related parameters, and (iv) characteristics of datasets. This research proposes a fast and accurate performance prediction approach to predict the calculation and communication time of an application running on a distributed computing platform. The major contribution of our prediction approach is that it can provide an accurate prediction of execution times for new datasets which have much larger sizes than the training datasets. Our approach consists of two models, i.e., a probabilistic self-learning model to predict calculation time and a simulation queuing model to predict network communication time. The combination of these two models provides data analysts a useful insight of optimal configuration of parallel resources (e.g., number of processes and number of cores) and application parameters setting. Janki Bhimani, Ningfang Mi, Miriam Leeser |
IPCCC | 1 |
| 2016 | Understanding performance of I/O intensive containerized applications for NVMe SSDsabstractOur cloud-based IT world is founded on hyper-visors and containers. Containers are becoming an important cornerstone, which is increasingly used day-by-day. Among different available frameworks, docker has become one of the major adoptees to use containerized platform in data centers and enterprise servers, due to its ease of deploying and scaling. Further more, the performance benefits of a lightweight container platform can be leveraged even more with a fast back-end storage like high performance SSDs. However, increase in number of simultaneously operating docker containers may not guarantee an aggregated performance improvement due to saturation. Thus, understanding performance bottleneck in a multi-tenancy docker environment is critically important to maintain application level fairness and perform better resource management. In this paper, we characterize the performance of persistent storage option (through data volume) for I/O intensive, dockerized applications. Our work investigates the impact on performance with increasing number of simultaneous docker containers in different workload environments. We provide, first of its kind study of I/O intensive containerized applications operating with NVMe SSDs. We show that 1) a six times better application throughput can be obtained, just by wise selection of number of containerized instances compared to single instance; and 2) for multiple application containers running simultaneously, an application throughput may degrade upto 50% compared to a stand-alone applications throughput, if good choice of application and workload is not made. We then propose novel design guidelines for an optimal and fair operation of both homogeneous and heterogeneous environments mixed with different applications and workloads. Janki Bhimani, Jingpei Yang, Zhengyu Yang 0001, Ningfang Mi, Qiumin Xu, Manu Awasthi, Rajinikanth Pandurangan, Vijay Balakrishnan |
IPCCC | 1 |
| 2016 | GReM: Dynamic SSD resource allocation in virtualized storage systems with heterogeneous IO workloadsabstractIn a shared virtualized storage system that runs VMs with heterogeneous IO demands, it becomes a problem for the hypervisor to cost-effectively partition and allocate SSD resources among multiple VMs. There are two straightforward approaches to solving this problem: equally assigning SSDs to each VM or managing SSD resources in a fair competition mode. Unfortunately, neither of these approaches can fully utilize the benefits of SSD resources, particularly when the workloads frequently change and bursty IOs occur from time to time. In this paper, we design a Global SSD Resource Management solution - GReM, which aims to fully utilize SSD resources as a second-level cache under the consideration of performance isolation. In particular, GReM takes dynamic IO demands of all VMs into consideration to split the entire SSD space into a long-term zone and a short-term zone, and cost-effectively updates the content of SSDs in these two zones. GReM is able to adaptively adjust the reservation for each VM inside the long-term zone based on their IO changes. GReM can further dynamically partition SSDs between the long- and short-term zones during runtime by leveraging the feedbacks from both cache performance and bursty workloads. Experimental results show that GReM can capture the cross-VM IO changes to make correct decisions on resource allocation, and thus obtain high IO hit ratio and low IO management costs, compared with both traditional and state-of-the-art caching algorithms. Zhengyu Yang 0001, Jianzhe Tai, Janki Bhimani, Ningfang Mi, Bo Sheng |
IPCCC | 3 |