VLDB 2026 Research / reviewers in the wild / expert
Subarna Chatterjee
dblp:139/7927
· DBLP profile ↗
15ranked-venue papers
8as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 4 · 3 first-author · 1 since 2021Computer networks · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Limousine: Blending Learned and Classical Indexes to Self-Design Larger-than-Memory Cloud Storage EnginesabstractWe present Limousine, a self-designing key-value storage engine, that can automatically morph to the near-optimal storage engine architecture shape given a workload, a cloud budget, and target performance. At its core, Limousine identifies the fundamental design principles of storage engines as combinations of learned and classical data structures that collaborate through algorithms for data storage and access. By unifying these principles over diverse hardware and three major cloud providers (AWS, GCP, and Azure), Limousine creates a massive design space of quindecillion (1048) storage engine designs the vast majority of which do not exist in literature or industry. Limousine contains a distribution-aware IO model to accurately evaluate any candidate design. Using these models, Limousine searches within the exhaustive design space to construct a navigable continuum of designs connected along a Pareto frontier of cloud cost and performance. If storage engines contain learned components, Limousine also introduces efficient lazy write algorithms to optimize the holistic read-write performance. Once the near-optimal design is decided for the given context, Limousine automatically materializes the corresponding design in Rust code. Using the YCSB benchmark, we demonstrate that storage engines automatically designed and generated by Limousine scale better by up to 3 orders of magnitude when compared with state-of-the-art industry-leading engines such as RocksDB, WiredTiger, FASTER, and Cosine, over diverse workloads, data sets, and cloud budgets. Subarna Chatterjee, Mark F. Pekala, Lev Kruglyak, Stratos Idreos |
Proc. ACM Manag. Data | 1 |
| 2022 | SNARF: A Learning-Enhanced Range FilterabstractWe present Sparse Numerical Array-Based Range Filters (SNARF), a learned range filter that efficiently supports range queries for numerical data. SNARF creates a model of the data distribution to map the keys into a bit array which is stored in a compressed form. The model along with the compressed bit array which constitutes SNARF are used to answer membership queries. We evaluate SNARF on multiple synthetic and real-world datasets as a stand-alone filter and by integrating it into RocksDB. For range queries, SNARF provides up to 50x better false positive rate than state-of-the-art range filters, such as SuRF and Rosetta, with the same space usage. We also evaluate SNARF in RocksDB as a filter replacement for filtering requests before they access on-disk data structures. For RocksDB, SNARF can improve the execution time of the system up to 10x compared to SuRF and Rosetta for certain read-only workloads. Kapil Vaidya, Tim Kraska, Subarna Chatterjee, Eric Knorr, Michael Mitzenmacher, Stratos Idreos |
Proc. VLDB Endow. | 3 |
| 2021 | Deep Learning: Systems and ResponsibilityabstractDeep learning enables numerous applications across diverse areas. Data systems researchers are also increasingly experimenting with deep learning to enhance data systems performance. We present a tutorial on deep learning, highlighting the data systems nature of neural networks as well as research opportunities for advancements through data management techniques. We focus on three critical aspects: (1) classic design tradeoffs in neural networks which we can enrich through a systems and data management perspective, e.g., thinking critically about storage, data movement, and computation; (2) classic design problems in data systems which we can reconsider with neural networks as a viable design option, e.g., to replace or help system components that make complex decisions such as database optimizers; and (3) essential considerations for responsible application of neural networks in critical human-facing problems in society and how these also link to data management and performance considerations. While these are seemingly a diverse set of rich topics, they are strongly interconnected through data management, and their combination offers rich opportunities for future research. Abdul Wasay, Subarna Chatterjee, Stratos Idreos |
SIGMOD Conference | 2 |
| 2021 | Cosine: A Cloud-Cost Optimized Self-Designing Key-Value Storage EngineabstractWe present a self-designing key-value storage engine, Cosine, which can always take the shape of the close to "perfect" engine architecture given an input workload, a cloud budget, a target performance, and required cloud SLAs. By identifying and formalizing the first principles of storage engine layouts and core key-value algorithms, Cosine constructs a massive design space comprising of sextillion (10 36 ) possible storage engine designs over a diverse space of hardware and cloud pricing policies for three cloud providers - AWS, GCP, and Azure. Cosine spans across diverse designs such as Log-Structured Merge-trees, B-trees, Log-Structured Hash-tables, in-memory accelerators for filters and indexes as well as trillions of hybrid designs that do not appear in the literature or industry but emerge as valid combinations of the above. Cosine includes a unified distribution-aware I/O model and a learned concurrency-aware CPU model that with high accuracy can calculate the performance and cloud cost of any possible design on any workload and virtual machines. Cosine can then search through that space in a matter of seconds to find the best design and materializes the actual code of the resulting storage engine design using a templated Rust implementation. We demonstrate that on average Cosine outperforms state-of-the-art storage engines such as write-optimized RocksDB, read-optimized WiredTiger, and very write-optimized FASTER by 53x, 25x, and 20x, respectively, for diverse workloads, data sizes, and cloud budgets across all YCSB core workloads and many variants. Subarna Chatterjee, Meena Jagadeesan, Wilson Qin, Stratos Idreos |
Proc. VLDB Endow. | 1 |
| 2021 | Big-Sensor-Cloud Infrastructure: A Holistic Prototype for Provisioning Sensors-as-a-ServiceabstractThe proposed work relates to the development ofBig-Sensor-Cloud Infrastructure(BSCI) that immensely enhances the usability and management of the physical sensor devices. Traditional Wireless Sensor Networks (WSNs) are manufactured in a proprietary, vendor-specific design. Thus, the renderability of WSNs is almost infeasible to people/organizations that do not own a network of their own. Thus, in the existing system, WSN-based applications are inaccessible to the naive-users or common people who do not own physical sensor devices. Recently, sensor-cloud infrastructure has been viewed as a substitute for traditional WSNs. However, with the increasing growth in the velocity, variety, and variability of data, the management becomes a serious concern and difficulty. Thus, existing systems are not able to capture, analyze, and control the present data efficiently, in real-time. BSCI is a distributed framework for “Big” sensor-data storage, processing, virtualization, leveraging, and efficient remote management. The methods of the proposed BSCI are persuasive as they are equipped with the ability to handle “Big” data with enormous heterogeneous data volumes (in zettabyte) generated with tremendous velocity. The framework interfaces between the physical and cyber worlds, thereby acquiring real-time data from the physical WSNs into the cloud platform. This data are processed and delivered to the end-users as a simple service – Sensors-as-a-Service (Se-aaS). BSCI completely maintains and manages the data and the metadata internally within its database. Multiple organizations with heterogeneous demand can be successfully served with Se-aaS through BSCI. From a user-perspective, BSCI is highly convenient as the users are completely abstracted from the underlying complex processing logic. This allows the naive users to envision the typical hardware sensor devices as simple accessible services like electricity, and water. Subarna Chatterjee, Arijit Roy 0002, Sanku Kumar Roy, Sudip Misra, Manmeet Singh Bhogal, Rachit Daga |
IEEE Trans. Cloud Comput. | 1 |
| 2020 | Rosetta: A Robust Space-Time Optimized Range Filter for Key-Value StoresabstractWe introduce Rosetta, a probabilistic range filter designed specifically for LSM-tree based key-value stores. The core intuition is that we can sacrifice filter probe time because it is not visible in end-to-end key-value store performance, which in turn allows us to significantly reduce the filter false positive rate for every level of the tree. Rosetta indexes all binary prefixes of a key using a hierarchically arranged set of Bloom filters. It then converts each range query into multiple probes, one for each non-overlapping binary prefix. Rosetta has the ability to track workload patterns and adopt a beneficial tuning for each individual LSM-tree run by adjusting the number of Bloom filters it uses and how memory is spread among them to optimize the FPR/CPU cost balance. We show how to integrate Rosetta in a full system, RocksDB, and we demonstrate that it brings as much as a 40x improvement compared to default RocksDB and 2-5x improvement compared to state-of-the-art range filters in a variety of workloads and across different levels of the memory hierarchy (memory, SSD, hard disk). We also show that, unlike state-of-the-art filters, Rosetta brings a net benefit in RocksDB's overall performance, i.e., it improves range queries without losing any performance for point queries. Siqiang Luo, Subarna Chatterjee, Rafael Ketsetsidis, Niv Dayan, Wilson Qin, Stratos Idreos |
SIGMOD Conference | 2 |
| 2019 | Optimal Data Center Scheduling for Quality of Service Management in Sensor-CloudabstractThe proposed work concentrates on the networking facets of sensor-cloud infrastructures-one of the first attempts of its kind. In a sensor-cloud, multiple sets of physical sensor nodes that are activated based on an application demand, in turn give rise to multiple distinct virtual sensors (VSs). The VSs are considered to span across multiple geographical regions; thereby, depositing the data (from each of the VS) to the closest cloud data center (DC). Quite obviously, multiple geospatial DCs get involve with an application. However, the principle of sensor-cloud is to store and conglomerate the data from various VSs, before they can be provisioned as Sensors-as-a-Service (Se-aaS). The assortation of data occurs within a single Virtual Machine (VM) (or in some cases multiple VMs) residing inside a particular DC. This work addresses the problem of scheduling a particular DC that congregates data from various VSs, and transmit the same to the end-user application. The work follows the general pairwise choice framework of the Optimal Decision Rule. The scheduling of the DC is performed under several network constraints, such as data migration cost, data delivery cost, and service delay of an application that ensures the preservation of the Quality-of-Service (QoS) and maintenance of the user satisfaction. The work quantifies the effective QoS of Se-aaS and determines an optimal decision rule for electing a particular DC. While arriving at a collective decision, the work incorporates the fallible decision making ability of a DC; thereby, excluding the loss of generality. Experimental results depict that the proposed algorithm for generating the optimal decision rule finds applicability in real-time cloud computing scenarios. Subarna Chatterjee, Sudip Misra, Samee Ullah Khan |
IEEE Trans. Cloud Comput. | 1 |
| 2018 | Experimental Study on the Performance and Resource Utilization of Data Streaming FrameworksabstractWith the advent of the Internet of Things (IoT), data stream processing have gained increased attention due to the ever-increasing need to process heterogeneous and voluminous data streams. This work addresses the problem of selecting a correct stream processing framework for a given application to be executed within a specific physical infrastructure. For this purpose, we focus on a thorough comparative analysis of three data stream processing platforms - Apache Flink, Apache Storm, and Twitter Heron (the enhanced version of Apache Storm), that are chosen based on their potential to process both streams and batches in real-time. The goal of the work is to enlighten the cloud-clients and the cloud-providers with the knowledge of the choice of the resource-efficient and requirement-adaptive streaming platform for a given application so that they can plan during allocation or assignment of Virtual Machines for application execution. For the comparative performance analysis of the chosen platforms, we have experimented using 8-node clusters on Grid5000 experimentation testbed and have selected a wide variety of applications ranging from a conventional benchmark to sensor-based IoT application and statistical batch processing application. In addition to the various performance metrics related to the elasticity and resource usage of the platforms, this work presents a comparative study of the “green-ness” of the streaming platforms by analyzing their power consumption - one of the first attempts of its kind. The obtained results are thoroughly analyzed to illustrate the functional behavior of these platforms under different computing scenarios. Subarna Chatterjee, Christine Morin |
CCGrid | 1 |
| 2018 | Assessment of the Suitability of Fog Computing in the Context of Internet of ThingsabstractThis work performs a rigorous, comparative analysis of the fog computing paradigm and the conventional cloud computing paradigm in the context of the Internet of Things (IoT), by mathematically formulating the parameters and characteristics of fog computing-one of the first attempts of its kind. With the rapid increase in the number of Internet-connected devices, the increased demand of real-time, low-latency services is proving to be challenging for the traditional cloud computing framework. Also, our irreplaceable dependency on cloud computing demands the cloud data centers (DCs) always to be up and running which exhausts huge amount of power and yield tons of carbon dioxide (CO2) gas. In this work, we assess the applicability of the newly proposed fog computing paradigm to serve the demands of the latency-sensitive applications in the context of IoT. We model the fog computing paradigm by mathematically characterizing the fog computing network in terms of power consumption, service latency, CO2emission, and cost, and evaluating its performance for an environment with high number of Internet-connected devices demanding real-time service. A case study is performed with traffic generated from the 100 highest populated cities being served by eight geographically distributed DCs. Results show that as the number of applications demanding real-time service increases, the fog computing paradigm outperforms traditional cloud computing. For an environment with 50 percent applications requesting for instantaneous, real-time services, the overall service latency for fog computing is noted to decrease by 50.09 percent. However, it is mentionworthy that for an environment with less percentage of applications demanding for low-latency services, fog computing is observed to be an overhead compared to the traditional cloud computing. Therefore, the work shows that in the context of IoT, with high number of latency-sensitive applications fog computing outperforms cloud computing. Subhadeep Sarkar 0001, Subarna Chatterjee, Sudip Misra |
IEEE Trans. Cloud Comput. | 2 |
| 2018 | Link-Quality Aware Path Selection in the Presence of Proactive Jamming in Fallible Wireless Sensor NetworksabstractIn this paper, we propose a mechanism to ensure the proper functioning of a wireless sensor network (WSN) in the presence of static and proactive jammer within the network. Existing research works have primarily focused on the detection of jammer node within the network and ameliorating its consequent effects on the network. However, these countermeasures to mitigate the effect of the jammer suffer from certain limitations as WSNs are primarily resource constrained and most of the countermeasures are computationally intensive. The objective of this paper is to prevent the disruption of the network in the presence of jamming by bypassing the jammed zone and setting up alternative paths. The alternative paths are chosen with the maximum link quality in order to maintain the quality of service of the network even after jamming. This paper proposes Link-quality Aware Path SElection (LAPSE) algorithm that chooses alternative paths based on the optimal link quality. LAPSE is based on optimal decision rule and its design considers the fallible nature of the nodes while choosing/rejecting a particular link. Finally, the performance of the proposed algorithm, LAPSE, is evaluated in terms of the network parameters-packet delivery rate, network throughput, transmission energy, node lifetime, and network lifetime. Results indicate that the performance of LAPSE is significantly better than the existing jamming avoidance algorithms. Prasenjit Bhavathankar, Subarna Chatterjee, Sudip Misra |
IEEE Trans. Commun. | 2 |
| 2017 | Dynamic Optimal Pricing for Heterogeneous Service-Oriented Architecture of Sensor-Cloud InfrastructureabstractThis paper proposes a dynamic and optimal pricing scheme for provisioning Sensors-as-a-Service (Se-aaS) [1] within the sensor-cloud infrastructure. Existing cloud pricing models are limited in terms of the homogeneity in service-types, and hence, are not compliant for the heterogeneous service oriented architecture of Se-aaS. We propose a new pricing model comprising of two components, applicable for Se-aaS architecture: pricing attributed to Hardware (pH) and pricing attributed to Infrastructure (pI). pH addresses the problem of pricing the physical sensor nodes subject to variable demand and utility of the end-users. It maximizes the profit incurred by every sensor owner, while keeping in mind the end-users' utility. pI mainly focuses on the pricing incurred due to the virtualization of resources. It takes into account the cost for the usage of the infrastructural resources, inclusive of the cost for maintaining virtualization within sensor-cloud. pI maximizes the profit of the sensor-cloud service provider (SCSP) by considering the user satisfaction. Simulation results depict improved performance of pH in comparison to the traditional hardware pricing algorithms, viz. PPM and Sprite, in terms of the residual energy, proximity to the base station (BS), received signal strength (RSS), overhead, and cumulative energy consumption. The results also show the tendency of the sensor-owners to converge to the end-user utility, but not exceed it. We also analyze the performance of pI. The results show the optimality in the profit incurred by SCSP and the user satisfaction. Subarna Chatterjee, Ranjana Ladia, Sudip Misra |
IEEE Trans. Serv. Comput. | 1 |
| 2016 | QoS estimation and selection of CSP in oligopoly environment for Internet of ThingsabstractThis work focuses on an automated selection of Cloud Service Provider (CSP) for a naive end-user in an IoT scenario. In traditional cloud computing model, the end-users are knowledgeable about the Virtual Machines (VMs) and are technically aware of their requirements in terms of the computing cores, processing abilities, and storage requirements. In case of IoT, the users are envisioned to be widespread from naive, unsophisticated people to even objects or things who are devoid of the required knowledge and expertise. Further, in IoT technology, multiple Cloud Service Providers (CSPs) may possess the potential of serving an IoT application. Therefore, it is required for the end-user to judiciously select a single CSP based on the maximum obtainable Quality of Service (QoS) from a CSP. This work proposes an algorithm QoS based Automated Selection of CSP (QASeC) for automated selection of a CSP from a set of nominated CSPs based on the maximum achievable QoS. The work identifies and models the QoS parameters for every CSP and defines a QoS utility metric for each CSP. Based on the metric, the work proposes an optimization for selection of the appropriate CSP and the cloud gateway associated with it. From the obtained results, we infer the suitability of QASeC in real-life IoT scenarios. Subarna Chatterjee, Sudip Misra |
WCNC | 1 |
| 2015 | Optimal composition of a virtual sensor for efficient virtualization within sensor-cloudabstractThe work focuses on optimal formation of virtual sensors (VSs) within a sensor-cloud infrastructure. Existing work on sensor-cloud have considered the formation of VS with the maximal set of compatible physical sensor nodes. However, as these underlying nodes are highly resource constrained, inefficient and redundant utilization of the nodes takes a toll on the entire performance of the cloud and the network. In this work, we propose algorithms for efficient virtualization of the physical sensor nodes and optimal composition of VSs - within the same geographic region (CoV-I) and spanning across multiple regions (CoV-II). Experimental results demonstrate that, compared to the existing strategy of maximal composition of VSs, CoV-I improves the cumulative energy consumption and the network lifetime by 34.9% and 61.04%, respectively, and CoV-II enhances the parameters by 68.4% and 29.59%, respectively. Subarna Chatterjee, Sudip Misra |
ICC | 1 |
| 2015 | QoS-aware sensor allocation for target tracking in sensor-cloud
Sudip Misra, Anuj Singh, Subarna Chatterjee, Amit Kumar Mandal |
Ad Hoc Networks | 3 |
| 2014 | Social choice considerations in cloud-assisted WBAN architecture for post-disaster healthcare: Data aggregation and channelization
Sudip Misra, Subarna Chatterjee |
Inf. Sci. | 2 |