VLDB 2026 Research / reviewers in the wild / expert
Chandranil Chakraborttii
dblp:321/4739 · also Chandranil Nil Chakraborttii
· DBLP profile ↗
13ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-4142-6089ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synthetic Data Generation for Storage Failure Prediction in Large-Scale Systems
Chandranil Chakraborttii, Ana Veroneze Solórzano, Devesh Tiwari |
CCGrid | 1 |
| 2026 | TAP++: a tier-aware transformer for adaptive prefetching system in hybrid memory systems
Chandranil Chakraborttii |
Neural Comput. Appl. | 1 |
| 2024 | MemFlex: A Hybrid Memory System to Boost Cost of Ownership in Data CentersabstractModern large-scale computing clusters face scalability challenges with traditional DRAM-based memory systems due to issues like increasing cell leakage current and reduced reliability. To overcome these limitations, alternative memory solutions have emerged, including 3D-stacked DRAM and emerging non-volatile memory (NVM) technologies. However, these alternatives are unlikely to fully replace DRAM due to capacity constraints and higher cost-per-bit. Hybrid memory systems, combining DRAM with NVM technologies, offer a cost-effective solution by leveraging the strengths of both memory types. Effective data placement decisions are crucial for optimizing hybrid memory systems. This paper introduces MemFlex, an machine learning (ML)-driven approach for migrating pages between memory tiers in a hybrid memory system based on predicted page lifetimes. MemFlex utilizes application-specific ML models to predict death-time ranges with high accuracy, guiding placement decisions to the appropriate storage tier. Evaluation results demonstrate MemFlex's superiority over state-of-the-art techniques, achieving an average performance improvement of 19% over evaluated baselines, with minimal performance degradation. This paper contributes to hybrid memory management research, leveraging real-world traces for evaluation, and introducing an ML based approach for optimized data placement decisions. Chandranil Chakraborttii, Mohammad Maminur Islam |
COMPSAC | 1 |
| 2024 | Improving Insurance Fraud Detection With Generated DataabstractInsurance fraud poses a significant threat to the in-dustry, with estimated annual losses of $40 billion. Traditional efforts to combat fraud have relied on identifying fraudulent traits, but machine learning (ML) techniques offer promising avenues for detection. However, ML methods face challenges due to data scarcity and class imbalance in fraud datasets, hindering model efficacy. Recent advancements in Generative Adversarial Networks (GANs), present a solution by generating synthetic data resembling real instances. Accurate representational data generation also offers an avenue of preserving data privacy. In this paper, we propose a novel approach leveraging GANs to generate accurate representational data of fraudulent transactions, addressing data scarcity and privacy concerns. By incorporating generated data into model training, we show predictive performance improvement by up to 24%. Our empirical results demonstrate the efficacy of the proposed method in augmenting fraud detection capabilities while preserving privacy, thus offering a promising direction for improving insurance fraud detection. Our research contributes to advancing fraud detection in the insurance domain, with potential applications in other sectors facing similar challenges. Future research directions include exploring other generative models and combining them with traditional ML techniques to further enhance fraud detection capabilities. Kiet Ha, Lucas Stowe, Chandranil Chakraborttii |
COMPSAC | 3 |
| 2024 | Towards Generating Surprising Content in 2D Platform GamesabstractSurprise is a key factor for driving engagement in video games. By incorporating unexpected elements, games can create a sense of excitement and curiosity, leading to a more immersive and enjoyable experience. In this paper, we propose a framework for generating levels in 2D platform games with embedded surprising content. Building upon the VCL (Violation of Expectation, Caught Off Guard, and Learning) model, we use it as a conceptual foundation to curate surprise by altering specific game elements (called metrics) during the level generation process. We developed a tile-based parametric level generator for Super Mario Bros., creating customized levels based on metrics such as linearity, enemy density, and pattern variation, among others. 393 participants in our study played 2 generated levels, with the first setting expectations and the second intentionally violating them by altering chosen metrics. By comparing player responses and gameplay data with the VCL model, we explore the phenomenon of surprise in games. Our findings reveal statistically significant correlations between certain metrics and player responses, teasing at the potential for automatically generating surprising levels in 2D platform games. Chandranil Chakraborttii, Lucas Ferreira |
FDG | 1 |
| 2023 | Enabling Multi-tenancy on SSDs with Accurate IO Interference ModelingabstractTechnological advancements in the past decades have substantially increased the capacity and performance of Solid State Drives (SSDs). Provisioning such high-capacity SSDs among tenants can reap multiple benefits, such as elevated performance, efficient resource utilization, and cost savings through reduced Total Cost of Ownership. However, workloads perform poorly when co-located with others on the same SSD due to IO Interference, potentially violating Service Level Objectives (SLOs). High overprovisioning can address the SLO issue, however, it entails low utilization. Prior works proposed Machine Learning (ML) techniques to predict SSD performance in the presence of interfering tenants for optimizing workload placement. However, we find that these works suffer from two notable limitations. First, previous ML models do not capture interference impact due to the non-uniform workload characteristics and SSD internals. Second, they fail to compute interference of an arbitrary number of workloads due to a lack of feature aggregation. As a result, these works still offer low utilization and can only enforce weak SLOs. To address these limitations, we propose a Gray-box feature representation and aggregation technique to capture the IO interference impact of multiple non-uniform workloads based on internal SSD characteristics. Our technique improves prediction accuracy by 12x (lower mean absolute error) over prior works, resulting in up to 60% higher resource utilization or enforcing up to 2.5× stricter SLOs. Lokesh N. Jaliminche, Chandranil Chakraborttii, Changho Choi, Heiner Litz |
SoCC | 2 |
| 2023 | Leveraging Temporality of Data to Improve Failure Predictions for Solid State Drives in Data CentersabstractThis paper introduces an unsupervised multivariate anomaly detection framework leveraging generative adversarial networks (GANs) to predict solid-state drive (SSD) failures. Recent prior research leveraged the spatial locality of drive failures, along with performance data, to improve drive failure predictions. However, this information is seldom available in most publicly available data sets and can be challenging to collect, requiring additional overhead. In this work, we show that drive failures have a strong temporal correlation that can be used to improve drive failure predictions. We analyze drive observations from over 30,000 SSDs from Google’s data center spanning six years and show that our GAN-based approach improves the performance of failed drive predictions by 12% compared to the state-of-the-art. In contrast to prior work, we utilize the difference in drive observations and leverage LSTMs (long short-term memory) with GANs to capture the temporal correlations in data. Additionally, we introduce a hierarchical prediction approach that can accurately predict replacements and infant mortality in SSDs with an accuracy of 79% and 98% respectively. Chandranil Chakraborttii, Jonas Boettner |
COMPSAC | 1 |
| 2023 | Towards data generation to alleviate privacy concerns for cybersecurity applicationsabstractWhile sharing of data is vital for learning progression and knowledge development, its full effectiveness is limited due to concerns about privacy and the presence of stringent regulations. This issue is particularly grave in the domain of cybersecurity applications where client data often comprises confidential and sensitive information. Furthermore, cybersecurity datasets tend to suffer from class imbalance, where data related to cyber attacks are rare compared to the benign conditions. Hence, performing machine learning (ML) tasks such as attack detection and classification becomes a challenging endeavor. Synthetic tabular data has emerged as a viable alternative to enable data sharing while satisfying regulatory and privacy constraints. In this paper, we present a methodology that utilizes the Intrusion Detection System (IDS) dataset to generate synthetic tabular representational data from raw dataset while addressing class imbalance issues during the data generation process. The methodology incorporates a feature selection process that identifies the most important features that help with accurate data generation, and demonstrates comparable performance using popular machine learning (ML) techniques on the anomaly detection task. The similarity between the original and generated datasets is evaluated using two metrics - distribution metric and data reduction metric - achieving up to 0.97 similarity score on the data reduction metric, outperforming a baseline approach that uses all input features by up to 11%. Dhiraj Ganji, Chandranil Chakraborttii |
COMPSAC | 2 |
| 2021 | Reducing write amplification in flash by death-time prediction of logical block addressesabstractFlash-based solid state drives lack support for in-place updates, and hence deploy a flash translation layer to absorb the writes. For this purpose, SSDs implement a log-structured storage system introducing garbage collection and write-amplification overheads. In this paper, we present a machine learning based approach for reducing write amplification in log structured file systems via death-time prediction of logical block addresses. We define death-time of a data element as the number of I/O writes before which the data element is overwritten. We leverage the sequential nature of I/O accesses to train lightweight, yet powerful, temporal convolutional network (TCN) based models to predict death-times of logical blocks in SSDs. We leverage the predicted death-times in designing ML-DT, a near-optimal data placement technique that minimizes write amplification (WA) in log structured storage systems. We compare our approach with three state-of-the-art data placement schemes and show that ML-DT achieves the lowest WA by utilizing the learnt I/O death-time patterns from real-world storage workloads. Our proposed approach results in up to 14% reduction in write amplification compared to the best baseline technique. Additionally, we present a mapping learning technique to test the applicability of our approach to new or unseen workloads and present a hyper-parameter sensitive study. Chandranil Chakraborttii, Heiner Litz |
SYSTOR | 1 |
| 2020 | Improving the accuracy, adaptability, and interpretability of SSD failure prediction modelsabstractFlash-based solid state drives represent an important storage tier in today's hyperscale data centers. Although solid state drives (SSDs) are relatively reliable, data center operators are interested in predicting future drive failures to administer drive replacement, data migration, and drive acquisition strategies. We analyze telemetry data from over 30,000 SSDs running live applications in Google's datacenters over a span of six years, for predicting and explaining SSD failures using machine learning techniques. We propose the use of 1-class isolation forest and autoencoder-based anomaly detection techniques for predicting previously unseen SSD failure types with high accuracy. We show that ignoring the minority class for training can improve the performance by up to 9.5% and if adaptability to dynamic environments is required, by up to 13%. Furthermore, this paper proposes to utilize 1-class autoencoders to enable model interpretability. In particular, our autoencoder-based approach enables reasoning about the causes that lead to SSD failures. Common to all approaches, we deploy a set of powerful feature selection techniques that improve the model performance by up to 1.3X and reduce training times by up to 1.8X. Chandranil Chakraborttii, Heiner Litz |
SoCC | 1 |
| 2020 | Learning I/O Access Patterns to Improve Prefetching in SSDs
Chandranil Chakraborttii, Heiner Litz |
ECML/PKDD (4) | 1 |
| 2018 | SSD QoS Improvements through Machine LearningabstractThe recent deceleration of Moore's law bespeaks new approaches for optimization of resources. Machine learning has been applied to a wide variety of problems across multiple domains; however, the space of machine learning research for storage optimization is only lightly explored. In this paper, we focus on learning IO access patterns with the aim of improving the performance of flash based devices. Flash based storage devices provide orders of magnitude better performance than HDDs, but they suffer from high tail latencies due to garbage collection (GC) which causes variable IO latency. In flash devices, GC is the method of relocating existing data and deleting stale data, in order to create empty blocks for new incoming data. By learning the temporal trends of IO accesses, we built workload specific regression models for predicting the future time when the SSD will be in GC mode. We tested our models on synthetic traces (random read/write mix with fixed blocksize) generated by FIO workload generator. For the purpose of determining when the SSD is in GC mode, we track I/O completion times and classify completions that take more than 10 times the median completion value as representing those times when the SSD is in GC mode. Experiments run on the SSD models we tested reveal that a GC phase usually last 400 ms and it happens every 7000 ms on average. Results show that our workload specific models are accurate in predicting the time of next GC mode, achieving RMSE score of 10.61. Chandranil Chakraborttii, Vikas Sinha, Heiner Litz |
SoCC | 1 |
| 2017 | Towards generative emotions in games based on cognitive modelingabstractProcedural Content Generation (PCG) accomplishes feats that once would be considered magic, creating near-infinite amounts of unique levels, worlds, objects and other content for games. Yet, despite this near magical quality, generated content is often found to have a sameness to it. After a short time it loses the interest of players. Many procedurally generated games, such as No Man's Sky, have disappointed customers in this sense. We argue that one important cause is generated content fails to create an emotional connection with players. Emotions help in keeping players engaged with game content [7] and thereby improve the gameplay experience. Among the many emotions players experience in games, surprise is perhaps the most important for PCG. Surprise intensifies other emotions [9] and it lies at the origin of humor, strategy and problem solving [15]. Thus, surprise helps to increase player enjoyment and engagement with games. So far, PCG in games has produced surprise by accident of chance. PCG systems which intentionally create surprising moments, in a controllable way, can play an important role in increasing engagement and interest in games. Chandranil Chakraborttii, Lucas Ferreira, E. James Whitehead Jr. |
FDG | 1 |