EDBT 2026 Demo / reviewers in the wild / expert
Rohan Basu Roy
dblp:246/6853
· DBLP profile ↗
27ranked-venue papers
13as first author
24since 2021 · last 2026
0000-0002-1082-9846ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 11 first-author · 23 since 2021Software engineering, systems software and programming languages · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LowCarb: Carbon-Aware Scheduling of Serverless FunctionsabstractServerless computing is observing rapid adoption in cloud computing platforms. Prior works have extensively focused on improving the performance of serverless computing platforms via “keeping alive” functions in memory proactively to lower the function execution latency, but the potential environmental sustainability aspects of such performance-enhancing strategies remain underexplored. This work highlights that serverless computing introduces unique carbon footprint sources and trade-offs between performance and sustainability. We present LowCarb, a novel reinforcement learning-based solution that co-optimizes serverless function performance and carbon footprint. LowCarb effectively quantifies and resolves the inherent conflict between performance and sustainability to achieve results within 15% of optimality. Rohan Basu Roy, Devesh Tiwari |
HPCA | 1 |
| 2026 | TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and EfficiencyabstractQuantum processors are being integrated into HPC ecosystems as co-processors, where compilation of quantum circuits into hardware-executable form determines both output fidelity and runtime. Current compilers use a fixed pass sequence and ignore the fact that optimal pass selection varies with circuit, hardware, and noise conditions. We present TuniQ, a reinforcement learning-based system that selects compilation passes at each pipeline stage, adapting to circuit, backend, and current noise profile. TuniQ introduces several novel design components like a dual-encoder for stage-aware representation, shaped rewards for cross-stage credit assignment, and dynamic action masking for valid compilation. Evaluated across diverse quantum workloads on multiple IBM Quantum Cloud processors, TuniQ improves fidelity and reduces compilation time over the state-of-the-art IBM Qiskit transpiler, generalizes across backends without retraining, and scales strongly to utility-scale circuits with growing advantage. Mohammad Abrarul Hasanat, Jason Ludmir, Tirthak Patel, Rohan Basu Roy |
ICS | 4 |
| 2026 | WaterSplit: Coordinated On-Site and Off-Site Water Allocation For Sustainable Datacenter Cooling
Yankai Jiang 0002, Raghavendra Kanakagiri, Rohan Basu Roy, Devesh Tiwari |
IPDPS | 3 |
| 2025 | DarwinGame: Playing Tournaments for Tuning Applications in Noisy Cloud Environments
Rohan Basu Roy, Vijay Gadepally, Devesh Tiwari |
ASPLOS (1) | 1 |
| 2025 | Water Footprint of Datacenter Applications: Methodological Implications of Manufacturing, Operational, and Decommissioning PhasesabstractRising computational demands have made cloud datacenters' water footprint a critical concern. We demonstrate how different water footprint accounting methodologies - incorporating operational, manufacturing, and decommissioning water consumption, impact measurements and highlight the need for methodology standardization for water-aware operations. Our analysis reveals opportunities for water-aware scheduling in datacenters by considering regional water variations and lifecycle impacts. Amit Samanta 0001, Yankai Jiang 0002, Ryan Stutsman, Rohan Basu Roy |
SoCC | 4 |
| 2025 | GridGreen: Integrating Serverless Computing in HPC Systems for Performance and SustainabilityabstractWe present GridGreen, a scheduling framework that improves the sustainability and performance of scientific workflow execution by integrating serverless computing with traditional on-premise high performance computing (HPC) clusters. GridGreen allocates workflow components across HPC and serverless environments leveraging spatio-temporal variation of carbon intensity and component execution characteristics. It incorporates component-level optimization, speculative pre-warming, I/O-aware data management, and fallback adaptation to jointly minimize carbon footprint and service time under user-defined cost constraints. Our evaluations on large-scale bioinformatics workflows across leadership-class HPC facilities and cloud-based serverless regions demonstrate that GridGreen achieves robust, cost-effective execution while improving carbon efficiency. Amit Samanta 0001, Ryan Stutsman, Rohan Basu Roy |
SoCC | 3 |
| 2025 | Bringing Differential Privacy to HPC: Privacy-Preserving Transformations of HPC TracesabstractMonitoring HPC systems yields valuable insights into user behavior, aiding resource management, collaborative research, and software design. However, privacy concerns raise the barrier for real-world HPC trace sharing between HPC facilities and researchers. Traditional anonymization methods fall short as user behavior remains identifiable. To address this, we propose a robust toolset for privacy protection of HPC traces using Differential Privacy (DP). Our toolset offers a set of DP algorithms, metrics, and visualizations to empower HPC operators to protect users' sensitive information under a privacy protection guarantee. We evaluated our toolset over real HPC systems traces for different parameters and data aggregations. Moreover, we show that machine learning models trained on privacy-preserved logs maintain accuracy compared to real data, which supports data publishing and sharing across different computing facilities. Ana Veroneze Solórzano, Rohan Basu Roy, Benjamin Schwaller, Sara Walton, Jim M. Brandt, Devesh Tiwari |
HPDC | 2 |
| 2025 | WaterWise: Co-optimizing Carbon- and Water-Footprint Toward Environmentally Sustainable Cloud ComputingabstractThe carbon and water footprint of large-scale computing systems poses serious environmental sustainability risks. In this study, we discover that, unfortunately, carbon and water sustainability are at odds with each other - and, optimizing one alone hurts the other. Toward that goal, we introduce, WaterWise, a novel job scheduler for parallel workloads that intelligently co-optimizes carbon and water footprint to improve the sustainability of geographically distributed data centers. Yankai Jiang 0002, Rohan Basu Roy, Raghavendra Kanakagiri, Devesh Tiwari |
PPoPP | 2 |
| 2025 | ThirstyFLOPS: Water Footprint Modeling and Analysis Toward Sustainable HPC SystemsabstractHigh-performance computing (HPC) systems are becoming increasingly water-intensive due to their reliance on water-based cooling and the energy used in power generation. However, the water footprint of HPC remains relatively underexplored—especially in contrast to the growing focus on carbon emissions. In this paper, we present ThirstyFLOPS - a comprehensive water footprint analysis framework for HPC systems. Our approach incorporates region-specific metrics, including Water Usage Effectiveness, Power Usage Effectiveness, and Energy Water Factor, to quantify water consumption using real-world data. Using four representative HPC systems – Marconi, Fugaku, Polaris, and Frontier – as examples, we provide implications for HPC system planning and management. We explore the impact of regional water scarcity and nuclear-based energy strategies on HPC sustainability. Our findings aim to advance the development of water-aware, environmentally responsible computing infrastructures. Yankai Jiang 0002, Raghavendra Kanakagiri, Rohan Basu Roy, Devesh Tiwari |
SC | 3 |
| 2025 | GreenMix: Energy-Efficient Serverless Computing via Randomized Sketching on Asymmetric Multi-CoresabstractGreenMix is motivated by the renewed interest in asymmetric multi-core processors and the emergence of the serverless computing model. Asymmetric multi-cores offer better energy and performance trade-offs by placing different core types on the same die. However, existing serverless scheduling techniques do not leverage these benefits. GreenMix is the first serverless work to reduce energy and serverless keep-alive costs while meeting QoS targets by leveraging asymmetric multi-cores. GreenMix employs randomized sketching, tailored for serverless execution and keep-alive, to perform within 10% of the optimal solution in terms of energy efficiency and keep-alive cost reduction. GreenMix’s effectiveness is demonstrated through evaluations on clusters of ARM big.LITTLE and Intel Alder Lake asymmetric processors. It outperforms competing state-of-the-art schedulers, offering a novel approach for energy-efficient serverless computing. Rohan Basu Roy, Tirthak Patel, Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari |
SC | 1 |
| 2024 | CodeCrunch: Improving Serverless Performance via Function Compression and Cost-Aware Warmup Location OptimizationabstractServerless computing has a critical problem of function cold starts. To minimize cold starts, state-of-the-art techniques predict function invocation times to warm them up. Warmed-up functions occupy space in memory and incur a keep-alive cost, which can become exceedingly prohibitive under bursty load. To address this issue, we design CodeCrunch, which introduces the concept of serverless function compression and exploits server heterogeneity to make serverless computing more efficient, especially under high memory pressure. Rohan Basu Roy, Tirthak Patel, Rohan Garg 0001, Devesh Tiwari |
ASPLOS (1) | 1 |
| 2024 | RainbowCake: Mitigating Cold-starts in Serverless with Layer-wise Container Caching and SharingabstractServerless computing has grown rapidly as a new cloud computing paradigm that promises ease-of-management, cost-efficiency, and auto-scaling by shipping functions via self-contained virtualized containers. Unfortunately, serverless computing suffers from severe cold-start problems---starting containers incurs non-trivial latency. Full container caching is widely applied to mitigate cold-starts, yet has recently been outperformed by two lines of research: partial container caching and container sharing. However, either partial container caching or container sharing techniques exhibit their drawbacks. Partial container caching effectively deals with burstiness while leaving cold-start mitigation halfway; container sharing reduces cold-starts by enabling containers to serve multiple functions while suffering from excessive memory waste due to over-packed containers. Hanfei Yu, Rohan Basu Roy, Christian Fontenot, Devesh Tiwari, Jian Li 0008, Hong Zhang 0025, Hao Wang 0022, Seung-Jong Park |
ASPLOS (1) | 2 |
| 2024 | The Hidden Carbon Footprint of Serverless ComputingabstractDue to the unique aspects of serverless computing like keep-alive and co-location of functions, it is challenging to account for its carbon footprint. This is the first work to introduce the need for systematic methodologies for carbon accounting in the serverless environment, propose new methodologies and in-depth analysis, and highlight how the carbon footprint estimation can vary based on the chosen methodology. It discusses how serverless-specific scheduling choices can impact the tradeoffs between performance and carbon footprint, with an aim toward standardizing methodological choices and identifying opportunities for future improvements. Rohan Basu Roy, Raghavendra Kanakagiri, Yankai Jiang 0002, Devesh Tiwari |
SoCC | 1 |
| 2024 | EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable ComputingabstractThis work introduces ECOLIFE, the first carbon-aware serverless function scheduler to co-optimize carbon footprint and performance. ECOLIFE builds on the key insight of intelligently exploiting multi-generation hardware to achieve high performance and lower carbon footprint. ECOLIFE designs multiple novel extensions to Particle Swarm Optimization (PSO) in the context of serverless execution environment to achieve high performance while effectively reducing the carbon footprint. Yankai Jiang 0002, Rohan Basu Roy, Baolin Li 0001, Devesh Tiwari |
SC | 2 |
| 2023 | ProPack: Executing Concurrent Serverless Functions Faster and CheaperabstractThe serverless computing model has been on the rise in recent years due to a lower barrier to entry and elastic scalability. However, our experimental evidence suggests that multiple serverless computing platforms suffer from serious performance inefficiencies when a high number of concurrent function instances are invoked, which is a desirable capability for parallel applications. To mitigate this challenge, this paper introduces ProPack, a novel solution that provides higher performance and yields cost savings for end users running applications with high concurrency. ProPack leverages insights obtained from experimental study to build a simple and effective analytical model that mitigates the scalability bottleneck. Our evaluation on multiple serverless platforms including AWS Lambda and Google confirms that ProPack can improve average performance by 85% and save cost by 66%. ProPack provides significant improvement (over 50%) over the state-of-the-art serverless workload manager such as Pywren, and is also, effective at mitigating the concurrency bottleneck for FuncX, a recent on-premise serverless execution platform for parallel applications. Rohan Basu Roy, Tirthak Patel, Richmond Liew, Yadu N. Babuji, Ryan Chard, Devesh Tiwari |
HPDC | 1 |
| 2023 | Toward Sustainable HPC: Carbon Footprint Estimation and Environmental Implications of HPC SystemsabstractThe rapid growth in demand for HPC systems has led to a rise in carbon footprint, which requires urgent intervention. In this work, we present a comprehensive analysis of the carbon footprint of highperformance computing (HPC) systems, considering the carbon footprint during both the hardware manufacturing and system operational stages. Our work employs HPC hardware component carbon footprint modeling, regional carbon intensity analysis, and experimental characterization of the system life cycle to highlight the importance of quantifying the carbon footprint of HPC systems. Baolin Li 0001, Rohan Basu Roy, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari |
SC | 2 |
| 2022 | IceBreaker: warming serverless functions better with heterogeneityabstractServerless computing, an emerging computing model, relies on "warming up" functions prior to its anticipated execution for faster and cost-effective service to users. Unfortunately, warming up functions can be inaccurate and incur prohibitively expensive cost during the warmup period (i.e., keep-alive cost). In this paper, we introduce IceBreaker, a novel technique that reduces the service time and the "keep-alive" cost by composing a system with heterogeneous nodes (costly and cheaper). IceBreaker does so by dynamically determining the cost-effective node type to warm up a function based on the function's time-varying probability of the next invocation. By employing heterogeneity, IceBreaker allows for more number of nodes under the same cost budget and hence, keeps more number of functions warm and reduces the wait time during high load. Our real-system evaluation confirms that IceBreaker reduces the overall keep-alive cost by 45% and execution time by 27% using representative serverless applications and industry-grade workload trace. IceBreaker is the first technique to employ and leverage the idea of mixing expensive and cheaper nodes to improve both service time and keep-alive cost for serverless functions -- opening up a new research avenue of serverless computing on heterogeneous servers for researchers and practitioners. Rohan Basu Roy, Tirthak Patel, Devesh Tiwari |
ASPLOS | 1 |
| 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and ImplicationsabstractProduction high-performance computing (HPC) systems are adopting and integrating GPUs into their design to accommodate artificial intelligence (AI), machine learning, and data visualization workloads. To aid with the design and operations of new and existing GPU-based large-scale systems, we provide a detailed characterization of system operations, job characteristics, user behavior, and trends on a contemporary GPU-accelerated production HPC system. Our insights indicate that the pre-mature phases in modern AI workflow take up significant GPU hours while underutilizing GPUs, which opens up the opportunity for a multi-tier system. Finally, we provide various potential recommendations and areas for future investment for system architects, operators, and users. Baolin Li 0001, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Basu Roy, Bill Bergeron, John T. Holodnak, Michael Houle 0001, Matthew Hubbell, Michael Jones 0001, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Julie Mullen, Andrew Prout, Benjamin Price, Albert Reuther, Antonio Rosa, Matthew L. Weiss, Charles Yee, Daniel Edelman, Allan Vanterpool, Anson Cheng, Vijay Gadepally, Devesh Tiwari |
HPCA | 8 |
| 2022 | Mashup: making serverless computing useful for HPC workflows via hybrid executionabstractThis work introduces Mashup, a novel strategy to leverage serverless computing model for executing scientific workflows in a hybrid fashion by taking advantage of both the traditional VM-based cloud computing platform and the emerging serverless platform. Mashup outperforms the state-of-the-art workflow execution engines by an average of 34% and 43% in terms of execution time reduction and cost reduction, respectively, for widely-used HPC workflows on the Amazon Cloud platform (EC2 and Lambda). Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Devesh Tiwari |
PPoPP | 1 |
| 2022 | DayDream: Executing Dynamic Scientific Workflows on Serverless Platforms with Hot StartsabstractHPC applications are increasingly being designed as dynamic workflows for the ease of development and scaling. This work demonstrates how the serverless computing model can be leveraged for efficient execution of complex, real-world scientific workflows, although serverless computing was not originally designed for executing scientific workflows. This work characterizes, quantifies, and improves the execution of three real-world, complex, dynamic scientific workflows: ExaFEL (workflow for investigating the molecular structures via X-Ray diffraction), Cosmoscout-Vr(workflow for large scale virtual reality simulation), and Core Cosmology Library (a cosmology workflow for investigating dark matter). The proposed technique, DayDream, employs the hot start mechanism for warming up the components of the workflows by decoupling the runtime environment from the component function code to mitigate cold start overhead. DayDream optimizes the service time and service cost jointly to reduce the service time by 45% and service cost by 23% over the state-of-the-art HPC workload manager. Rohan Basu Roy, Tirthak Patel, Devesh Tiwari |
SC | 1 |
| 2021 | Operating Liquid-Cooled Large-Scale Systems: Long-Term Monitoring, Reliability Analysis, and Efficiency MeasuresabstractThe past decade has seen a rise in the use of liquid cooling due to its energy efficiency. While many previous works have helped make progress toward improving data center cooling, a vast majority of them perform studies on a small system over a short span. The computer systems and HPC community lacks a long-term study highlighting the challenges and solutions in operating a liquid-cooled large-scale data center. We conduct the first detailed characterization of a petascale supercomputer, Mira, over a span of six years. The study is enabled by systematic monitoring of the environmental metrics, and discusses new research avenues, including coolant monitor failures. Rohan Basu Roy, Tirthak Patel, Rajkumar Kettimuthu, William E. Allcock, Paul M. Rich, Adam Scovel, Devesh Tiwari |
HPCA | 1 |
| 2021 | SATORI: Efficient and Fair Resource Partitioning by Sacrificing Short-Term Benefits for Long-Term Gains*abstractMulti-core architectures have enabled data centers to increasingly co-locate multiple jobs to improve resource utilization and lower the operational cost. Unfortunately, naively co-locating multiple jobs may lead to only a modest increase in system throughput. Worse, some users may observe proportionally higher performance degradation compared to other users co-located on the same physical multi-core system. SATORI is a novel strategy to partition multi-core architectural resources to achieve two conflicting goals simultaneously: increasing system throughput and achieving fairness among the co-located jobs. Rohan Basu Roy, Tirthak Patel, Devesh Tiwari |
ISCA | 1 |
| 2021 | Bliss: auto-tuning complex applications using a pool of diverse lightweight learning modelsabstractAs parallel applications become more complex, auto-tuning becomes more desirable, challenging, and time-consuming. We propose, Bliss, a novel solution for auto-tuning parallel applications without requiring apriori information about applications, domain-specific knowledge, or instrumentation. Bliss demonstrates how to leverage a pool of Bayesian Optimization models to find the near-optimal parameter setting 1.64× faster than the state-of-the-art approaches. Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Devesh Tiwari |
PLDI | 1 |
| 2021 | RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instancesabstractDeep learning model inference is a key service in many businesses and scientific discovery processes. This paper introduces Ribbon, a novel deep learning inference serving system that meets two competing objectives: quality-of-service (QoS) target and cost-effectiveness. The key idea behind Ribbon is to intelligently employ a diverse set of cloud computing instances (heterogeneous instances) to meet the QoS target and maximize cost savings. Ribbon devises a Bayesian Optimization-driven strategy that helps users build the optimal set of heterogeneous instances for their model inference service needs on cloud computing platforms - and, Ribbon demonstrates its superiority over existing approaches of inference serving systems using homogeneous instance pools. Ribbon saves up to 16% of the inference service cost for different learning models including emerging deep learning recommender system models and drug-discovery enabling models. Baolin Li 0001, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Karen Gettings, Devesh Tiwari |
SC | 2 |
| 2020 | Experimental evaluation of NISQ quantum computers: error measurement, characterization, and implicationsabstractNoisy Intermediate-Scale Quantum (NISQ) computers are being increasingly used for executing early-stage quantum programs to establish the practical realizability of existing quantum algorithms. These quantum programs have uses cases in the realm of high-performance computing ranging from molecular chemistry and physics simulations to addressing NP-complete optimization problems. However, NISQ devices are prone to multiple types of errors, which affect the fidelity and reproducibility of the program execution. As the technology is still primitive, our understanding of these quantum machines and their error characteristics is limited. To bridge that understanding gap, this is the first work to provide a systematic and rich experimental evaluation of IBM Quantum Experience (QX) quantum computers of different scales and topologies. Our experimental evaluation uncovers multiple important and interesting aspects of benchmarking and evaluating quantum program on NISQ machines. We have open-sourced our experimental framework and dataset to help accelerate the evaluation of quantum computing systems. Tirthak Patel, Abhay Potharaju, Baolin Li 0001, Rohan Basu Roy, Devesh Tiwari |
SC | 4 |
| 2020 | UREQA: Leveraging Operation-Aware Error Rates for Effective Quantum Circuit Mapping on NISQ-Era Quantum Computers
Tirthak Patel, Baolin Li 0001, Rohan Basu Roy, Devesh Tiwari |
USENIX ATC | 3 |
| 2019 | Design of a Computationally Economical Image Classifier using Generic FeaturesabstractIn this paper, we propose an image classification technique which uses a simple autoencoder with a regularizer. Nowadays, Convolutional Neural Networks (CNN) are primarily used for image classification. Our method can be used for image classification with much reduced requirement of computational capability than a complex CNN which has a huge number of degrees of freedom. Here, the terms simple and complex, respectively, correspond to the simplicity and the complexity of a network in terms of the number of learnable parameters (degrees of freedom) and the number of hidden layers. This technique uses features extracted from a pretrained CNN, trained on a completely different dataset. Genetic algorithm solves for the optimal hyperparameters of the pretrained CNN. It is observed that these features serve as important and robust parameters for the training of the autoencoder, as a final average image classification accuracy improvement of about 17.45% is observed with the inclusion of these features. We use a pretrained CNN on MNIST dataset and classify images of several other benchmark datasets. We utilize different classifiers for image classification based on features extracted from the autoencoder and repeat each of the experiments a number of times with different random initialization of the classifier and the weight matrix of the autoencoder. We also perform experiments by pretraining the CNN with different datasets. Our results show a notable image classification accuracy and a significant reduction of training time with respect to a complex CNN. Rohan Basu Roy, Anisha Halder, Amit Konar, Atulya K. Nagar |
CEC | 1 |