EDBT 2026 Demo / reviewers in the wild / expert
Baolin Li 0001
dblp:72/2660-1
· DBLP profile ↗
14ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0001-9778-1023ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GreenMix: Energy-Efficient Serverless Computing via Randomized Sketching on Asymmetric Multi-CoresabstractGreenMix is motivated by the renewed interest in asymmetric multi-core processors and the emergence of the serverless computing model. Asymmetric multi-cores offer better energy and performance trade-offs by placing different core types on the same die. However, existing serverless scheduling techniques do not leverage these benefits. GreenMix is the first serverless work to reduce energy and serverless keep-alive costs while meeting QoS targets by leveraging asymmetric multi-cores. GreenMix employs randomized sketching, tailored for serverless execution and keep-alive, to perform within 10% of the optimal solution in terms of energy efficiency and keep-alive cost reduction. GreenMix’s effectiveness is demonstrated through evaluations on clusters of ARM big.LITTLE and Intel Alder Lake asymmetric processors. It outperforms competing state-of-the-art schedulers, offering a novel approach for energy-efficient serverless computing. Rohan Basu Roy, Tirthak Patel, Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari |
SC | 3 |
| 2024 | Sprout: Green Generative AI with Carbon-Efficient LLM InferenceabstractThe rapid advancement of generative AI has heightened environmental concerns, particularly regarding carbon emissions.Our framework, SPROUT, addresses these challenges by reducing the carbon footprint of inference in large language models (LLMs).SPROUT introduces "generation directives" to guide the autoregressive generation process, achieving a balance between ecological sustainability and high-quality outputs.By employing a strategic optimizer for directive assignment and a novel offline quality evaluator, SPROUT reduces the carbon footprint of generative LLM inference by over 40% in real-world evaluations, using the Llama model and global electricity grid data.This work is crucial as the rising interest in inference time compute scaling laws amplifies environmental concerns, emphasizing the need for eco-friendly AI solutions. Baolin Li 0001, Yankai Jiang 0002, Vijay Gadepally, Devesh Tiwari |
EMNLP | 1 |
| 2024 | Interpretable Analysis of Production GPU Clusters Monitoring Data via Association Rule MiningabstractModern high-performance computing (HPC) and cloud computing systems are integrating powerful GPUs to accelerate increasingly demanding deep learning workloads. To improve cluster efficiency and better understand user behavior and job characteristics, system operators will collect operational data for trace analysis. However, previous efforts on these system logs have lacked the interpretability aspect, and there is no systematic approach that can be widely applied to different datacenter traces and return interpretable results. In this work, we propose a workflow to discover hidden association relation-ships between collected features of system jobs. The outcome of our analysis approach yields useful association rules that can be directly interpreted into operational insights. Using this approach, we have conducted case studies using the traces of three large-scale multi-tenant GPU clusters running production machine learning workloads. We have focused on the observations of GPU underutilization and job failures, revealing the possible reasons for these job behaviors and suggesting solutions to mitigate them. Our case studies have demonstrated the feasibility of our interpretable analysis workflow, which can be widely used by more HPC and cloud computing system operators. Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari |
IPDPS | 1 |
| 2024 | EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable ComputingabstractThis work introduces ECOLIFE, the first carbon-aware serverless function scheduler to co-optimize carbon footprint and performance. ECOLIFE builds on the key insight of intelligently exploiting multi-generation hardware to achieve high performance and lower carbon footprint. ECOLIFE designs multiple novel extensions to Particle Swarm Optimization (PSO) in the context of serverless execution environment to achieve high performance while effectively reducing the carbon footprint. Yankai Jiang 0002, Rohan Basu Roy, Baolin Li 0001, Devesh Tiwari |
SC | 3 |
| 2023 | Sustainable Supercomputing for AI: GPU Power Capping at HPC ScaleabstractAs research and deployment of AI grows, the computational burden to support and sustain its progress inevitably does too. To train or fine-tune state-of-the-art models in NLP, computer vision, etc., some form of AI hardware acceleration is virtually a requirement. Recent large language models require considerable resources to train and deploy, resulting in significant energy usage, potential carbon emissions, and massive demand for GPUs and other hardware accelerators. However, this surge carries large implications for energy sustainability at the HPC/datacenter level. In this paper, we study the effects of power-capping GPUs at a research supercomputing center on GPU temperature and power draw; we show significant decreases in both temperature and power draw, reducing power consumption and potentially improving hardware life-span, with minimal impact on job performance. To our knowledge, our work is the first to conduct and make available a detailed analysis of the effects of GPU power-capping at the supercomputing scale. We hope our work will inspire HPCs/datacenters to further explore, evaluate, and communicate the impact of power-capping AI hardware accelerators for more sustainable AI. Dan Zhao 0007, Siddharth Samsi, Joseph McDonald, Baolin Li 0001, David Bestor, Michael Jones 0001, Devesh Tiwari, Vijay Gadepally |
SoCC | 4 |
| 2023 | Kairos: Building Cost-Efficient Machine Learning Inference Systems with Heterogeneous Cloud ResourcesabstractOnline inference is becoming a key service product for many businesses, deployed in cloud platforms to meet customer demands. Despite their revenue-generation capability, these services need to operate under tight Quality-of-Service (QoS) and cost budget constraints. This paper introduces KAIROS, a novel runtime framework that maximizes the query throughput while meeting QoS target and a cost budget. KAIROS designs and implements novel techniques to build a pool of heterogeneous compute hardware without online exploration overhead, and distribute inference queries optimally at runtime. Our evaluation using industry-grade machine learning (ML) models shows that KAIROS yields up to 2x the throughput of an optimal homogeneous solution, and outperforms state-of-the-art schemes by up to 70%, despite advantageous implementations of the competing schemes to ignore their exploration overhead. Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari |
HPDC | 1 |
| 2023 | Toward Sustainable HPC: Carbon Footprint Estimation and Environmental Implications of HPC SystemsabstractThe rapid growth in demand for HPC systems has led to a rise in carbon footprint, which requires urgent intervention. In this work, we present a comprehensive analysis of the carbon footprint of highperformance computing (HPC) systems, considering the carbon footprint during both the hardware manufacturing and system operational stages. Our work employs HPC hardware component carbon footprint modeling, regional carbon intensity analysis, and experimental characterization of the system life cycle to highlight the importance of quantifying the carbon footprint of HPC systems. Baolin Li 0001, Rohan Basu Roy, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari |
SC | 1 |
| 2023 | Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference ServiceabstractThis paper presents a solution to the challenge of mitigating carbon emissions from hosting large-scale machine learning (ML) inference services. ML inference is critical to modern technology products, but it is also a significant contributor to carbon footprint. We introduce, Clover, a carbon-friendly ML inference service runtime system that balances performance, accuracy, and carbon emissions through mixed-quality models and GPU resource partitioning. Our experimental results demonstrate that Clover is effective in substantially reducing carbon emissions while maintaining high accuracy and meeting service level agreement (SLA) targets. Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari |
SC | 1 |
| 2022 | MISO: exploiting multi-instance GPU capability on multi-tenant GPU clustersabstractGPU technology has been improving at an expedited pace in terms of size and performance, empowering HPC and AI/ML researchers to advance the scientific discovery process. However, this also leads to inefficient resource usage, as most GPU workloads, including complicated AI/ML models, are not able to utilize the GPU resources to their fullest extent - encouraging support for GPU multi-tenancy. We propose MISO, a technique to exploit the Multi-Instance GPU (MIG) capability on the latest NVIDIA datacenter GPUs (e.g., A100, H100) to dynamically partition GPU resources among co-located jobs. MISO's key insight is to use the lightweight, more flexible Multi-Process Service (MPS) capability to predict the best MIG partition allocation for different jobs, without incurring the overhead of implementing them during exploration. Due to its ability to utilize GPU resources more efficiently, MISO achieves 49% and 16% lower average job completion time than the unpartitioned and optimal static GPU partition schemes, respectively. Baolin Li 0001, Tirthak Patel, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari |
SoCC | 1 |
| 2022 | Do Temperature and Humidity Exposures Hurt or Benefit Your SSDs?abstractSSDs are becoming mainstream data storage de-vices, replacing HDDs in most data centers, consumer goods, and IoT gadgets. In this work, we ask an uncharted research question: What is the environmental conditions' impact on SSD performance? To answer it, we systematically measure, quantify, and characterize the impact of various commonly changing envi-ronmental conditions such as temperature and humidity on the performance of SSDs. Our experiments and analysis uncover that exposure to changes in temperature and humidity can significantly affect SSD performance. Adnan Maruf, Sashri Brahmakshatriya, Baolin Li 0001, Devesh Tiwari, Gang Quan, Janki Bhimani |
DATE | 3 |
| 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and ImplicationsabstractProduction high-performance computing (HPC) systems are adopting and integrating GPUs into their design to accommodate artificial intelligence (AI), machine learning, and data visualization workloads. To aid with the design and operations of new and existing GPU-based large-scale systems, we provide a detailed characterization of system operations, job characteristics, user behavior, and trends on a contemporary GPU-accelerated production HPC system. Our insights indicate that the pre-mature phases in modern AI workflow take up significant GPU hours while underutilizing GPUs, which opens up the opportunity for a multi-tier system. Finally, we provide various potential recommendations and areas for future investment for system architects, operators, and users. Baolin Li 0001, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Basu Roy, Bill Bergeron, John T. Holodnak, Michael Houle 0001, Matthew Hubbell, Michael Jones 0001, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Julie Mullen, Andrew Prout, Benjamin Price, Albert Reuther, Antonio Rosa, Matthew L. Weiss, Charles Yee, Daniel Edelman, Allan Vanterpool, Anson Cheng, Vijay Gadepally, Devesh Tiwari |
HPCA | 1 |
| 2021 | RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instancesabstractDeep learning model inference is a key service in many businesses and scientific discovery processes. This paper introduces Ribbon, a novel deep learning inference serving system that meets two competing objectives: quality-of-service (QoS) target and cost-effectiveness. The key idea behind Ribbon is to intelligently employ a diverse set of cloud computing instances (heterogeneous instances) to meet the QoS target and maximize cost savings. Ribbon devises a Bayesian Optimization-driven strategy that helps users build the optimal set of heterogeneous instances for their model inference service needs on cloud computing platforms - and, Ribbon demonstrates its superiority over existing approaches of inference serving systems using homogeneous instance pools. Ribbon saves up to 16% of the inference service cost for different learning models including emerging deep learning recommender system models and drug-discovery enabling models. Baolin Li 0001, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Karen Gettings, Devesh Tiwari |
SC | 1 |
| 2020 | Experimental evaluation of NISQ quantum computers: error measurement, characterization, and implicationsabstractNoisy Intermediate-Scale Quantum (NISQ) computers are being increasingly used for executing early-stage quantum programs to establish the practical realizability of existing quantum algorithms. These quantum programs have uses cases in the realm of high-performance computing ranging from molecular chemistry and physics simulations to addressing NP-complete optimization problems. However, NISQ devices are prone to multiple types of errors, which affect the fidelity and reproducibility of the program execution. As the technology is still primitive, our understanding of these quantum machines and their error characteristics is limited. To bridge that understanding gap, this is the first work to provide a systematic and rich experimental evaluation of IBM Quantum Experience (QX) quantum computers of different scales and topologies. Our experimental evaluation uncovers multiple important and interesting aspects of benchmarking and evaluating quantum program on NISQ machines. We have open-sourced our experimental framework and dataset to help accelerate the evaluation of quantum computing systems. Tirthak Patel, Abhay Potharaju, Baolin Li 0001, Rohan Basu Roy, Devesh Tiwari |
SC | 3 |
| 2020 | UREQA: Leveraging Operation-Aware Error Rates for Effective Quantum Circuit Mapping on NISQ-Era Quantum Computers
Tirthak Patel, Baolin Li 0001, Rohan Basu Roy, Devesh Tiwari |
USENIX ATC | 2 |