Siddharth Samsi

dblp:15/7615 · also Sid Samsi · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 8 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 GreenMix: Energy-Efficient Serverless Computing via Randomized Sketching on Asymmetric Multi-Cores
abstract
GreenMix is motivated by the renewed interest in asymmetric multi-core processors and the emergence of the serverless computing model. Asymmetric multi-cores offer better energy and performance trade-offs by placing different core types on the same die. However, existing serverless scheduling techniques do not leverage these benefits. GreenMix is the first serverless work to reduce energy and serverless keep-alive costs while meeting QoS targets by leveraging asymmetric multi-cores. GreenMix employs randomized sketching, tailored for serverless execution and keep-alive, to perform within 10% of the optimal solution in terms of energy efficiency and keep-alive cost reduction. GreenMix’s effectiveness is demonstrated through evaluations on clusters of ARM big.LITTLE and Intel Alder Lake asymmetric processors. It outperforms competing state-of-the-art schedulers, offering a novel approach for energy-efficient serverless computing.
Rohan Basu Roy, Tirthak Patel, Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
SC4
2024 Interpretable Analysis of Production GPU Clusters Monitoring Data via Association Rule Mining
abstract
Modern high-performance computing (HPC) and cloud computing systems are integrating powerful GPUs to accelerate increasingly demanding deep learning workloads. To improve cluster efficiency and better understand user behavior and job characteristics, system operators will collect operational data for trace analysis. However, previous efforts on these system logs have lacked the interpretability aspect, and there is no systematic approach that can be widely applied to different datacenter traces and return interpretable results. In this work, we propose a workflow to discover hidden association relation-ships between collected features of system jobs. The outcome of our analysis approach yields useful association rules that can be directly interpreted into operational insights. Using this approach, we have conducted case studies using the traces of three large-scale multi-tenant GPU clusters running production machine learning workloads. We have focused on the observations of GPU underutilization and job failures, revealing the possible reasons for these job behaviors and suggesting solutions to mitigate them. Our case studies have demonstrated the feasibility of our interpretable analysis workflow, which can be widely used by more HPC and cloud computing system operators.
Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
IPDPS2
2023 Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
abstract
As research and deployment of AI grows, the computational burden to support and sustain its progress inevitably does too. To train or fine-tune state-of-the-art models in NLP, computer vision, etc., some form of AI hardware acceleration is virtually a requirement. Recent large language models require considerable resources to train and deploy, resulting in significant energy usage, potential carbon emissions, and massive demand for GPUs and other hardware accelerators. However, this surge carries large implications for energy sustainability at the HPC/datacenter level. In this paper, we study the effects of power-capping GPUs at a research supercomputing center on GPU temperature and power draw; we show significant decreases in both temperature and power draw, reducing power consumption and potentially improving hardware life-span, with minimal impact on job performance. To our knowledge, our work is the first to conduct and make available a detailed analysis of the effects of GPU power-capping at the supercomputing scale. We hope our work will inspire HPCs/datacenters to further explore, evaluate, and communicate the impact of power-capping AI hardware accelerators for more sustainable AI.
Dan Zhao 0007, Siddharth Samsi, Joseph McDonald, Baolin Li 0001, David Bestor, Michael Jones 0001, Devesh Tiwari, Vijay Gadepally
SoCC2
2023 Kairos: Building Cost-Efficient Machine Learning Inference Systems with Heterogeneous Cloud Resources
abstract
Online inference is becoming a key service product for many businesses, deployed in cloud platforms to meet customer demands. Despite their revenue-generation capability, these services need to operate under tight Quality-of-Service (QoS) and cost budget constraints. This paper introduces KAIROS, a novel runtime framework that maximizes the query throughput while meeting QoS target and a cost budget. KAIROS designs and implements novel techniques to build a pool of heterogeneous compute hardware without online exploration overhead, and distribute inference queries optimally at runtime. Our evaluation using industry-grade machine learning (ML) models shows that KAIROS yields up to 2x the throughput of an optimal homogeneous solution, and outperforms state-of-the-art schemes by up to 70%, despite advantageous implementations of the competing schemes to ignore their exploration overhead.
Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
HPDC2
2023 Toward Sustainable HPC: Carbon Footprint Estimation and Environmental Implications of HPC Systems
abstract
The rapid growth in demand for HPC systems has led to a rise in carbon footprint, which requires urgent intervention. In this work, we present a comprehensive analysis of the carbon footprint of highperformance computing (HPC) systems, considering the carbon footprint during both the hardware manufacturing and system operational stages. Our work employs HPC hardware component carbon footprint modeling, regional carbon intensity analysis, and experimental characterization of the system life cycle to highlight the importance of quantifying the carbon footprint of HPC systems.
Baolin Li 0001, Rohan Basu Roy, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
SC4
2023 Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference Service
abstract
This paper presents a solution to the challenge of mitigating carbon emissions from hosting large-scale machine learning (ML) inference services. ML inference is critical to modern technology products, but it is also a significant contributor to carbon footprint. We introduce, Clover, a carbon-friendly ML inference service runtime system that balances performance, accuracy, and carbon emissions through mixed-quality models and GPU resource partitioning. Our experimental results demonstrate that Clover is effective in substantially reducing carbon emissions while maintaining high accuracy and meeting service level agreement (SLA) targets.
Baolin Li 0001, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
SC2
2022 MISO: exploiting multi-instance GPU capability on multi-tenant GPU clusters
abstract
GPU technology has been improving at an expedited pace in terms of size and performance, empowering HPC and AI/ML researchers to advance the scientific discovery process. However, this also leads to inefficient resource usage, as most GPU workloads, including complicated AI/ML models, are not able to utilize the GPU resources to their fullest extent - encouraging support for GPU multi-tenancy. We propose MISO, a technique to exploit the Multi-Instance GPU (MIG) capability on the latest NVIDIA datacenter GPUs (e.g., A100, H100) to dynamically partition GPU resources among co-located jobs. MISO's key insight is to use the lightweight, more flexible Multi-Process Service (MPS) capability to predict the best MIG partition allocation for different jobs, without incurring the overhead of implementing them during exploration. Due to its ability to utilize GPU resources more efficiently, MISO achieves 49% and 16% lower average job completion time than the unpartitioned and optimal static GPU partition schemes, respectively.
Baolin Li 0001, Tirthak Patel, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
SoCC3
2022 AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications
abstract
Production high-performance computing (HPC) systems are adopting and integrating GPUs into their design to accommodate artificial intelligence (AI), machine learning, and data visualization workloads. To aid with the design and operations of new and existing GPU-based large-scale systems, we provide a detailed characterization of system operations, job characteristics, user behavior, and trends on a contemporary GPU-accelerated production HPC system. Our insights indicate that the pre-mature phases in modern AI workflow take up significant GPU hours while underutilizing GPUs, which opens up the opportunity for a multi-tier system. Finally, we provide various potential recommendations and areas for future investment for system architects, operators, and users.
Baolin Li 0001, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Basu Roy, Bill Bergeron, John T. Holodnak, Michael Houle 0001, Matthew Hubbell, Michael Jones 0001, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Julie Mullen, Andrew Prout, Benjamin Price, Albert Reuther, Antonio Rosa, Matthew L. Weiss, Charles Yee, Daniel Edelman, Allan Vanterpool, Anson Cheng, Vijay Gadepally, Devesh Tiwari
HPCA3
2020 SEVIR : A Storm Event Imagery Dataset for Deep Learning Applications in Radar and Satellite Meteorology
abstract
Modern deep learning approaches have shown promising results in meteorological applications like precipitation nowcasting, synthetic radar generation, front detection and several others. In order to effectively train and validate these complex algorithms, large and diverse datasets containing high-resolution imagery are required. Petabytes of weather data, such as from the Geostationary Environmental Satellite System (GOES) and the Next-Generation Radar (NEXRAD) system, are available to the public; however, the size and complexity of these datasets is a hindrance to developing and training deep models. To help address this problem, we introduce the Storm EVent ImagRy (SEVIR) dataset - a single, rich dataset that combines spatially and temporally aligned data from multiple sensors, along with baseline implementations of deep learning models and evaluation metrics, to accelerate new algorithmic innovations. SEVIR is an annotated, curated and spatio-temporally aligned dataset containing over 10,000 weather events that each consist of 384 km x 384 km image sequences spanning 4 hours of time. Images in SEVIR were sampled and aligned across five different data types: three channels (C02, C09, C13) from the GOES-16 advanced baseline imager, NEXRAD vertically integrated liquid mosaics, and GOES-16 Geostationary Lightning Mapper (GLM) flashes. Many events in SEVIR were selected and matched to the NOAA Storm Events database so that additional descriptive information such as storm impacts and storm descriptions can be linked to the rich imagery provided by the sensors. We describe the data collection methodology and illustrate the applications of this dataset with two examples of deep learning in meteorology: precipitation nowcasting and synthetic weather radar generation. In addition, we also describe a set of metrics that can be used to evaluate the outputs of these models. The SEVIR dataset and baseline implementations of selected applications are available for download.
Mark S. Veillette, Siddharth Samsi, Christopher J. Mattioli
NeurIPS2
2017 Learning by doing, High Performance Computing education in the MOOC era
Julie Mullen, Chansup Byun, Vijay Gadepally, Siddharth Samsi, Albert Reuther, Jeremy Kepner
J. Parallel Distributed Comput.4
2008 Experiences from Cyberinfrastructure Development for Multi-user Remote Instrumentation
abstract
Computer-controlled scientific instruments such as electron microscopes, spectrometers, and telescopes are expensive to purchase and maintain. Also, they generate large amounts of raw and processed data that has to be annotated and archived. Cyber-enabling these instruments and their data sets using remote instrumentation cyberinfrastructures can improve user convenience and significantly reduce costs. In this paper, we discuss our experiences in gathering technical and policy requirements of remote instrumentation for research and training purposes. Next, we describe the cyberinfrastructure solutions we are developing for supporting related multi-user workflows. Finally, we present our solution-deployment experiences in the form of case studies. The case studies cover both technical issues (bandwidth provisioning, collaboration tools, data management, system security) and policy issues (service level agreements, use policy, usage billing). Our experiences suggest that developing cyberinfrastructures for remote instrumentation requires: (a) understanding and overcoming multi-disciplinary challenges, (b) developing reconfigurable-and-integrated solutions, and (c) close collaborations between instrument labs, infrastructure providers, and application developers.
Prasad Calyam, Abdul Kalash, Neil Ludban, Sowmya Gopalan, Siddharth Samsi, Karen A. Tomko, David E. Hudak, Ashok K. Krishnamurthy 0001
eScience5
2007 Survey of Parallel MATLAB Techniques and Applications to Signal and Image Processing
abstract
We present a survey of modern parallel MATLAB techniques. We concentrate on the most promising and well supported techniques with an emphasis in SIP applications. Some of these methods require writing explicit code to perform inter-processor communication while others hide the complexities of communication and computation by using higher level programming interfaces. We cover each approach with special emphasis given to performance and productivity issues.
Ashok K. Krishnamurthy 0001, John Nehrbass, Juan Carlos Chaves, Siddharth Samsi
ICASSP (4)4