Seetharami Seelam

dblp:19/1624 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
5since 2021 · last 2027
0000-0002-7595-3477ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2027 Closed-loop calculations of electronic structure on a quantum processor and a classical supercomputer at full scale
abstract
Quantum computers must operate in concert with classical computers to deliver on the promise of quantum advantage for practical problems. To achieve that, it is important to understand how quantum and classical computing can interact together, and how one can characterize the scalability and efficiency of hybrid quantum–classical workflows. So far, early experiments with quantum-centric supercomputing workflows have been limited in scale and complexity. Here, we use a Heron quantum processor deployed on premises with the entire supercomputer Fugaku to perform the largest computation of electronic structure involving quantum and classical high-performance computing. We design a closed-loop workflow between the quantum processors and 152,064 classical nodes of Fugaku, to approximate the electronic structure of chemistry models beyond the reach of exact diagonalization, with accuracy comparable to some all-classical approximation methods. Our work pushes the limits of the integration of quantum and classical high-performance computing, showcasing computational resource orchestration at the largest scale possible for current classical supercomputers.
Tomonori Shirakawa, Javier Robledo Moreno, Toshinari Itoko, Vinay Tripathi, Kento Ueda, Yukio Kawashima, Lukas Broers, William M. Kirby, Himadri Pathak, Hanhee Paik, Miwako Tsuji, Yuetsu Kodama, Mitsuhisa Sato, Constantinos Evangelinos, Seetharami Seelam, Robert Walkup, Seiji Yunoki, Mario Motta, Petar Jurcevic, Hiroshi Horii, Antonio Mezzacapo
Future Gener. Comput. Syst.15
2025 Vela: A Virtualized LLM Training System with GPU Direct RoCE
abstract
Vela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark.
Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam
ASPLOS (2)25
2025 Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
abstract
We present the design and implementation of a new lifetime-aware tensor offloading framework for GPU memory expansion using low-cost PCIe-based solid-state drives (SSDs). Our framework, TERAIO, is developed explicitly for large language model (LLM) training with multiple GPUs and multiple SSDs. Its design is driven by our observation that the active tensors take only a small fraction (1.7% on average) of allocated GPU memory in each LLM training iteration, the inactive tensors are usually large and will not be used for a long period of time, creating ample opportunities for offloading/prefetching tensors to/from slow SSDs without stalling the GPU training process. TERAIO accurately estimates the lifetime (active period of time in GPU memory) of each tensor with the profiling of the first few iterations in the training process. With the tensor lifetime analysis, TERAIO will generate an optimized tensor offloading/prefetching plan and integrate it into the compiled LLM program via PyTorch. TERAIO has a runtime tensor migration engine to execute the offloading/prefetching plan via GPUDirect storage, which allows direct tensor migration between GPUs and SSDs for alleviating the CPU bottleneck and maximizing the SSD bandwidth utilization. In comparison with state-of-the-art studies such as ZeRO-Offload and ZeRO-Infinity, we show that TERAIO improves the training performance of various LLMs by 1.47× on average, and achieves 80.7% of the ideal performance assuming unlimited GPU memory.
Yirui Eric Zhou, Apoorve Mohan, I-Hsin Chung, Seetharami Seelam, Jian Huang 0006
NeurIPS6
2024 2024 IEEE International Conference on Cloud Computing, Message from the Chairs
abstract
We are delighted to welcome all participants to the 2024 IEEE International Conference on Cloud Computing (CLOUD 2024), which is taking place in the beautiful city of Shenzhen, China, from July 7th to 13th.
Tevfik Kosar, Krishnan Venkateswaran, Shangguang Wang, Seetharami Seelam, Santonu Sarkar, Xuanzhe Liu
CLOUD4
2023 Chic-sched: a HPC Placement-Group Scheduler on Hierarchical Topologies with Constraints
abstract
Efficient placement of advanced HPC and AI workloads with application constraints is raising challenges for resource schedulers on shared infrastructures, such as the Cloud. In this work, we propose a novel Constraints- and Heuristics-based scheduler on HIerarchical Topologies for High-Performance Computing workloads in the Cloud (chic-sched, for short). Our heuristics-based algorithm enables placement across multiple levels in a network hierarchy with loosely specified constraints, and it works without retries by providing suboptimal placements to minimize placement failures. This allows for fast scheduling at scale, and the O(N log N) complexity enables placement decisions within tens of milliseconds for groups of hundreds of virtual machines (VM). We introduce a new and simple metric to quantify the goodness of group placements. With this metric, in terms of deviation from ideal placements, we show that chic-sched is 20-50% better than the common bestFit or worstFit algorithms in all scenarios of two-level placements with spreading and packing constraints. We evaluate chic-sched with publicly available VM-request traces from a production Cloud, and, comparing against bestFit, we show that it achieves 8% lower placement failure rates and more than 40% better placement locality. Finally, to quantify the goodness of constraints-based placements, we conduct experiments with a realistic MPI workload on synthetically allocated VM clusters in a public cloud. We measure a 9% performance improvement over an adverse placement in a scenario where our heuristics-based scheduler would return a good, but not perfect, placement.
Laurent Schares, Asser N. Tantawi, Pavlos Maniotis, Ming-Hung Chen, Claudia Misale, Seetharami Seelam, Hao Yu 0008
IPDPS6
2019 ConfAdvisor: A Performance-centric Configuration Tuning Framework for Containers on Kubernetes
abstract
Configuration tuning of software is often a good option to improve application performance without any application code modifications. Although we can casually change configurations, it is not easy to apply optimal configurations, as optimal configurations require deep knowledge of the underlying system. This is problematic because applications with suboptimal configuration result in poor performance. As container and container management systems have emerged as an application platform on the cloud, configuration tuning becomes even more challenging because containers add more complexity to the application performance. We need to consider not only fundamental misconfiguration but also container image verification, deployment configuration, application characteristics awareness based on metrics and logs. Although previous knowledge regarding how we should tune configurations for a system software is sometimes available, knowledge about performance tuning practices is neither normalized nor reusable to expand on any advice for misconfiguration to the containers. Even in the cloud-native environment, there is no centralized service to deliver knowledge continuously to application containers nor a framework to develop a misconfiguration fix rule for a container throughout its lifetime. In this paper, we propose a performance-centric configuration tuning framework for containers on Kubernetes, named ConfAdvisor, that enables containers to achieve a higher performance by validating various misconfigurations adaptively. ConfAdivsor gives config tuning advice to application containers, images, and Kubernetes specs and also provides a development framework to build configuration validation rules. We present the design of ConfAdvisor and provide several case studies to tune application containers in the real world.
Tatsuhiro Chiba, Rina Nakazawa, Hiroshi Horii, Sahil Suneja, Seetharami Seelam
IC2E5
2017 Taming Performance Degradation of Containers in the Case of Extreme Memory Overcommitment
abstract
The efficiency of datacenters is important consideration for cloud service providers to make their datacenters always ready for fulfilling the increasing demand for computing resources. Container-based virtualization is one approach to improving efficiency by reducing the overhead of virtualization. Resource overcommitment is another approach, but cloud providers tend to make conservative allocations of resources because there is no good understanding of the relationship between physical resource overcommitment and its impact on performance. This paper presents a quantitative study of performance degradation of containerized workloads due to memory overcommitment and a technique to mitigate it. We focused on physical memory overcommitment, where the sum of the working set memory is larger than the physical memory. We drove a small fraction of Docker containers at a high load level and the rest of them at a very low load level to emulate a common usage pattern of cloud datacenters. Detailed measurements revealed it is difficult to predict how many additional containers can be launched before thrashing hurts performance. We show that tuning the per-container swappiness of heavily loaded containers is effective for launching a larger number of containers and that it achieves an overcommitment of about three times.
Rina Nakazawa, Kazunori Ogata, Seetharami Seelam, Tamiya Onodera
CLOUD3
2015 Introduction to special issue on High Performance Computing Architectures and Systems
Jia Hu 0001, Seetharami Seelam, Laurent Lefèvre
J. Comput. Syst. Sci.2
2014 Introduction to special issue on embedded systems architecture and applications
Jia Hu 0001, Jens Palsberg, Seetharami Seelam, Marco Di Natale, Lei (Chris) Liu
J. Syst. Archit.3
2013 Extreme scale computing: Modeling the impact of system noise in multi-core clustered systems
Seetharami Seelam, Liana L. Fong, Asser N. Tantawi, John Lewars, John Divirgilio, Kevin J. Gildea
J. Parallel Distributed Comput.1
2012 Partitioned Parallel Job Scheduling for Extreme Scale Computing
David Brelsford, George Chochia, Nathan Falk, Kailash Marthi, Ravindra Sure, Norman Bobroff, Liana L. Fong, Seetharami Seelam
JSSPP8
2012 Evaluation of Multi-core Scalability Bottlenecks in Enterprise Java Workloads
abstract
The increasing number of cores integrated into modern processors is blurring the line between supercomputers and enterprise-grade servers. Therefore, the same attention to lock contention bottlenecks must be given to Java-based business workloads as it is given to massively parallel, high-performance computing applications, especially when it comes to characterizing global trends that would ease the transition of today's code base to tomorrow's parallel configurations. This paper first presents the characteristics of a typical Java-based business application software stack and examines the locking contentions that can appear at each level of that stack. Second, it presents scalability evaluation of three enterprise-grade, Java-based workloads and details the lock contention founds. Third, it summarizes the results of our findings, emphasizing the need for a streamlined methodology for lock-contention analysis of enterprise Java workloads.
Xavier Guerin, Wei Tan 0001, Seetharami Seelam, Parijat Dube
MASCOTS4
2012 vPFS: Bandwidth virtualization of parallel storage systems
abstract
Existing parallel file systems are unable to differentiate I/Os requests from concurrent applications and meet per-application bandwidth requirements. This limitation prevents applications from meeting their desired Quality of Service (QoS) as high-performance computing (HPC) systems continue to scale up. This paper presents vPFS, a new solution to address this challenge through a bandwidth virtualization layer for parallel file systems. vPFS employs user-level parallel file system proxies to interpose requests between native clients and servers and to schedule parallel I/Os from different applications based on configurable bandwidth management policies. vPFS is designed to be generic enough to support various scheduling algorithms and parallel file systems. Its utility and performance are studied with a prototype which virtualizes PVFS2, a widely used parallel file system. Enhanced proportional sharing schedulers are enabled based on the unique characteristics (parallel striped I/Os) and requirement (high throughput) of parallel storage systems. The enhancements include new threshold- and layout-driven scheduling synchronization schemes which reduce global communication overhead while delivering total-service fairness. An experimental evaluation using typical HPC benchmarks (IOR, NPB BTIO) shows that the throughput overhead of vPFS is small (;96% of target sharing ratio) for competing applications with diverse I/O patterns.
Yiqi Xu, Dulcardo Arteaga, Ming Zhao 0002, Yonggang Liu 0004, Renato J. O. Figueiredo, Seetharami Seelam
MSST6
2012 Experiences in building and scaling an enterprise application on multicore systems
abstract
SUMMARY Even though Java is the de facto programming language for enterprise applications, there exist only a limited number of Java‐based benchmarks to understand the performance on emerging multicore systems. To bridge this gap, this paper presents a report generation benchmark that is developed on top of Open Source Apache Geronimo's DayTrader benchmark. Report generation and rendering is at the heart of many enterprise business analytics and business intelligence software products, and it is used by many enterprise applications. We evaluate the performance scalability of this benchmark on a state‐of‐the‐art Power7 multicore system with 8 Power7 cores and 32 hardware threads. The benchmark throughput scales linearly up to eight hardware threads, but beyond that point, the throughput falls sharply. Significant locking in the Java class libraries for non‐shared objects results in this performance drop. Splitting the locks on these shared classes results in near linear scaling from eight to 32 threads and improved the throughput by 80%. We also show that the Linux operating system load balancing could result in a degraded application performance in hardware multithreaded systems and simultaneous‐multithreads‐aware task scheduling results in uniform core‐resource utilization as well as improved application performance. Copyright © 2011 John Wiley & Sons, Ltd.
Seetharami Seelam, Parijat Dube, Megumi Ito, Deniz Binay, Michael Dawson 0001, Pramod Nagaraja, Graeme Johnson, Liana L. Fong, Michel Hack, Xiaoqiao Meng, Li Zhang 0002
Concurr. Comput. Pract. Exp.1
2011 Characterization of System Services and Their Performance Impact in Multi-core Nodes
abstract
The performance of parallel applications on large scale systems is shown to disproportionately degrade due to interference from system services. This interference from system services is also known as jitter. However, there is limited understanding of sources and patterns of jitter on multi-core systems. In this paper, we identify and characterize jitter sources in terms of their amplitude and execution interval distributions on multi-core IBM Power systems with UNIX-based general purpose operating systems: AIX and Linux. Our analysis shows that there are various kinds of jitter sources and their execution varies drastically between different cores and between hardware threads within each core for practical reasons. This in-depth knowledge of jitter events is leveraged to devise effective approaches to mitigate the jitter impact on application performance in large scale systems. Moreover, such knowledge would provide useful insights to a new generation of operating system designs such as multikernel or satellite kernel for multicore systems.
Seetharami Seelam, Liana L. Fong, John Lewars, John Divirgilio, Brian F. Veale, Kevin J. Gildea
IPDPS1
2010 Extreme scale computing: Modeling the impact of system noise in multicore clustered systems
abstract
System noise or Jitter is the activity of hardware, firmware, operating system, runtime system, and management software events. It is shown to disproportionately impact application performance in current generation large-scale clustered systems running general-purpose operating systems (GPOS). Jitter mitigation techniques such as co-scheduling jitter events across operating systems improve application performance but their effectiveness on future petascale systems is unknown. To understand if existing co-scheduling solutions enable scalable petascale performance, we construct two complementary jitter models based on detailed analysis of system noise from the nodes of a large-scale system running a GPOS. We validate these two models using experimental data from a system consisting of 128 GPOS instances with 4096 CPUs. Based on our models, we project a minimum slowdown of 2.1%, 5.9%, and 11.5% for applications executing on a similar one petaflop system running 1024 GPOS instances and having global synchronization operations once every 1000 msec, 100 msec, and 10 msec, respectively. Our projections indicate that additional system noise mitigation techniques are required to contain the impact of jitter on multi-petaflop systems, especially for tightly synchronized applications.
Seetharami Seelam, Liana L. Fong, Asser N. Tantawi, John Lewars, John Divirgilio, Kevin J. Gildea
IPDPS1