EDBT 2026 Demo / reviewers in the wild / expert
Seetharami Seelam
dblp:19/1624
· DBLP profile ↗
16ranked-venue papers
4as first author
5since 2021 · last 2027
0000-0002-7595-3477ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Closed-loop calculations of electronic structure on a quantum processor and a classical supercomputer at full scaleabstractQuantum computers must operate in concert with classical computers to deliver on the promise of quantum advantage for practical problems. To achieve that, it is important to understand how quantum and classical computing can interact together, and how one can characterize the scalability and efficiency of hybrid quantum–classical workflows. So far, early experiments with quantum-centric supercomputing workflows have been limited in scale and complexity. Here, we use a Heron quantum processor deployed on premises with the entire supercomputer Fugaku to perform the largest computation of electronic structure involving quantum and classical high-performance computing. We design a closed-loop workflow between the quantum processors and 152,064 classical nodes of Fugaku, to approximate the electronic structure of chemistry models beyond the reach of exact diagonalization, with accuracy comparable to some all-classical approximation methods. Our work pushes the limits of the integration of quantum and classical high-performance computing, showcasing computational resource orchestration at the largest scale possible for current classical supercomputers. Tomonori Shirakawa, Javier Robledo Moreno, Toshinari Itoko, Vinay Tripathi, Kento Ueda, Yukio Kawashima, Lukas Broers, William M. Kirby, Himadri Pathak, Hanhee Paik, Miwako Tsuji, Yuetsu Kodama, Mitsuhisa Sato, Constantinos Evangelinos, Seetharami Seelam, Robert Walkup, Seiji Yunoki, Mario Motta, Petar Jurcevic, Hiroshi Horii, Antonio Mezzacapo |
Future Gener. Comput. Syst. | 15 |
| 2025 | Vela: A Virtualized LLM Training System with GPU Direct RoCEabstractVela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark. Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam |
ASPLOS (2) | 25 |
| 2025 | Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect StorageabstractWe present the design and implementation of a new lifetime-aware tensor offloading
framework for GPU memory expansion using low-cost PCIe-based solid-state
drives (SSDs). Our framework, TERAIO, is developed explicitly for large language
model (LLM) training with multiple GPUs and multiple SSDs. Its design is driven
by our observation that the active tensors take only a small fraction (1.7% on
average) of allocated GPU memory in each LLM training iteration, the inactive
tensors are usually large and will not be used for a long period of time, creating
ample opportunities for offloading/prefetching tensors to/from slow SSDs without
stalling the GPU training process. TERAIO accurately estimates the lifetime (active
period of time in GPU memory) of each tensor with the profiling of the first few
iterations in the training process. With the tensor lifetime analysis, TERAIO will
generate an optimized tensor offloading/prefetching plan and integrate it into the
compiled LLM program via PyTorch. TERAIO has a runtime tensor migration
engine to execute the offloading/prefetching plan via GPUDirect storage, which
allows direct tensor migration between GPUs and SSDs for alleviating the CPU
bottleneck and maximizing the SSD bandwidth utilization. In comparison with
state-of-the-art studies such as ZeRO-Offload and ZeRO-Infinity, we show that
TERAIO improves the training performance of various LLMs by 1.47× on average,
and achieves 80.7% of the ideal performance assuming unlimited GPU memory. Yirui Eric Zhou, Apoorve Mohan, I-Hsin Chung, Seetharami Seelam, Jian Huang 0006 |
NeurIPS | 6 |
| 2024 | 2024 IEEE International Conference on Cloud Computing, Message from the ChairsabstractWe are delighted to welcome all participants to the 2024 IEEE International Conference on Cloud Computing (CLOUD 2024), which is taking place in the beautiful city of Shenzhen, China, from July 7th to 13th. Tevfik Kosar, Krishnan Venkateswaran, Shangguang Wang, Seetharami Seelam, Santonu Sarkar, Xuanzhe Liu |
CLOUD | 4 |
| 2023 | Chic-sched: a HPC Placement-Group Scheduler on Hierarchical Topologies with ConstraintsabstractEfficient placement of advanced HPC and AI workloads with application constraints is raising challenges for resource schedulers on shared infrastructures, such as the Cloud. In this work, we propose a novel Constraints- and Heuristics-based scheduler on HIerarchical Topologies for High-Performance Computing workloads in the Cloud (chic-sched, for short). Our heuristics-based algorithm enables placement across multiple levels in a network hierarchy with loosely specified constraints, and it works without retries by providing suboptimal placements to minimize placement failures. This allows for fast scheduling at scale, and the O(N log N) complexity enables placement decisions within tens of milliseconds for groups of hundreds of virtual machines (VM). We introduce a new and simple metric to quantify the goodness of group placements. With this metric, in terms of deviation from ideal placements, we show that chic-sched is 20-50% better than the common bestFit or worstFit algorithms in all scenarios of two-level placements with spreading and packing constraints. We evaluate chic-sched with publicly available VM-request traces from a production Cloud, and, comparing against bestFit, we show that it achieves 8% lower placement failure rates and more than 40% better placement locality. Finally, to quantify the goodness of constraints-based placements, we conduct experiments with a realistic MPI workload on synthetically allocated VM clusters in a public cloud. We measure a 9% performance improvement over an adverse placement in a scenario where our heuristics-based scheduler would return a good, but not perfect, placement. Laurent Schares, Asser N. Tantawi, Pavlos Maniotis, Ming-Hung Chen, Claudia Misale, Seetharami Seelam, Hao Yu 0008 |
IPDPS | 6 |
| 2019 | ConfAdvisor: A Performance-centric Configuration Tuning Framework for Containers on KubernetesabstractConfiguration tuning of software is often a good option to improve application performance without any application code modifications. Although we can casually change configurations, it is not easy to apply optimal configurations, as optimal configurations require deep knowledge of the underlying system. This is problematic because applications with suboptimal configuration result in poor performance. As container and container management systems have emerged as an application platform on the cloud, configuration tuning becomes even more challenging because containers add more complexity to the application performance. We need to consider not only fundamental misconfiguration but also container image verification, deployment configuration, application characteristics awareness based on metrics and logs. Although previous knowledge regarding how we should tune configurations for a system software is sometimes available, knowledge about performance tuning practices is neither normalized nor reusable to expand on any advice for misconfiguration to the containers. Even in the cloud-native environment, there is no centralized service to deliver knowledge continuously to application containers nor a framework to develop a misconfiguration fix rule for a container throughout its lifetime. In this paper, we propose a performance-centric configuration tuning framework for containers on Kubernetes, named ConfAdvisor, that enables containers to achieve a higher performance by validating various misconfigurations adaptively. ConfAdivsor gives config tuning advice to application containers, images, and Kubernetes specs and also provides a development framework to build configuration validation rules. We present the design of ConfAdvisor and provide several case studies to tune application containers in the real world. Tatsuhiro Chiba, Rina Nakazawa, Hiroshi Horii, Sahil Suneja, Seetharami Seelam |
IC2E | 5 |
| 2017 | Taming Performance Degradation of Containers in the Case of Extreme Memory OvercommitmentabstractThe efficiency of datacenters is important consideration for cloud service providers to make their datacenters always ready for fulfilling the increasing demand for computing resources. Container-based virtualization is one approach to improving efficiency by reducing the overhead of virtualization. Resource overcommitment is another approach, but cloud providers tend to make conservative allocations of resources because there is no good understanding of the relationship between physical resource overcommitment and its impact on performance. This paper presents a quantitative study of performance degradation of containerized workloads due to memory overcommitment and a technique to mitigate it. We focused on physical memory overcommitment, where the sum of the working set memory is larger than the physical memory. We drove a small fraction of Docker containers at a high load level and the rest of them at a very low load level to emulate a common usage pattern of cloud datacenters. Detailed measurements revealed it is difficult to predict how many additional containers can be launched before thrashing hurts performance. We show that tuning the per-container swappiness of heavily loaded containers is effective for launching a larger number of containers and that it achieves an overcommitment of about three times. Rina Nakazawa, Kazunori Ogata, Seetharami Seelam, Tamiya Onodera |
CLOUD | 3 |
| 2015 | Introduction to special issue on High Performance Computing Architectures and Systems
Jia Hu 0001, Seetharami Seelam, Laurent Lefèvre |
J. Comput. Syst. Sci. | 2 |
| 2014 | Introduction to special issue on embedded systems architecture and applications
Jia Hu 0001, Jens Palsberg, Seetharami Seelam, Marco Di Natale, Lei (Chris) Liu |
J. Syst. Archit. | 3 |
| 2013 | Extreme scale computing: Modeling the impact of system noise in multi-core clustered systems
Seetharami Seelam, Liana L. Fong, Asser N. Tantawi, John Lewars, John Divirgilio, Kevin J. Gildea |
J. Parallel Distributed Comput. | 1 |
| 2012 | Partitioned Parallel Job Scheduling for Extreme Scale Computing
David Brelsford, George Chochia, Nathan Falk, Kailash Marthi, Ravindra Sure, Norman Bobroff, Liana L. Fong, Seetharami Seelam |
JSSPP | 8 |
| 2012 | Evaluation of Multi-core Scalability Bottlenecks in Enterprise Java WorkloadsabstractThe increasing number of cores integrated into modern processors is blurring the line between supercomputers and enterprise-grade servers. Therefore, the same attention to lock contention bottlenecks must be given to Java-based business workloads as it is given to massively parallel, high-performance computing applications, especially when it comes to characterizing global trends that would ease the transition of today's code base to tomorrow's parallel configurations. This paper first presents the characteristics of a typical Java-based business application software stack and examines the locking contentions that can appear at each level of that stack. Second, it presents scalability evaluation of three enterprise-grade, Java-based workloads and details the lock contention founds. Third, it summarizes the results of our findings, emphasizing the need for a streamlined methodology for lock-contention analysis of enterprise Java workloads. Xavier Guerin, Wei Tan 0001, Seetharami Seelam, Parijat Dube |
MASCOTS | 4 |
| 2012 | vPFS: Bandwidth virtualization of parallel storage systemsabstractExisting parallel file systems are unable to differentiate I/Os requests from concurrent applications and meet per-application bandwidth requirements. This limitation prevents applications from meeting their desired Quality of Service (QoS) as high-performance computing (HPC) systems continue to scale up. This paper presents vPFS, a new solution to address this challenge through a bandwidth virtualization layer for parallel file systems. vPFS employs user-level parallel file system proxies to interpose requests between native clients and servers and to schedule parallel I/Os from different applications based on configurable bandwidth management policies. vPFS is designed to be generic enough to support various scheduling algorithms and parallel file systems. Its utility and performance are studied with a prototype which virtualizes PVFS2, a widely used parallel file system. Enhanced proportional sharing schedulers are enabled based on the unique characteristics (parallel striped I/Os) and requirement (high throughput) of parallel storage systems. The enhancements include new threshold- and layout-driven scheduling synchronization schemes which reduce global communication overhead while delivering total-service fairness. An experimental evaluation using typical HPC benchmarks (IOR, NPB BTIO) shows that the throughput overhead of vPFS is small (;96% of target sharing ratio) for competing applications with diverse I/O patterns. Yiqi Xu, Dulcardo Arteaga, Ming Zhao 0002, Yonggang Liu 0004, Renato J. O. Figueiredo, Seetharami Seelam |
MSST | 6 |
| 2012 | Experiences in building and scaling an enterprise application on multicore systemsabstractSUMMARY Even though Java is the de facto programming language for enterprise applications, there exist only a limited number of Java‐based benchmarks to understand the performance on emerging multicore systems. To bridge this gap, this paper presents a report generation benchmark that is developed on top of Open Source Apache Geronimo's DayTrader benchmark. Report generation and rendering is at the heart of many enterprise business analytics and business intelligence software products, and it is used by many enterprise applications. We evaluate the performance scalability of this benchmark on a state‐of‐the‐art Power7 multicore system with 8 Power7 cores and 32 hardware threads. The benchmark throughput scales linearly up to eight hardware threads, but beyond that point, the throughput falls sharply. Significant locking in the Java class libraries for non‐shared objects results in this performance drop. Splitting the locks on these shared classes results in near linear scaling from eight to 32 threads and improved the throughput by 80%. We also show that the Linux operating system load balancing could result in a degraded application performance in hardware multithreaded systems and simultaneous‐multithreads‐aware task scheduling results in uniform core‐resource utilization as well as improved application performance. Copyright © 2011 John Wiley & Sons, Ltd. Seetharami Seelam, Parijat Dube, Megumi Ito, Deniz Binay, Michael Dawson 0001, Pramod Nagaraja, Graeme Johnson, Liana L. Fong, Michel Hack, Xiaoqiao Meng, Li Zhang 0002 |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | Characterization of System Services and Their Performance Impact in Multi-core NodesabstractThe performance of parallel applications on large scale systems is shown to disproportionately degrade due to interference from system services. This interference from system services is also known as jitter. However, there is limited understanding of sources and patterns of jitter on multi-core systems. In this paper, we identify and characterize jitter sources in terms of their amplitude and execution interval distributions on multi-core IBM Power systems with UNIX-based general purpose operating systems: AIX and Linux. Our analysis shows that there are various kinds of jitter sources and their execution varies drastically between different cores and between hardware threads within each core for practical reasons. This in-depth knowledge of jitter events is leveraged to devise effective approaches to mitigate the jitter impact on application performance in large scale systems. Moreover, such knowledge would provide useful insights to a new generation of operating system designs such as multikernel or satellite kernel for multicore systems. Seetharami Seelam, Liana L. Fong, John Lewars, John Divirgilio, Brian F. Veale, Kevin J. Gildea |
IPDPS | 1 |
| 2010 | Extreme scale computing: Modeling the impact of system noise in multicore clustered systemsabstractSystem noise or Jitter is the activity of hardware, firmware, operating system, runtime system, and management software events. It is shown to disproportionately impact application performance in current generation large-scale clustered systems running general-purpose operating systems (GPOS). Jitter mitigation techniques such as co-scheduling jitter events across operating systems improve application performance but their effectiveness on future petascale systems is unknown. To understand if existing co-scheduling solutions enable scalable petascale performance, we construct two complementary jitter models based on detailed analysis of system noise from the nodes of a large-scale system running a GPOS. We validate these two models using experimental data from a system consisting of 128 GPOS instances with 4096 CPUs. Based on our models, we project a minimum slowdown of 2.1%, 5.9%, and 11.5% for applications executing on a similar one petaflop system running 1024 GPOS instances and having global synchronization operations once every 1000 msec, 100 msec, and 10 msec, respectively. Our projections indicate that additional system noise mitigation techniques are required to contain the impact of jitter on multi-petaflop systems, especially for tightly synchronized applications. Seetharami Seelam, Liana L. Fong, Asser N. Tantawi, John Lewars, John Divirgilio, Kevin J. Gildea |
IPDPS | 1 |