EDBT 2026 Demo / reviewers in the wild / expert
Srinivasan Subramaniyan
dblp:268/1747
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-5848-5667ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DySM: Dynamic Scaling of GPU Streaming Multiprocessor in Spatially Shared Real-Time Embedded GPU SystemsabstractMany of today’s real-time embedded systems are increasingly relying on GPUs for AI-related computing. However, existing GPU scheduling solutions for spatially shared GPU systems are still mostly open-loop and rely on worst-case execution time (WCET) estimation for offline schedulability analysis, which cannot adapt to online workload variations. Although adaptive scheduling has been proposed to handle runtime execution time variations, prior approaches target either CPU or time-slicing GPUs, where only one task can execute on the GPU within a time slice. In contrast, spatial sharing enables concurrent kernel execution via Streaming Multiprocessor (SM) partitioning, allowing better GPU resource utilization. Therefore, new adaptive solutions must be designed for spatially shared GPU systems. In this paper, we propose DySM, a closed-loop response time control algorithm for spatially shared GPUs in soft real-time systems. In face of runtime workload variations, DySM leverages dynamic SM scaling to control task response times with low runtime overhead. To model GPU resource contention among tasks, we analytically derive a multi-input-multi-output (MIMO) system model that captures the impact of SM scaling on the response times of different tasks. Based on this model, DySM is designed using feedback control theory for guaranteed system stability and control accuracy. Experimental results on an Nvidia GPU testbed demonstrate that DySM outperforms state-of-the-art solutions by providing runtime real-time guarantees. Compared to the best-performing baseline, DySM can reduce the deadline miss ratio by up to 90.93%. Srinivasan Subramaniyan |
ECRTS | 1 |
| 2026 | CATS: Correlation-aware Task Scheduling for GPU Power Optimization in AI Data CentersabstractGPUs have been increasingly deployed in the data centers of big IT companies to enhance their AI/ML infrastructures. Since GPUs typically consume significantly more power than CPUs, it is crucial to optimize the power consumption of GPU data centers. Task consolidation has been demonstrated to be an effective way to reduce GPU power consumption by consolidating ML tasks onto a smaller set of GPUs and putting unused GPUs and servers into sleep. Unfortunately, existing work on GPU sharing and consolidation assumes that the GPU utilization of each ML task can be approximated as a constant during consolidation. This is in contrast to our analysis of real-world traces, which shows the GPU utilization of ML workloads fluctuates significantly over time. Srinivasan Subramaniyan |
ICS | 1 |
| 2025 | Power Capping of GPU Servers for Machine Learning Inference OptimizationabstractPower capping, which is an essential component of power oversubscription, has been widely used in data centers to host more servers than allowed by the capacity of their power infrastructures, in order to avoid expensive power upgrade and reduce capital expenses. Traditionally, power capping is performed mainly with CPU frequency and voltage scaling, which cannot be directly applied to the GPU servers that are commonly deployed in today’s data centers, because GPUs can have much higher power consumption than CPUs. Recently proposed GPU power capping solutions are designed for a single GPU and so cannot be used on GPU servers that have a host CPU and multiple GPUs to process machine learning (ML) workloads. Hence, a joint power capping solution must be designed to coordinate the host CPU and all the GPUs in a server for optimizing ML inference performance. Srinivasan Subramaniyan |
ICPP | 2 |
| 2025 | SEEB-GPU: Early-Exit Aware Scheduling and Batching for Edge GPU InferenceabstractThe deployment of deep neural networks (DNNs) on edge devices is becoming increasingly common in latency-sensitive applications such as autonomous driving, real-time video analytics, and augmented reality. However, modern DNNs are rapidly growing in complexity, and edge GPUs often lack the computational resources available in cloud counterparts. This leads to increased inference latency and challenges in meeting strict Service Level Agreements (SLAs). Srinivasan Subramaniyan, Rudra Joshi, Marco Brocanelli |
SEC | 1 |
| 2025 | Exploiting ML Task Correlation in the Minimization of Capital Expense for GPU Data CentersabstractEfficiently scheduling Machine Learning (ML) training tasks in a GPU data center presents a significant research challenge. Existing solutions commonly schedule such tasks based on their demanded GPU utilization, but simply assume that the GPU utilization of each task can be approximated as a constant number (e.g., by using the peak value), even though the ML training tasks commonly have their GPU utilization varying significantly over time. Using a constant number to schedule tasks can result in an overestimation of the needed GPU count and, therefore, a high capital expense for GPU purchases. To address this, we design CorrGPU, a correlation-aware GPU scheduling algorithm that considers the utilization correlation among different tasks to minimize the number of needed GPUs in a data center. CorrGPU is designed based on a key observation from the analysis of real ML traces that different tasks do not have their GPU utilization peak at exactly the same time. As a result, if the correlations among tasks are considered in scheduling, more tasks can be scheduled onto the same GPUs, without extending the training duration beyond the desired due time. For a GPU data center to be constructed based on an estimated ML workload, CorrGPU can help the operators purchase a smaller number of GPUs, thus minimizing their capital expense. Our hardware testbed results demonstrate CorrGPU's potential to reduce the number of GPUs needed. Our simulation results on real-world ML traces also show that CorrGPU outperforms several state-of-the-art solutions by reducing capital expense by 20.88 %. Srinivasan Subramaniyan |
IPCCC | 1 |
| 2025 | FC-GPU: Feedback Control GPU Scheduling for Real-time Embedded SystemsabstractGPUs have recently been adopted in many real-time embedded systems. However, existing GPU scheduling solutions are mostly open-loop and rely on the estimation of worst-case execution time (WCET). Although adaptive solutions, such as feedback control scheduling, have been previously proposed to handle this challenge for CPU-based real-time tasks, they cannot be directly applied to GPU, because GPUs have different and more complex architectures and so schedulable utilization bounds cannot apply to GPUs yet. In this article, we propose FC-GPU, the first Feedback Control GPU scheduling framework for real-time embedded systems. To model the GPU resource contention among tasks, we analytically derive a multi-input-multi-output (MIMO) system model that captures the impacts of task rate adaptation on the response times of different tasks. Building on this model, we design a MIMO controller that dynamically adjusts task rates based on measured response times. Our extensive hardware testbed results on an Nvidia RTX 3090 GPU and an AMD MI-100 GPU demonstrate that FC-GPU can provide better real-time performance even when the task execution times significantly increase at runtime. Srinivasan Subramaniyan |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | Latency-Guaranteed Co-Location of Inference and Training for Reducing Data Center ExpensesabstractToday's data centers often need to run various machine learning (ML) applications with stringent SLO (Service-Level Objective) requirements, such as inference latency. To that end, data centers prefer to 1) over-provision the number of servers used for inference processing and 2) isolate them from other servers that run ML training, despite both use GPUs extensively, to minimize possible competition of computing resources. Those practices result in a low GPU utilization and thus a high capital expense. Hence, if training and inference jobs can be safely co-located on the same GPUs with explicit SLO guarantees, data centers could flexibly run fewer training jobs when an inference burst arrives and run more afterwards to increase GPU utilization, reducing their capital expenses. In this paper, we propose GPUColo, a two-tier co-location solution that provides explicit ML inference SLO guarantees for co-located GPUs. In the outer tier, we exploit GPU spatial sharing to dynamically adjust the percentage of active GPU threads allocated to spatially co-located inference and training processes, so that the inference latency can be guaranteed. Because spatial sharing can introduce considerable overheads and thus cannot be conducted at a fine time granularity, we design an inner tier that puts training jobs into periodic sleep, so that the inference jobs can quickly get more GPU resources for more prompt latency control. Our hardware testbed results show that GPUColo can precisely control the inference latency to the desired SLO, while maximizing the throughput of the training jobs co-located on the same GPUs. Our large-scale simulation with a 57-day real-world data center trace (6500 GPUs) also demonstrates that GPU Colo enables latency-guaranteed inference and training co-location. Consequently, it allows 74.9 % of GPUs to be saved for a much lower capital expense. Guoyu Chen, Srinivasan Subramaniyan |
ICDCS | 2 |
| 2020 | Gbit/s Non-Binary LDPC Decoders: High-Throughput using High-Level SpecificationsabstractIt is commonly perceived that an HLS specification targeted for FPGAs cannot provide throughput performance in par with equivalent RTL descriptions. In this work we developed a complex design of a non-binary LDPC decoder, that although hard to generalise, shows that HLS provides sufficient architectural refinement options. They allow attaining performance above CPU- and GPU-based ones and excel at providing a faster design cycle when compared to RTL development. Oscar Ferraz, Srinivasan Subramaniyan, Joseph R. Cavallaro, Gabriel Falcão Paiva Fernandes, Madhura Purnaprajna |
FCCM | 2 |