VLDB 2026 Research / reviewers in the wild / expert
Sahil Tyagi
dblp:206/7277
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0007-8314-4745ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training
Sahil Tyagi, Feiyi Wang |
CCGrid | 1 |
| 2025 | OmniLearn: A Framework for Distributed Deep Learning Over Heterogeneous ClustersabstractDeep learning systems are optimized for clusters with homogeneous resources. However, heterogeneity is prevalent in computing infrastructure across edge, cloud and HPC. When training neural networks using stochastic gradient descent techniques on heterogeneous resources, performance degrades due to stragglers and stale updates. In this work, we develop an adaptive batch-scaling framework calledOmniLearnto mitigate the effects of heterogeneity in distributed training. Our approach is inspired by proportional controllers to balance computation across heterogeneous servers, and works under varying resource availability. By dynamically adjusting worker mini-batches at runtime,OmniLearnreduces training time by 14-85%. We also investigate asynchronous training, where our techniques improve accuracy by up to 6.9%. Sahil Tyagi, Prateek Sharma 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | GraVAC: Adaptive Compression for Communication-Efficient Distributed DL TrainingabstractDistributed data-parallel (DDP) training improves overall application throughput as multiple devices train on a subset of data and aggregate updates to produce a globally shared model. The periodic synchronization at each iteration incurs considerable overhead, exacerbated by the increasing size and complexity of state-of-the-art neural networks. Although many gradient compression techniques propose to reduce communication cost, the ideal compression factor that leads to maximum speedup or minimum data exchange remains an open-ended problem since it varies with the quality of compression, model size and structure, hardware, network topology and bandwidth. We propose GraVAC, a framework to dynamically adjust compression factor throughout training by evaluating model progress and assessing gradient information loss associated with compression. GraVAC works in an online, black-box manner without any prior assumptions about a model or its hyperparameters, while achieving the same or better accuracy than dense SGD (i.e., no compression) in the same number of iterations/epochs. As opposed to using a static compression factor, GraVAC reduces end-to-end training time for ResNet101, VGG16 and LSTM by 4.32×, 1.95× and 6.67× respectively. Compared to other adaptive schemes, our framework provides 1.94× to 5.63× overall speedup. Sahil Tyagi, D. Martin Swany |
CLOUD | 1 |
| 2023 | Flexible Communication for Optimal Distributed Learning over Unpredictable NetworksabstractGradient compression alleviates expensive communication in distributed deep learning by sending fewer values and its corresponding indices, typically via Allgather (AG). Training with high compression ratio (CR) achieves high accuracy like DenseSGD, but has lower parallel scaling due to high communication cost (i.e., parallel efficiency). Using lower CRs improves parallel efficiency by lowering synchronization cost, but degrades model accuracy as well (statistical efficiency). Further, speedup attained with different models and CRs also varies with network latency, effective bandwidth and collective op used for aggregation. In many cases, collectives like Allreduce (AR) have lower cost than AG to exchange the same amount of data. In this paper, we propose an AR-compatible Top k compressor that is bandwidth-optimal and thus performs better than AG in certain network configurations. We develop a flexible communication strategy that switches between AG and AR based on which collective is optimal in the current settings, and model the pareto-relationship between parallel and statistical efficiency as a multiobjective optimization (MOO) problem to dynamically adjust CR and accelerate training while still converging to high accuracy. Sahil Tyagi, D. Martin Swany |
IEEE Big Data | 1 |
| 2023 | Scavenger: A Cloud Service For Optimizing Cost and Performance of ML TrainingabstractCloud computing platforms can provide the compu-tational resources required for training large machine learning models such as deep neural networks. While the pay-as-you- go nature of cloud virtual machines (VMs) makes it easy to spin-up large clusters for training models, it can also lead to ballooning costs. The 100s of virtual machine sizes provided by cloud platforms also makes it extremely challenging to select the “right” cloud cluster configuration for training. Furthermore, the training time and cost of distributed model training is highly sensitive to the cluster configurations, and presents a large and complex tradeoff-space. In this paper, we develop principled and practical techniques for optimizing the training time and cost of distributed ML model training on the cloud. Our key insight is that both the parallel and statistical efficiency must be considered when selecting the optimum job configuration parameters such as the number of workers and the batch size. By combining conventional parallel scaling concepts and new insights into SGD noise, we develop models for estimating the time and cost on different cluster configurations. Using the repetitive nature of training and our performance models, our Scavenger cloud service can search for optimum cloud configurations in a black-box, online manner. Our approach reduces training times by 2 x and costs by more than 50 %. Our performance models are accurate to within 2 %, and our search imposes only a 10% overhead compared to an ideal oracle- based approach. Sahil Tyagi, Prateek Sharma 0001 |
CCGrid | 1 |
| 2023 | Accelerating Distributed ML Training via Selective SynchronizationabstractIn distributed training, deep neural networks (DNNs) are launched over multiple workers concurrently and aggregate their local updates on each step in bulk-synchronous parallel (BSP) training. However, BSP does not linearly scale-out due to high communication cost of aggregation. To mitigate this overhead, alternatives like Federated Averaging (FedAvg) and Stale-Synchronous Parallel (SSP) either reduce synchronization frequency or eliminate it altogether, usually at the cost of lower final accuracy. In this paper, we present SelSync, a practical, low-overhead method for DNN training that dynamically chooses to incur or avoid communication at each step either by calling the aggregation op or applying local updates based on their significance. We propose various optimizations as part of SelSync to improve convergence in the context of semi-synchronous training. Our system converges to the same or better accuracy than BSP while reducing training time by up to 14×. Sahil Tyagi, D. Martin Swany |
CLUSTER | 1 |
| 2022 | ScaDLES: Scalable Deep Learning over Streaming data at the EdgeabstractDistributed deep learning (DDL) training systems are designed for cloud and data-center environments that assumes homogeneous compute resources, high network bandwidth, sufficient memory and storage, a s w ell a s independent and identically distributed (IID) data across all nodes. However, these assumptions don’t necessarily apply on the edge, especially when training neural networks on streaming data in an online manner. Computing on the edge suffers from both systems and statistical heterogeneity. Systems heterogeneity is attributed to differences in compute resources and bandwidth specific to each device, while statistical heterogeneity comes from unbalanced and skewed data on the edge. Different streaming-rates among devices can be another source of heterogeneity when dealing with streaming data. If the streaming rate is lower than training batch-size, device needs to wait until enough samples have streamed in before performing a single iteration of stochastic gradient descent (SGD). Thus, low-volume streams act like stragglers slowing down devices with high-volume streams in synchronous training. On the other hand, data can accumulate quickly in the buffer if the streaming rate is too high and the devices can’t train at line-rate. In this paper, we introduce ScaDLES to efficiently train on streaming data at the edge in an online fashion, while also addressing the challenges of limited bandwidth and training with non-IID data. We empirically show that ScaDLES converges up to 3.29× faster compared to conventional distributed SGD. Sahil Tyagi, D. Martin Swany |
IEEE Big Data | 1 |
| 2021 | Cost-Effective Sharing of Streaming Dataflows for IoT ApplicationsabstractInternet of Things (IoT) applications are often designed as dataflows that analyze sensor data in real-time to make decisions. Stream processing systems likeApache Stormexecute these on Cloud infrastructure. As IoT applications within shared data environments like smart cities grow, they will duplicate tasks like pre-processing and analytics. This offers the opportunity to collaboratively reuse the outputs of overlapping dataflows, improving the resource efficiency on Clouds. We proposedataflow reuse algorithmsthat when given a submitted dataflow, identify the intersection of reusable tasks and streams from existing dataflows to form amerged dataflow, with guaranteed equivalence of their output streams. Algorithms to unmerge dataflows when they are removed, and defragment partially reused dataflows are also proposed. We implement these algorithms for the Storm fast-data platform, and validate their performance and resource savings using 86 real and synthetic dataflows from eScience and IoT domains. Our reuse strategies reduce the number of running tasks by 34–45 percent and the cumulative CPU usage by 29–63 percent. Including defragmentation of incremental dataflows achieves a monetary savings on Cloud resources of 36–44 percent compared to dataflows without reuse, and has limited redeployment overheads. Shilpa Chaturvedi, Sahil Tyagi, Yogesh L. Simmhan |
IEEE Trans. Cloud Comput. | 2 |
| 2019 | Anomaly Detection over Streaming Data: Indy500 Case StudyabstractSports racing is attracting billions of audiences each year. It is powered and transformed by the latest data analysis technologies, from race car design, driving skill improvements to audience engagement on social media. However, most of the data processing are off-line and retrospective analysis. The emerging real-time data analysis from the Internet of Things (IoT) result in fast data streams generated from distributed sensors. Applying advanced Machine Learning/Artificial Intelligence over such data streams to discover new information, predict future insights and make control decision is a crucial process. In this paper, we start by articulating racing car big data characteristics and present time-critical anomaly detection of the racing cars with the real-time sensors of cars and the tracks from actual racing events. We build a scalable system infrastructure based on neuro-morphic Hierarchical Temporal Memory Algorithm (HTM) algorithm and Storm stream processing engine. By courtesy of historical Indy500 racing logs, evaluation experiments on this prototype system demonstrate good performance in terms of anomaly detection accuracy and service level objective (SLO) of latency for a real-world streaming application. Chathura Widanage, Sahil Tyagi, Ravi Teja, Bo Peng 0011, Supun Kamburugamuve, Dan Baum, Dayle Smith, Judy Qiu, Jon Koskey |
CLOUD | 3 |
| 2017 | Collaborative Reuse of Streaming Dataflows in IoT ApplicationsabstractDistributed Stream Processing Systems (DSPS) like Apache Storm and Spark Streaming enable composition of continuous dataflows that execute persistently over data streams. They are used by Internet of Things (IoT) applications to analyze sensor data from Smart City cyber-infrastructure, and make active utility management decisions. As the ecosystem of such IoT applications that leverage shared urban sensor streams continue to grow, applications will perform duplicate pre-processing and analytics tasks. This offers the opportunity to collaboratively reuse the outputs of overlapping dataflows, thereby improving the resource efficiency. In this paper, we propose dataflow reuse algorithms that given a submitted dataflow, identifies the intersection of reusable tasks and streams from a collection of running dataflows to form a merged dataflow. Similar algorithms to unmerge dataflows when they are removed are also proposed. We implement these algorithms for the popular Apache Storm DSPS, and validate their performance and resource savings for 35 synthetic dataflows based on public OPMW workflows with diverse arrival and departure distributions, and on 21 real IoT dataflows from RIoTBench. We see that our Reuse algorithms reduce the count of running tasks by 38 - 46% for the two workloads, and a reduction in cumulative CPU usage of 36-51%, that can result in real cost savings on Cloud resources. Shilpa Chaturvedi, Sahil Tyagi, Yogesh L. Simmhan |
eScience | 2 |