EDBT 2026 Demo / reviewers in the wild / expert
Sudipta Saha Shubha
dblp:229/5547
· DBLP profile ↗
12ranked-venue papers
9as first author
9since 2021 · last 2026
0009-0002-9284-507XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaGen: Workload-Adaptive Cluster Scheduler for Latency-Optimal LLM Inference ServingabstractThe inference workloads of Large Language Models (LLMs) pose significant latency and cost challenges due to increasing model sizes and demand for real-time responses. Existing cluster schedulers for multi-instance LLM serving primarily focus on load balancing to optimize memory usage, which is insufficient for workloads with diverse request characteristics. In such cases, the compute layout — the arrangement of tokens across iterations within each instance—plays a crucial role in determining latency. We propose AdaGen, a workload-adaptive cluster scheduler that minimizes latency and thus maximizes SLO attainment by optimizing compute layouts across instances. AdaGen employs a multi-step scheduling strategy: it first classifies requests based on prefill and decode lengths, then balances load, and finally performs selective distributed execution across instances. Each step incrementally refines the scheduling based on the compute layouts derived from the decision of the previous step. To avoid the overhead of actual execution to generate the layouts, AdaGen introduces a novel simulation-based estimator. Extensive experiments using production workloads show that AdaGen achieves up to 3.6× higher SLO attainment and 2× better cost-efficiency compared to the existing systems, while ensuring scalability. Sudipta Saha Shubha, Ayush Goel, Diman Zad Tootaghaj, Khaled Diab 0001, Hardik Soni 0001, K. K. Ramakrishnan, Puneet Sharma 0001, Haiying Shen |
EuroSys | 1 |
| 2025 | CIS: Checkpointed Inference for Data Drift-Resilient Model Serving at Edge ServersabstractSmall deep learning models deployed at edge servers suffer from decreasing accuracy due to data drift and hence require continual learning, which leads to resource competition with inference execution, decreasing accuracy and latency service-level-objective (SLO) fulfillment. Previous methods fail to maximize both accuracy and SLO fulfillment. To address this problem, we propose a Checkpointed Inference-based system for accurate and SLO-guaranteed data drift-resilient model Serving (CIS). CIS incorporates checkpointed inference - maximizing accuracy by continuously switching to intermediately retrained models during inference. However, model switching introduces significant inference latency overhead. To mitigate this, first, CIS proposes a lightweight allocation and placement scheduler to minimize switching time. Second, during inference execution, CIS reorders retraining samples to reduce switching frequency. Third, it temporarily reallocates GPU space from retraining tasks to inference tasks to address request queuing issue caused by the switching, with minimal impact on accuracy. Trace-driven experiments show that CIS achieves up to 25.1% higher accuracy without affecting SLO fulfillment, and requires 4× lower GPU cost compared to the existing method in achieving similar or higher accuracy. Sudipta Saha Shubha, Haiying Shen, Ganesh Ananthanarayanan |
SoCC | 1 |
| 2024 | An Accurate and Efficient Clustered Federated Learning for Mobile Edge DevicesabstractFederated Learning (FL) has been increasingly used in various edge device applications. Cluster-wise separate Federated Learning (CFL) is an effective means of addressing the low accuracy issue resulting from data heterogeneity in FL. In CFL, nearby edge devices form a cluster and each cluster trains a separate machine learning (ML) model. As many edge devices (e.g., cellular phones) are mobile, a device may leave its cluster during training, which degrades the training accuracy. However, the existing works on CFL do not consider the mobility of the devices. We propose enabling such devices to keep participating in the trainings for its previous clusters and also the current cluster. However, due to constrained resources and high workload of multiple CFLs, such devices may become stragglers, affecting training time and accuracy. Also, cluster environment changes and less-visited clusters will generate low model accuracy. In this paper, we first conducted experimental analysis to verify the motivation and illustrate the problems. Then, to address the problems, we propose Clustered Federated Learning for Mobile edge devices (CFLM). CFLM decreases training time by sharing training data and computation results between multiple trainings. Additionally, CFLM increases accuracy by handling dynamic environment and by increasing the accuracy of lessvisited clusters. Our extensive evaluations on both CPU and GPU devices show that CFLM decreases training time by up to 68% and increases accuracy by up to 18% compared to the existing works. Sudipta Saha Shubha, Haiying Shen |
SEC | 1 |
| 2024 | USHER: Holistic Interference Avoidance for Resource Optimized ML Inference
Sudipta Saha Shubha, Haiying Shen, Anand Padmanabha Iyer |
OSDI | 1 |
| 2023 | Accurate and Efficient Distributed COVID-19 Spread Prediction based on a Large-Scale Time-Varying People Mobility GraphabstractCompared to previous epidemics, COVID-19 spreads much faster in people gatherings. Thus, we need not only more accurate epidemic spread prediction considering the people gatherings but also more time-efficient prediction for taking actions (e.g., allocating medical equipments) in time. Motivated by this, we analyzed a time-varying people mobility graph of the United States (US) for one year and the effectiveness of previous methods in handling time-varying graphs. We identified several factors that influence COVID-19 spread and observed that some graph changes are transient, which degrades the effectiveness of the previous graph repartitioning and replication methods in distributed graph processing since they generate more time overhead than saved time. Based on the analysis, we propose an accurate and time-efficient Distributed Epidemic Spread Prediction system (DESP). First, DESP incorporates the factors into a previous prediction model to increase the prediction accuracy. Second, DESP conducts repartitioning and replication only when a graph change is stable for a certain time period (predicted using machine learning) to ensure the operation improves time-efficiency. We conducted extensive experiments on Amazon AWS based on real people movement datasets. Experimental results show DESP reduces communication time by up to 52%, while enhancing accuracy by up to 24% compared to existing methods. Sudipta Saha Shubha, Shohaib Mahmud, Haiying Shen, Geoffrey C. Fox, Madhav V. Marathe |
IPDPS | 1 |
| 2023 | AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersabstractVarious audio and video applications rely on multiple deep neural network (DNN) models deployed on edge servers to conduct inference with ms-level latency service-level-objectives (SLOs). To avoid accuracy decreases caused by data drift, continual retraining is necessary. However, this poses a challenge for GPU resource allocation to satisfy the tight SLOs while maintaining high accuracy in this scenario. There has been no research devoted to tackling this issue. In this paper, we conducted trace-based experimental analysis in this particular scenario, which shows that different models have varying degrees of impact from data drift, incremental retraining (proposed by us that retrains certain samples before inference) and early-exit model structures can help increase accuracy, and the interdependencies among tasks may lead to significant CPU-GPU memory communications. Leveraging these unique observations, we propose a data drift Adaptive scheduler for accurate and SLO-guaranteed Inference serving at edge servers (AdaInf). AdaInf uses incremental retraining and allocates GPU amount among applications based on their SLOs. For each application, it splits GPU time between retraining and inference to satisfy its SLO, and then allocates GPU time among retraining tasks based on their impact degrees. In addition, AdaInf proposes strategies that leverage the job features in this scenario to reduce the impact of CPU-GPU memory communications on latency. Our real trace-driven experimental evaluation shows that AdaInf can increase accuracy by up to 21% and reduce SLO violations by up to 54% compared to existing methods. Achieving similar accuracy as AdaInf requires 4× more GPU resources on the edge server for the existing method. Sudipta Saha Shubha, Haiying Shen |
SIGCOMM | 1 |
| 2022 | Trustworthy Distributed Deep Neural Network Training in an Edge Device NetworkabstractWith the increased usage of edge devices having local computation capabilities, deep neural network (DNN) training in a network of edge devices becomes promising. Several recent works have proposed fully edge-based distributed training systems for situations when the communication to cloud is unstable or intermittent. However, such distributed systems become vulnerable when there are untrusted devices that launch data and model poisoning attacks during training, deteriorating the accuracy of the DNN model. To handle this challenge, we propose a Trustworthy distributed system for Machine learning training in an edge device network (TrustMe). TrustMe realizes both data and model parallelisms. It detects the untrusted devices producing illegitimate outputs. Next, it reassigns the training tasks of the untrusted devices to other trusted devices in such a way that the reassignment and the training that is restarted after the reassignment require minimal time. Our container-based emulation and real device experiments demonstrate that TrustMe achieves up to 12% higher accuracy and 45% less training time compared to existing methods in the presence of untrusted devices. Sudipta Saha Shubha, Haiying Shen |
IEEE Big Data | 1 |
| 2021 | Challenges of Distributed Computing for Pandemic Spread Prediction based on Large-Scale Human Interaction DataabstractPandemic like COVID-19 poses severe challenges to public health and causes great damages to human lives and economy. Computational epidemiology allows the authorities (e.g., federal, state, city governments) to predict the future states of pandemic and take preventive measures based on that. However, due to widespread human mobility, epidemic prediction on a small-scale (e.g., city) may not produce effective results. Therefore, to obtain a clearer picture of pandemics, authorities often rely on large-scale human mobility data. Processing such a large graph dataset is highly computation intensive, and thus generates long latency. Distributed computing, in which the graph is partitioned and processed by multiple servers in parallel, is a solution to this problem. However, as human mobility changes over time, the graph varies over time, which may make the previous graph partition not effective anymore in limiting the communication overhead. In this paper, we study the existing works in the literature for processing large-scale human mobility graph data to predict pandemic spread. We conducted comprehensive experiments based on real-world large-scale human mobility data that covers the entire state of Virginia, USA. Based on the experimental evaluations, we present our findings and discuss the possible future research directions for the aforementioned challenge. Sudipta Saha Shubha, Shohaib Mahmud, Haiying Shen |
CLOUD | 1 |
| 2021 | A Diverse Noise-Resilient DNN Ensemble Model on Edge Devices for Time-Series DataabstractMany applications such as healthcare and transportation on edge devices will use deep neural network (DNN) prediction based on time-series data collected by the devices. However, the existence of noises in the on-device sensors negatively impacts the sensing output of the DNN models. The state-of-the-art time-series based DNN approaches can deal with Gaussian noise but cannot effectively handle other types of noises in spite of the existence of different types of noises such as shot, burst, transient noises, and their combination. In this paper, we propose an ensemble-based DNN model, namely E-Sense, which consists of different expert models for different noises and shows higher prediction accuracy. Since an edge device may have limited resources to run a large DNN model, we further propose a novel searching-based model compression method called E-Comp that uses knowledge distillation to compress E-Sense to a smaller DNN model while maintaining the accuracy. Our real experiments on live sensor data and trace-driven experiments on three real traces show that E-Sense outperforms other methods in accuracy, and E-Comp reduces 27% inference time without sacrificing accuracy compared with other DNN compression methods. We also distributed our source code. Sudipta Saha Shubha, Tanmoy Sen, Haiying Shen, Matthew Normansell |
SECON | 1 |
| 2020 | Bangla Voice Command Recognition in end-to-end System Using Topic Modeling based Contextual RescoringabstractIn this work, we perform contextual rescoring using multi-label topic modeling to improve the performance of an End-to-End Bangla voice command recognition system. We use a hybrid of Connectionist Temporal Classification (CTC) and Attention mechanism in our End-to-End architecture. We use Recurrent Neural Network (RNN) as language model and La-beled LDA (Latent Dirichlet allocation) for contextual rescoring. Our experiments show that our rescoring method reduces Word Error Rate (WER) from 16.7% to 12.8% in Bangla voice command recognition task when the relevant context is provided. The system does not lose any performance when irrelevant context is provided. Nafis Sadeq, Shafayat Ahmed, Sudipta Saha Shubha, Md. Nahidul Islam, Muhammad Abdullah Adnan |
ICASSP | 3 |
| 2020 | Preparation of Bangla Speech Corpus from Publicly Available Audio & TextabstractAutomatic speech recognition systems require large annotated speech corpus. The manual annotation of a large corpus is very difficult. In this paper, we focus on the automatic preparation of a speech corpus for Bangladeshi Bangla. We have used publicly available Bangla audiobooks and TV news recordings as audio sources. We designed and implemented an iterative algorithm that takes as input a speech corpus and a huge amount of raw audio (without transcription) and outputs a much larger speech corpus with reasonable confidence. We have leveraged speaker diarization, gender detection, etc. to prepare the annotated corpus. We also have prepared a synthetic speech corpus for handling out-of-vocabulary word problems in Bangla language. Our corpus is suitable for training with Kaldi. Experimental results show that the use of our corpus in addition to the Google Speech corpus (229 hours) significantly improves the performance of the ASR system. Shafayat Ahmed, Nafis Sadeq, Sudipta Saha Shubha, Md. Nahidul Islam, Muhammad Abdullah Adnan, Mohammad Zuberul Islam |
LREC | 3 |
| 2018 | Maximizing heterogeneous coverage in over and under provisioned visual sensor networks
Abdullah Al Zishan, Imtiaz Karim, Sudipta Saha Shubha, Ashikur Rahman |
J. Netw. Comput. Appl. | 3 |