Alind Khare

dblp:211/0360 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-4649-9022ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 84% Kernel, tree and ensemble methods · 9% Time series and sequential data · 7%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
1.322024
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-device Inference · ECCV (79) 2024
CompOFA - Compound Once-For-All Networks for Faster Multi-Platform Deployment · ICLR 2021
Cloud and datacenter computing
inference serving
0.912025
SuperServe: Fine-Grained Inference Serving for Unpredictable Workloads · NSDI 2025
Machine learning › Efficient and distributed learning
federated learning
0.812024
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-device Inference · ECCV (79) 2024
Machine learning › Efficient and distributed learning › federated learning › federated AutoML
federated neural architecture search
0.812024
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-device Inference · ECCV (79) 2024
Machine learning › Efficient and distributed learning
on-device inference
0.812024
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-device Inference · ECCV (79) 2024
Machine learning › Efficient and distributed learning
inference serving
0.722025
HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units · KDD 2020
SuperServe: Fine-Grained Inference Serving for Unpredictable Workloads · NSDI 2025
Machine learning › Efficient and distributed learning › inference efficiency
cost-aware inference
0.612022
UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification · NeurIPS 2022
Machine learning › Time series and sequential data › time series modeling
dynamic prediction
0.612022
UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification · NeurIPS 2022
Machine learning › Efficient and distributed learning › adaptive computation
model cascading
0.612022
UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification · NeurIPS 2022
Machine learning › Kernel, tree and ensemble methods › classifier combination
multi-stage classification
0.612022
UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification · NeurIPS 2022
Machine learning › Efficient and distributed learning
model compression
0.512021
CompOFA - Compound Once-For-All Networks for Faster Multi-Platform Deployment · ICLR 2021
Medical and health informatics
clinical decision support
0.412020
HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units · KDD 2020
Machine learning › Kernel, tree and ensemble methods
model ensemble
0.112020
HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units · KDD 2020

Methods — techniques the papers use, named apart from their topics

neural architecture search · 1.3latency-aware scheduling · 0.9ensemble selection · 0.9federated learning · 0.8uncertainty estimation · 0.6cascade of classifiers · 0.6compound once-for-all · 0.5
YearPublicationVenuePosition
2025 ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
abstract
Large multimodal models (LMMs) demonstrate impressive capabilities in understanding images, videos, and audio beyond text. However, efficiently serving LMMs in production environments poses significant challenges due to their complex model architectures and heterogeneous characteristics across their multi-stage inference pipelines and modalities.
Haoran Qiu, Anish Biswas, Jayashree Mohan, Alind Khare, Esha Choukse, Íñigo Goiri, Zeyu Zhang 0005, Haiying Shen, Chetan Bansal, Ramachandran Ramjee, Rodrigo Fonseca
SoCC5
2025 SuperServe: Fine-Grained Inference Serving for Unpredictable Workloads
Alind Khare, Dhruv Garg, Sukrit Kalra, Snigdha Grandhi, Ion Stoica, Alexey Tumanov
NSDI1
2024 DεpS: Delayed ε-Shrinking for Faster Once-for-All Training
Aditya Annavajjala, Alind Khare, Animesh Agrawal, Igor Fedorov, Hugo Latapie, Myungjin Lee, Alexey Tumanov
ECCV (89)2
2024 SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-device Inference
Alind Khare, Animesh Agrawal, Aditya Annavajjala, Payman Behnam, Myungjin Lee, Hugo Latapie, Alexey Tumanov
ECCV (79)1
2022 UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification
abstract
Machine Learning (ML) research has focused on maximizing the accuracy of predictive tasks. ML models, however, are increasingly more complex, resource intensive, and costlier to deploy in resource-constrained environments. These issues are exacerbated for prediction tasks with sequential classification on progressively transitioned stages with “happens-before” relation between them.We argue that it is possible to “unfold” a monolithic single multi-class classifier, typically trained for all stages using all data, into a series of single-stage classifiers. Each single- stage classifier can be cascaded gradually from cheaper to more expensive binary classifiers that are trained using only the necessary data modalities or features required for that stage. UnfoldML is a cost-aware and uncertainty-based dynamic 2D prediction pipeline for multi-stage classification that enables (1) navigation of the accuracy/cost tradeoff space, (2) reducing the spatio-temporal cost of inference by orders of magnitude, and (3) early prediction on proceeding stages. UnfoldML achieves orders of magnitude better cost in clinical settings, while detecting multi- stage disease development in real time. It achieves within 0.1% accuracy from the highest-performing multi-class baseline, while saving close to 20X on spatio- temporal cost of inference and earlier (3.5hrs) disease onset prediction. We also show that UnfoldML generalizes to image classification, where it can predict different level of labels (from coarse to fine) given different level of abstractions of a image, saving close to 5X cost with as little as 0.4% accuracy reduction.
Yanbo Xu, Alind Khare, Glenn Matlin, Monish Ramadoss, Rishikesan Kamaleswaran, Chao Zhang 0014, Alexey Tumanov
NeurIPS2
2021 CompOFA - Compound Once-For-All Networks for Faster Multi-Platform Deployment
Manas Sahni, Shreya Varshini, Alind Khare, Alexey Tumanov
ICLR3
2020 HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units
abstract
Deep learning models have achieved expert-level performance in healthcare with an exclusive focus on training accurate models. However, in many clinical environments such as intensive care unit (ICU), real-time model serving is equally if not more important than accuracy, because in ICU patient care is simultaneously more urgent and more expensive. Clinical decisions and their timeliness, therefore, directly affect both the patient outcome and the cost of care. To make timely decisions, we argue the underlying serving system must be latency-aware. To compound the challenge, health analytic applications often require a combination of models instead of a single model, to better specialize individual models for different targets, multi-modal data, different prediction windows, and potentially personalized predictions. To address these challenges, we propose HOLMES---an online model ensemble serving framework for healthcare applications. HOLMES dynamically identifies the best performing set of models to ensemble for highest accuracy, while also satisfying sub-second latency constraints on end-to-end prediction. We demonstrate that HOLMES is able to navigate the accuracy/latency tradeoff efficiently, compose the ensemble, and serve the model ensemble pipeline, scaling to simultaneously streaming data from 100 patients, each producing waveform data at 250~Hz. HOLMES outperforms the conventional offline batch-processed inference for the same clinical task in terms of accuracy and latency (by order of magnitude). HOLMES is tested on risk prediction task on pediatric cardio ICU data with above 95% prediction accuracy and sub-second latency on 64-bed simulation.
Shenda Hong, Yanbo Xu, Alind Khare, Satria Priambada, Kevin O. Maher, Alaa Aljiffry, Jimeng Sun 0001, Alexey Tumanov
KDD3
2017 Distributed Algorithm for High-Utility Subgraph Pattern Mining Over Big Data Platforms
abstract
Frequent subgraph pattern mining (FSM) finds subgraph patterns that occur in a graph database with a frequency that is more than a given threshold. In FSM, the notion of occurrence captures the presence or absence of a node and an edge in a binary fashion and considers relevance of each edge or node as same. However, an edge or a node may have different relevancy score. Therefore, the utility of a pattern should be defined using the relevance score of participating edges or nodes. This paper defines the utility notion of a pattern using this idea and presents algorithms to mine high-utility patterns from a given graph database. A significant issue in high-utility pattern mining is that the antimonotonic property no longer holds contrary to the FSM. Hence pruning of the search space becomes a daunting task. To address this issue, we incorporate a function to estimate an upper-bound utility of a pattern object that also satisfies the anti-monotonic property. This paper presents three optimization heuristics for the solution on a distributed platform, namely, a novel use of bloom filter to avoid exploration of non-candidates, avoidance of sending database information with each pattern, and avoidance of sending pattern embeddings with each pattern. The experimental study on Apache Spark shows the effectiveness of our proposed optimization strategies.
Alind Khare, Vikram Goyal, Srikanth Baride, Sushil K. Prasad, Michael McDermott, Dhara Shah
HiPC1