EDBT 2026 Demo / reviewers in the wild / expert
Aakash Sharma
dblp:173/0802
· DBLP profile ↗
11ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dally: A Network-Placement Sensitive Cluster Scheduler for Deep Learning
Aakash Sharma, Vivek M. Bhasi, Sonali Singh, Mahmut T. Kandemir, George Kesidis, Chita R. Das |
IEEE Big Data | 1 |
| 2024 | Paldia: Enabling SLO-Compliant and Cost-Effective Serverless Computing on Heterogeneous HardwareabstractAmong the variety of applications (apps) being deployed on serverless platforms, apps such as Machine Learning (ML) inference serving can achieve better performance from leveraging accelerators like GPUs. Yet, major serverless providers, despite having GPU-equipped servers, do not offer GPU support for their serverless functions. Given that serverless functions are deployed on various generations of CPUs already, extending this to various (typically more expensive) GPU generations can offer providers a greater range of hardware to serve incoming requests according to the functions and request traffic. Here, providers are faced with the challenge of selecting hardware to reach a well-proportioned trade-off point between cost and performance. While recent works have attempted to address this, they often fail to do so as they overlook optimization opportunities arising from intelligently leveraging existing GPU sharing mechanisms. To address this point, we devise a heterogeneous serverless framework, PALDIA, which uses a prudent Hardware selection policy to acquire capable, cost-effective hardware and perform intelligent request scheduling on it to yield high performance and cost savings. Specifically, our scheduling algorithm employs hybrid spatio-temporal GPU sharing that intelligently trades off job queueing delays and interference to allow the chosen cost-effective hardware to also be highly performant. We extensively evaluate PALDIA using 16 ML inference workloads with real-world traces on a 6 node heterogeneous cluster. Our results show that PALDIA significantly outperforms state-of-the-art works in terms of Service Level Objective (SLO) compliance (up to 13.3% more) and tail latency (up to ∼50% less), with cost savings up to 86%. Vivek M. Bhasi, Aakash Sharma, Shruti Mohanty, Mahmut T. Kandemir, Chita R. Das |
IPDPS | 2 |
| 2024 | Towards SLO-Compliant and Cost-Effective Serverless Computing on Emerging GPU ArchitecturesabstractServerless platforms are supporting an increasing variety of applications (apps). Among these, apps such as Machine Learning (ML) inference serving can benefit significantly from leveraging accelerators like GPUs. Yet, major serverless providers, despite having GPU-equipped servers, do not offer GPU support for their serverless functions. While recent works have attempted to bridge this gap, they are agnostic to the capabilities of new-generation GPUs, thereby, overlooking several performance optimization opportunities. Vivek M. Bhasi, Aakash Sharma, Jashwant Raj Gunasekaran, Ashutosh Pattnaik, Mahmut T. Kandemir, Chita R. Das |
Middleware | 2 |
| 2023 | Stash: A Comprehensive Stall-Centric Characterization of Public Cloud VMs for Distributed Deep LearningabstractDeep neural networks (DNNs) are increasingly popular owing to their ability to solve complex problems such as image recognition, autonomous driving, and natural language processing. Their growing complexity coupled with the use of larger volumes of training data (to achieve acceptable accuracy) has warranted the use of GPUs and other accelerators. Such accelerators are typically expensive, with users having to pay a high upfront cost to acquire them. For infrequent use, users can, instead, leverage the public cloud to mitigate the high acquisition cost. However, with the wide diversity of hardware instances (particularly GPU instances) available in public cloud, it becomes challenging for a user to make an appropriate choice from a cost/performance standpoint. In this work, we try to address this problem by (i) introducing a comprehensive distributed deep learning (DDL) profiler Stash, which determines the various execution stalls that DDL suffers from, and (ii) using Stash to extensively characterize various public cloud GPU instances by running popular DNN models on them. Specifically, it estimates two types of communication stalls, namely, interconnect and network stalls, that play a dominant role in DDL execution time. Stash is implemented on top of prior work, DS-analyzer, that computes only the CPU and disk stalls. Using our detailed stall characterization, we list the advantages and shortcomings of public cloud GPU instances for users to help them make an informed decision(s). Our characterization results indicate that the more expensive GPU instances may not be the most performant for all DNN models and that AWS can sometimes sub-optimally allocate hardware interconnect resources. Specifically, the intra-machine interconnect can introduce communication overheads of up to 90% of DNN training time and the network-connected instances can suffer from up to 5× slowdown compared to training on a single instance. Furthermore, (iii) we also model the impact of DNN macroscopic features such as the number of layers and the number of gradients on communication stalls, and finally, (iv) we briefly discuss a cost comparison with existing work. Aakash Sharma, Vivek M. Bhasi, Sonali Singh, Jashwant Raj Gunasekaran, Subrata Mitra, Mahmut T. Kandemir, George Kesidis, Chita R. Das |
ICDCS | 1 |
| 2023 | Capturing Nutrition Data for Sports: Challenges and Ethical Issues
Aakash Sharma, Katja Pauline Czerwinska, Dag Johansen, Håvard D. Johansen |
MMM (1) | 1 |
| 2022 | Cypress: input size-sensitive container provisioning and request scheduling for serverless platformsabstractThe growing popularity of the serverless platform has seen an increase in the number and variety of applications (apps) being deployed on it. The majority of these apps process user-provided input to produce the desired results. Existing work in the area of input-sensitive profiling has empirically shown that many such apps have input size-dependent execution times which can be determined through modelling techniques. Nevertheless, existing serverless resource management frameworks are agnostic to the input size-sensitive nature of these apps. We demonstrate in this paper that this can potentially lead to container over-provisioning and/or end-to-end Service Level Objective (SLO) violations. To address this, we propose Cypress, an input size-sensitive resource management framework, that minimizes the containers provisioned for apps, while ensuring a high degree of SLO compliance. We perform an extensive evaluation of Cypress on top of a Kubernetes-managed cluster using 5 apps from the AWS Serverless Application Repository and/or Open-FaaS Function Store with real-world traces and varied input size distributions. Our experimental results show that Cypress spawns up to 66% fewer containers, thereby, improving container utilization and saving cluster-wide energy by up to 2.95X and 23%, respectively, versus state-of-the-art frameworks, while remaining highly SLO-compliant (up to 99.99%). Vivek M. Bhasi, Jashwant Raj Gunasekaran, Aakash Sharma, Mahmut T. Kandemir, Chita R. Das |
SoCC | 3 |
| 2021 | CASH: A Credit Aware Scheduling for Public Cloud PlatformsabstractDistributed data processing frameworks such as Hadoop, Tez, Spark, and Flink are exclusively used by public cloud tenants for executing large scale data analytics applications in various domains including but not limited to content management, financial sector, healthcare etc. These frameworks slice a job into a number of smaller tasks, which are then executed by a job scheduler on a multi-node compute cluster. While making scheduling decisions, the State-of-art schedulers employed in these frameworks assume hardware resources such as CPU, disk I/O and network I/O to offer a fixed service rate. However, in a public cloud environment, many of these resources are associated with burstable service rates. More specifically, the resources offer a guaranteed baseline service rate with an option to burst above their baseline rate by expending accumulated burst credits. Being unaware about this underlying hardware burstability, schedulers tend to make sub-optimal task placement decisions, thereby adversely affecting the job completion times, leading to higher deployment costs.In this paper, we propose CASH, a burst credit aware scheduler, which is cognizant about the burst credits associated with the individual hardware resources in the public cloud cluster. Through coarse grained task annotations depicting the burst credit demand of individual tasks and dynamically monitoring the credits for the underlying resources, CASH performs optimal task placement decisions. We prototype CASH on YARN, Hadoop, and Tez, and extensively evaluate it using both batch and streaming workloads. Our experimental results with CASH show CPU-credit based instances, like AWS T3, are a viable cost effective alternative when compared to self-managed offerings like Amazon EMR, for running large scale batch workloads. Furthermore, we demonstrate that CASH can accelerate streaming SQL queries on a large Hive database by up to 39.4% , leading to public cloud cost savings by up to 22%. Aakash Sharma, Saravanan Dhakshinamurthy, George Kesidis, Chita R. Das |
CCGRID | 1 |
| 2021 | Designing a Service for Compliant Sharing of Sensitive Research Data
Aakash Sharma, Thomas Bye Nilsen, Sivert Johansen, Dag Johansen, Håvard D. Johansen |
CRiSIS | 1 |
| 2017 | Dynamic Map Update Protocol for Highly Automated Driving Vehicles
Florian Jomrich, Aakash Sharma, Tobias Rueckelt, Daniel Burgstahler, Doreen Böhnstedt |
VEHITS | 2 |
| 2016 | Extracting PICO Sentences from Clinical Trial Reports using Supervised Distant SupervisionabstractSystematic reviews underpin Evidence Based Medicine (EBM) by addressing precise clinical questions via comprehensive synthesis of all relevant published evidence. Authors of systematic reviews typically define a Population/Problem, Intervention, Comparator, and Outcome (a PICO criteria) of interest, and then retrieve, appraise and synthesize results from all reports of clinical trials that meet these criteria. Identifying PICO elements in the full-texts of trial reports is thus a critical yet time-consuming step in the systematic review process. We seek to expedite evidence synthesis by developing machine learning models to automatically extract sentences from articles relevant to PICO elements. Collecting a large corpus of training data for this task would be prohibitively expensive. Therefore, we derive distant supervision (DS) with which to train models using previously conducted reviews. DS entails heuristically deriving 'soft' labels from an available structured resource. However, we have access only to unstructured, free-text summaries of PICO elements for corresponding articles; we must derive from these the desired sentence-level annotations. To this end, we propose a novel method -- supervised distant supervision (SDS) -- that uses a small amount of direct supervision to better exploit a large corpus of distantly labeled instances by learning to pseudo-annotate articles using the available DS. We show that this approach tends to outperform existing methods with respect to automated PICO extraction. Byron C. Wallace, Joël Kuiper, Aakash Sharma, Mingxi (Brian) Zhu, Iain James Marshall |
J. Mach. Learn. Res. | 3 |
| 2015 | Fuzzy match index for scale-invariant feature transform (SIFT) features with application to face recognition with weak supervisionabstractA fuzzy match index for scale‐invariant feature transform (SIFT) features is proposed in this study that cumulatively involves all the test SIFT keypoints in the decision‐making process. The new fuzzy SIFT classifier is adapted successfully for robust face recognition from complex backgrounds without using any face cropping tools and using only a single training template. The further incorporation of entropy weights ensures that the facial features have a greater role in the soft decision‐making as compared with the background features. The highlights of the authors’ work are: (i) The development of a novel highly efficient fuzzy SIFT descriptor matching tool; (ii) incorporation of feature entropy weights to highlight the contribution of facial features; (iii) application to robust face recognition from uncropped images having diverse backgrounds with a single template for each subject. The authors thus allow for weak supervision of the face recognition experiment and obtain high accuracy for 20 subjects of the CALTECH‐256 face database, 133 subjects of the labelled faces for the wild dataset and 994 subjects of the FERET database, with state‐of‐the‐art comparisons indicating the supremacy of the authors’ approach. Seba Susan, Aakash Sharma, Shikhar Verma, Siddhant Jain |
IET Image Process. | 3 |