Aditya Dhakal

dblp:209/8665 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-8297-8525ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 4 since 2021Computer networks · 3 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Variability-Guided Performance Optimization
abstract
The past few decades have seen software and hardware growing more heterogeneous and layered in abstractions. This trend produced many benefits for hiding complexity and increasing efficiency and modularity. But it also makes reasoning about performance and identifying its underlying factors more challenging because of the presence of performance variability. Moreover, performance variability can prevent synchronous applications from scaling and server applications from meeting service-level agreements.
Eitan Frachtenberg, Viyom Mittal, Mohammed Baydoun, Aditya Dhakal, Izzat El Hajj, Dejan S. Milojicic
ICPE4
2026 Are We There Yet? Predicting if Executing Applications are Near Completion
Mohammad Sonji, Mohammed Baydoun, Safaa Diab, Amir Nassereldine, Pedro Bruel, Aditya Dhakal, Rolando P. Hong Enriquez, Gourav Rattihalli, Diman Zad Tootaghaj, Gallig Renaud, Barbara M. Chapman, Fatima K. Abu Salem, Eitan Frachtenberg, Dejan S. Milojicic, Izzat El Hajj
ICPE6
2024 Opportunistic Energy-Aware Scheduling for Container Orchestration Platforms Using Graph Neural Networks
abstract
Reducing the energy consumption of data centers is critical to meeting international climate goals and lowering operation costs. Container orchestration platforms can help counteract this trend by optimally placing applications across the infrastructure to increase resource utilization and reduce energy consumption. But platforms in use today are still energy-agnostic and do not offer any insights into energy consumption. In this paper, we present a monitoring framework and a new modeling approach for resource usage in data centers. The model captures heterogeneous hardware and software and acts as input for a Graph Neural Network (GNN) to predict power consumption. Based on this model, we derive a set of container scheduling algorithms that opportunistically schedule applications based on the estimated energy impact of incoming containers. Our results show that the GNN-based prediction model is very accurate and achieves an average RMSE (Root Mean Square Error) of 7.5%. We have implemented a custom scheduler to demonstrate the benefits of using our prediction, and our scheduler can decrease energy consumption on average by 6.2% without any code changes for the application and without increasing workload completion time compared to the default Kubernetes scheduler.
Philipp Raith, Gourav Rattihalli, Aditya Dhakal, Sai Rahul Chalamalasetti, Dejan S. Milojicic, Eitan Frachtenberg, Stefan Nastic, Schahram Dustdar
CCGrid3
2024 Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads
Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma 0001, Aleksandar Kuzmanovic
USENIX ATC3
2024 Quantum optimization algorithms: Energetic implications
abstract
Summary Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. In this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least to more energy efficient than their classical counterparts; why feedback‐based quantum optimization is energy‐inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of 20. Finally, we use the NchooseK high‐level programming model to run optimization problems on both gate‐based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate‐based counterparts.
Rolando P. Hong Enriquez, Rosa M. Badia, Barbara M. Chapman, Kirk Bresniker, Scott Pakin, Alok Mishra 0002, Pedro Bruel, Aditya Dhakal, Gourav Rattihalli, Ninad Hogade, Eitan Frachtenberg, Dejan S. Milojicic
Concurr. Comput. Pract. Exp.8
2024 D-STACK: High Throughput DNN Inference by Effective Multiplexing and Spatio-Temporal Scheduling of GPUs
abstract
Hardware accelerators such as GPUs are required for real-time, low latency inference with Deep Neural Networks (DNN). Providing inference services in the cloud can be resource intensive, and effectively utilizing accelerators in the cloud is important. Spatial multiplexing of the GPU, while limiting the GPU resources (GPU%) to each DNN to the right amount, leads to higher GPU utilization and higher inference throughput. Right-sizing the GPU for each DNN the optimal batching of requests to balance throughput and service level objectives (SLOs), and maximizing throughput by appropriately scheduling DNNs are still significant challenges.This article introduces a dynamic and fair spatio-temporal scheduler (D-STACK) for multiple DNNs to run in the GPU concurrently. We develop and validate a model that estimates the parallelism each DNN can utilize and a lightweight optimization formulation to find an efficient batch size for each DNN. Our holistic inference framework provides high throughput while meeting application SLOs. We compare D-STACK with other GPU multiplexing and scheduling methods (e.g., NVIDIA Triton, Clipper, Nexus), using popular DNN models. Our controlled experiments with multiplexing several popular DNN models achieve up to$1.6\times$improvement in GPU utilization and up to$4\times$improvement in inference throughput.
Aditya Dhakal, Sameer G. Kulkarni, K. K. Ramakrishnan
IEEE Trans. Cloud Comput.1
2023 Fine-Grained Heterogeneous Execution Framework with Energy Aware Scheduling
abstract
The growing convergence of high-performance, data analytics, and machine-learning applications is increasingly pushing computing systems toward heterogeneous processors and specialized hardware accelerators. Hardware heterogeneity, in turn, leads to finer-grained workflows. State-of-the-art server-less computing resource managers do not currently provide efficient scheduling of such fine-grained tasks on systems with heterogeneous CPUs and specialized hardware accelerators (e.g., GPUs and FPGAs). Working with fine-grained tasks presents an opportunity for more efficient energy use via new scheduling models. Our proposed scheduler enables technologies like Nvidia's Multi-Process Service (MPS) to pack multiple fine-grained tasks on GPUs efficiently. Its advantages include better co-location of jobs and better sharing of hardware resources such as GPUs that were not previously possible on container orchestration systems. We propose a Kubernetes-native energy-aware scheduler that integrates with our heterogeneous framework. Combining fine-grained resource scheduling on heterogeneous hardware and energy-aware scheduling results in up to 17.6% improvement in makespan, up to 20.16% reduction in energy consumption for CPU workloads, and up to 58.15% improvement in makespan, and up to 28.92% reduction in energy consumption for GPU workloads.
Gourav Rattihalli, Ninad Hogade, Aditya Dhakal, Eitan Frachtenberg, Rolando P. Hong Enriquez, Pedro Bruel, Alok Mishra 0002, Dejan S. Milojicic
CLOUD3
2023 Kernel-as-a-Service: A Serverless Programming Model for Heterogeneous Hardware Accelerators
abstract
With the slowing of Moore's law and decline of Dennard scaling, computing systems increasingly rely on specialized hardware accelerators in addition to general-purpose compute units. Increased hardware heterogeneity necessitates disaggregating applications into workflows of fine-grained tasks that run on a diverse set of CPUs and accelerators. Current accelerator delivery models cannot support such applications efficiently, as (1) the overhead of managing accelerators erases performance benefits for fine-grained tasks; (2) exclusive accelerator use per task leads to underutilization; and (3) specialization increases complexity for developers.
Tobias Pfandzelter, Aditya Dhakal, Eitan Frachtenberg, Sai Rahul Chalamalasetti, Darel Emmot, Ninad Hogade, Rolando P. Hong Enriquez, Gourav Rattihalli, David Bermbach, Dejan S. Milojicic
Middleware2
2022 SLAM-share: visual simultaneous localization and mapping for real-time multi-user augmented reality
abstract
Augmented reality (AR) devices perform visual simultaneous localization and mapping (SLAM) to map the real world and localize themselves in it, enabling them to render the virtual holograms appropriately. Current multi-user AR platforms fall short in that they only allow asymmetric sharing of this SLAM information, resulting in multiple "secondary" devices viewing holograms placed by a single "primary" device, instead of equal participation. The goal of this work is to enable all AR devices to participate equally, by constructing a common global map to which all AR devices can contribute. However, doing so with low latency and high accuracy is challenging on resource-constrained mobile devices. This work proposes an appropriate partitioning between clients and a server to achieve high-throughput, low latency, multi-user SLAM. In our system, SLAM-Share, the edge server performs the complex SLAM computations so that the client devices need only perform lightweight operations. The server utilizes shared memory and efficient map merging to build and update a global map from different clients. It also exploits the parallelism of GPU processing to achieve high-performance tracking. Evaluations show that SLAM-Share is able to achieve significant tracking speedups (up to 50% reduction compared to alternative approaches), maintain good localization accuracy, and merge and update maps within 200 ms.
Aditya Dhakal, Xukan Ran, Yunshu Wang, Jiasi Chen, K. K. Ramakrishnan
CoNEXT1
2022 Slice-Tune: a system for high performance DNN autotuning
abstract
Autotuning DNN models prior to their deployment is an essential but time-consuming task. Using expensive (and power-hungry) GPU and TPU accelerators efficiently is also key. Since DNNs do not always use a GPU fully, spatial multiplexing of multiple models can provide just the right amount of GPU resources for each DNN. We find that a DNN model tuned with the maximum GPU resources has higher inference latency if less GPU resources are available at inference time. We present methods to tune a DNN model, so that we provide the right amount of accelerator resources during tuning. Thus, even when a wide range of GPU resources are available at inference time, the tuned model achieves low inference latency. Further, existing autotuning frameworks take a long time to tune a model due to inefficient utilization of the client and server-side CPU and GPU. Our system, Slice-Tune, improves several autotuning frameworks to efficiently use system resources by re-thinking the partitioning of tasks between the client and server (where models are profiled on the server GPU), in a Kubernetes environment. We increase parallelism during tuning by sharding the tuning model across multiple tuning application instances, providing concurrent tuning of different operators of a model. We also scale server instances to achieve better GPU multiplexing. Slice-Tune reduces DNN autotuning time in a single GPU and in GPU clusters. Slice-Tune decreases DNN autotuning time by up to 75%, and increase autotuning throughput by a factor of 5, across 3 different autotuning frameworks (TVM, Ansor, and Chameleon).
Aditya Dhakal, K. K. Ramakrishnan, Sameer G. Kulkarni, Puneet Sharma 0001, Junguk Cho
Middleware1
2021 Primitives Enhancing GPU Runtime Support for Improved DNN Performance
abstract
Deep neural networks (DNNs) are increasingly used for real-time inference, requiring low latency, but require significant computational power as they continue to increase in complexity. Edge clouds promise to offer lower latency due to their proximity to end users and having powerful accelerators like GPUs to provide the computation power needed for DNNs. But it is also important to ensure that the edge-cloud resources are utilized well. For this, multiplexing several DNN models through spatial sharing of the GPU can substantially improve edge-cloud resource usage. Typical GPU runtime environments have significant interactions with the CPU, to transfer data to the GPU, for CPU-GPU synchronization on inference task completions, etc. These result in overheads. We present a DNN inference framework with a set of software primitives that reduce the overhead for DNN inference, increase GPU utilization and improve performance, with lower latency and higher throughput. Our first primitive uses the GPU DMA effectively, reducing the CPU cycles spent to transfer the data to the GPU. A second primitive uses asynchronous ‘events' for faster task completion notification. GPU runtimes typically preclude fine-grained user control on GPU resources, causing long GPU downtimes when adjusting resources. Our third primitive supports overlapping of model-loading and execution, thus allowing GPU resource re-allocation with very little GPU idle time. Our other primitives increase inference throughput by improving scheduling and processing more requests. Overall, our primitives decrease inference latency by more than 35% and increase DNN throughput by 2-3x.
Aditya Dhakal, Sameer G. Kulkarni, K. K. Ramakrishnan
CLOUD1
2020 GSLICE: controlled spatial sharing of GPUs for a scalable inference platform
abstract
The increasing demand for cloud-based inference services requires the use of Graphics Processing Unit (GPU). It is highly desirable to utilize GPU efficiently by multiplexing different inference tasks on the GPU. Batched processing, CUDA streams and Multi-process-service (MPS) help. However, we find that these are not adequate for achieving scalability by efficiently utilizing GPUs, and do not guarantee predictable performance.
Aditya Dhakal, Sameer G. Kulkarni, K. K. Ramakrishnan
SoCC1
2020 Machine Learning at the Edge: Efficient Utilization of Limited CPU/GPU Resources by Multiplexing
abstract
Edge clouds can provide very responsive services for end-user devices that require more significant compute capabilities than they have. But edge cloud resources such as CPUs and accelerators such as GPUs are limited and must be shared across multiple concurrently running clients. However, multiplexing GPUs across applications is challenging. Further, edge servers are likely to require considerable amounts of streaming data to be processed. Getting that data from the network stream to the GPU can be a bottleneck, limiting the amount of work GPUs do. Finally, the lack of prompt notification of job completion from GPU also results in ineffective GPU utilization. We propose a framework that addresses these challenges in the following manner. We utilize spatial sharing of GPUs to multiplex the GPU more efficiently. While spatial sharing of GPU can increase GPU utilization, the uncontrolled spatial sharing currently available with state-of-the-art systems such as CUDA-MPS can cause interference between applications, resulting in unpredictable latency. Our framework utilizes controlled spatial sharing of GPU, which limits the interference across applications. Our framework uses the GPU DMA engine to offload data transfer to GPU, therefore preventing CPU from being bottleneck while transferring data from the network to GPU. Our framework uses the CUDA event library to have timely, low overhead GPU notifications. Preliminary experiments show that we can achieve low DNN inference latency and improve DNN inference throughput by a factor of ~ 1.4.
Aditya Dhakal, Sameer G. Kulkarni, K. K. Ramakrishnan
ICNP1
2019 NetML: An NFV Platform with Efficient Support for Machine Learning Applications
abstract
Real-time applications such as autonomous and connected cars, surveillance, and online learning applications have to train on streaming data. They require low-latency, high throughput machine learning (ML) functions resident in the network and in the cloud to perform learning and inference. NFV on edge cloud platforms can provide support for these applications by having heterogeneous computing including GPUs and other accelerators to offload ML-related computation. GPUs provide the necessary speedup for performing learning and inference to meet the needs of these latency sensitive real-time applications. Supporting ML inference and learning efficiently for streaming data in NFV platforms has several challenges. In this paper, we present a framework, NetML, that runs existing ML applications on an heterogeneous NFV platform that includes both CPUs and GPUs. NetML efficiently transfers the appropriate packet payload to the GPU, minimizing overheads, avoiding locks, and avoiding CPU-based data copies. Additionally, NetML minimizes latency by maximizing overlap between the data movement and GPU computation. We evaluate the efficiency of our approach for training and inference using popular object detection algorithms on our platform. NetML reduces the latency for inferring images by more than 20% and increases the training throughput by 30% while reducing CPU utilization compared to other state-of-the-art alternatives.
Aditya Dhakal, K. K. Ramakrishnan
NetSoft1
2017 Machine learning at the network edge for automated home intrusion monitoring
abstract
Monitoring of residences and businesses can be effectively performed using machine learning algorithms. As sensors and devices used for monitoring become more complex, having humans process the information to detect intrusions would be expensive and difficult to scale. We propose an automated home/business monitoring system which resides on edge servers performing online learning on streaming data coming from homes and businesses in the neighborhood. The edge servers run Open-NetVM, a Network Function Virtualization (NFV) platform, and host multiple machine learning applications instantiated on demand. This enables us to serve a set of customers in the neighborhood on a timely basis, permitting customization and learning of the behavior of each home. We combine the results of the multiple classifiers, with each classifier examining a distinct feature related to a distinct sensor, to finally infer whether the entry is a normal one or an intrusion. Our results show that our system is able to classify intrusions better than basing the decision on a single classifier, thus reducing false alarms. We have also shown that our system can effectively scale and monitor thousands of homes.
Aditya Dhakal, K. K. Ramakrishnan
ICNP1