EDBT 2026 Demo / reviewers in the wild / expert
Mehdi Ghasemi 0003
dblp:30/10826-3
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-3947-5639ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uncertainty-Aware RL-Based Scheduling of Multi-DNN Workloads on Edge MPSoCsabstractEmerging Machine Learning (ML) workloads, particularly those deployed at the edge, increasingly rely on a network of Deep Neural Networks (DNNs), where each model is tailored to a specific task. The primary challenge is efficiently utilizing all available resources simultaneously on heterogeneous multi-processor system-on-chips (MPSoCs) when executing these computationally demanding multi-DNN workloads. This becomes even more complicated because inference behavior can vary with batch size, input content, and runtime dynamics. This paper presents a novel framework to schedule the execution of multi-DNN workloads on a heterogeneous MPSoC with the goal of minimizing the latency of the workload execution. The proposed approach uses a reinforcement learning (RL) framework to handle three major sources of variability, including differences in DNN execution and batching due to hardware heterogeneity, workload fluctuations driven by input content, and runtime execution unpredictability caused by system-level constraints. The RL method integrates a graph neural network (GNN) and a pointer-based policy network to adapt to stochastic execution behavior. GNN-based embeddings enable scalable and efficient scheduling decisions across diverse workload conditions. The proposed approach achieves up to 6.2× and 2.0× performance improvement over CAMDNN and HEFT, respectively, when evaluated on a Qualcomm RB5 development kit. Soroush Heidari, Mehdi Ghasemi 0003, Sarma B. K. Vrudhula |
SEC | 2 |
| 2025 | Energy-Efficient, Delay-Constrained Edge Computing of a Network of DNNsabstractThis paper presents a novel approach for executing the inference of a network of pre-trained deep neural networks (DNNs) on commercial-off-the-shelf devices that are deployed at the edge. The problem is to partition the computation of the DNNs between an energy-constrained and performance-limited edge device$\boldsymbol{\mathcal{E}}$, and an energy-unconstrained, higher performance device$\boldsymbol{\mathcal{C}}$, referred to as thecloudlet, with the objective of minimizing the energy consumption of$\boldsymbol{\mathcal{E}}$subject to a deadline constraint. The proposed partitioning algorithm takes into account the performance profiles of executing DNNs on the devices, the power consumption profiles, and the variability in the delay of the wireless channel. The algorithm is demonstrated on a platform that consists of an NVIDIA Jetson Nano as the edge device$\boldsymbol{\mathcal{E}}$and a Dell workstation with a Titan Xp GPU as the cloudlet. Experimental results show significant improvements both in terms of energy consumption of$\boldsymbol{\mathcal{E}}$and processing delay of the application. Additionally, it is shown how the energy-optimal solution is changed when the deadline constraint is altered. Moreover, the overhead of decision-making for our proposed method is significantly lower than the state-of-the-art Integer Linear Programming (ILP) solutions. Mehdi Ghasemi 0003, Soroush Heidari, Younggeun Kim 0001, Carole-Jean Wu, Sarma B. K. Vrudhula |
IEEE Trans. Computers | 1 |
| 2024 | Elastic Execution of Multi-Tenant DNNs on Heterogeneous Edge MPSoCsabstractThe growing complexity of machine learning (ML) tasks drives the rapid deployment of multi-tenant ML workloads at the edge presenting unique challenges due to the variable computational demands and strict latency requirements. This paper introduces a holistic elastic scheduler, EMERALD, designed to optimize the execution of multi-tenant machine learning (ML) workloads on heterogeneous edge (Multiprocessor System on Chip) MPSoCs under strict runtime constraints. EMERALD employs input resolution scaling to dynamically adjust the computational demands of deep neural networks (DNNs), thereby enhancing the ability to meet stringent latency requirements while maintaining high accuracy. The scheduler consists of two main components: a local greedy scheduler and a global scheduler. The local scheduler actively manipulates input resolution in response to deadline violations, selecting the resolutions that minimally impact accuracy and maximally reduce response time. The global scheduler, an Integer Linear Programming (ILP)based scheduler, fine-tunes the decisions of the local scheduler by considering factors such as DNN dependencies, scene complexity, hardware heterogeneity, and the trade-offs between accuracy and makespan associated with input scaling adjustments. This hierarchical approach allows EMERALD to effectively balance computational efficiency and accuracy, significantly reducing missed deadlines—achieving 11x and 12.3x fewer missed deadlines compared to CAMDNN and HEFT, respectively, in scenarios demanding 30 frames per second. The results underscore the critical role of adaptive input scaling in managing the complexities of edge-based ML deployments. Soroush Heidari, Mehdi Ghasemi 0003, Younggeun Kim 0001, Carole-Jean Wu, Sarma B. K. Vrudhula |
SEC | 2 |
| 2022 | CAMDNN: Content-Aware Mapping of a Network of Deep Neural Networks on Edge MPSoCsabstractMachine Learning (ML) workloads are increasingly deployed at the edge. Enabling efficient inference execution while considering model and system heterogeneity remains challenging, especially for ML tasks built with a network of DNNs. The challenge is to maximize the utilization of all available resources on the multiprocessor system on a chip (MPSoC) at the same time. This becomes even more complicated because the optimal mapping for the network of DNNs can vary with input batch sizes and scene complexity. In this paper, a holistic hierarchical scheduling framework is presented to optimize the execution time for a network of DNN models on an edge MPSoC at runtime, considering varying input characteristics. The framework consists of a local and a global scheduler. The local scheduler maps individual DNNs in the inference pipeline to the best-performing hardware unit while the global scheduler customizes an Integer Linear Programming (ILP) solution to instantiate DNN remapping. To minimize scheduler runtime overhead, an imitation learning (IL) based scheduler is used that approximates the ILP solutions. The proposed scheduling framework (CAMDNN) was implemented on a Qualcomm Robotic RB5 platform. CAMDNN resulted in lower execution time of up to 32% than HEFT, and by factors of 6.67X, 5.6X and 2.17X than the CPU-only, GPU-only and Central Queue schedulers. Soroush Heidari, Mehdi Ghasemi 0003, Younggeun Kim 0001, Carole-Jean Wu, Sarma B. K. Vrudhula |
IEEE Trans. Computers | 2 |
| 2022 | EdgeWise: Energy-efficient CNN Computation on Edge Devices under Stochastic Communication DelaysabstractThis article presents a framework to enable the energy-efficient execution of convolutional neural networks (CNNs) on edge devices. The framework consists of a pair of edge devices connected via a wireless network: a performance and energy-constrained deviceDas the first recipient of data and an energy-unconstrained deviceNas an accelerator forD. DeviceDdecides on-the-fly how to distribute the workload with the objective of minimizing its energy consumption while accounting for the inherent uncertainty in network delay and the overheads involved in data transfer. These challenges are tackled by adopting the data-driven modeling framework of Markov Decision Processes, whereby an optimal policy is consulted byDinO(1) time to make layer-by-layer assignment decisions. As a special case, a linear-time dynamic programming algorithm is also presented for finding optimal layer assignment at once, under the assumption that the network delay is constant throughout the execution of the application. The proposed framework is demonstrated on a platform comprised of a Raspberry PI 3 asDand an NVIDIA Jetson TX2 asN. An average improvement of 31% and 23% in energy consumption is achieved compared to the alternatives of executing the CNNs entirely onDandN. Two state-of-the-art methods were also implemented and compared with the proposed methods. Mehdi Ghasemi 0003, Daler N. Rakhmatov, Carole-Jean Wu, Sarma B. K. Vrudhula |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2021 | Energy-Efficient Mapping for a Network of DNN Models at the EdgeabstractThis paper describes a novel framework for executing a network of trained deep neural network (DNN) models on commercial-off-the-shelf devices that are deployed in an IoT environment. The scenario consists of two devices connected by a wireless network: a user-end device (U), which is a low-end, energy and performance-limited processor, and a cloudlet (C), which is a substantially higher performance and energy-unconstrained processor. The goal is to distribute the computation of the DNN models between U and C to minimize the energy consumption of U while taking into account the variability in the wireless channel delay and the performance overhead of executing models in parallel. The proposed framework was implemented using an NVIDIA Jetson Nano for U and a Dell workstation with Titan Xp GPU as C. Experiments demonstrate significant improvements both in terms of energy consumption of U and processing delay. Mehdi Ghasemi 0003, Soroush Heidari, Younggeun Kim 0001, Aaron Lamb, Carole-Jean Wu, Sarma B. K. Vrudhula |
SMARTCOMP | 1 |