EDBT 2026 Demo / reviewers in the wild / expert
Seyed Morteza Nabavinejad
dblp:175/7611
· DBLP profile ↗
12ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0002-5123-6318ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CarbonDIS: Carbon-Aware DNN Inference Scheduling on Heterogeneous GPUsabstractDeploying deep neural network (DNN) inference applications in data centers and clouds to empower various services is on the increase. The significant power consumption of these applications contributes to the carbon emissions of underlying infrastructures. Therefore, minimizing the carbon emissions of these applications leads to a reduced carbon footprint of data centers and clouds. This paper introduces CarbonDIS, a carbon-aware inference scheduler designed to maintain the latency of DNN inference applications while minimizing their carbon emissions by leveraging heterogeneous GPUs. CarbonDIS considers an inference architecture where high-end and low-end GPUs are used in tandem to serve inference jobs. By leveraging the varying computing capability and power consumption of heterogeneous GPUs, CarbonDIS can effectively balance the performance and carbon emissions. Evaluation using three types of GPUs, 10 DNN models, and real-world carbon intensity traces show that CarbonDIS can find a trade-off between carbon footprint and latency of the jobs and reduce the carbon emissions by 16% compared with a performance-centric approach. Seyed Morteza Nabavinejad, Shubbhi Taneja, Tian Guo 0001 |
CCGrid | 1 |
| 2024 | MediatorDNN: Contention Mitigation for Co-Located DNN Inference JobsabstractWith the increase in computing power of cutting-edge hardware platforms, it is a common practice to run multiple jobs on a single machine for improved resource utilization and throughput. However, this leads to inevitable resource contention among co-located jobs, impacting their performance. The resource contention can worsen due to fluctuations in resource utilization of jobs caused by variations in their input workload. To tackle the co-location contention for DNN inference jobs, we propose MediatorDNN, which considers contention and resource utilization variation when co-locating DNN inference jobs. It profiles each DNN, monitors microarchitectural metrics such as memory bandwidth and cache access pattern, and high-level resource utilization like CPU utilization. Based on profiling results and leveraging Modern Portfolio Theory (MPT), MediatorDNN decides on the co-location of jobs. Experimental results with various DNNs on two hardware platforms show that MediatorDNN improves throughput by up to 108% (21 % on average) compared to an approach only considering contention and ignoring resource utilization variation. Seyed Morteza Nabavinejad, Sherief Reda, Tian Guo 0001 |
CLOUD | 1 |
| 2023 | Mobility-Aware Computation Offloading in Edge Computing Using Machine LearningabstractCloudlets are resource-rich computing infrastructures of edge computing that are located at physical proximity of users to provide one-hop, high-bandwidth wireless access to additional computational resources. They enable computation offloading for user applications, which compensates for the resource limitation of user devices by providing ultra-low latency processing for their applications. Although the computation capability of user devices is dramatically augmented by offloading, spatio-temporal uncertainties due to user mobility and changes in application specifications bring the most challenging obstacles in deciding where to offload to provide minimum latency. In this paper, we focus on these challenges by designing efficient offloading approaches that take into account these uncertainties and dynamics in order to minimize the turnaround time of the applications, which is constituted by offloading latency, migration delay, and execution time. We first formulate this NP-hard problem as an integer programming model to obtain optimal offloading decisions. We tackle its intractability by designing two novel offloading approaches, called S-OAMC and G-OAMC, that fully assign applications to cloudlets by considering their expected future locations and specifications predicted by Matrix Completion, a machine learning method. S-OAMC is a sampling-based approximation dynamic programming approach that enhances scalability and obtains near-optimal solutions. G-OAMC is a fast greedy-based approach for finding low-turnaround time offloading decisions. We conduct extensive experiments to assess the performance of our proposed approaches. The results show that S-OAMC and G-OAMC lead to near-optimal turnaround time in a reasonable time, and they both obtain low migration rates. Erfan Farhangi Maleki, Lena Mashayekhy, Seyed Morteza Nabavinejad |
IEEE Trans. Mob. Comput. | 3 |
| 2022 | Inference Time Reduction of Deep Neural Networks on Embedded Devices: A Case StudyabstractFrom object detection to semantic segmentation, deep learning has achieved many groundbreaking results in recent years. However, due to the increasing complexity, the execution of neural networks on embedded platforms is greatly hindered. This has motivated the development of several neural network minimisation techniques, amongst which pruning has gained a lot of focus. In this work, we perform a case study on a series of methods with the goal of finding a small model that could run fast on embedded devices. First, we suggest a simple, but effective, ranking criterion for filter pruning called Mean Weight. Then, we combine this new criterion with a threshold-aware layer-sensitive filter pruning method, called T-sensitive pruning, to gain high accuracy. Further, the pruning algorithm follows a structured filter pruning approach that removes all selected filters and their dependencies from the DNN model, leading to less computations, and thus low inference time in lower-end CPUs. To validate the effectiveness of the proposed method, we perform experiments on three different datasets (with 3, 101, and 1000 classes) and two different deep neural networks (i.e., SICK-Net and MobileNet V1). We have obtained speedups of up to 13x on lower-end CPUs (Armv8) with less than 1% drop in accuracy. This satisfies the goal of transferring deep neural networks to embedded hardware while attaining a good trade-off between inference time and accuracy. Isma-Ilou Sadou, Seyed Morteza Nabavinejad, Zhonghai Lu, Masoumeh Ebrahimi |
DSD | 2 |
| 2022 | Coordinated Batching and DVFS for DNN Inference on GPU AcceleratorsabstractEmploying hardware accelerators to improve the performance and energy-efficiency of DNN applications is on the rise. One challenge of using hardware accelerators, including the GPU-based ones, is that their performance is limited by internal and external factors, such as power caps. A common approach to meet the power cap constraint is using the Dynamic Voltage Frequency Scaling (DVFS) technique. However, the functionally of this technique is limited and platform-dependent. To tackle this challenge, we propose a new control knob, which is the size of input batches fed to the GPU accelerator in DNN inference applications. We first evaluate the impact of batch size on power consumption and performance of DNN inference. Then, we introduce the design and implementation of a fast and lightweight runtime system, called BatchDVFS. Dynamic batching is implemented inBatchDVFSto adaptively change the batch size, and hence, trade-off throughput with power consumption. It employs an approach based on binary search to find the proper batch size within a short period of time. Combining dynamic batching with the DVFS technique,BatchDVFScan control the power consumption in wider ranges, and hence, yield higher throughput in the presence of power caps. To find near-optimal solution for long-running jobs that can afford a relatively significant profiling overhead, compared withBatchDVFSoverhead, we also design an approach, called BOBD, that employs Bayesian Optimization to wisely explore the vast state space resulted by combination of the batch size and DVFS solutions. Conducting several experiments using a modern GPU and several DNN models and input datasets, we show that ourBatchDVFScan significantly surpass the techniques solely based on DVFS or batching, regarding throughput (up to 11.2x and 2.2x, respectively), while successfully meeting the power cap. Seyed Morteza Nabavinejad, Sherief Reda, Masoumeh Ebrahimi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | BatchSizer: Power-Performance Trade-off for DNN InferenceabstractGPU accelerators can deliver significant improvement for DNN processing; however, their performance is limited by internal and external parameters. A well-known parameter that restricts the performance of various computing platforms in real-world setups, including GPU accelerators, is the power cap imposed usually by an external power controller. A common approach to meet the power cap constraint is using the Dynamic Voltage Frequency Scaling (DVFS) technique. However, the functionally of this technique is limited and platform-dependent. To improve the performance of DNN inference on GPU accelerators, we propose a new control knob, which is the size of input batches fed to the GPU accelerator in DNN inference applications. After evaluating the impact of this control knob on power consumption and performance of GPU accelerators and DNN inference applications, we introduce the design and implementation of a fast and lightweight runtime system, called BatchSizer. This runtime system leverages the new control knob for managing the power consumption of GPU accelerators in the presence of the power cap. Conducting several experiments using a modern GPU and several DNN models and input datasets, we show that our BatchSizer can significantly surpass the conventional DVFS technique regarding performance (up to 29%), while successfully meeting the power cap. Seyed Morteza Nabavinejad, Sherief Reda, Masoumeh Ebrahimi |
ASP-DAC | 1 |
| 2021 | Profit Maximization of Big Data Jobs in Cloud Using Stochastic OptimizationabstractReserved instances offered by cloud providers make it possible to reserve resources and computing capacity for a specific period of time. One should pay for all the hours of that time interval; in exchange, the hourly rate is significantly lower than on-demand instances. Reserved Instances can significantly reduce the monetary cost of resources needed to process big data applications in cloud. However, purchases of these instances are non-refundable, and hence, one should be able to estimate the required resources prior to purchase to avoid over-payment. It becomes important especially when the results obtained by big data job has monetary value, such as business intelligence applications. But, estimating the resource demand of big data processing jobs is hard because of numerous factors that affect them such as data locality, data skew, stragglers, internal settings of big data processing framework, interference among instances, instances availability, etc. To maximize the profit of processing such big data jobs in cloud considering fluctuating nature of their resource demand, as well as reserved instances limitations, we propose Reserved Instances Stochastic Allocation (RISA) approach. Using historical traces of resource demand of big data jobs submitted by user, RISA leverages stochastic optimization to determine the amount of resources needed to be reserved for that user to maximize the profit. Our evaluation using real-world traces shows that RISA can increase the net profit by up to 10x, compared to previous approaches. RISA can also find solutions as close as 2 percent to the best possible solution. Seyed Morteza Nabavinejad, Maziar Goudarzi |
IEEE Trans. Cloud Comput. | 1 |
| 2020 | ApproxDNN: Incentivizing DNN Approximation in CloudabstractService providers leverage discounted prices of reserved instances offered by cloud providers to amortize their operational costs. They reserve a certain number of instances to cover a significant portion of their computing resource requirements, and further employ on-demand instances to cover remaining requirements not satisfied by the reserved instances. Because of the higher price of on-demand instances, service providers seek to lower their usage to minimize operational costs. In this work, we propose ApproxDNN approach for Machine Learning as a Service to reduce operational costs of service providers by incentivizing approximate results, based on the capabilities of cutting-edge GPUs and a discounted pricing model. When the deadlines of jobs submitted by users are very tight, a service provider might not be able to execute all of them on reserved instances under the default precision. In such cases, Ap- proxDNN leverages the reduced-precision instructions to reduce the execution time of the jobs with slight reduction in their final accuracy, and consequently, to minimize the employment of on- demand instances. To incentivize users to accept the approximate results of reduced-precision instructions, ApproxDNN offers them a discounted price for the service based on a newly designed pricing model. Our proposed pricing model of ApproxDNN guarantees lower or equal cost for service providers compared to the conventional method that solely depends on employment of on-demand instances in case of the reserved instance shortage. We employ real-world traces to conduct an extensive set of experiments and evaluate the performance of our proposed approach. The results show that ApproxDNN reduces the cost of service providers by 18%, while never exceeding the cost of the conventional method and slightly affecting the accuracy by 0.14%. Seyed Morteza Nabavinejad, Lena Mashayekhy, Sherief Reda |
CCGRID | 1 |
| 2019 | Faster MapReduce Computation on Clouds Through Better Performance EstimationabstractProcessing Big Data in cloud is on the increase. An important issue for efficient execution of Big Data processing jobs on a cloud platform is selecting the best fitting virtual machine (VM) configuration(s) among the miscellany of choices that cloud providers offer. Wise selection of VM configurations can lead to better performance, cost and energy consumption. Therefore, it is crucial to explore the available configurations and opt for the best ones that well suit each MapReduce application. Profiling the given application on all the configurations is costly, time and energy consuming. An alternative is to run the application on a subset of configurations (sample configurations) and estimate its performance on other configurations based on the obtained values by sample configurations. We show that the choice of these sample configurations highly affects accuracy of later estimations. Our Smart Configuration Selection (SCS) scheme chooses better representatives from among all configurations by once-off analysis of given performance figures of the benchmarks so as to increase the accuracy of estimations of missing values, and consequently, to more accurately choose the configuration providing the highest performance. The results show that the SCS choice of sample configurations is very close to the best choice, and can reduce estimation error to 11.58 percent from the original 19.72 percent of random configuration selection. More importantly, using SCS estimations in a makespan minimization algorithm improves the execution time by up to 36.03 percent compared with random sample selection. Seyed Morteza Nabavinejad, Maziar Goudarzi |
IEEE Trans. Cloud Comput. | 1 |
| 2018 | QoR-aware power capping for approximate big data processingabstractTo limit the peak power consumption of a cluster, a centralized power capping system typically assigns power caps to the individual servers, which are then enforced using local capping controllers. Consequently, the performance and throughput of the servers are affected, and the runtime of jobs is extended as a result. We observe that servers in big data processing clusters often execute big data applications that have different tolerance for approximate results. To mitigate the impact of power capping, we propose a new power-Capping aware resource manager for Approximate Big data processing (CAB) that takes into consideration the minimum Quality-of-Result (QoR) of the jobs. We use industry-standard feedback power capping controllers to enforce a power cap quickly, while, simultaneously modifying the resource allocations to various jobs based on their progress rate, target minimum QoR, and the power cap such that the impact of capping on runtime is minimized. Based on the applied cap and the progress rates of jobs, CAB dynamically allocates the computing resources (i.e., number of cores and memory) to the jobs to mitigate the impact of capping on the finish time. We implement CAB in Hadoop-2.7.3 and evaluate its improvement over other methods on a state-of-the-art 28-core Xeon server. We demonstrate that CAB minimizes the impact of power capping on runtime by up to 39.4% while meeting the minimum QoR constraints. Seyed Morteza Nabavinejad, Xin Zhan, Maziar Goudarzi, Sherief Reda |
DATE | 1 |
| 2016 | Energy efficiency in cloud-based MapReduce applications through better performance estimation
Seyed Morteza Nabavinejad, Maziar Goudarzi |
DATE | 1 |
| 2016 | The Memory Challenge in Reduce Phase of MapReduce ApplicationsabstractMapReduce has become a popular paradigm for Big Data processing. Each MapReduce Application has two phases: Map and Reduce. Each phase consist of several tasks in a defaulted sequence of processes. It is common place to determine the number of Map tasks equal to the number of data blocks in the input data. However, there is no specific rule for determining the number of Reduce tasks based on the amount of intermediate data generated by Map tasks or the specifications of machines that execute the tasks. Since the Reduce tasks bring the data into memory for processing, this may lead to inefficient execution of application and even application failure because of memory shortage or temporary consumption. In this work, we first evaluate this challenge and show its problematic significance. To address this challenge, we propose a Mnemonic approach. Mnemonic leverages a profiling mechanism to detect the application behavior regarding intermediate data generation. It first decides the amount of memory to be dedicated to each Reduce slot. Then it determines the number of Reduce tasks based on the gathered information through profiling and the decided size of memory for Reduce slots. Experimental results using PUMA benchmark suit indicates that our proposed memory-aware approach can 1) completely remove the likelihood of application failure due to out of memory error and 2) decrease the execution time of Reduce phase up to 58.27, 79.36, and 88.79 percent compared with Memory Oblivious, Fine Grain 1, and Fine Grain 2 approaches, respectively. Seyed Morteza Nabavinejad, Maziar Goudarzi, Shirin Mozaffari |
IEEE Trans. Big Data | 1 |