VLDB 2026 Research / reviewers in the wild / expert
Mohanad Odema
dblp:216/1646
· DBLP profile ↗
14ranked-venue papers
9as first author
14since 2021 · last 2025
0000-0002-0828-949XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 8 first-author · 12 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving PerceptionabstractWe study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems from how chiplets technology is becoming integral to emerging vehicular architectures, providing a cost-effective tradeoff between performance, modularity, and customization; and from perception models being the most computationally demanding workloads in a autonomous driving system. Using the Tesla Autopilot perception pipeline as a case study, we first breakdown its constituent models and profile their performance on different chiplet accelerators. From the insights, we propose a novel scheduling strategy to efficiently deploy perception workloads on multi-chip AI accelerators. Our experiments using a standard DNN performance simulator, MAESTRO, show our approach realizes 82% and 2.8 × increase in throughput and processing engines utilization compared to monolithic accelerator designs. Mohanad Odema, Hyoukjun Kwon, Mohammad Abdullah Al Faruque |
DATE | 1 |
| 2025 | DisCovHAR: Contrastive Attention for Human Activity Recognition Under Distribution ShiftsabstractAdvances in Internet of Things (IoT) wearable sensors and edge-artificial intelligence (Edge-AI) have enabled practical realizations of machine learning (ML)-enabled mobile sensing applications like human activity recognition (HAR). The effective deployment of these data-driven models necessitates learning robust representations capable of handling prevalent distribution shifts (DS), including new users, device positions, rotations, and more. In that respect, contrastive learning (CL) has shown promise in learning transformation-invariant features, outperforming traditional HAR methods. However, recent findings reveal that the contrastive loss induces shrinkage and expansion of the feature space which may limit the generalization capacity of the model. To address this, we propose DisCovHAR, a contrastive attention method to selectively apply the contrastive loss to a subset of the feature space through the transformer encoder attention mechanism. Extensive experiments on three HAR datasets (DSADS, PAMAP2, and USCHAD) demonstrate its superiority over state-of-the-art methods. Specifically, our approach yields up to 4.47% and 7.82% average accuracy improvements in subject-wise and position-wise generalization settings. Furthermore, DisCovHAR demonstrates up to 5.07% increased robustness compared to prior methods under multivariate distribution shift scenarios. Mohanad Odema, Mohammad Abdullah Al Faruque |
IEEE Internet Things J. | 2 |
| 2024 | SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module AcceleratorsabstractEmerging multi-model workloads with heavy models like recent large language models significantly increased the compute and memory demands on hardware. To address such increasing demands, designing a scalable hardware architecture became a key problem. Among recent solutions, the 2.5D silicon interposer multi-chip module (MCM)-based AI accelerator has been actively explored as a promising scalable solution due to their significant benefits in the low engineering cost and composability. However, previous MCM accelerators are based on homogeneous architectures with fixed dataflow, which encounter major challenges from highly heterogeneous multi-model work-loads due to their limited workload adaptivity. Therefore, in this work, we explore the opportunity in the heterogeneous dataflow MCM AI accelerators. We identify the scheduling of multi-model workload on heterogeneous dataflow MCM AI accelerator is an important and challenging problem due to its significance and scale, which reaches$\mathbf{O}(10^{56})$even for a two-model workload on 6×6 chiplets. We develop a set of heuristics to navigate the huge scheduling space and codify them into a scheduler, SCAR, with advanced techniques such as inter-chiplet pipelining. Our evaluation on ten multi-model workload scenarios for datacenter multitenancy and AR/VR use-cases has shown the efficacy of our approach, achieving on average 27.6% and 29.6% less energy-delay product (EDP) for the respective applications settings compared to homogeneous baselines. Mohanad Odema, Hyoukjun Kwon, Mohammad Abdullah Al Faruque |
MICRO | 1 |
| 2024 | PrivyNAS: Privacy-Aware Neural Architecture Search for Split Computing in Edge-Cloud SystemsabstractSplit Computing has become a prominent resource-efficient method to enable machine learning (ML) applications on user-constrained edge devices, where compute-intensive ML workloads can be delegated to remote cloud servers for processing. However, the exposure of data to the cloud service providers as such raises privacy alarms due to the possible leakage of sensitive user information. On a relevant note, typical deep neural network (DNN) design frameworks do not take into account model splitting and its complications at the early design stages of DNNs. Thus, a natural question arises on how to bridge this gap and optimize the DNN design process such that split computing operations can meet the requirements of accuracy, performance, and privacy. In this paper, we strive to address this question through adopting a privacy-by-design approach, where privacy is characterized either as a constraint or an objective to realize privacy-aware models tailored for split computing. Using the -differential privacy standard for our case study, we conduct intensive empirical analysis on the relation between architectural parameters and intrinsic privacy budgets, and propose PrivyNAS – a privacy-aware Neural Architecture Search framework for split computing. On the CIFAR-10 dataset, our approach has demonstrated promising results in providing DNN architectures that balance the required design trade-offs. Mohanad Odema, Mohammad Abdullah Al Faruque |
IEEE Internet Things J. | 1 |
| 2023 | Map-and-Conquer: Energy-Efficient Mapping of Dynamic Neural Nets onto Heterogeneous MPSoCsabstractHeterogeneous MPSoCs comprise diverse processing units of varying compute capabilities. To date, the mapping strategies of neural networks (NNs) onto such systems are yet to exploit the full potential of processing parallelism, made possible through both the intrinsic NNs’ structure and underlying hardware composition. In this paper, we propose a novel framework to effectively map NNs onto heterogeneous MPSoCs in a manner that enables them to leverage the underlying processing concurrency. Specifically, our approach identifies an optimal partitioning scheme of the NN along its ‘width’ dimension, which facilitates deployment of concurrent NN blocks onto different hardware computing units. Additionally, our approach contributes a novel scheme to deploy partitioned NNs onto the MPSoC as dynamic multi-exit networks for additional performance gains. Our experiments on a standard MPSoC platform have yielded dynamic mapping configurations that are 2.1x more energy-efficient than the GPU-only mapping while incurring 1.7x less latency than DLA-only mapping. Halima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Smaïl Niar, Mohammad Abdullah Al Faruque |
DAC | 2 |
| 2023 | SEO: Safety-Aware Energy Optimization Framework for Multi-Sensor Neural Controllers at the EdgeabstractRuntime energy management has become quintessential for multi-sensor autonomous systems at the edge for achieving high performance given the platform constraints. Typical for such systems, however, is to have their controllers designed with formal guarantees on safety that precede in priority such optimizations, which in turn limits their application in real settings. In this paper, we propose a novel energy optimization framework that is aware of the autonomous system’s safety state, and leverages it to regulate the application of energy optimization methods so that the system’s formal safety properties are preserved. In particular, through the formal characterization of a system’s safety state as a dynamic processing deadline, the computing workloads of the underlying models can be adapted accordingly. For our experiments, we model two popular runtime energy optimization methods, offloading and gating, and simulate an autonomous driving system (ADS) use-case in the CARLA simulation environment with performance characterizations obtained from the standard Nvidia Drive PX2 ADS platform. Our results demonstrate that through a formal awareness of the perceived risks in the test case scenario, energy efficiency gains are still achieved (reaching 89.9%) while maintaining the desired safety properties. Mohanad Odema, James Ferlez, Yasser Shoukry, Mohammad Abdullah Al Faruque |
DAC | 1 |
| 2023 | HADAS: Hardware-Aware Dynamic Neural Architecture Search for Edge Performance ScalingabstractDynamic neural networks (DyNNs) have become viable techniques to enable intelligence on resource-constrained edge devices while maintaining computational efficiency. In many cases, the implementation of DyNNs can be sub-optimal due to its underlying backbone architecture being developed at the design stage independent of both: (i) potential support for dynamic computing, e.g. early exiting, and (ii) resource efficiency features of the underlying hardware, e.g., dynamic voltage and frequency scaling (DVFS). Addressing this, we present HADAS, a novel Hardware-Aware Dynamic Neural Architecture Search framework that realizes DyNN architectures whose backbone, early exiting features, and DVFS settings have been jointly optimized to maximize performance and resource efficiency. Our experiments using the CIFAR-100 dataset and a diverse set of edge computing platforms have shown that HADAS can elevate dynamic models' energy efficiency by up to 57% for the same level of accuracy scores. Our code is available at https://github.com/HalimaBouzidi/HADAS Halima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Mohammad Abdullah Al Faruque, Smaïl Niar |
DATE | 2 |
| 2023 | Testudo: Collaborative Intelligence for Latency-Critical Autonomous SystemsabstractEdge computing is to be widely adopted for autonomous systems (ASs) applications as compute-intensive processing tasks can be offloaded to compute-capable servers located at the edge of the network infrastructure. Given the critical nature of numerous AS applications, their tasks are mostly governed by strict execution deadlines to alleviate any safety concerns from delayed responses. Although wireless link uncertainty has prompted recent works to designate redundant local execution as an offloading fail-safe to ensure these deadlines are met, frequent invocation of such fail-safe mechanisms can potentially undermine the extent of performance gains from offloading. In this article, we thoroughly analyze how redundant execution overheads can influence the overall performance. Then, we present TESTUDO, a methodology to optimize the energy consumption for latency-sensitive AS applications employing collaborative edge computing. Primarily, our methodology encompasses two main stages: 1) designing processing pipelines supporting optimal offloading points and fail-safe integration using modular design techniques and 2) developing a context-aware adaptive runtime solution based on deep reinforcement learning to adapt the mode of operation according to the wireless network status. Our experiments for end-to-end control and object detection use-cases have shown that TESTUDO achieved energy gains reaching up to 31% and 13.4% (15.9% and 5.3% on average) for the former and latter, respectively, while incurring little-to-no degradation in prediction scores (< 1% change) from state-of-the-art strategies. Mohanad Odema, Marco Levorato, Mohammad Abdullah Al Faruque |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | MaGNAS: A Mapping-Aware Graph Neural Architecture Search Framework for Heterogeneous MPSoC DeploymentabstractGraph Neural Networks (GNNs) are becoming increasingly popular for vision-based applications due to their intrinsic capacity in modeling structural and contextual relations between various parts of an image frame. On another front, the rising popularity of deep vision-based applications at the edge has been facilitated by the recent advancements in heterogeneous multi-processor Systems on Chips (MPSoCs) that enable inference under real-time, stringent execution requirements. By extension, GNNs employed for vision-based applications must adhere to the same execution requirements. Yet contrary to typical deep neural networks, the irregular flow of graph learning operations poses a challenge to running GNNs on such heterogeneous MPSoC platforms. In this paper, we propose a novel unified design-mapping approach for efficient processing of vision GNN workloads on heterogeneous MPSoC platforms. Particularly, we develop MaGNAS, a mapping-aware Graph Neural Architecture Search framework. MaGNAS proposes a GNN architectural design space coupled with prospective mapping options on a heterogeneous SoC to identify model architectures that maximize on-device resource efficiency. To achieve this, MaGNAS employs a two-tier evolutionary search to identify optimal GNNs and mapping pairings that yield the best performance trade-offs. Through designing a supernet derived from the recent Vision GNN (ViG) architecture, we conducted experiments on four (04) state-of-the-art vision datasets using both ( i ) a real hardware SoC platform (NVIDIA Xavier AGX) and ( ii ) a performance/cost model simulator for DNN accelerators. Our experimental results demonstrate that MaGNAS is able to provide 1.57 × latency speedup and is 3.38 × more energy-efficient for several vision datasets executed on the Xavier MPSoC vs. the GPU-only deployment while sustaining an average 0.11% accuracy reduction from the baseline. Mohanad Odema, Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, Mohammad Abdullah Al Faruque |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2022 | Romanus: Robust Task Offloading in Modular Multi-Sensor Autonomous Driving SystemsabstractDue to the high performance and safety requirements of self-driving applications, the complexity of modern autonomous driving systems (ADS) has been growing, instigating the need for more sophisticated hardware which could add to the energy footprint of the ADS platform. Addressing this, edge computing is poised to encompass self-driving applications, enabling the compute-intensive autonomy-related tasks to be offloaded for processing at compute-capable edge servers. Nonetheless, the intricate hardware architecture of ADS platforms, in addition to the stringent robustness demands, set forth complications for task offloading which are unique to autonomous driving. Hence, we present ROMANUS, a methodology for robust and efficient task offloading for modular ADS platforms with multi-sensor processing pipelines. Our methodology entails two phases: (i) the introduction of efficient offloading points along the execution path of the involved deep learning models, and (ii) the implementation of a runtime solution based on Deep Reinforcement Learning to adapt the operating mode according to variations in the perceived road scene complexity, network connectivity, and server load. Experiments on the object detection use case demonstrated that our approach is 14.99% more energy-efficient than pure local execution while achieving a 77.06% reduction in risky behavior from a robust-agnostic offloading baseline. Mohanad Odema, Mohammad Abdullah Al Faruque |
ICCAD | 2 |
| 2021 | Energy-Aware Design Methodology for Myocardial Infarction Detection on Low-Power Wearable DevicesabstractMyocardial Infarction (MI) is a heart disease that damages the heart muscle and requires immediate treatment. Its silent and recurrent nature necessitates real-time continuous monitoring of patients. Nowadays, wearable devices are smart enough to perform on-device processing of heartbeat segments and report any irregularities in them. However, the small form factor of wearable devices imposes resource constraints and requires energy-efficient solutions to satisfy them. In this paper, we propose a design methodology to automate the design space exploration of neural network architectures for MI detection. This methodology incorporates Neural Architecture Search (NAS) using Multi-Objective Bayesian Optimization (MOBO) to render Pareto optimal architectural models. These models minimize both detection error and energy consumption on the target device. The design space is inspired by Binary Convolutional Neural Networks (BCNNs) suited for mobile health applications with limited resources. The models' performance is validated using the PTB diagnostic ECG database from PhysioNet. Moreover, energy-related measurements are directly obtained from the target device in a typical hardware-in-the-loop fashion. Finally, we benchmark our models against other related works. One model exceeds state-of-the-art accuracy on wearable devices (reaching 91.22%), whereas others trade off some accuracy to reduce their energy consumption (by a factor reaching 8.26x). Mohanad Odema, Mohammad Abdullah Al Faruque |
ASP-DAC | 1 |
| 2021 | LENS: Layer Distribution Enabled Neural Architecture Search in Edge-Cloud HierarchiesabstractEdge-Cloud hierarchical systems employing intelligence through Deep Neural Networks (DNNs) endure the dilemma of workload distribution within them. Previous solutions proposed to distribute workloads at runtime according to the state of the surroundings, like the wireless conditions. However, such conditions are usually overlooked at design time. This paper addresses this issue for DNN architectural design by presenting a novel methodology, LENS, which administers multi-objective Neural Architecture Search (NAS) for two-tiered systems, where the performance objectives are refashioned to consider the wireless communication parameters. From our experimental search space, we demonstrate that LENS improves upon the traditional solution’s Pareto set by 76.47% and 75% with respect to the energy and latency metrics, respectively. Mohanad Odema, Berken Utku Demirel, Mohammad Abdullah Al Faruque |
DAC | 1 |
| 2021 | EExNAS: Early-Exit Neural Architecture Search Solutions for Low-Power Wearable DevicesabstractEquipping wearable devices with intelligence is essential for promoting mobile healthcare applications. However, challenges remain due to the resource limitations of these devices. In this work, we introduce EExNAS, a methodology for designing high-performance and resource-efficient dynamic Neural Architecture solutions for wearable devices. The methodology incorporates a platform-aware Neural Architecture Search (NAS) that accounts for energy efficiency at runtime through an Early-Exit (EEx) option. We showcase our methodology’s merit across 2 wearable applications, Myocardial Infarction (MI) detection and Human Activity Recognition (HAR). Solutions from EExNAS are compared against those from related works in terms of accuracy and performance. For MI detection, our final solutions with EEx capability could reach 98.54% accuracy on the PTB ECG dataset. Mohanad Odema, Mohammad Abdullah Al Faruque |
ISLPED | 1 |
| 2021 | SAGE: A Split-Architecture Methodology for Efficient End-to-End Autonomous Vehicle ControlabstractAutonomous vehicles (AV) are expected to revolutionize transportation and improve road safety significantly. However, these benefits do not come without cost; AVs require large Deep-Learning (DL) models and powerful hardware platforms to operate reliably in real-time, requiring between several hundred watts to one kilowatt of power. This power consumption can dramatically reduce vehicles’ driving range and affect emissions. To address this problem, we propose SAGE: a methodology for selectively offloading the key energy-consuming modules of DL architectures to the cloud to optimize edge, energy usage while meeting real-time latency constraints. Furthermore, we leverage Head Network Distillation (HND) to introduce efficient bottlenecks within the DL architecture in order to minimize the network overhead costs of offloading with almost no degradation in the model’s performance. We evaluate SAGE using an Nvidia Jetson TX2 and an industry-standard Nvidia Drive PX2 as the AV edge, devices and demonstrate that our offloading strategy is practical for a wide range of DL models and internet connection bandwidths on 3G, 4G LTE, and WiFi technologies. Compared to edge-only computation, SAGE reduces energy consumption by an average of 36.13% , 47.07% , and 55.66% for an AV with one low-resolution camera, one high-resolution camera, and three high-resolution cameras, respectively. SAGE also reduces upload data size by up to 98.40% compared to direct camera offloading. Arnav Vaibhav Malawade, Mohanad Odema, Sebastien Lajeunesse-DeGroot, Mohammad Abdullah Al Faruque |
ACM Trans. Embed. Comput. Syst. | 2 |