Naram Mhaisen

dblp:265/7849 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-0211-2666ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DRONE-RL: Dynamic reinforcement learning for online navigation of UAVs in evolving environments
Noor Khial, Mhd Saria Allahham, Naram Mhaisen, Loay Ismail, Mohamed Abdalla Mabrok, Amr Mohamed 0001
Knowl. Based Syst.3
2025 Multi-Target Path Planning with Probabilistic Detection in Cluttered Environments
abstract
Autonomous Unmanned Aerial Vehicles (UAVs) offer substantial advantages for tasks such as surveillance, disaster management, and environmental monitoring, where human intervention can be risky. With advancements in their agility and autonomy, UAVs are becoming essential for critical tasks in combat, reconnaissance, wildfire monitoring, and disaster search and rescue. This paper addresses a key challenge in UAV path planning: efficiently visiting multiple unknown mobile targets in complex, obstacle-filled environments. We leverage the Deep Deterministic Policy Gradient (DDPG) framework to continuously control UAV movement to enable effective obstacle avoidance and sequential target visitation. Our approach allows the UAV to learn the unknown distribution of mobile targets and determine optimal paths while navigating around obstacles. With limited environment information, the agent receives rewards based on the confidence of detecting targets within its observation field. We validate the effectiveness of our method through comparison with an optimal benchmark that assumes perfect knowledge of target mobility and obstacle locations. Results indicate that increasing target numbers significantly impacts the agent's performance by requiring additional training time. Moreover, heavily cluttered environments reduce mission success rates for target visitation.
Noor Khial, Naram Mhaisen, Loay Ismail, Mohamed Abdalla Mabrok, Amr Mohamed 0001
ICC2
2025 On the Dynamic Regret of Following the Regularized Leader: Optimism with History Pruning
abstract
We revisit the Follow the Regularized Leader (FTRL) framework for Online Convex Optimization (OCO) over compact sets, focusing on achieving dynamic regret guarantees. Prior work has highlighted the framework’s limitations in dynamic environments due to its tendency to produce "lazy" iterates. However, building on insights showing FTRL’s ability to produce "agile" iterates, we show that it can indeed recover known dynamic regret bounds through optimistic composition of future costs and careful linearization of past costs, which can lead to pruning some of them. This new analysis of FTRL against dynamic comparators yields a principled way to interpolate between greedy and agile updates and offers several benefits, including refined control over regret terms, optimism without cyclic dependence, and the application of minimal recursive regularization akin to AdaFTRL. More broadly, we show that it is not the "lazy" projection style of FTRL that hinders (optimistic) dynamic regret, but the decoupling of the algorithm’s state (linearized history) from its iterates, allowing the state to grow arbitrarily. Instead, pruning synchronizes these two when necessary.
Naram Mhaisen, George Iosifidis
ICML1
2025 An online learning framework for UAV search mission in adversarial environments
Noor Khial, Naram Mhaisen, Mohamed Abdalla Mabrok, Amr Mohamed 0001
Expert Syst. Appl.2
2025 Slicing for AI: An Online Learning Framework for Network Slicing Supporting AI Services
abstract
The forthcoming 6G networks will embrace a new realm of AI-driven services that requires innovative network slicing strategies, namely slicing for AI, which involves the creation of customized network slices to meet Quality of Service (QoS) requirements of diverse AI services. This poses challenges due to time-varying dynamics of users’ behavior and mobile networks. Thus, this paper proposes an online learning framework to determine the allocation of computational and communication resources to AI services, to optimize their accuracy as one of their unique key performance indicators (KPIs), while abiding by resources, learning latency, and cost constraints. We define a problem of optimizing the total accuracy while balancing conflicting KPIs, prove its NP-hardness, and propose an online learning framework for solving it in dynamic environments. We present a basic online solution and two variations employing a pre-learning elimination method for reducing the decision space to expedite the learning. Furthermore, we propose a biased decision space subset selection by incorporating prior knowledge to enhance the learning speed without compromising performance and present two alternatives of handling the selected subset. Our results depict the efficiency of the proposed solutions in converging to the optimal decisions, while reducing decision space and improving time complexity. Additionally, our solution outperforms State-of-the-Art techniques in adapting to diverse environmental dynamics and excels under varying levels of resource availability.
Menna Helmy, Alaa Awad, Naram Mhaisen, Amr Mohamed 0001, Aiman Erbad
IEEE Trans. Netw. Serv. Manag.3
2024 Online Caching With no Regret: Optimistic Learning via Recommendations
abstract
The design of effective online caching policies is an increasingly important problem for content distribution networks, online social networks and edge computing services, among other areas. This paper proposes a new algorithmic toolbox for tackling this problem through the lens ofoptimisticonline learning. We build upon the Follow-the-Regularized-Leader (FTRL) framework, which is developed further here to include predictions for the file requests, and we design online caching algorithms for bipartite networks with pre-reserved or dynamic storage subject to time-average budget constraints. The predictions are provided by a content recommendation system that influences the users viewing activity and hence can naturally reduce the caching network's uncertainty about future requests. We also extend the framework to learn and utilize the best request predictor in cases where many are available. We prove that the proposed optimistic learning caching policies can achievesub-zeroperformance loss (regret) for perfect predictions, and maintain the sub-linear regret bound$O(\sqrt{T})$, which is the best achievable bound for policies that do not use predictions, even for arbitrary-bad predictions. The performance of the proposed algorithms is evaluated with detailed trace-driven numerical tests.
Naram Mhaisen, George Iosifidis, Douglas J. Leith
IEEE Trans. Mob. Comput.1
2023 Federated Learning for Online Resource Allocation in Mobile Edge Computing: A Deep Reinforcement Learning Approach
abstract
Federated learning (FL) is increasingly considered to circumvent the disclosure of private data in mobile edge computing (MEC) systems. Training with large data can enhance FL learning accuracy, which is associated with non-negligible energy use. Scheduled edge devices with small data save energy but decrease FL learning accuracy due to a reduction in energy consumption. A trade-off between the energy consumption of edge devices and the learning accuracy of FL is formulated in this proposed work. The FL-enabled twin-delayed deep deterministic policy gradient (FL-TD3) framework is proposed as a solution to the formulated problem because its state and action spaces are large in a continuous domain. This framework provides the maximum accuracy ratio of FL divided by the device’s energy consumption. A comparison of the numerical results with the state-of-the-art demonstrates that the ratio has been improved significantly.
Kai Li 0002, Naram Mhaisen, Wei Ni 0001, Eduardo Tovar, Mohsen Guizani
WCNC3
2023 Reinforcement Learning for Intelligent Healthcare Systems: A Review of Challenges, Applications, and Open Research Issues
abstract
The rise of chronic disease patients and the pandemic pose immediate threats to healthcare expenditure and mortality rates. This calls for transforming healthcare systems away from one-on-one patient treatment into intelligent health systems, leveraging the recent advances of Internet of Things and smart sensors. Meanwhile, reinforcement learning (RL) has witnessed an intrinsic breakthrough in solving a variety of complex problems for distinct applications and services. Thus, this article presents a comprehensive survey of the recent models and techniques of RL that have been developed/used for supporting Intelligent-healthcare (I-health) systems. It can guide the readers to deeply understand the state-of-the-art regarding the use of RL in the context of I-health. Specifically, we first present an overview of the I-health systems’ challenges, architecture, and how RL can benefit these systems. We then review the background and mathematical modeling of different RL, deep RL (DRL), and multiagent RL models. We highlight important guidelines on how to select the appropriate RL model for a given problem, and provide quantitative comparisons, showing the results of deploying key RL models in two scenarios that can be followed in monitoring applications. After that, we conduct an in-depth literature review on RL’s applications in I-health systems, covering edge intelligence, smart core network, and dynamic treatment regimes. Finally, we highlight emerging challenges and future research directions to enhance RL’s success in I-health systems, which opens the door for exploring some interesting and unsolved problems.
Alaa Awad, Naram Mhaisen, Amr Mohamed 0001, Aiman Erbad, Mohsen Guizani
IEEE Internet Things J.2
2022 Communication-efficient hierarchical federated learning for IoT heterogeneous systems with imbalanced data
abstract
Federated Learning (FL) is a distributed learning methodology that allows multiple nodes to cooperatively train a deep learning model, without the need to share their local data. It is a promising solution for telemonitoring systems that demand intensive data collection, for detection, classification, and prediction of future events, from different locations while maintaining a strict privacy constraint. Due to privacy concerns and critical communication bottlenecks, it can become impractical to send the FL updated models to a centralized server. Thus, this paper studies the potential of hierarchical FL in Internet of Things (IoT) heterogeneous systems. In particular, we propose an optimized solution for user assignment and resource allocation over hierarchical FL architecture for IoT heterogeneous systems. This work focuses on a generic class of machine learning models that are trained using gradient-descent-based schemes while considering the practical constraints of non-uniformly distributed data across different users. We evaluate the proposed system using two real-world datasets, and we show that it outperforms state-of-the-art FL solutions. Specifically, our numerical results highlight the effectiveness of our approach and its ability to provide 4–6% increase in the classification accuracy, with respect to hierarchical FL schemes that consider distance-based user assignment. Furthermore, the proposed approach could significantly accelerate FL training and reduce communication overhead by providing 75–85% reduction in the communication rounds between edge nodes and the centralized server, for the same model accuracy.
Alaa Awad, Naram Mhaisen, Amr Mohamed 0001, Aiman Erbad, Mohsen Guizani, Zaher Dawy, Wassim Nasreddine
Future Gener. Comput. Syst.2
2022 Exploring Deep-Reinforcement-Learning-Assisted Federated Learning for Online Resource Allocation in Privacy-Preserving EdgeIoT
abstract
Federated learning (FL) has been increasingly considered to preserve data training privacy from eavesdropping attacks in mobile-edge computing-based Internet of Things (EdgeIoT). On the one hand, the learning accuracy of FL can be improved by selecting the IoT devices with large data sets for training, which gives rise to a higher energy consumption. On the other hand, the energy consumption can be reduced by selecting the IoT devices with small data sets for FL, resulting in a falling learning accuracy. In this article, we formulate a new resource allocation problem for privacy-preserving EdgeIoT to balance the learning accuracy of FL and the energy consumption of the IoT device. We propose a new FL-enabled twin-delayed deep deterministic policy gradient (FL-DLT3) framework to achieve the optimal accuracy and energy balance in a continuous domain. Furthermore, long short-term memory (LSTM) is leveraged in FL-DLT3 to predict the time-varying network state while FL-DLT3 is trained to select the IoT devices and allocate the transmit power. Numerical results demonstrate that the proposed FL-DLT3 achieves fast convergence (less than 100 iterations) while the FL accuracy-to-energy consumption ratio is improved by 51.8% compared to the existing state-of-the-art benchmark.
Kai Li 0002, Naram Mhaisen, Wei Ni 0001, Eduardo Tovar, Mohsen Guizani
IEEE Internet Things J.3
2021 Rational Contracts: Data-driven Service Provisioning in Blockchain-powered Systems
abstract
Smart Contracts (SCs), which are software programs that run on blockchain platforms, provide appealing security guarantees characterized by decentralized, autonomous, and verifiable execution. On the other hand, Service Provisioning (SP) systems (i.e., systems that assign users to service providers in a way that maximizes the global utility) have been leveraging SCs to provide trust and transparency features. Such features are obtained by deploying the SP’s assignment criteria as an SC on the blockchain. However, deploying optimal assignment criteria as SCs does not guarantee the best performance over time since the blockchain participants join and leave flexibly, and their load varies with time, potentially deeming the initial assignment sub-optimal. Furthermore, modifying the criteria manually by a third party at every variation in the blockchain jeopardizes the autonomous and independent execution promised by SCs. Thus, in this paper, we propose the use of online learning SCs that leverage the chained data to continuously self-tune their assignment criteria and maintain maximum utility. We show that the proposed data-driven method can achieve high performance on the multi-stage assignment problem. We also compare the proposed approach to multiple assignment algorithms as well as planning methods. Results show a significant performance advantage over heuristics and better adaptability to the dynamic nature of blockchain networks compared to planning techniques.
Naram Mhaisen, Amr Mohamed 0001, Aiman Erbad, Mohsen Guizani
ICC1
2020 To chain or not to chain: A reinforcement learning approach for blockchain-enabled IoT monitoring applications
Naram Mhaisen, Noora Fetais, Aiman Erbad, Amr Mohamed 0001, Mohsen Guizani
Future Gener. Comput. Syst.1