Waleed Ahsan

dblp:226/2324 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-3550-7077ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Resource Allocation for Semantic Aware Relay Networks using Multi-Agent Reinforcement Learning
abstract
The variety and volume of real-time video streaming services make network resources like power, frequency, and time critical to the user experience. Semantic communication, by transmitting only meaningful data, significantly reduces the load on the network, leading to faster transmission times and reduced latency-both crucial for enhancing user experience. In this paper, we develop a context-aware resource allocation framework for a semantic-aware regenerative AI-based Unmanned Aerial Vehicle (UAV) setup that accommodates both conventional and semantic communication users. Specifically, we define a semantic feature pooling mechanism, upon which a novel Quality of Experience (QoE) model is proposed. For dynamic network environments, we formulate a long-term resource allocation problem by maximizing the expected rewards. Each UAV is modeled as a learning agent, with each resource allocation solution corresponding to an action taken by the UAVs. We then develop a Multi-Agent Reinforcement Learning (MARL) framework in which each agent discovers its optimal strategy based on local observations. More specifically, we propose an agent-independent method where all agents execute a decision algorithm independently while sharing a common structure based on Q-learning. The proposed framework also facilitates intelligent traffic steering by dynamically adjusting resource distribution based on real-time context. Finally, simulation results demonstrate the effectiveness and superiority of the proposed method in terms of overall QoE.
Waleed Ahsan, Chuan Heng Foh, Yi Ma 0002
VTC2025-Spring1
2024 Reinforcement Learning Based Age of Information Minimization for Downlink NOMA Systems
abstract
In this paper, we develop learning based strategies to optimize the age of information (AoI) based resource allocation for downlink non-orthogonal multiple access (NOMA) systems. Considering dynamic network loads, we design state-action-reward-state-action (SARSA) Q-learning for small-scale networks, as it provides more reliable policy than traditional Q-learning. For large-scale networks, we design a deep deterministic policy gradient (DDPG) cooperative agent system, as DDPG is able to efficiently handle huge state-action spaces with deep neural networks (DNNs). With the aid of the designed algorithms, the entire AoI centric dynamic resource allocations are updated using actor- critic based for DNNs and the SARSA Q-table for large/small scale networks. We propose a multi-objective reward system based on the sum rate, mean AoI, and max power, which helps to improve the update of the DNNs weights in DDPG and Q-table values in SARSA. The simulation outcomes prove that a) Proposed schemes are able to find efficient optimal allocation policies for all considered types of network traffic; b) The proposed multi-objective reward function can assist reinforcement learning agents for long-term communications by converging within 100 episodes; and c) The long-term average throughput performance of the NOMA AoI system is better than the conventional orthogonal multiple access AoI system.
Waleed Ahsan, Anees Ahsan, Faiqa Bibi
GLOBECOM1
2022 A Reliable Reinforcement Learning for Resource Allocation in Uplink NOMA-URLLC Networks
abstract
In this paper, we propose a deep state-action-reward-state-action (SARSA)$\lambda $learning approach for optimising the uplink resource allocation in non-orthogonal multiple access (NOMA) aided ultra-reliable low-latency communication (URLLC). To reduce the mean decoding error probability in time-varying network environments, this work designs a reliable learning algorithm for providing a long-term resource allocation, where the reward feedback is based on the instantaneous network performance. With the aid of the proposed algorithm, this paper addresses three main challenges of the reliable resource sharing in NOMA-URLLC networks: 1) user clustering; 2) Instantaneous feedback system; and 3) Optimal resource allocation. All of these designs interact with the considered communication environment. Lastly, we compare the performance of the proposed algorithm with conventional Q-learning and SARSA Q-learning algorithms. The simulation outcomes show that: 1) Compared with the traditional Q learning algorithms, the proposed solution is able to converge within 200 episodes for providing as low as$10^{-2}$long-term mean error; 2) NOMA assisted URLLC outperforms traditional OMA systems in terms of decoding error probabilities; and 3) The proposed feedback system is efficient for the long-term learning process.
Waleed Ahsan, Wenqiang Yi, Yuanwei Liu, Arumugam Nallanathan
IEEE Trans. Wirel. Commun.1
2021 Reliable Reinforcement Learning Based NOMA Schemes for URLLC
abstract
In this paper, we propose a deep state-action-reward-state-action (SARSA)$A$learning approach for optimising the uplink resource allocation in non-orthogonal multiple access (NOMA) aided ultra-reliable low-latency communication (URLLC). To reduce the mean decoding error probability in time-varying network environments, this work designs a reliable learning algorithm for providing a long-term resource allocation, where the reward feedback is based on the instantaneous network performance. With the aid of the proposed algorithm, this paper addresses three main challenges of the reliable resource sharing in NOMA-URLLC networks: 1) Dynamic user clustering; 2) Instantaneous feedback system; and 3) Optimal resource allocation. All of these designs interact with the considered communication environment. The simulation outcomes show that: 1) Compared with the traditional Q learning algorithm, the proposed solution converges faster and obtains better performance; 2) NOMA assisted URLLC outperforms traditional OMA systems in terms of decoding error probabilities; and 3) The dynamic feedback system is efficient for the long-term learning process.
Waleed Ahsan, Wenqiang Yi, Yuanwei Liu, Arumugam Nallanathan
GLOBECOM1
2021 Resource Allocation in Uplink NOMA-IoT Networks: A Reinforcement-Learning Approach
abstract
Non-orthogonal multiple access (NOMA) exploits the potential of the power domain to enhance the connectivity for the Internet of Things (IoT). Due to time-varying communication channels, dynamic user clustering is a promising method to increase the throughput of NOMA-IoT networks. This article develops an intelligent resource allocation scheme for uplink NOMA-IoT communications. To maximise the average performance of sum rates, this work designs an efficient optimization approach based on two reinforcement learning algorithms, namely deep reinforcement learning (DRL) and SARSA-learning. For light traffic, SARSA-learning is used to explore the safest resource allocation policy with low cost. For heavy traffic, DRL is used to handle traffic-introduced huge variables. With the aid of the considered approach, this work addresses two main problems of fair resource allocation in NOMA techniques: 1) allocating users dynamically and 2) balancing resource blocks and network traffic. We analytically demonstrate that the rate of convergence is inversely proportional to network sizes. Numerical results show that: 1) Compared with the optimal benchmark scheme, the proposed DRL and SARSA-learning algorithms have lower complexity with acceptable accuracy and 2) NOMA-enabled IoT networks outperform the conventional orthogonal multiple access based IoT networks in terms of system throughput.
Waleed Ahsan, Wenqiang Yi, Zhijin Qin, Yuanwei Liu, Arumugam Nallanathan
IEEE Trans. Wirel. Commun.1
2018 Clustering algorithm for internet of vehicles (IoV) based on dragonfly optimizer (CAVDO)
Farhan Aadil, Waleed Ahsan, Zahoor-Ur Rehman, Peer Azmat Shah, Seungmin Rho, Irfan Mehmood
J. Supercomput.2