Malik Aqeel Anwar

dblp:154/3556 · also Aqeel Anwar · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2022
0000-0001-6768-058XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2022 FRL-FI: Transient Fault Analysis for Federated Reinforcement Learning-Based Navigation Systems
abstract
Swarm intelligence is being increasingly deployed in autonomous systems, such as drones and unmanned vehicles. Federated reinforcement learning (FRL), a key swarm intelligence paradigm where agents interact with their own environments and cooperatively learn a consensus policy while preserving privacy, has recently shown potential advantages and gained popularity. However, transient faults are increasing in the hardware system with continuous technology node scaling and can pose threats to FRL systems. Meanwhile, conventional redundancy-based protection methods are challenging to deploy on resource-constrained edge applications. In this paper, we experimentally evaluate the fault tolerance of FRL navigation systems at various scales with respect to fault models, fault locations, learning algorithms, layer types, communication intervals, and data types at both training and inference stages. We further propose two cost-effective fault detection and recovery techniques that can achieve up to$3.3\times$improvement in resilience with$<2.7\%$overhead in FRL systems.
Zishen Wan, Malik Aqeel Anwar, Abdulrahman Mahmoud, Yu-Shun Hsiao, Vijay Janapa Reddi, Arijit Raychowdhury
DATE2
2022 RAPID-RL: A Reconfigurable Architecture with Preemptive-Exits for Efficient Deep-Reinforcement Learning
abstract
Present-day Deep Reinforcement Learning (RL) systems show great promise towards building intelligent agents surpassing human-level performance. However, the computational complexity associated with the underlying deep neural networks (DNNs) leads to power-hungry implementations. This makes deep RL systems unsuitable for deployment on resource-constrained edge devices. To address this challenge, we propose a reconfigurable architecture with preemptive exits for effi-cient deep RL (RAPID-RL). RAPID-RL enables conditional activation of DNN layers based on the difficulty level of inputs. This allows to dynamically adjust the compute effort during inference while maintaining competitive performance. We achieve this by augmenting a deep Q-network (DQN) with side-branches capable of generating intermediate predictions along with an associated confidence score. We also propose a novel training methodology for learning the actions and branch confidence scores in a dynamic RL setting. Our experiments evaluate the proposed framework for Atari 2600 gaming tasks and a realistic Drone navigation task on an open-source drone simulator (PEDRA). We show that RAPID-RL incurs 0.34 × (0.25 ×) number of operations (OPS) while maintaining performance above 0.88 × (0.91 ×) on Atari (Drone navigation) tasks, compared to a baseline-DQN without any side-branches. The reduction in OPS leads to fast and efficient inference, proving to be highly beneficial for the resource-constrained edge where making quick decisions with minimal compute is essential.
Adarsh Kosta, Malik Aqeel Anwar, Priyadarshini Panda, Arijit Raychowdhury, Kaushik Roy 0001
ICRA2
2021 Analyzing and Improving Fault Tolerance of Learning-Based Navigation Systems
abstract
Learning-based navigation systems are widely used in autonomous applications, such as robotics, unmanned vehicles and drones. Specialized hardware accelerators have been proposed for high-performance and energy-efficiency for such navigational tasks. However, transient and permanent faults are increasing in hardware systems and can catastrophically violate tasks safety. Meanwhile, traditional redundancy-based protection methods are challenging to deploy on resource-constrained edge applications. In this paper, we experimentally evaluate the resilience of navigation systems with respect to algorithms, fault models and data types from both RL training and inference. We further propose two efficient fault mitigation techniques that achieve $2 \times$ success rate and 39% quality-of-flight improvement in learning-based navigation systems.
Zishen Wan, Malik Aqeel Anwar, Yu-Shun Hsiao, Vijay Janapa Reddi, Arijit Raychowdhury
DAC2
2021 A decentralized policy gradient approach to multi-task reinforcement learning
abstract
We develop a mathematical framework for solving multi-task reinforcement learning (MTRL) problems based on a type of policy gradient method. The goal in MTRL is to learn a common policy that operates effectively in different environments; these environments have similar (or overlapping) state spaces, but have different rewards and dynamics. We highlight two fundamental challenges in MTRL that are not present in its single task counterpart, and illustrate them with simple examples. We then develop a decentralized entropyregularized policy gradient method for solving the MTRL problem, and study its finite-time convergence rate. We demonstrate the effectiveness of the proposed method using a series of numerical experiments. These experiments range from small-scale "GridWorld" problems that readily demonstrate the trade-offs involved in multi-task learning to large-scale problems, where common policies are learned to navigate an airborne drone in multiple (simulated) environments.
Sihan Zeng, Malik Aqeel Anwar, Thinh T. Doan 0001, Arijit Raychowdhury, Justin K. Romberg
UAI2
2019 Transfer and Online Reinforcement Learning in STT-MRAM Based Embedded Systems for Autonomous Drones
abstract
In this paper we present an algorithm-hardware co-design for camera-based autonomous flight in small drones. We show that the large write-latency and write-energy for nonvolatile memory (NVM) based embedded systems makes them unsuitable for real-time reinforcement learning (RL). We address this by performing transfer learning (TL) on meta-environments and RL on the last few layers of a deep convolutional network. While the NVM stores the meta-model from TL, an on-die SRAM stores the weights of the last few layers. Thus all the real-time updates via RL are carried out on the SRAM arrays. This provides us with a practical platform with comparable performance as end-to-end RL and 83.4% lower energy per image frame.
Insik Yoon, Malik Aqeel Anwar, Titash Rakshit, Arijit Raychowdhury
DATE2
2014 Acoustic sensor network relative self-calibration using joint TDOA and DOA with unknown beacon positions
abstract
In this paper, we propose an efficient self-calibration mechanism, achieving localization as well as orientation, for wireless acoustic sensor network. Time difference of arrival (TDOA) based on wireless and acoustic signals, while direction of arrival (DOA) based on acoustic only, are used jointly to achieve acoustic sensor node calibration. Each sensor node is equipped with three microphones and a wireless transceiver. Sensor nodes are calibrated by using two beacon positions with respect to a reference node whose position is assumed to be known. Our approach requires a single moving beacon (MB) equipped with RF and acoustic signal sources. Once a node has been calibrated (in terms of location and orientation) it can be considered as a valid reference for the remaining uncalibrated nodes. We have developed hardware platform to validate the proposed calibration mechanism. Performance results show the effectiveness and usefulness of the proposed mechanism.
Malik Aqeel Anwar, Hammad Hassan, Hasan Maqbool, Akif Rehman, Muhammad Tahir 0002
WCNC1