VLDB 2026 Research / reviewers in the wild / expert
Arslan Munir
dblp:47/710
· DBLP profile ↗
40ranked-venue papers
19as first author
7since 2021 · last 2026
0000-0002-3126-8945ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 8 first-author · 1 since 2021Computer networks · 5 · 4 first-author · 1 since 2021Security and privacy · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition
Hayat Ullah, Muhammad Ali Shafique, Arslan Munir |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Design and Development of an Unmanned Orchard Spraying Robot with Two Degrees of FreedomabstractAutomated and efficient spraying mechanisms are imperative to cope with the scarcity of human resources in agriculture, reduce the wastage of pesticides, and improve agricultural yield. In this paper, we present an indigenous solution for spraying of pesticides on different agricultural crops particularly fruit orchards in order to avoid the wastage of pesticide liquid and to prevent the health hazards of the pesticides faced by the agricultural workers utilizing a targeted robotic spraying technique. The implementation of this technique in agricultural procedures has shown a positive impact through the reduction of wastage, enhancement of food security, and optimization of resource utilization. The spraying robot consists of an unmanned ground vehicle (UGV) and a robotic spraying manipulator with two degrees of freedom capable of moving in azimuth and elevation orientations. The spraying mechanism is mounted over the UGV which is capable of moving over all kinds of rough agricultural terrains. The operator is able to control the UGV and the spraying manipulator through a LabView-based computer terminal connected to a joystick via a wireless radio frequency (RF) link at a range of approximately six hundred meters. In case of long distances, the UGV is equipped with a night-vision camera to continue the spraying campaigns irrespective of the lighting conditions. The efficiency of the robot is tested in practical environments and has shown effective results without causing harm to the health of the operator. Sardar Ali Abbas, Waqar Shahid Qureshi, Arslan Munir |
IECON | 3 |
| 2023 | AI-Driven Salient Soccer Events Recognition Framework for Next-Generation IoT-Enabled EnvironmentsabstractThe salient event recognition of soccer matches in the next-generation Internet of Things (Nx-IoT) environment aims to analyze the performance of players/teams by the sports analytics and managerial staff. The embedded Nx-IoT devices carried by the soccer players during the match capture and transmit data to an artificial intelligence (AI)-assisted computing platform. The interconnectivity of data acquisition devices with an AI-assisted computing platform in the Nx-IoT environment will not only allow the spectators to track the formation of their favorite players during a soccer match but will also enable the managerial staff to evaluate the players’ performance in the soccer match as well as in practice sessions. This Nx-IoT-enabled salient event detection feature can be provided to spectators and sports’ managerial staff as a financial technology (FinTech) service. In this article, we propose an efficient deep-learning-based framework for multiperson salient soccer event recognition in IoT-enabled FinTech. The proposed framework performs event recognition in three steps: 1) frames preprocessing; 2) frame-level discriminative features extraction; and 3) high-level events recognition in soccer videos. Moreover, we introduce a new soccer video events (SVE) data set containing videos of six salient events of soccer games. To provide a strong baseline, we evaluate our newly created SVE data set using different traditional machine learning and deep learning algorithms. We also perform event recognition on untrimmed soccer videos using our proposed framework and compare the results with state-of-the-art methods. The obtained results validate the suitability of our proposed framework for salient event recognition in Nx-IoT environments. Khan Muhammad 0001, Hayat Ullah, Mohammad S. Obaidat, Amin Ullah, Arslan Munir, Victor Hugo C. de Albuquerque |
IEEE Internet Things J. | 5 |
| 2022 | PUF-RAKE: A PUF-Based Robust and Lightweight Authentication and Key Establishment ProtocolabstractPhysically unclonable functions (PUFs) bind a device's identity to its physical hardware and thus, can be employed for device identification, authentication and cryptographic key generation. However, PUFs are susceptible to modeling attacks if a number of PUFs’ challenge-response pairs (CRPs) are exposed to the adversary. Furthermore, many of the embedded devices requiring authentication and inter-device communication in a real-time environment/system have stringent resource and low latency requirements, and thus require a lightweight authentication and key establishment mechanism to quickly realize an authenticated and secure connection. We propose PUF-RAKE, a PUF-based lightweight, highly reliable authentication and key establishment scheme. The proposed scheme enhances the reliability of PUF as well as alleviates the resource constraints by employing error correction in the server instead of the device as well as removing cryptographic hashing required by earlier PUF-based protocols. The proposed PUF-RAKE is robust against masquerade, brute force, replay, and modeling attacks. In PUF-RAKE, we introduce an inexpensive yet secure stream authentication scheme inside the device which authenticates the server before the underlying PUF can be invoked. This prevents an adversary from brute forcing the device's PUF to acquire CRPs essentially locking out the device from unauthorized model generation. Additionally, we also introduce a lightweight CRP obfuscation mechanism involving XOR and shuffle operations. The security of PUF-RAKE has been formally verified. A prototype of the protocol has been implemented on two Xilinx Zynq 7000 system-on-chips with one present on Xilinx zc706 evaluation board and the other present on the Avnet Zedboard. Observations, security analysis and results verify that the PUF-RAKE is secure against a probabilistic polynomial time adversary under both the unauthenticated link and authenticated link adversarial models while providing ∼99% reliable authentication. In addition, PUF-RAKE provides a reduction of 60 and 72 percent for look-up tables (LUTs) and register count, respectively, in the programmable logic (PL) part of the Zynq 7000 as compared to a recently proposed approach while providing additional advantages. Mahmood Azhar Qureshi, Arslan Munir |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | Multi-Feature Extraction with Ensemble Network for Tracing Chronic Retinal DisordersabstractThe retina manifests a vital role in tracing chronic retinal disorders. It is located near the optic nerve that transforms the captured light into neural signals. The most prominent chronic eye diseases exhibit themselves in the retina. The analysis of a retina for detecting disease symptoms is quite challenging. Most of the prior methods developed using shallow and deep learning algorithms primarily emphasized single feature extraction for disease diagnosis. The underlying article has designed an ensemble network for extracting multiple retinal features using a single comprehensive platform. It includes a set of models that reflect feature-based needs to prevent intensity loss, micro-vessels overlap, and data redundancy. The proposed method has experimented with prominent benchmark datasets developed for vessels tree, optic disc/cup, and arteries/veins extraction. It is also compared with other methods and achieved promising results. Our platform is helpful for physicians to trace the variations in the retina of subjects facing chronic retinal disorders. Muhammad Zubair Khan, Yugyung Lee, Arslan Munir, Muazzam Ali Khan |
BIBM | 3 |
| 2021 | Detecting Chronic Vascular Damage with Attention-Guided Neural SystemabstractThe retinal vasculature has a vital role in predicting chronic diabetic and hypertensive retinopathy. Recently, the advent of deep learning algorithms has brought a revolution in ocular disease prediction. The researchers frequently design complex and intricate techniques to efficiently segment vessels, micro-vessels and achieve better response on publicly available benchmark datasets. This article has designed an attention-guided neural system to extract vascular tree and distinguish it in arteries and veins. The proposed learning protocol with a minimalist approach can compete with state-of-the-art work without a performance compromise. Our method has achieved a promising response on numerous retinal image datasets. The pitfall of previously proposed work is also addressed through the self-defined assessment criteria. The in-depth analysis highlights that the underlying problem is unsolved for unseen data with different distribution than training. Our method is cross-validated to report the performance loss by keeping diversity in data selection. The technique is further applied for the arteries and veins extraction. Our effort can be adapted as an efficient vision-critical platform to scan and localize retinal damage and diagnose the disease symptoms early to prevent vision impairment. Muhammad Zubair Khan, Yugyung Lee, Arslan Munir, Muazzam Ali Khan |
BIBM | 3 |
| 2021 | Design and Evaluation of a Reconfigurable ECU Architecture for Secure and Dependable Automotive CPSabstractThe next generation of automobiles integrate a multitude of electronic control units (ECUs) to implement various automotive control and infotainment applications. However, recent works have demonstrated that these pervasively computerized modern automobiles are susceptible to security attacks that could compromise the physical safety of the driver and/or passengers. In this paper, we propose a novel ECU architecture for automotive cyber-physical systems (CPS) that simultaneously integrates both security and dependability primitives in the design with negligible performance, energy, and resources overhead. We implement our proposed ECU architecture on Xilinx Automotive (XA) Spartan-6 FPGA. We demonstrate the effectiveness of our proposed architecture using a steer-by-wire (SBW) application over controller area network (CAN) with flexible data rate (CAN FD) as a case study. We also optimize and implement a prior secure and dependable automotive work on NXP quad-core iMX6Q SABRE automotive board. We quantify the performance, energy, and error resilience of our proposed architecture for the SBW case study. Results reveal that our proposed architecture can attain a speedup of 47.9× while consuming 2.4× lesser energy than the optimized SABRE board implementation of security and dependability primitives. We further perform a comparative analysis of prior designs and the proposed ECU architecture for different in-vehicle networks, viz., CAN, CAN FD, and FlexRay. Results verify the feasibility as well as the superiority of the proposed ECU over other prior designs in terms of response time, energy efficiency, and error resilience. Bikash Poudel, Arslan Munir |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | OTMS: A Novel Radar-Based Mapping System for Automotive ApplicationsabstractCurrent automotive radar systems are designed to detect objects within an area around the vehicle and provide alerts or other driving assistance features to the driver. The main goal of these driving assistance systems is to increase safety of not only the driver but also increase safety for pedestrians on and around roadways. In this paper, we propose an on-board terrain mapping system (OTMS), which is an automotive radar-based system that extends the automotive safety goals to include mapping and vehicle-to-everything (V2X) communication. The proposed system utilizes a 94 GHz 3-dimensional (3D) pulsed radar to sense the environment around the vehicle and create a 3D model of the environment that can be sampled into a global database for mapping purposes. Objects are also detected and tracked within the model and can be used to provide assistive driving to the user even in adverse weather conditions. The model generated by the system can then be stored locally on the automobile or transmitted to other vehicles or infrastructure for use in mapping, navigation, and/or safety applications. The results validate that the OTMS is a valid system for providing safety and mapping functions for advanced driver assistance systems (ADAS). Justin Bode, Arslan Munir |
CCNC | 2 |
| 2020 | PUF-IPA: A PUF-based Identity Preserving Protocol for Internet of Things AuthenticationabstractPhysically unclonable functions (PUFs) can be used for Internet of things (IoT) based identification, authentication and authorization. However, PUF based authentication systems are vulnerable to various attacks including, but not limited to, replay and modeling attacks. In this paper, we propose PUF-IPA, a PUF-based identity-preserving protocol for IoT device authentication. The PUF-IPA provides stronger resilience against security attacks as compared to previous approaches assuming a threat model where adversary can conduct not only passive or active attacks during authentication phase but can also breach the server storing PUF credentials. The proposed PUF-IPA is robust against brute force, replay, and modeling attacks. In PUF-IPA, no partial/full challenge-response pairs (CRPs) or soft models associated to a PUF within a device are stored, generated, transmitted, or received by the server during authentication events. The PUF-IPA improves the PUF response accuracy by enabling self-checking. Results reveal that the PUF-IPA improves the PUF response accuracy from 89% to 98% without the use of hardware-expensive error correction codes. Mahmood Azhar Qureshi, Arslan Munir |
CCNC | 2 |
| 2020 | NeuroMAX: A High Throughput, Multi-Threaded, Log-Based Accelerator for Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) require high throughput hardware accelerators for real time applications owing to their huge computational cost. Most traditional CNN accelerators rely on single core, linear processing elements (PEs) in conjunction with 1D dataflows for accelerating convolution operations. This limits the maximum achievable ratio of peak throughput per PE count to unity. Most of the past works optimize their dataflows to attain close to a 100% hardware utilization to reach this ratio. In this paper, we introduce a high throughput, multi-threaded, log-based PE core. The designed core provides a 200% increase in peak throughput per PE count while only incurring a 6% increase in area overhead compared to a single, linear multiplier PE core with same output bit precision. We also present a 2D weight broadcast dataflow which exploits the multi-threaded nature of the PE cores to achieve a high hardware utilization per layer for various CNNs. The entire architecture, which we refer to as NeuroMAX, is implemented on Xilinx Zynq 7020 SoC at 200 MHz processing clock. Detailed analysis is performed on throughput, hardware utilization, area and power breakdown, and latency to show performance improvement compared to previous FPGA and ASIC designs. Mahmood Azhar Qureshi, Arslan Munir |
ICCAD | 2 |
| 2020 | Design and Analysis of Secure and Dependable Automotive CPS: A Steer-by-Wire Case StudyabstractThe next generation of automobiles (also known as cybercars) will increasingly incorporate electronic control units (ECUs) in novel automotive control applications. Recent work has demonstrated the vulnerability of modern car control systems to security attacks that directly impacts the cybercar's physical safety and dependability. In this paper, we provide an integrated approach for the design of secure and dependable automotive cyber-physical systems (CPS) using a case study: a steer-by-wire (SBW) application over controller area network (CAN). The challenge is to embed both security and dependability over CAN while ensuring that the real-time constraints of the automotive CPS are not violated. Our approach enables early design feasibility analysis of automotive CPS by embedding essential security primitives (i.e., confidentiality, integrity, and authentication) over CAN subject to the real-time constraints imposed by the desired quality of service and behavioral reliability. Our method leverages multicore ECUs for providing fault tolerance by redundant multi-threading (RMT) and also further enhances RMT for quick error detection and correction. We quantify the error resilience of our approach and evaluate the interplay of performance, fault tolerance, security, and scalability for our SBW case study. Arslan Munir, Farinaz Koushanfar |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2019 | TrolleyMod v1.0: An Open-Source Simulation and Data-Collection Platform for Ethical Decision Making in Autonomous VehiclesabstractThis paper presents TrolleyMod v1.0, an open-source platform based on the CARLA simulator for the collection of ethical decision-making data for autonomous vehicles. This platform is designed to facilitate experiments aiming to observe and record human decisions and actions in high-fidelity simulations of ethical dilemmas that occur in the context of driving. Targeting experiments in the class of trolley problems, TrolleyMod provides a seamless approach to creating new experimental settings and environments with the realistic physics-engine and the high-quality graphical capabilities of CARLA and the Unreal Engine. Also, TrolleyMod provides a straightforward interface between the CARLA environment and Python to enable the implementation of custom controllers, such as deep reinforcement learning agents. The results of such experiments can be used for sociological analyses, as well as the training and tuning of value-aligned autonomous vehicles based on social values that are inferred from observations. Vahid Behzadan, James Minton, Arslan Munir |
AIES | 3 |
| 2019 | PUF-RLA: A PUF-Based Reliable and Lightweight Authentication Protocol Employing Binary String ShufflingabstractPhysically unclonable functions (PUFs) can be employed for device identification, authentication, secret key storage, and other security tasks. However, PUFs are susceptible to modeling attacks if a number of PUFs' challenge-response pairs (CRPs) are exposed to the adversary. Furthermore, many of the embedded devices requiring authentication have stringent resource constraints and thus require a lightweight authentication mechanism. We propose PUF-RLA, a PUF-based lightweight, highly reliable authentication scheme employing binary string shuffling. The proposed scheme enhances the reliability of PUF as well as alleviates the resource constraints by employing error correction in the server instead of the device without compromising the security. The proposed PUF-RLA is robust against brute force, replay, and modeling attacks. In PUF-RLA, we introduce an inexpensive yet secure stream authentication scheme inside the device which authenticates the server before the underlying PUF can be invoked. This prevents an adversary from brute forcing the device's PUF to acquire CRPs essentially locking out the device from unauthorized model generation. Additionally, we also introduce a lightweight CRP obfuscation mechanism involving XOR and shuffle operations. Results and security analysis verify that the PUF-RLA is secure against brute force, replay, and modeling attacks, and provides 99% reliable authentication. In addition, PUF-RLA provides a reduction of 63% and 74% for look-up tables (LUTs) and register count, respectively, in FPGA compared to a recently proposed approach while providing additional authentication advantages. Mahmood Azhar Qureshi, Arslan Munir |
ICCD | 2 |
| 2019 | A Framework for Distributed Deep Neural Network Training with Heterogeneous Computing PlatformsabstractDeep neural network (DNN) training is generally performed by cloud computing platforms. However, cloud-based training has several problems such as network bottleneck, server management cost, and privacy. To overcome these problems, one of the most promising solutions is distributed DNN model training which trains the model with not only high-performance servers but also low-end power-efficient mobile edge or user devices. However, due to the lack of a framework which can provide an optimal cluster configuration (i.e., determining which computing devices participate in DNN training tasks), it is difficult to perform efficient DNN model training considering DNN service providers' preferences such as training time or energy efficiency. In this paper, we introduce a novel framework for distributed DNN training that determines the best training cluster configuration with available heterogeneous computing resources. Our proposed framework utilizes pre-training with a small number of training steps and estimates training time, power, energy, and energy-delay product (EDP) for each possible training cluster configuration. Based on the estimated metrics, our framework performs DNN training for the remaining steps with the chosen best cluster configurations depending on DNN service providers' preferences. Our framework is implemented in TensorFlow and evaluated with three heterogeneous computing platforms and five widely used DNN models. According to our experimental results, in 76.67% of the cases, our framework chooses the best cluster configuration depending on DNN service providers' preferences with only a small training time overhead. Bontak Gu, Joonho Kong, Arslan Munir, Younggeun Kim 0001 |
ICPADS | 3 |
| 2018 | Design and Evaluation of a PVT Variation-Resistant TRNG CircuitabstractOn-chip true random number generators (TRNGs) can be designed by using traditional CMOS technology or by using more recent nanoscale devices and technologies. In CMOS technology, TRNGs are designed by harnessing random physical variations (e.g., thermal/supply/telegraph noise, jitter, latch metastability, etc.). Since CMOS technologies are slowly saturating in development, more recently, nanoscale devices and technologies, such as memristor, magnetic tunnel junction, carbon nanotubes, graphene, etc., are being used to design TRNG circuits. An ideal TRNG circuit is expected to generate a sequence of random bits with very high bit-entropy and zero correlation among the generated bitstreams. However, increasing variations in the fabrication process and the sensitivity of transistors to operating conditions (e.g., voltage, and temperature (PVT)) have a significant impact on bit-entropy of TRNGs designed in deep nanometer technologies. Furthermore, PVT variations can be exploited by an adversary as effective tools to attack TRNGs. To mitigate these issues, we propose three probabilistic circuits: probability booster, probability dropper, and probability stabilizer. We use these circuits to build our proposed TRNG circuit. We also propose a stochastic model of our proposed TRNG circuit. To validate our proposed model, we have simulated our TRNG circuit using PSpice simulator using 65nm and 28nm processes. Results reveal that our proposed TRNG can generate random numbers with bit-entropy that always lies in the range [0.998, 1] at a data rate of 16 Mbps and is robust against PVT variations. Using the NIST SP 800-22 test suite for randomness, we demonstrate that the output of the proposed TRNG circuit is statistically random with 99% confidence levels. Bikash Poudel, Arslan Munir |
ICCD | 2 |
| 2018 | A Design Space Exploration Methodology for Parameter Optimization in Multicore ProcessorsabstractThe need for application-specific design of multicore/manycore processing platforms is evident with computing systems finding use in diverse application domains. In order to tailor multicore/manycore processors for application specific requirements, a multitude of processor design parameters have to be tuned accordingly which involves rigorous and extensive design space exploration over large search spaces. In this paper, we propose an efficient methodology for design space exploration. We evaluate our methodology over two search spaces small and large, using a cycle-accurate simulator (ESESC) and a standard set of PARSEC and SPLASH-2 benchmarks. For the smaller design space, we compare results obtained from our design space exploration methodology with results obtained from fully exhaustive search. The results show that solution quality obtained from our methodology are within 1.35 - 3.69 percent of the results obtained from fully exhaustive search while only exploring 2.74 - 3 percent of the design space. For larger design space, we compare solution quality of different results obtained by varying the number of tunable processor design parameters included in the exhaustive search phase of our methodology. The results show that including more number of tunable parameters in the exhaustive search phase of our methodology greatly improves solution quality. Prasanna Kansakar, Arslan Munir |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Design and comparative evaluation of GPGPU- and FPGA-based MPSoC ECU architectures for secure, dependable, and real-time automotive CPSabstractIn this paper, we propose and implement two electronic control unit (ECU) architectures for real-time automotive cyber-physical systems that incorporate security and dependability primitives with low resources and energy overhead. These ECUs architectures follow the multiprocessor system-on-chip (MPSoC) design paradigm wherein the ECUs have multiple heterogeneous processing engines with specific functionalities. The first architecture, GED, leverages an ARM-based application processor and a GPGPU-based co-processor. The second architecture, RED, integrates an ARM based application processor with a FPGA-based co-processor. We quantify and compare temporal performance, energy, and error resilience of our proposed architectures for a steer-by-wire case study over CAN, CAN FD, and FlexRay in-vehicle networks. Hardware implementation results reveal that RED and GED can attain a speedup of 31.7× and 1.8×, respectively, while consuming 1.75× and 2× less energy, respectively, than contemporary ECU architectures. Bikash Poudel, Naresh Kumar Giri, Arslan Munir |
ASAP | 3 |
| 2017 | Design and evaluation of a novel ECU architecture for secure and dependable automotive CPSabstractIn this paper, we propose a new architecture for automotive ECUs that incorporates security and dependability primitives with negligible performance, energy, and resources overhead. We implement our proposed ECU architecture on Xilinx Automotive (XA) Spartan-6 FPGA. We demonstrate the effectiveness of our proposed architecture using a steer-by-wire (SBW) application over controller area network with flexible data rate (CAN FD) as a case study that considers major security-relevant use cases and constraints. We also optimize and implement a prior secure and dependable automotive work on the NXP quad-core iMX6Q SABRE automotive board. We further quantify performance, energy, and error resilience of our proposed architecture for our SBW case study. Results reveal that our proposed architecture can attain a speedup of 47.93× while consuming 2.4× lesser energy than optimized SABRE board implementation. Bikash Poudel, Arslan Munir |
CCNC | 2 |
| 2017 | An efficient computation offloading architecture for the Internet of Things (IoT) devicesabstractProliferation of the connected Internet of things (IoT) devices and applications like augmented reality have resulted in a paradigm shift in computation requirement and power management of these devices. Furthermore, processing enormous amounts of data generated by ubiquitous IoT devices and meeting real-time deadline requirements of novel IoT applications exacerbate the challenges in IoT design. To address these challenges, in this paper, we propose a computation offloading architecture to process the huge amount of data generated by IoT devices while simultaneously meeting the real-time deadlines of IoT applications. In our proposed architecture, a resource-constrained IoT device requests a relatively resourceful computing device (e.g., a personal computer) in the same local network for computation offloading. Additionally, in our proposed computation offloading architecture, both client and server devices tune their tunable parameters, such as operating frequency and number of active cores, to meet the application's real-time deadline requirements. We compare our proposed computation offloading architecture with contemporary computation offloading models that use cloud computing. Experimental results verify that our proposed architecture provides a performance improvement of 21.4% on average as compared to cloud-based computation offloading schemes. Raj Mani Shukla, Arslan Munir |
CCNC | 2 |
| 2017 | Evolving side-channel resistant reconfigurable hardware for elliptic curve cryptographyabstractWe propose to use a genetic algorithm to evolve novel reconfigurable hardware to implement elliptic curve cryptographic combinational logic circuits. Elliptic curve cryptography offers high security-level with a short key length making it one of the most popular public-key cryptosystems. Furthermore, there are no known sub-exponential algorithms for solving the elliptic curve discrete logarithm problem. These advantages render elliptic curve cryptography attractive for incorporating in many future cryptographic applications and protocols. However, elliptic curve cryptography has proven to be vulnerable to non-invasive side-channel analysis attacks such as timing, power, visible light, electromagnetic, and acoustic analysis attacks. In this paper, we use a genetic algorithm to address this vulnerability by evolving combinational logic circuits that correctly implement elliptic curve cryptographic hardware that is also resistant to simple timing and power analysis attacks. Using a fitness function composed of multiple objectives - maximizing correctness, minimizing propagation delays and minimizing circuit size, we can generate correct combinational logic circuits resistant to non-invasive, side channel attacks. To the best of our knowledge, this is the first work to evolve a cryptography circuit using a genetic algorithm. We implement evolved circuits in hardware on a Xilinx Kintex-7 FPGA. Results reveal that the evolutionary algorithm can successfully generate correct, and side-channel resistant combinational circuits with negligible propagation delay. Bikash Poudel, Sushil J. Louis, Arslan Munir |
CEC | 3 |
| 2016 | D2CyberSoft: A design automation tool for soft error analysis of Dependable CybercarsabstractNext generation of automobiles (also known as cybercars) will escalate the proliferation of electronic control units (ECUs) to implement novel distributed control applications such as steer-by-wire (SBW), brake-by-wire, etc. Although the use of electronic embedded systems improves performance, driving comfort, safety, and economy for the customer; this electronic control of automotive systems makes these systems susceptible to soft errors, which can significantly reduce the availability of these systems. Enhancing reliability and availability at a minimum additional cost is a major challenge in automotive systems design. In this paper, we propose D2CyberSoft-a design automation tool for cybercars that facilitate cybercar designers in selecting designs with minimum cost overhead to make safety-critical automotive functions less vulnerable to soft errors. D2CyberSoft aids designers by providing built-in models, easy to specify inputs, and easy to interpret outputs. We elaborate Markov models that form the basis of D2CyberSoft using SBW as a case study. We further provide evaluation insights obtained from D2CyberSoft. Arslan Munir, Farinaz Koushanfar |
CCNC | 1 |
| 2016 | Design and performance analysis of secure and dependable cybercars: A steer-by-wire case studyabstractThe next generation of automobiles (also known as cybercars) will increasingly incorporate electronic control units (ECUs) in novel automotive control applications. Recent work has demonstrated vulnerability of modern car control systems to security attacks that directly impact the cybercar's physical safety and dependability. In this paper, we provide an integrated approach for the design of secure and dependable cybercars using a case study: a steer-by-wire (SBW) application over controller area network (CAN). The challenge is to embed both security and dependability over CAN while ensuring that the real-time constraints of the cybercar applications are not violated. Our approach enables early design feasibility analysis by embedding essential security primitives (i.e., confidentiality, integrity, and authentication) over CAN subject to the real-time constraints imposed by the desired quality of service and behavioral reliability. Our method leverages multi-core ECUs for providing fault-tolerance by redundant multi-threading (RMT) and also further enhances RMT for quick error detection. We quantify the error resilience of our approach and evaluate the interplay of performance, fault-tolerance, security, and scalability for our SBW case study. Arslan Munir, Farinaz Koushanfar |
CCNC | 1 |
| 2015 | Fine-Grained Voltage Boosting for Improving Yield in Near-Threshold Many-Core ProcessorsabstractProcess variation is a major impediment in optimizing yield, energy, and performance in near-threshold many-core processors. In this paper, we present a comprehensive analysis on yield losses in near-threshold many-core processors. Based on our analysis, we propose energy-efficient yield improvement techniques for near-threshold many-core processors: SRAM cell arrays and Wordline driver voltage Boosting (SWBoost) and Cache voltage Boosting (CBoost). Results reveal that SWBoost and CBoost improve a chip yield by up to 66% and 83%, respectively. Furthermore, runtime energy overheads of SWBoost and CBoost are only 0.46% and 0.54%, respectively, which are much lower than conventional voltage boosting techniques. Joonho Kong, Arslan Munir, Farinaz Koushanfar |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | Modeling and Analysis of Fault Detection and Fault Tolerance in Wireless Sensor NetworksabstractTechnological advancements in communications and embedded systems have led to the proliferation of Wireless Sensor Networks (WSNs) in a wide variety of application domains. These application domains include but are not limited to mission-critical (e.g., security, defense, space, satellite) or safety-related (e.g., health care, active volcano monitoring) systems. One commonality across all WSN application domains is the need to meet application requirements (e.g., lifetime, reliability). Many application domains require that sensor nodes be deployed in harsh environments, such as on the ocean floor or in an active volcano, making these nodes more prone to failures. Sensor node failures can be catastrophic for critical or safety-related systems. This article models and analyzes fault detection and fault tolerance in WSNs. To determine the effectiveness and accuracy of fault detection algorithms, we simulate these algorithms using ns-2. We investigate the synergy between fault detection and fault tolerance and use the fault detection algorithms’ accuracies in our modeling of Fault-Tolerant (FT) WSNs. We develop Markov models for characterizing WSN reliability and Mean Time to Failure (MTTF) to facilitate WSN application-specific design. Results obtained from our FT modeling reveal that an FT WSN composed of duplex sensor nodes can result in as high as a 100% MTTF increase and approximately a 350% improvement in reliability over a Non-Fault-Tolerant (NFT) WSN. The article also highlights future research directions for the design and deployment of reliable and trustworthy WSNs. Arslan Munir, Joseph Antoon, Ann Gordon-Ross |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2014 | D2Cyber: A design automation tool for dependable cybercarsabstractNext generation of automobiles (also known as cybercars) will increasingly incorporate electronic control units (ECUs) to implement various safety-critical functions such as x-by-wire (e.g., steer-by-wire (SBW), brake-by-wire). ISO 26262 specifies automotive safety integrity levels (ASILs) to signify the criticality associated with a function. Meeting a design's ASIL requirements at a minimum additional cost is a major challenge in cybercar design. In this paper, we propose D2Cyber - a design automation tool for cybercars that facilitates designers in selecting dependable designs by providing built-in models, easy to specify inputs, and easy to interpret outputs. D2Cyber considers the effects of temperature, electronics quality grade, and design lifetime in cybercar's design space exploration for determining a cost-effective solution and also advises on the attainable ASIL from a given design. We elaborate Markov models that form the basis of D2Cyber using SBW as a case study. We further provide evaluation insights obtained from D2Cyber. Arslan Munir, Farinaz Koushanfar |
DATE | 1 |
| 2014 | A queueing theoretic approach for performance evaluation of low-power multi-core embedded systems
Arslan Munir, Ann Gordon-Ross, Sanjay Ranka, Farinaz Koushanfar |
J. Parallel Distributed Comput. | 1 |
| 2014 | Multi-Core Embedded Wireless Sensor Networks: Architecture and ApplicationsabstractTechnological advancements in the silicon industry, as predicted by Moore's law, have enabled integration of billions of transistors on a single chip. To exploit this high transistor density for high performance, embedded systems are undergoing a transition from single-core to multi-core. Although a majority of embedded wireless sensor networks (EWSNs) consist of single-core embedded sensor nodes, multi-core embedded sensor nodes are envisioned to burgeon in selected application domains that require complex in-network processing of the sensed data. In this paper, we propose an architecture for heterogeneous hierarchical multi-core embedded wireless sensor networks (MCEWSNs) as well as an architecture for multi-core embedded sensor nodes used in MCEWSNs. We elaborate several compute-intensive tasks performed by sensor networks and application domains that would especially benefit from multi-core embedded sensor nodes. This paper also investigates the feasibility of two multi-core architectural paradigms-symmetric multiprocessors (SMPs) and tiled many-core architectures (TMAs)-for MCEWSNs. We compare and analyze the performance of an SMP (an Intel-based SMP) and a TMA (Tilera's TILEPro64) based on a parallelized information fusion application for various performance metrics (e.g., runtime, speedup, efficiency, cost, and performance per watt). Results reveal that TMAs exploit data locality effectively and are more suitable for MCEWSN applications that require integer manipulation of sensor data, such as information fusion, and have little or no communication between the parallelized tasks. To demonstrate the practical relevance of MCEWSNs, this paper also discusses several state-of-the-art multi-core embedded sensor node prototypes developed in academia and industry. We further discuss research challenges and future research directions for MCEWSNs. Arslan Munir, Ann Gordon-Ross, Sanjay Ranka |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | CyCAR'2013: first international academic workshop on security, privacy and dependability for cybervehiclesabstractThe next generation of automobiles (also known as CyberVehicles) will increasingly incorporate electronic control units in novel automotive control applications. Recent work has demonstrated vulnerability of modern automotive control systems to security attacks that directly impact CyberVehicles' physical safety and dependability. The First International Academic Workshop on Security, Privacy and Dependability for CyberVehicles (CyCAR'13) focuses on security and privacy topics in CyberVehicles that are within the scope of ACM Conference on Computer and Communications Security (CCS). Specifically, the workshop targets issues related to security and privacy issues in computerized, complex, and connected modern vehicles as well as their complex supply chains. This workshop offers an opportunity to trigger the transfer of the accumulated knowledge by the ACM CCS community to the car industry while taking into account typical automotive constraints such as interoperability, reliability, dependability, quality, resource constraints and/or complex supply chain. Arslan Munir, Farinaz Koushanfar, Hervé Seudie, Ahmad-Reza Sadeghi |
CCS | 1 |
| 2013 | High-performance optimizations on tiled many-core embedded systems: a matrix multiplication case study
Arslan Munir, Farinaz Koushanfar, Ann Gordon-Ross, Sanjay Ranka |
J. Supercomput. | 1 |
| 2012 | Online algorithms for wireless sensor networks dynamic optimizationabstractTechnological advancements in wireless communications and embedded systems have led to the proliferation of wireless sensor network (WSN) applications, each with varying application requirements (i.e., lifetime, throughput, reliability, etc.). Sensor node tunable parameters enable WSN designers to specialize/tune a sensor node to meet application requirements, but however, parameter tuning is a challenging process that requires designer expertise to consider sensor node complexities and changing environmental stimuli. In this paper, we develop lightweight, online optimization algorithms for sensor node parameter tuning, which enables dynamic optimizations to meet application requirements and adapt to changing environmental stimuli. Results reveal that our online optimizations quickly converge to a near optimal solution using minimal computational and storage resources, and are thus amenable for implementation on resource and energy-constrained sensor nodes. Arslan Munir, Ann Gordon-Ross, Susan Lysecky, Roman L. Lysecky |
CCNC | 1 |
| 2012 | Dynamic phase-based tuning for embedded systems using phase distance mappingabstractPhase-based tuning specializes a system's tunable parameters to the varying runtime requirements of an application's different phases of execution to meet optimization goals. Since the design space for tunable systems can be very large, one of the major challenges in phase-based tuning is determining the best configuration for each phase without incurring significant tuning overhead (e.g., energy and/or performance) during design space exploration. In this paper, we propose phase distance mapping, which directly determines the best configuration for a phase, thereby eliminating design space exploration. Phase distance mapping applies the correlation between a known phase's characteristics and best configuration to determine a new phase's best configuration based on the new phase's characteristics. Experimental results verify that our phase distance mapping approach determines configurations within 3% of the optimal configurations on average and yields an energy delay product savings of 26% on average. Tosiron Adegbija, Ann Gordon-Ross, Arslan Munir |
ICCD | 3 |
| 2012 | Parallelized benchmark-driven performance evaluation of SMPs and tiled multi-core architectures for embedded systemsabstractWith Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multi-core to exploit this high transistor density for high performance. However, there exists a plethora of multi-core architectures and the suitability of these multi-core architectures for different embedded domains (e.g., distributed, real-time, reliability-constrained) requires investigation. Despite the diversity of embedded domains, one of the critical applications in many embedded domains (especially distributed embedded domains) is information fusion. Furthermore, many other applications consist of various kernels, such as Gaussian elimination (used in network coding), that dominate the execution time. In this paper, we evaluate two embedded systems multi-core architectural paradigms: symmetric multiprocessors (SMPs) and tiled multi-core architectures (TMAs). We base our evaluation on a parallelized information fusion application and benchmarks that are used as building blocks in applications for SMPs and TMAs. We compare and analyze the performance of an Intel-based SMP and Tilera's TILEPro64 TMA based on our parallelized benchmarks for the following performance metrics: runtime, speedup, efficiency, cost, scalability, and performance per watt. Results reveal that TMAs are more suitable for applications requiring integer manipulation of data with little communication between the parallelized tasks (e.g., information fusion) whereas SMPs are more suitable for applications with floating point computations and a large amount of communication between processor cores. Arslan Munir, Ann Gordon-Ross, Sanjay Ranka |
IPCCC | 1 |
| 2012 | An MDP-Based Dynamic Optimization Methodology for Wireless Sensor NetworksabstractWireless sensor networks (WSNs) are distributed systems that have proliferated across diverse application domains (e.g., security/defense, health care, etc.). One commonality across all WSN domains is the need to meet application requirements (i.e., lifetime, responsiveness, etc.) through domain specific sensor node design. Techniques such as sensor node parameter tuning enable WSN designers to specialize tunable parameters (i.e., processor voltage and frequency, sensing frequency, etc.) to meet these application requirements. However, given WSN domain diversity, varying environmental situations (stimuli), and sensor node complexity, sensor node parameter tuning is a very challenging task. In this paper, we propose an automated Markov Decision Process (MDP)-based methodology to prescribe optimal sensor node operation (selection of values for tunable parameters such as processor voltage, processor frequency, and sensing frequency) to meet application requirements and adapt to changing environmental stimuli. Numerical results confirm the optimality of our proposed methodology and reveal that our methodology more closely meets application requirements compared to other feasible policies. Arslan Munir, Ann Gordon-Ross |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | High-Performance Energy-Efficient Multicore Embedded ComputingabstractWith Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multicore to exploit this high-transistor density for high performance. Embedded systems differ from traditional high-performance supercomputers in that power is a first-order constraint for embedded systems; whereas, performance is the major benchmark for supercomputers. The increase in on-chip transistor density exacerbates power/thermal issues in embedded systems, which necessitates novel hardware/software power/thermal management techniques to meet the ever-increasing high-performance embedded computing demands in an energy-efficient manner. This paper outlines typical requirements of embedded applications and discusses state-of-the-art hardware/software high-performance energy-efficient embedded computing (HPEEC) techniques that help meeting these requirements. We also discuss modern multicore processors that leverage these HPEEC techniques to deliver high performance per watt. Finally, we present design challenges and future research directions for HPEEC system development. Arslan Munir, Sanjay Ranka, Ann Gordon-Ross |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Markov Modeling of Fault-Tolerant Wireless Sensor NetworksabstractTechnological advancements in communications and embedded systems have led to the proliferation of wireless sensor networks (WSNs) in a wide variety of application domains. One commonality across all WSN application domains is the need to meet application requirements (e.g., lifetime, reliability, etc.). Many application domains require that sensor nodes be deployed in harsh environments (e.g., ocean floor, active volcanoes), making these sensor nodes more prone to failures. Unfortunately, sensor node failures can be catastrophic for critical or safety related systems. To improve reliability in such systems, we propose a fault-tolerant sensor node model for applications with high reliability requirements. We develop Markov models for characterizing WSN reliability and MTTF (Mean Time to Failure) to facilitate WSN application-specific design. Results show that our proposed fault-tolerant model can result in as high as a 100% MTTF increase and approximately a 350% improvement in reliability over a non-fault-tolerant WSN. Results also highlight the significance of a robust fault detection algorithm to leverage the benefits of fault-tolerant WSNs. Arslan Munir, Ann Gordon-Ross |
ICCCN | 1 |
| 2011 | A queueing theoretic approach for performance evaluation of low-power multi-core embedded systemsabstractWith Moore's law supplying billions of transistors on-chip, embedded systems are undergoing a transition from single-core to multi-core to exploit this high transistor density for high performance. However, the optimal layout of these multiple cores along with the memory subsystem (caches and main memory) to satisfy power, area, and often stringent real-time constraints is a challenging design endeavor. The short time-to-market constraint of embedded systems exacerbates this design challenge and necessitates the architectural modeling of embedded systems to reduce the time-to-market by expediting target applications to device/architecture mapping. In this paper, we present a queueing theoretic approach for modeling multi-core embedded systems that provides a quick and inexpensive performance evaluation both in terms of time and resources as compared to the development of multi-core simulators and running benchmarks on these simulators. We also calculate chip area and power consumption for different multi-core embedded architectures with a varying number of processor cores and cache configurations to provide a comparative analysis of multicore embedded architectures in terms of performance, area, and power consumption. Our performance and power results indicate that multi-core embedded system architectures that leverage shared last-level caches (LLCs) provide the best LLC performance per watt but may introduce main memory response time and throughput bottlenecks for high cache miss rates, whereas architectures leveraging a hybrid of private and shared LLCs alleviate main memory bottlenecks at the expense of reduced performance per watt. Arslan Munir, Ann Gordon-Ross, Sanjay Ranka |
ICCD | 1 |
| 2010 | A lightweight dynamic optimization methodology for wireless sensor networksabstractTechnological advancements in embedded systems due to Moore's law have lead to the proliferation of wireless sensor networks (WSNs) in different application domains (e.g. defense, health care, surveillance systems) with different application requirements (e.g. lifetime, reliability). Many commercial-off-the-shelf (COTS) sensor nodes can be specialized to meet these requirements using tunable parameters (e.g. voltage, frequency) to specialize the operating state. Since a sensor node's performance depends greatly on environmental stimuli, dynamic optimizations enable sensor nodes to automatically determine their operating state in-situ. However, dynamic optimization methodology development given a large design space and resource constraints (memory and computational) is a very challenging task. In this paper, we propose a lightweight dynamic optimization methodology that intelligently selects initial tunable parameter values to produce a high-quality initial operating state in one-shot for time-critical or highly constrained applications. Further operating state improvements are made using an efficient greedy exploration algorithm, achieving optimal or near-optimal operating states while exploring only 0.04% of the design space on average. Arslan Munir, Ann Gordon-Ross, Susan Lysecky, Roman L. Lysecky |
WiMob | 1 |
| 2010 | SIP-Based IMS Signaling Analysis for WiMax-3G Interworking ArchitecturesabstractThe third-generation partnership project (3GPP) and 3GPP2 have standardized the IP multimedia subsystem (IMS) to provide ubiquitous and access network-independent IP-based services for next-generation networks via merging cellular networks and the Internet. The application layer Session Initiation Protocol (SIP), standardized by 3GPP and 3GPP2 for IMS, is responsible for IMS session establishment, management, and transformation. The IEEE 802.16 worldwide interoperability for microwave access (WiMax) promises to provide high data rate broadband wireless access services. In this paper, we propose two novel interworking architectures to integrate WiMax and third-generation (3G) networks. Moreover, we analyze the SIP-based IMS registration and session setup signaling delay for 3G and WiMax networks with specific reference to their interworking architectures. Finally, we explore the effects of different WiMax-3G interworking architectures on the IMS registration and session setup signaling delay. Arslan Munir, Ann Gordon-Ross |
IEEE Trans. Mob. Comput. | 1 |
| 2007 | LCSCW2: an architecture for IP multimedia subsystemsabstractThe future fourth generation wireless heterogeneous networks aim to integrate various wireless access technologies and to support the IMS (IP multimedia subsystem) sessions. In this paper, we propose the Loosely Coupled Satellite-Cellular-WiMax-WLAN (LCSCW2) interworking architecture. The LCSCW2 architecture uses the loosely coupling approach and integrates the satellite networks, 3G wireless networks, WiMax, and WLANs. It can support IMS sessions and provide global coverage. The LCSCW2 architecture facilitates independent deployment and traffic engineering of various access networks. We also propose an analytical model to determine the associate cost for the signaling and data traffic for inter-system communication in the LCSCW2 architecture. The cost analysis includes the transmission, processing, and queueing costs at various entities. Numerical results are presented for different arrival rates and session lengths. Arslan Munir, Vincent W. S. Wong 0001 |
QSHINE | 1 |
| 2007 | Interworking Architectures for IP Multimedia Subsystems
Arslan Munir, Vincent W. S. Wong 0001 |
Mob. Networks Appl. | 1 |