Ayesha Afzal

dblp:220/2370 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Meta-learning meets transformers: A novel approach to enterprise network intrusion detection
Ali Haider Khan, Kaleem Razzaq Malik, Ayesha Afzal, Jianqiang Li 0002
Expert Syst. Appl.4
2026 Exploring metrics for analyzing dynamic behavior in MPI programs via a coupled-oscillator model
abstract
We propose a novel, lightweight, and physically inspired approach to modeling the dynamics of parallel distributed-memory programs. Inspired by the Kuramoto model, we represent MPI processes as coupled oscillators with topology-aware interactions, custom coupling potentials, and stochastic noise. The resulting system of nonlinear ordinary differential equations opens a path to modeling key performance phenomena of parallel programs, including synchronization, delay propagation and decay, bottlenecks, and self-desynchronization. This paper introduces interaction potentials to describe memory- and compute-bound workloads and employs multiple quantitative metrics – such as an order parameter, synchronization entropy, phase gradients, and phase differences – to evaluate phase coherence and disruption. We also investigate the role of local noise and show that moderate noise can accelerate resynchronization in scalable applications. Our simulations align qualitatively with MPI trace data, showing the potential of physics-informed abstractions to predict performance patterns, which offers a new perspective for performance modeling and software-hardware co-design in parallel computing.
Ayesha Afzal, Georg Hager, Gerhard Wellein
Parallel Comput.1
2024 Demo: Orchflow: Orchestration and Management of IoT-Centric Distributed Workflows
abstract
The evolution of edge and cloud computing infrastructures has opened avenues for developing Internet-centered distributed applications characterized by adaptability, evolvability, and emergence. These applications, such as knowledge-driven distributed workflows are dynamically orchestrated and managed using resources across enterprise networks, cloud data centers, and Internet of Things (IoT) devices. Unlike traditional business processes and scientific workflows, knowledge-driven workflows dynamically evolve and adapt by responding to the environmental context, current execution status, and specific parameters of the case at hand, which are not predictable beforehand. In this demonstration, we present the Orchflow system, designed for dynamic orchestration and management of adaptive IoT-centric workflows. Orchflow enables flexible and location-aware resource selection, iterative and incremental binding, and dynamic deployment of services by considering factors such as spatio-temporal requirements of workflow tasks, resource availability status, underlying service infrastructure constraints, and real-time processing requirements within the workflow.
Sehrish Amjad, Ahmed Akhtar, Basit Shafiq, Shafay Shamail, Ayesha Afzal, Jaideep Vaidya
ICDCS6
2024 Orchestration and Management of Adaptive IoT-Centric Distributed Applications
abstract
Current Internet of Things (IoT) devices provide a diverse range of functionalities, ranging from measurement and dissemination of sensory data observation, to computation services for real-time data stream processing. In extreme situations such as emergencies, a significant benefit of IoT devices is that they can help gain a more complete situational understanding of the environment. However, this requires the ability to utilize IoT resources while taking into account location, battery life, and other constraints of the underlying edge and IoT devices. A dynamic approach is proposed for orchestration and management of distributed workflow applications using services available in cloud data centers, deployed on servers, or IoT devices at the network edge. Our proposed approach is specifically designed for knowledge-driven business process workflows that are adaptive, interactive, evolvable and emergent. A comprehensive empirical evaluation shows that the proposed approach is effective and resilient to situational changes.
Sehrish Amjad, Ahmed Akhtar, Ayesha Afzal, Basit Shafiq, Jaideep Vaidya, Shafay Shamail, Omer F. Rana
IEEE Internet Things J.4
2024 Blockchain Based Auditable Access Control for Business Processes With Event Driven Policies
abstract
The use of blockchain technology has been proposed to provide auditable access control for individual resources. Unlike the case where all resources are owned by a single organization, this work focuses on distributed applications such as business processes and distributed workflows. These applications are often composed of multiple resources/services that are subject to the security and access control policies of different organizational domains. Here, blockchains provide an attractive decentralized solution to provide auditability. However, the underlying access control policies may have event-driven constraints and can be overlapping in terms of the component conditions/rules as well as events. Existing work cannot handle event-driven constraints and does not sufficiently account for overlaps leading to significant overhead in terms of cost and computation time for evaluating authorizations over the blockchain. In this work, we propose an automata-theoretic approach for generating a cost-efficient composite access control policy. We reduce this composite policy generation problem to the standard weighted set cover problem. We show that the composite policy correctly captures all the local access control policies and reduces the policy evaluation cost over the blockchain. We have implemented the initial prototype of our approach using Ethereum as the underlying blockchain and empirically validated the effectiveness and efficiency of our approach. Ablation studies were conducted to determine the impact of changes in individual service policies on the overall cost.
Ahmed Akhtar, Masoud Barati, Basit Shafiq, Omer F. Rana, Ayesha Afzal, Jaideep Vaidya, Shafay Shamail
IEEE Trans. Dependable Secur. Comput.5
2023 Making applications faster by asynchronous execution: Slowing down processes or relaxing MPI collectives
Ayesha Afzal, Georg Hager, Stefano Markidis, Gerhard Wellein
Future Gener. Comput. Syst.1
2023 RL-IoT: Reinforcement Learning-Based Routing Approach for Cognitive Radio-Enabled IoT Communications
abstract
Internet of Things (IoT) devices are widely being used in various smart applications and being equipped with cognitive radio (CR) capabilities for dynamic spectrum allocation. Our objectives in this work are to achieve higher data rates and minimize end-to-end routing delays in CR-enabled IoT communication in order to maximize throughput. We propose a reinforcement learning (RL)-based routing approach in the cognitive radio network (CRN)-based IoT environment. The idea is to add the channel selection decision capability to the network layer in order to minimize packet collisions as well as end-to-end delay (EED). We perform a comprehensive performance evaluation of the proposed RL-IoT routing mechanism by simulating the cognitive radio-enabled Internet of Things (CR-IoT) communication environment in the cognitive radio cognitive network (CRCN) simulator and comparing the network performance achieved by our proposed mechanism with that of the recent AODV-based routing mechanism for IoT (AODV-IoT), ELD-CRN, and SpEED-IoT routing approaches. Our evaluation results show that the RL-IoT model performs better than existing approaches in terms of average data rate, throughput, packet collision, and EED.
Tauqeer Safdar Malik, Kaleem Razzaq Malik, Ayesha Afzal, Lei Wang 0005, Houbing Song, Nadir Shah
IEEE Internet Things J.3
2023 The Role of Idle Waves, Desynchronization, and Bottleneck Evasion in the Performance of Parallel Programs
abstract
The performance of highly parallel applications on distributed-memory systems is influenced by many factors. Analytic performance modeling techniques aim to provide insight into performance limitations and are often the starting point of optimization efforts. However, coupling analytic models across the system hierarchy (socket, node, network) fails to encompass the intricate interplay between the program code and the hardware, especially when execution and communication bottlenecks are involved. In this paper we investigate the effect ofbottleneck evasionand how it can lead to automatic overlap of communication overhead with computation. Bottleneck evasion leads to a gradual loss of the initial bulk-synchronous behavior of a parallel code so that its processes become desynchronized. This occurs most prominently in memory-bound programs, which is why we choose memory-bound benchmark and application codes, specifically an MPI-augmented STREAM Triad, sparse matrix-vector multiplication, and a collective-avoiding Chebyshev filter diagonalization code to demonstrate the consequences of desynchronization on two different supercomputing platforms. We investigate the role of idle waves as possible triggers for desynchronization and show the impact of automatic asynchronous communication for a spectrum of code properties and parameters, such as saturation point, matrix structures, domain decomposition, and communication concurrency. Our findings reveal how eliminating synchronization points (such as collective communication or barriers) precipitates performance improvements that go beyond what can be expected by simply subtracting the overhead of the collective from the overall runtime.
Ayesha Afzal, Georg Hager, Gerhard Wellein
IEEE Trans. Parallel Distributed Syst.1
2023 Collaborative Business Process Fault Resolution in the Services Cloud
abstract
The emergence of cloud and edge computing has enabled rapid development and deployment of Internet-centric distributed applications. There are many platforms and tools that can facilitate users to develop distributed business process (BP) applications by composing relevant service components in a plug and play manner. However, there is no guarantee that a BP application developed in this way is fault-free. In this paper, we formalize the problem of collaborative BP fault resolution which aims to utilize information from existing fault-free BPs that use similar services to resolve faults in a user developed BP. We present an approach based on association analysis of pairwise transformations between a faulty BP and existing BPs to identify the smallest possible set of transformations to resolve the fault(s) in the user developed BP. An extensive experimental evaluation over both synthetically generated faulty BPs and real BPs developed by users shows the effectiveness of our approach.
Muhammad Adeel Zahid, Basit Shafiq, Jaideep Vaidya, Ayesha Afzal, Shafay Shamail
IEEE Trans. Serv. Comput.4
2022 BP-DEBUG: A Fault Debugging and Resolution Tool for Business Processes
abstract
Cloud computing and Internet-ware software paradigm have enabled rapid development of distributed business process (BP) applications. Several tools are available to facilitate automated/ semi-automated development and deployment of such distributed BPs by orchestrating relevant service components in a plug-and-play fashion. However, the BPs developed using such tools are not guaranteed to be fault-free. In this demonstration, we present a tool called BP-DEBUG for debugging and automated repair of faulty BPs. BP-DEBUG implements our Collaborative Fault Resolution (CFR) approach that utilizes the knowledge of existing BPs with a similar set of web services fault detection and resolution in a given user BP. Essentially, CFR attempts to determine any semantic and structural differences between a faulty BP and related BPs and computes a minimum set of transformations which can be used to repair the faulty BP. Demo url: https://youtu.be/mf49oSekLOA.
Muhammad Adeel Zahid, Basit Shafiq, Shafay Shamail, Ayesha Afzal, Jaideep Vaidya
ICDCS4
2022 An Integrated Framework for Fault Resolution in Business Processes
abstract
Cloud and edge-computing based platforms have enabled rapid development of distributed business process (BP) applications in a plug and play manner. However, these platforms do not provide the needed capabilities for identifying or repairing faults in BPs. Faults in BP may occur due to errors made by BP designers because of their lack of understanding of the underlying component services, misconfiguration of these services, or incorrect/incomplete BP workflow specifications. Such faults may not be discovered at design or development stage and may occur at runtime. In this paper, we present a unified framework for automated fault resolution in BPs. The proposed framework employs a novel and efficient fault resolution approach that extends the generate-and-validate program repair approach. In addition, we propose a hybrid approach that performs fault resolution by analyzing a faulty BP in isolation as well as by comparing with other BPs using similar services. This hybrid approach results in improved accuracy and broader coverage of fault types. We also perform an extensive experimental evaluation to compare the effectiveness of the proposed approach using a dataset of 208 faulty BPs.
Muhammad Adeel Zahid, Ahmed Akhtar, Basit Shafiq, Shafay Shamail, Ayesha Afzal, Jaideep Vaidya
ICWS5
2022 Addressing White-box Modeling and Simulation Challenges in Parallel Computing
abstract
No abstract available.
Ayesha Afzal, Gerhard Wellein, Georg Hager
SIGSIM-PADS1
2022 Analytic performance model for parallel overlapping memory-bound kernels
abstract
Abstract Complex applications running on multicore processors show a rich performance phenomenology. The growing number of cores per ccNUMA domain complicates performance analysis of memory‐bound code since system noise, load imbalance, or task‐based programming models can lead to thread desynchronization. Hence, the simplifying assumption that all cores execute the same loop can not be upheld. Motivated by observations on plain and modified versions of the HPCG benchmark, we construct a performance model of execution of memory‐bound loop kernels. It can predict the memory bandwidth share per kernel on a memory contention domain depending on the number of active cores and which other workload the kernel is paired with. The only code features required are the single‐thread memory request fraction per kernel, which is directly related to the single‐thread memory bandwidth, and its saturated bandwidth. The former can either be measured directly or predicted using the Execution‐Cache‐Memory performance model. The computational intensity of the kernels and the detailed structure of the code is of no significance. We validate our model on Intel Broadwell, Intel Cascade Lake, and AMD Rome processors pairing various streaming and stencil kernels. The error in predicting the bandwidth share per kernel is less than 8%.
Ayesha Afzal, Georg Hager, Gerhard Wellein
Concurr. Comput. Pract. Exp.1
2022 A Framework for Dynamic Composition and Management of Emergency Response Processes
abstract
An emergency response process outlines the workflow of different activities that need to be performed in response to an emergency. Effective emergency response requires communication and coordination with the operational systems belonging to different collaborating organizations. Therefore, it is necessary to establish information sharing and system-level interoperability among the diverse operational systems. Unlike typical e-government processes that are well structured and have a well-defined outcome, emergency response processes are knowledge-centric and their workflow structure and execution may evolve as the incident unfolds. It is impractical to define static plans and response process workflows for every possible situation. Instead, a dynamic response should be adaptable to the changing situation. We present an integrated approach that facilitates the dynamic composition of an executable response process. The proposed approach employs ontology-based reasoning to determine the default actions and resource requirements for the given incident and to identify relevant response organizations based on their jurisdictional and mutual aid agreement rules. The Web service APIs of the identified response organizations are then used to generate an executable response process that evolves dynamically. The proposed approach is implemented and experimentally validated using an example scenario derived from the FEMA Hazardous Materials Tabletop Exercises Manual.
Abeer Elahraf, Ayesha Afzal, Ahmed Akhtar, Basit Shafiq, Jaideep Vaidya, Shafay Shamail, Nabil R. Adam
IEEE Trans. Serv. Comput.2
2021 ASSEMBLE: Attribute, Structure and Semantics Based Service Mapping Approach for Collaborative Business Process Development
abstract
Development of a Business Process (BP) is a challenging task for small and medium enterprises (SMEs) which often do not have adequate resources for design, coding, and management of their BPs. Knowledge of existing BPs of related organizations can be exploited for collaborative BP development. However, syntactic and semantic heterogeneity among the Web service operations of BPs across organizations is a major obstacle to such collaborative BP development. In this paper, we propose an approach for collaborative BP development that exploits the attribute and structural similarity of related BPs as well as the semantic information including preconditions and postconditions of operations, to compute a mapping between the available service operations of the user organization and the BP operations of other organizations. We experimentally evaluate the approach with real world data from e-commerce sales BPs and demonstrate its effectiveness.
Ayesha Afzal, Basit Shafiq, Shafay Shamail, Abeer Elahraf, Jaideep Vaidya, Nabil R. Adam
IEEE Trans. Serv. Comput.1
2020 BP-Com: A Service Mapping Tool for Rapid Development of Business Processes
abstract
Business Process (BP) composition is a challenging task for small and medium organizations that do not have sufficient resources for design, coding, and management of their BPs. Cloud infrastructure and service-oriented middleware can be leveraged for rapid development and deployment of BPs of such organizations. BP development in the cloud-based environment can be done by exploiting the knowledge of existing BPs of related organizations. In this demonstration, we present the BP- Com tool which is a Web-based interactive system that enables efficient development of BPs in the cloud. BP-Com implements our service mapping approach called ASSEMBLE that utilizes the attribute, structural and semantics information of service operations of existing BPs in a given domain to help a user organization to compose its BP. Given a collection of related BPs and available service operations of a user organization, BP-Com computes a mapping between the available service operations of the user organization and the BP operations of other organizations. The results of operation mapping are presented to the user for refinement and customization of the generated BP workflow. Executable BP code is then generated in standard BPEL language, which can be deployed on any process execution engine on the user organization's site or on the cloud.
Ayesha Afzal, Muhammad Adeel Zahid, Ahmad Akhtar, Basit Shafiq, Shafay Shamail, Abeer Elahraf, Jaideep Vaidya, Nabil R. Adam
ICDCS1
2020 Blockchain Based Auditable Access Control for Distributed Business Processes
abstract
The use of blockchain technology has been proposed to provide auditable access control for individual resources. However, when all resources are owned by a single organization, such expensive solutions may not be needed. In this work we focus on distributed applications such as business processes and distributed workflows. These applications are often composed of multiple resources/services that are subject to the security and access control policies of different organizational domains. Here, blockchains can provide an attractive decentralized solution to provide auditability. However, the underlying access control policies may be overlapping in terms of the component conditions/rules, and simply using existing solutions would result in repeated evaluation of user's authorization separately for each resource, leading to significant overhead in terms of cost and computation time over the blockchain. To address this challenge, we propose an approach that formulates a constraint optimization problem to generate an optimal composite access control policy. This policy is in compliance with all the local access control policies and minimizes the policy evaluation cost over the blockchain. The developed smart contract(s) can then be deployed to the blockchain, and used for access control enforcement. We also discuss how the access control enforcement can be audited using a game-theoretic approach to minimize cost. We have implemented the initial prototype of our approach using Ethereum as the underlying blockchain and experimentally validated the effectiveness and efficiency of our approach.
Ahmad Akhtar, Basit Shafiq, Jaideep Vaidya, Ayesha Afzal, Shafay Shamail, Omer F. Rana
ICDCS4
2019 Propagation and Decay of Injected One-Off Delays on Clusters: A Case Study
abstract
Analytic, first-principles performance modeling of distributed-memory applications is difficult due to a wide spectrum of random disturbances caused by the application and the system. These disturbances (commonly called “noise”) run contrary to the assumptions about regularity that one usually employs when constructing simple analytic models. Despite numerous efforts to quantify, categorize, and reduce such effects, a comprehensive quantitative understanding of their performance impact is not available, especially for long, one-off delays of execution periods that have global consequences for the parallel application. In this work, we investigate various traces collected from synthetic benchmarks that mimic real applications on simulated and real message-passing systems in order to pin-point the mechanisms behind delay propagation. We analyze the dependence of the propagation speed of “idle waves,” i.e., propagating phases of inactivity, emanating from injected delays with respect to the execution and communication properties of the application, study how such delays decay under increased noise levels, and how they interact with each other. We also show how fine-grained noise can make a system immune against the adverse effects of propagating idle waves. Our results contribute to a better understanding of the collective phenomena that manifest themselves in distributed-memory parallel applications.
Ayesha Afzal, Georg Hager, Gerhard Wellein
CLUSTER1
2019 ClusterCockpit - A web application for job-specific performance monitoring
abstract
Monitoring is a common component of HPC system software. Up to now, monitoring focused mainly on health checking and system level performance as well as on job scheduler information and was targeted towards system administrators. Recently job-specific performance monitoring based on hardware performance counter metrics has gained attention at academic HPC computing centers. HPC is becoming a mainstream tool that is also used by non-HPC experts, and HPC centers see a demand to check for pathological jobs and jobs with large optimization potential. The possibility to measure hardware performance counter data with negligible overhead allows assessment of efficient resource utilization and detection of pathological jobs. Pathological jobs are, e.g. jobs with errors in the batch script, jobs which do not terminate, jobs with severe load imbalance, or jobs that do not use any resources. This paper introduces ClusterCockpit, a web front-end tailor-made tool for job-specific performance monitoring. While many recent job-specific performance monitoring efforts concentrate on the measurement and data collection layers, ClusterCockpit provides a modern user interface targeted towards performance analysts as well as application users.
Jan Eitzinger, Thomas Gruber 0007, Ayesha Afzal, Thomas Zeiser, Gerhard Wellein
CLUSTER3
2018 Solving Maxwell's Equations with Modern C++ and SYCL: A Case Study
abstract
In scientific computing, unstructured meshes are a crucial foundation for the simulation of real-world physical phenomena. Compared to regular grids, they allow resembling the computational domain with a much higher accuracy, which in turn leads to more efficient computations. There exists a wealth of supporting libraries and frameworks that aid programmers with the implementation of applications working on such grids, each built on top of existing parallelization technologies. However, many approaches require the programmer to introduce a different programming paradigm into their application or provide different variants of the code. SYCL is a new programming standard providing a remedy to this dilemma by building on standard C++ 17 with its so-called single-source approach: Programmers write standard C++ code and expose parallelism using C++ 17 keywords. The application is then transformed into a concrete implementation by the SYCL implementation. By encapsulating the OpenCL ecosystem, different SYCL implementations enable not only the programming of CPUs but also of heterogeneous platforms such as GPUs or other devices. For the first time, this paper showcases a SY CL-based solver for the nodal Discontinuous Galerkin method for Maxwell's equations on unstructured meshes. We compare our solution to a previous C-based implementation with respect to programmability and performance on heterogeneous platforms.
Ayesha Afzal, Christian Schmitt 0003, Samer Alhaddad, Yevgen Grynko, Jürgen Teich, Jens Förstner, Frank Hannig
ASAP1
2018 OpenCL-Based FPGA Design to Accelerate the Nodal Discontinuous Galerkin Method for Unstructured Meshes
abstract
The exploration of FPGAs as accelerators for scientific simulations has so far mostly been focused on small kernels of methods working on regular data structures, for example in the form of stencil computations for finite difference methods. In computational sciences, often more advanced methods are employed that promise better stability, convergence, locality and scaling. Unstructured meshes are shown to be more effective and more accurate, compared to regular grids, in representing computation domains of various shapes. Using unstructured meshes, the discontinuous Galerkin method preserves the ability to perform explicit local update operations for simulations in the time domain. In this work, we investigate FPGAs as target platform for an implementation of the nodal discontinuous Galerkin method to find time-domain solutions of Maxwell's equations in an unstructured mesh. When maximizing data reuse and fitting constant coefficients into suitably partitioned on-chip memory, high computational intensity allows us to implement and feed wide data paths with hundreds of floating point operators. By decoupling off-chip memory accesses from the computations, high memory bandwidth can be sustained, even for the irregular access pattern required by parts of the application. Using the Intel/Altera OpenCL SDK for FPGAs, we present different implementation variants for different polynomial orders of the method. In different phases of the algorithm, either computational or bandwidth limits of the Arria 10 platform are almost reached, thus outperforming a highly multithreaded CPU implementation by around 2x.
Tobias Kenter, Gopinath Mahale, Samer Alhaddad, Yevgen Grynko, Christian Schmitt 0003, Ayesha Afzal, Frank Hannig, Jens Förstner, Christian Plessl
FCCM6