Ali Zahir

dblp:87/8068 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0003-4657-4475ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 SAS: Speculative Locality Aware Scheduling for I/O intensive scientific analysis in clouds
Ali Zahir, Ashiq Anjum, Satish Narayana Srirama, Rajkumar Buyya
Future Gener. Comput. Syst.1
2023 Toward Optimal Softcore Carry-aware Approximate Multipliers on Xilinx FPGAs
abstract
Domain-specific accelerators for signal processing, image processing, and machine learning are increasingly being implemented on SRAM-based field-programmable gate arrays (FPGAs). Owing to the inherent error tolerance of such applications, approximate arithmetic operations, in particular, the design of approximate multipliers, have become an important research problem. Truncation of lower bits is a widely used approximation approach; however, analyzing and limiting the effects of carry-propagation due to this approximation has not been explored in detail yet. In this article, an optimized carry-aware approximate radix-4 Booth multiplier design is presented that leverages the built-in slice look-up tables (LUTs) and carry-chain resources in a novel configuration. The proposed multiplier simplifies the computation of the upper and lower bits and provides significant benefits in terms of FPGA resource usage (LUTs saving 38.5%–42.9%), Power Delay Product (PDP saving 49.4%–53%), performance metric (LUTs × critical path delay (CPD) × PDP saving 68.9%–73.1%) and errors (70% improvement in mean relative error distance) compared to the latest state-of-the-art designs. Therefore, the proposed designs are an attractive choice to implement multiplication on FPGA-based accelerators.
Muhammad Awais Khan 0002, Ali Zahir, Syed Ayaz Ali Shah, Pedro Reviriego, Anees Ullah, Nasim Ullah, Adam Khan, Hazrat Ali
ACM Trans. Embed. Comput. Syst.2
2021 Towards Low Latency and Resource-Efficient FPGA Implementations of the MUSIC Algorithm for Direction of Arrival Estimation
abstract
The estimation of the Direction of Arrival (DoA) is one of the most critical parameters for target recognition, identification and classification. MUltiple SIgnal Classification (MUSIC) is a powerful technique for DoA estimation. The algorithm requires complex mathematical operations like the computation of the covariance matrix for the input signals, eigenvalue decomposition and signal peak search. All these signal processing operations make real-time and resource-efficient implementation of the MUSIC algorithm on Field Programmable Gate Arrays (FPGAs) a challenge. In this paper, a novel design approach is proposed for the FPGA-implementation of the MUSIC algorithm. This approach enables a significant reduction in both FPGA resources and latency. In more detail, the proposed design enables the estimation of DoA in real-time scenarios in 2μsec with 30% to 50% fewer resources as compared to existing techniques.
Uzma M. Butt, Shoab A. Khan, Anees Ullah, Abdul Khaliq, Pedro Reviriego, Ali Zahir
IEEE Trans. Circuits Syst. I Regul. Pap.6
2020 A Survey of Interpretability of Machine Learning in Accelerator-based High Energy Physics
abstract
Data intensive studies in the domain of accelerator-based High Energy Physics, HEP, have become increasingly more achievable due to the emergence of machine learning with high-performance computing and big data technologies. In recent years, the intricate nature of physics tasks and data has prompted the use of more complex learning methods. To accurately identify physics of interest, and draw conclusions against proposed theories, it is crucial that these machine learning predictions are explainable. For it is not enough to accept an answer based on accuracy alone, but it is important in the process of physics discovery to understand exactly why an output was generated. That is, completeness of a solution is required. In this paper, we survey the application of machine learning methods to a variety of accelerator-based tasks in a bid to understand what role interpretability plays within this area. The main contribution of this paper is to promote the need for explainable artificial intelligence, XAI, for the future of machine learning in HEP.
Danielle Turvill, Lee Barnby, Bo Yuan 0004, Ali Zahir
BDCAT4
2020 FracTCAM: Fracturable LUTRAM-Based TCAM Emulation on Xilinx FPGAs
abstract
In this brief, we present FracTCAM, an efficient methodology for ternary content addressable memory (TCAM) emulation on Xilinx field-programmable gate arrays (FPGAs) by leveraging primitive architectural resources. The proposed methodology exploits the fracturable nature of lookup table random access memories (LUTRAMs) and built-in slice flip-flops for deeper pipelining. Multiple slices can be combined together to build deeper and wider TCAMs using ANDing operations. This results in TCAM implementations that achieve lower resources utilization, lower delay, and power consumption. A comparison with the existing schemes shows that FracTCAM consistently achieves the best performance per area (PA) and performance per area per watt (PAW).
Ali Zahir, Shadan Khan Khattak, Anees Ullah, Pedro Reviriego, Fahad Bin Muslim, Waleed Ahmad
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Multiple Hash Matching Units (MHMU): An Algorithmic Ternary Content Addressable Memory Design for Field Programmable Gate Arrays
abstract
As applications and user requirements are constantly evolving, there is a need to provide flexible networks that are able to process packets at high speed. One of the basic functions used for packet processing is matching a key formed by some fields of the incoming packet header against a set of stored rules. This is done for example to determine the next hop of a packet or to apply security checks on a firewall. In many cases, the stored rules have do not care bits as that enables a more flexible and compact representation of the rules. Therefore, the matching can be done in hardware using Ternary Content Addressable Memories (TCAMs). However, TCAMs pose several problems in many implementations. For example, for ASICs they require much more circuit area and power than standard SRAMs. On the other hand, designs based on programmable logic such as Field Programmable Gate Arrays (FPGAs) can only use the blocks provided by the FPGA that do not typically include TCAMs. In this last case, a TCAM can be emulated using the FPGA logic resources but with a large cost. To reduce the cost of implementing TCAMs, a number of algorithmic solutions have been proposed and are known as Algorithmic TCAMs or A-TCAMs. Most of those schemes target either software or ASIC implementations. In this paper we present Multiple Hash Matching Units (MHMU) an A-TCAM solution targeted towards FPGA implementations. The proposed scheme exploits the massive parallelism of FPGAs to implement many hash based matching units that use the embedded block RAM memories of the FPGA. The proposed MHMU scheme has been mapped to a Xilinx series 7 FPGA to check its efficiency in terms of resource usage and its scalability. To validate the effectiveness of MHMU, a simple configuration has been tested with Classbench generated sets of rules. The results show that the MHMU is able to consistently accommodate sets with several tens of thousands of rules with large keys.
Pedro Reviriego, Salvatore Pontarelli, Anees Ullah, Ali Zahir, Giuseppe Bianchi 0001
HPSR4
2010 Enhanced Generic Information Services Using Mobile Messaging
Muhammad Saleem 0002, Ali Zahir, Yasir Ismail, Bilal Saeed
GPC2