EDBT 2026 Demo / reviewers in the wild / expert
Afzal Ahmad
dblp:209/1558
· DBLP profile ↗
10ranked-venue papers
7as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAPO: Design Structure-Aware Pass Ordering for HLS via Contrastive and Reinforcement LearningabstractHigh-Level Synthesis (HLS) tools are widely adopted in FPGA-based domain-specific accelerator design. However, existing tools rely on fixed optimization strategies inherited from software compilations, limiting their effectiveness. Tailoring optimization strategies to specific designs requires deep semantic understanding, accurate hardware metric estimation, and advanced search algorithms - capabilities that current approaches lack.We propose DAPO, a design structure-aware pass ordering framework that extracts program semantics from control and data flow graphs, employs contrastive learning to generate rich embeddings, and leverages an analytical model for accurate hardware metric estimation. These components jointly guide a reinforcement learning agent to discover design-specific optimization strategies. Evaluations on standard HLS benchmarks demonstrate that our end-to-end flow delivers 1.67× speedup on pragma-free designs and a 2.36× speedup on designs with pragmas over Vitis HLS with comparable resource usage. Jinming Ge, Linfeng Du, Likith Anaparty, Shangkun Li, Tingyuan Liang, Afzal Ahmad, Vivek Chaturvedi, Sharad Sinha, Zhiyao Xie, Jiang Xu 0001, Wei Zhang 0012 |
DATE | 6 |
| 2025 | Automated Design Space Exploration in High-Level Physical SynthesisabstractImplementing HLS accelerators on large-scale multi-die FPGAs presents significant challenges. To address this, researchers have proposed High-Level Physical Synthesis (HLPS), which co-optimizes high-level synthesis and physical design to improve achievable frequency. However, existing HLPS techniques suffer from unstable and inconsistent quality of results (QoRs), largely due to the vast number of parameters that need to be selected by the user in an ad-hoc way. As a result, achieving satisfactory solutions still requires substantial manual effort and expertise in low-level circuit design.We propose a robust and practical design space exploration (DSE) framework that enhances the reliability and QoRs of HLPS by automating the iterative parameter tuning process. Informed by metrics extracted from physical implementation outcomes, the framework applies tailored heuristics to refine HLPS parameters, enabling consistent and automated timing closure. In evaluations with large-scale, real-world designs implemented on representative multi-die devices, our framework achieves an average frequency of 311.06 MHz, reaching 2.42× the frequency of the AMD Vitis/Vivado toolchain (128.48 MHz) and 1.67× that of the leading academic solutions (186.21 MHz). Linfeng Du, Jason Lau, Yuze Chi, Yutong Xie 0011, Chunyou Su, Afzal Ahmad, Zifan He, Jake Ke, Jinming Ge, Jason Cong, Wei Zhang 0012, Licheng Guo |
ICCAD | 7 |
| 2024 | Accel-NASBench: Sustainable Benchmarking for Accelerator-Aware NASabstractOne of the primary challenges impeding the progress of Neural Architecture Search (NAS) is its extensive reliance on exorbitant computational resources. NAS benchmarks aim to simulate runs of NAS experiments at zero cost, remediating the need for extensive compute. However, existing NAS benchmarks use synthetic datasets and model proxies that make simplified assumptions about the characteristics of these datasets and models, leading to unrealistic evaluations. We present a technique that allows searching for training proxies that reduce the cost of benchmark construction by significant margins, making it possible to construct realistic NAS benchmarks for large-scale datasets. Using this technique, we construct an open-source bi-objective NAS benchmark for the ImageNet2012 dataset combined with the on-device performance of accelerators, including GPUs, TPUs, and FPGAs. Through extensive experimentation with various NAS optimizers and hardware platforms, we show that the benchmark is accurate and allows searching for state-of-the-art hardware-aware models at zero cost. Afzal Ahmad, Linfeng Du, Zhiyao Xie, Wei Zhang 0012 |
DAC | 1 |
| 2024 | Fast and Practical Strassen's Matrix Multiplication using FPGAsabstractMatrix multiplication is a cornerstone operation in a wide array of scientific fields, including machine learning and computer graphics. The standard algorithm for matrix multiplication has a complexity of $O\left(n^{3}\right)$ for $n \times n$ matrices. Strassen’s algorithm improves this to $O\left(n^{2.807}\right)$, but its practicality is limited for small to medium matrix sizes due to the large number of additions it introduces. This paper presents a novel FPGA-based implementation of Strassen’s algorithm that achieves superior speed over an optimized General Matrix Multiply (GeMM) implementation for matrices as small as n = 256. Our design, tested extensively on two high-performance FPGA accelerators (Alveo U50 and U280) across various data types, matches or surpasses the performance of a highly optimized baseline across a range of matrix sizes. Afzal Ahmad, Linfeng Du |
FPL | 1 |
| 2023 | PertNAS: Architectural Perturbations for Memory-Efficient Neural Architecture SearchabstractDifferentiable Neural Architecture Search (NAS) relies on aggressive weight-sharing to reduce its search cost. This leads to GPU-memory bottlenecks that hamper the algorithm’s scalability. To resolve these bottlenecks, we propose a perturbations-based evolutionary approach that significantly reduces the memory cost while largely maintaining the efficiency benefits of weight-sharing. Our approach makes minute changes to compact neural architectures and measures their impact on performance. In this way, it extracts high-quality motifs from the search space. We utilize these perturbations to perform NAS in compact models evolving over time to traverse the search space. Our method disentangles GPU-memory consumption from search space size, offering exceptional scalability to large search spaces. Results show competitive accuracy on multiple benchmarks, including CIFAR10, ImageNet2012, and NASBench-301. Specifically, our approach improves accuracy on ImageNet and NASBench-301 by 0.3% and 0.87%, respectively. Furthermore, the memory consumption of search is reduced by roughly 80% against state-of-the-art weight-shared differentiable NAS works while achieving a search time of only 6 GPU hours. Afzal Ahmad, Zhiyao Xie, Wei Zhang 0012 |
DAC | 1 |
| 2021 | Autonomous Aerial Swarming in GNSS-denied Environments with High Obstacle DensityabstractThe compact flocking of relatively localized Un-manned Aerial Vehicles (UAVs) in high obstacle density areas is discussed in this paper. The presented work tackles realistic scenarios in which the environment map is not known apriori and the use of a global localization system and communication infrastructure is difficult due to the presence of obstacles. To achieve flocking in such a constrained environment, we propose a fully decentralized, bio-inspired control law that uses only onboard sensor data for safe flocking through the environment without any communication with other agents. In the proposed approach, each UAV agent uses onboard sensors to self-localize and estimate the relative position of other agents in its local reference frame. The usability and performance of the proposed approach were verified and evaluated using various experiments in a realistic robotic simulator and a natural forest. The presented experiments also validate the utility of onboard relative localization for autonomous multi-UAV applications in the absence of global localization information and communication. Afzal Ahmad, Viktor Walter, Pavel Petrácek, Matej Petrlík, Tomás Báca, David Zaitlík, Martin Saska |
ICRA | 1 |
| 2020 | Accelerating Tiny YOLOv3 using FPGA-Based Hardware/Software Co-DesignabstractConvolutional Neural Networks (CNNs) are influencing major breakthroughs in computer vision by achieving unprecedented accuracy on tasks such as image classification, object detection, landmark detection and semantic segmentation. Owing to high computational complexity of most modern CNN architectures, graphical processing units (GPUs) are being utilized to achieve real-time performance albeit at a high energy cost. Consequently, Field Programmable Gate Arrays (FPGAs) based hardware accelerators are also making their way as they demonstrate GPU-like performance with significantly lower energy consumption that is well-suited for embedded vision applications. In this paper, we employ Hardware/Software Co-Design approach to accelerate Tiny YOLOv3 - an efficient CNN architecture for object detection - by designing a hardware accelerator for convolution, the most complex operation involved in the CNNs. Experimental results show significant performance gains, in the range of 3.9× to 21.3×, over previous implementations of efficient object detection algorithms. Afzal Ahmad, Muhammad Adeel Pasha, Ghulam Jilani Raza |
ISCAS | 1 |
| 2020 | FFConv: An FPGA-based Accelerator for Fast Convolution Layers in Convolutional Neural NetworksabstractImage classification is known to be one of the most challenging problems in the domain of computer vision. Significant research is being done on developing systems and algorithms improving accuracy, performance, area, and power consumption for related problems. Convolutional Neural Networks (CNNs) have shown to give outstanding accuracies for problems such as image classification, object detection, and semantic segmentation. While CNNs are pioneering the development of high accuracy systems, their excessive computational complexity presents a barrier for a more permeated deployment. Although Graphical Processing Units (GPUs), due to their massively parallel architecture, have shown to give performance orders of magnitude better than general purpose processors, the former are limited by their high power consumption and generality. Consequently, Field Programmable Gate Arrays (FPGAs) are being explored to implement CNN architectures, as they also provide massively parallel logic resources but with a relatively lower power consumption than GPUs. In this article, we present FFConv, an efficient FPGA-based fast convolutional layer accelerator for CNNs. We design a pipelined, high-throughput convolution engine based on the Winograd minimal filtering (also called Fast Convolution) algorithms for computing the convolutional layers of three popular CNN architectures: VGG16, Alexnet, and Shufflenet. We implement our accelerator on a Virtex-7 FPGA platform where we exploit the computational parallelization to the maximum while exploring optimizations aimed at improving performance. The resultant design loses only 0.43%, 0.47%, and 0.61% Top-1 classification accuracy for VGG16, Alexnet, and Shufflenet-v1, respectively, while significantly improving throughput, resource, and power efficiency compared to previous state-of-the-art designs. Afzal Ahmad, Muhammad Adeel Pasha |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2019 | Towards Design Space Exploration and Optimization of Fast Algorithms for Convolutional Neural Networks (CNNs) on FPGAsabstractConvolutional Neural Networks (CNNs) have gained widespread popularity in the field of computer vision and image processing. Due to huge computational requirements of CNNs, dedicated hardware-based implementations are being explored to improve their performance. Hardware platforms such as Field Programmable Gate Arrays (FPGAs) are widely being used to design parallel architectures for this purpose. In this paper, we analyze Winograd minimal filtering or fast convolution algorithms to reduce the arithmetic complexity of convolutional layers of CNNs. We explore a complex design space to find the sets of parameters that result in improved throughput and power-efficiency. We also design a pipelined and parallel Winograd convolution engine that improves the throughput and power-efficiency while reducing the computational complexity of the overall system. Our proposed designs show up to 4.75× and 1.44× improvements in throughput and power-efficiency, respectively, in comparison to the state-of-the-art design while using approximately 2.67× more multipliers. Furthermore, we obtain savings of up to 53.6% in logic resources compared with the state-of-the-art implementation. Afzal Ahmad, Muhammad Adeel Pasha |
DATE | 1 |
| 2017 | Multiple Beacon Based Robust Cooperative Spectrum Sensing in MIMO Cognitive Radio Networks under CSI UncertaintyabstractThis paper presents multiple beacon vectors based robust detection schemes for cooperative spectrum sensing (CSS) in multiple-input multiple-output (MIMO) cognitive radio (CR) networks under channel state information (CSI) uncertainty. The inaccuracies in the estimate of the CSI are modeled as the standard ellipsoidal uncertainty set. We develop a multiple beacon vector based linear discriminant framework to obtain robust detectors for the problem of primary user detection in MIMO cognitive radio networks under ellipsoidal CSI uncertainty. Next, we employ this framework to develop signaling scheme based application specific detectors, namely the antipodal signaling based robust detector and on-off signaling based robust detector, along with their closed form expressions. Further, for the two signalling schemes, we even present allied detectors that are advantageous under low signal-to-noise ratio (SNR) conditions. Simulation results demonstrate a superior detection performance of the proposed robust detection schemes in comparison to the uncertainty agnostic matched filter detector for CSS in MIMO cognitive radio networks. Adarsh Patel, Afzal Ahmad, Rajeev Tripathi |
VTC Spring | 2 |