EDBT 2026 Demo / reviewers in the wild / expert
Tayyeb Mahmood
dblp:39/9350
· DBLP profile ↗
7ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-8853-305XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAPER-AI accelerator: a systolic array-based power-efficient reconfigurable AI acceleratorabstractDeep learning (DL) accelerators are critical for handling the growing computational demands of modern neural networks. Systolic array (SA)-based accelerators consist of a 2D mesh of processing elements (PEs) working cooperatively to accelerate matrix multiplication. The power efficiency of such accelerators is of primary importance, especially considering the edge AI regime. This work presents the SAPER-AI accelerator, an SA accelerator with power intent specified via a unified power format representation in a simplified manner with negligible microarchitectural optimization effort. Our proposed accelerator switches off rows and columns of PEs in a coarse-grained manner, thus leading to SA microarchitecture complying with the varying computational requirements of modern DL workloads. Our analysis demonstrates enhanced power efficiency ranging between 10% and 25% for the best case 32×32 and 64×64 SA designs, respectively. Additionally, the power delay product (PDP) exhibits a progressive improvement of around 6% for larger SA sizes. Moreover, a performance comparison between the MobileNet and ResNet50 models indicates generally better SA performance for the ResNet50 workload. This is due to the more regular convolutions portrayed by ResNet50 that are more favored by SAs, with the performance gap widening as the SA size increases. Fahad Bin Muslim, Kashif Inayat, Muhammad Zain Siddiqi, Safiullah Khan, Tayyeb Mahmood, Ihtesham Ul Islam |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2025 | AGD: Analytic Gradient Descent for Discrete Optimization in EDA and its Use to Gate SizingabstractIn electronic design automation (EDA), simulation models are often non-differentiable, and many design choices are discrete. As a result, greedy optimization methods based on numerical gradients are widely used, although they often lead to suboptimal solutions. In contrast, analytical methods may provide better solutions but require significant research effort. Reinforcement learning (RL) has been employed to address this problem; however, RL also suffers from notorious sample inefficiency, which is exaggerated in EDA because data sampling in EDA is very expensive due to slow simulations. This article proposes an alternative to RL for EDA, namely analytic gradient descent (AGD). Our method starts with a differentiable performance model, which can be either a learned surrogate or a static model. It then applies transformations similar to Shannon decomposition for each design variable in the performance model. Finally, one design option for each variable is selected using a one-hot variable, which is trained via a straight-through estimator (STE) through gradient descent. We demonstrate AGD on the well-known gate sizing problem using both a learned surrogate and a static model across 20 industrial benchmark circuits. Our experimental results show that the proposed method can outperform a several-decade-old commercial tool in the gate sizing task for 19 out of the 20 circuits. Phuoc Pham, Tae-Min Park 0001, Sung-Hyuk Cho, Tayyeb Mahmood, Joon-Sung Yang, Jaeyong Chung |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2024 | FPGA-assisted Design Space Exploration of Parameterized AI Accelerators: A Quickloop Approach
Kashif Inayat, Fahad Bin Muslim, Tayyeb Mahmood, Jaeyong Chung |
J. Syst. Archit. | 3 |
| 2023 | Quickloop: An Efficient, FPGA-Accelerated Exploration of Parameterized DNN AcceleratorsabstractQuickloop is a design-space exploration (DSE) framework of parameterized RTL generators, their software stack, and their simulation on FPGA. FPGAs are recently accelerating RTL simulations due to their rapid turnaround times (TAT), compared to ASIC. However, this TAT is still restrictive in DSE. We adopt a data-driven approach to optimize Quickloop's TAT and leverage this framework to extensively search the design space of an open source DNN accelerator. We show that our approach effectively slashes the TAT by above 30%, compared to conventional toolflow. Tayyeb Mahmood, Kashif Inayat, Jaeyong Chung |
PACT | 1 |
| 2015 | Ensuring Cache Reliability and Energy Scaling at Near-Threshold Voltage With MachoabstractNanoscale process variations in conventional SRAM cells are known to limit voltage scaling in microprocessor caches. Recently, a number of novel cache architectures have been proposed which substitute faulty words of one cache line with healthy words of others, to tolerate these failures at low voltages. These schemes rely on the fault maps to identify faulty words, inevitably increasing the chip area. Besides, the relationship between word sizes and the cache failure rates is not well studied in these works. In this paper, we analyze the word substitution schemes by employing Fault Tree Model and Collision Graph Model. A novel cache architecture (Macho) is then proposed based on this model. Macho is dynamically reconfigurable and is locally optimized (tailored to local fault density) using two algorithms: 1) a graph coloring algorithm for moderate fault densities and 2) a bipartite matching algorithm to support high fault densities. An adaptive matching algorithm enables on-demand reconfiguration of Macho to concentrate available resources on cache working sets. As a result, voltage scaling down to 400 mV is possible, tolerating bit failure rates reaching 1 percent (one failure in every 100 cells). This near-threshold voltage (NTV) operation achieves 44 percent energy reduction in our simulated system (CPU+DRAM models) with a 1 MB L2 cache. Tayyeb Mahmood, Seokin Hong, Soontae Kim |
IEEE Trans. Computers | 1 |
| 2013 | Macho: A failure model-oriented adaptive cache architecture to enable near-threshold voltage scalingabstractRecent interest in CMOS voltage scaling has produced a class of cache architectures which tolerate parametric SRAM failures at low voltage by substituting faulty words of one cache line with healthy words of another line. These caches rely on the fault maps (which grow reciprocally with smaller word sizes) for fault identification. Therefore, the benefits of cache voltage scaling must be rigorously investigated against the cost of their fault map overheads, especially in large caches. This paper reviews the word substitution caches and develops their parametric failure model. Our developed model leads to a non-intrusive and reconfigurable cache (Macho) which can be locally optimized (based on local fault density) by two graph-based algorithms. Specifically, our adaptive matching algorithm increases effective cache capacity by dynamically concentrating healthy cache blocks into active cache sets. Macho enables voltage scaling down to 400mV by tolerating high SRAM-failure rates (≥ 1%) and achieves better energy reduction (44%) than other substitution caches with similar area overheads. Tayyeb Mahmood, Soontae Kim, Seokin Hong |
HPCA | 1 |
| 2011 | Realizing near-true voltage scaling in variation-sensitive l1 caches via fault buffersabstractVoltage scaling can be applied to cache memories to reduce their energy consumptions. However, reduced supply voltage to the cache memories increases defective SRAM cells due to process variations, which will decrease their yields and performance nullifying the benefits of voltage scaling. To mitigate this problem, we propose a fault buffer-based scheme for L1 caches. Faults are identified and isolated at the granularity of individual words in the L1 caches. Actively used faulty cache words are allocated in the fault buffers dynamically. The fault buffers are organized as multiple banks for low cost implementation and can be reconfigured dynamically to reflect varying performance demands of programs. This dynamic scheme is shown to be more energy- and area-efficient than, and to be performing comparably to the previously proposed static schemes. Tayyeb Mahmood, Soontae Kim |
CASES | 1 |