EDBT 2026 Demo / reviewers in the wild / expert
Endri Taka
dblp:252/7612
· DBLP profile ↗
10ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0000-5136-7580ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen™ AI NPUsabstractThe high computational and memory demands of modern deep learning (DL) workloads have led to the development of specialized hardware devices from cloud to edge, such as AMD's Ryzen™ AI XDNA™ NPUs. Optimizing general matrix multiplication (GEMM) algorithms for these architectures is critical for improving DL workload performance. To this end, this paper presents a common systematic methodology to optimize GEMM workloads across the two current NPU generations, namely XDNA and XDNA2. Our implementations exploit the unique architectural features of AMD's NPUs and address key performance bottlenecks at the system level. End-to-end performance evaluation across various GEMM sizes demonstrates state-of-the-art throughput of up to 6.76 TOPS (XDNA) and 38.05 TOPS (XDNA2) for 8-bit integer (int8) precision. Similarly, for brain floating-point (bf16) precision, our GEMM implementations attain up to 3.14 TOPS (XDNA) and 14.71 TOPS (XDNA2). This work provides significant insights into key performance aspects of optimizing GEMM workloads on Ryzen AI NPUs. Endri Taka, André Rösti, Joseph Melber, Pranathi Vasireddy, Kristof Denolf, Diana Marculescu |
FPGA | 1 |
| 2026 | From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIRabstractGeneral-purpose compilers abstract away parallelism, locality, and synchronization, limiting their effectiveness on modern spatial architectures. As modern computing architectures increasingly rely on fine-grained control over data movement, execution order, and compute placement for performance, compiler infrastructure must provide explicit mechanisms for orchestrating compute and data to fully exploit such architectures. We introduce MLIR-AIR, a novel, open source compiler stack built on MLIR that bridges the semantic gap between high-level workloads and fine-grained spatial architectures such as AMD’s NPUs. MLIR-AIR defines the AIR dialect, which provides structured representations for asynchronous and hierarchical operations across compute and memory resources. AIR primitives allow the compiler to orchestrate spatial scheduling, distribute computation across hardware regions, and overlap communication with computation without relying on ad hoc runtime coordination or manual scheduling. We demonstrate MLIR-AIR’s capabilities through two case studies: matrix multiplication and the multi-head attention block from the LLaMA 2 model. For matrix multiplication, MLIR-AIR achieves up to 78.7% compute efficiency and generates implementations with performance almost identical to state-of-the-art, hand-optimized matrix multiplication written using the lower-level, close-to-metal MLIR-AIE framework. For multi-head attention, we demonstrate that the AIR interface supports fused implementations using approximately 150 lines of code, enabling tractable expression of complex workloads with efficient mapping to spatial hardware. MLIR-AIR transforms high-level structured control flow into spatial programs that efficiently utilize the compute fabric and memory hierarchy of an NPU, leveraging asynchronous execution, tiling, and communication overlap through compiler-managed scheduling. Erwei Wang, Samuel Bayliss, Andra Bisca, Zachary Blair, Sangeeta Chowdhary, Kristof Denolf, Jeff Fifield, Brandon Freiberger, Erika Hunhoff, Phil James-Roxby, Jack Lo, Joseph Melber, Stephen Neuendorffer, Eddie Richter, André Rösti, Javier Setoain, Gagandeep Singh 0002, Endri Taka, Pranathi Vasireddy, Zhewen Yu, Niansong Zhang, Jinming Zhuang |
ACM Trans. Reconfigurable Technol. Syst. | 18 |
| 2025 | Performance Analysis of GEMM Workloads on the AMD Versal PlatformabstractAMD Versal is a new heterogeneous computing hardware architecture comprised of adaptive intelligence (AI) engines, programmable logic, and a processing system. General Matrix Multiplication (GEMM) is the fundamental building block of modern deep learning (DL) applications such as ChatGPT, and GEMM workloads can be mapped onto Versal in different ways, each with distinct trade-offs. This paper presents a thorough analysis of GEMM workloads of different shapes and sizes, showcasing performance artifacts associated with the AMD Versal architecture. Focusing on the unique aspects of the Versal architecture, multiple research questions related to performance scaling, sensitivity, and efficiency are explored. This paper aims to assist developers in the FPGA community looking to implement GEMM on AMD Versal by providing guidelines and insights for enhancing performance. Kaustubh Manohar, Venkata Guru Prasanth Mulleti, Curt John Bansil, Endri Taka, Aman Arora 0001 |
FPGA | 4 |
| 2025 | Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI AccelerationabstractFPGA architectures have recently been enhanced to meet the substantial computational demands of modern deep neural networks (DNNs). To this end, both FPGA vendors and academic researchers have proposed in-fabric blocks that perform efficient tensor computations. However, these blocks are primarily optimized for dense computation, while most DNNs exhibit sparsity. To address this limitation, we propose incorporating structured sparsity support into FPGA architectures. We architect 2D systolic in-fabric blocks, named systolic sparse tensor (SST) slices, that support multiple degrees of sparsity to efficiently accelerate a wide variety of DNNs. SSTs support dense operation, 2:4 (50%) and 1:4 (75%) sparsity, as well as a new 1:3 (66.7%) sparsity level to further increase flexibility. When demonstrating on general matrix multiplication (GEMM) accelerators, which are the heart of most current DNN accelerators, our sparse SST-based designs attain up to 5× higher FPGA frequency and 10.9× lower area, compared to traditional FPGAs. Moreover, evaluation of the proposed SSTs on state-of-the-art sparse ViT and CNN models exhibits up to 3.52× speedup with minimal area increase of up to 13.3%, compared to dense in-fabric acceleration. Endri Taka, Ning-Chi Huang, Chi-Chih Chang, Kai-Chiang Wu, Aman Arora 0001, Diana Marculescu |
FPGA | 1 |
| 2025 | GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI EnginesabstractGeneral matrix-matrix multiplication (GEMM) is a fundamental operation in machine learning (ML) applications. We present the first comprehensive performance acceleration of GEMM workloads on AMD's second-generation AIE-ML architecture, which is specifically optimized for ML applications. Compared to AI-Engine (AIE), AIE-ML offers increased compute throughput and larger on-chip memory capacity. We propose a novel design that maximizes AIE-ML memory utilization, incorporates custom buffer placement within the AIE-ML and staggered kernel placement across the AIE-ML array, significantly reducing performance bottlenecks such as memory stalls and routing congestion, resulting in improved performance and efficiency compared to the default AMD's compiler. We evaluate the performance benefits of our design at three levels: single AIEML, pack of AIE-ML's and the complete AIE-ML array. GAMA achieves state-of-the-art performance, delivering up to$\mathbf{1 6 5}$TOPS (85% of peak) for int8 precision and 83 TBFLOPS (86% of peak) for bfloat16 precision GEMM workloads. Our solution achieves$8.7 \%, 9 \%, 39 \%$and 53.6 % higher peak throughput efficiency compared to the state-of-the-art AIE frameworks AMA, MAXEVA, ARIES and CHARM, respectively. Kaustubh Manohar, Endri Taka, Aman Arora 0001 |
FPL | 2 |
| 2025 | Performance Analysis of GEMM Workloads on the AMD Versal PlatformabstractAMD Versal is a new heterogeneous computing hardware architecture comprised of adaptive intelligence (AI) engines, programmable logic, and a processing system. General Matrix Multiplication (GEMM) is the fundamental building block of modern deep learning (DL) applications such as ChatGPT, and GEMM workloads can be mapped onto Versal in different ways, each with distinct trade-offs. This paper presents a thorough analysis of GEMM workloads of different shapes and sizes, showcasing performance artifacts associated with the AMD Versal architecture. Focusing on the unique aspects of the Versal architecture, multiple research questions related to performance scaling, sensitivity, and efficiency are explored. This paper aims to assist FPGA developers looking to implement GEMM on AMD Versal by providing insights for enhancing performance. Kaustubh Manohar, Venkata Guru Prashanth Mulleti, Curt John Bansil, Endri Taka, Aman Arora 0001 |
ISPASS | 4 |
| 2024 | Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAsabstractFPGAs are a promising platform for accelerating Deep Learning (DL) applications, due to their high performance, low power consumption, and reconfigurability. Recently, the leading FPGA vendors have enhanced their architectures to more efficiently support the computational demands of DL workloads. However, the two most prominent AI-optimized FPGAs, i.e., AMD/Xilinx Versal ACAP and Intel Stratix 10 NX, employ significantly different architectural approaches. This paper presents novel systematic frameworks to optimize the performance of General Matrix Multiplication (GEMM), a fundamental operation in DL workloads, by exploiting the unique and distinct architectural characteristics of each FPGA. Our evaluation on GEMM workloads for int8 precision shows up to 77 and 68 TOPs (int8) throughput, with up to 0.94 and 1.35 TOPs/W energy efficiency for Versal VC1902 and Stratix 10 NX, respectively. This work provides insights and guidelines for optimizing GEMM-based applications on both platforms, while also delving into their programmability trade-offs and associated challenges. Endri Taka, Dimitrios Gourounas, Andreas Gerstlauer, Diana Marculescu, Aman Arora 0001 |
FCCM | 1 |
| 2022 | Improving the performance of RISC-V softcores on FPGA by exploiting PVT variability and DVFSabstractImproving RISC-V processors becomes important in a plethora of applications, many of which rely exclusively on FPGA fabric to achieve custom HW/SW co-processing. Our approach is to improve softcore implementations by also accounting for the PVT peculiarities of each underlying FPGA. The proposed method bases on a custom DVFS technique to overcome PVT-induced guardbands, in-the-field. We evaluate the potential gains of an example RISC-V HDL core on Zynq MPSoC while varying multiple parameters, i.e., Voltage, Frequency, SW benchmarks, and RISC-V configurations. Our exploration indicates up to 75–149% throughput increase and/or 40% power decrease, vs STA, along with a need for careful tuning of RISC-V memory size. Endri Taka, George Lentaris, Dimitrios Soudris |
ISCAS | 1 |
| 2021 | Process Variability Analysis in Interconnect, Logic, and Arithmetic Blocks of 16-nm FinFET FPGAsabstractIn the current work, we study the process variability of logic, interconnect, and arithmetic/DSP resources in commercial 16-nm FPGAs. We create multiple, soft-macro sensors for each distinct resource under evaluation, and we deploy them across the FPGA fabric to measure intra-die variation, as well as across multiple FPGAs to measure inter-die variation. The derived results are used to create device-signature variability maps characterizing the distribution of variability across the die. Our study includes decoupling of variability to systematic and stochastic parts, exploration of variability under various voltage and temperature conditions and correlation analysis between the variability maps of the different resources. Furthermore, we scrutinize the impact of variability on the performance of actual test circuits and correlate the retrieved results with the sensor-based maps. Our experimental results on four Zynq XCZU7EV FPGAs showed significant intra- and inter-die variability, up to 7.8% and 8.9%, respectively, with a small increase under certain operating conditions. The correlation analysis demonstrated a strong correlation between the logic and arithmetic resources, whereas the interconnects showed a slightly weaker correlation in specific devices. Finally, a relatively moderate correlation was calculated between the variability maps and performance of test circuits due their dissimilar operating behavior versus our sensors. Endri Taka, Konstantinos Maragos 0001, George Lentaris, Dimitrios Soudris |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2019 | Analysis of Performance Variation in 16nm FinFET FPGA DevicesabstractProcess variability is a challenging fabrication issue impacting, mainly, the reliability and performance of chips. Variability is already present in current technology nodes and is expected to become even more significant in the future. In this work, we focus on the study of performance variation in 16nm FinFET FPGAs. We devise a comprehensive assessment methodology based on multiple programmable sensors with diverse resource and delay characteristics. Additionally, we consider various voltage and temperature conditions and decouple variability to systematic and stochastic. The experimental results on Zynq XCZU7EV show up to 7.3% intra-die variation increasing to 9.9% for certain operating conditions. Our approach demonstrates that logic and interconnect resources present different variability, slightly uncorrelated, which highlights the necessity and way towards more sophisticated mitigation methods/tools. Konstantinos Maragos 0001, Endri Taka, George Lentaris, Ioannis Stratakos, Dimitrios Soudris |
FPL | 2 |