Trevor E. Pogue

dblp:259/8396 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0002-6791-3758ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Karatsuba Matrix Multiplication and Its Efficient Custom Hardware Implementations
abstract
While the Karatsuba algorithm reduces the complexity of large integer multiplication, the extra additions required minimize its benefits for smaller integers of more commonly-used bitwidths. In this work, we propose the extension of the scalar Karatsuba multiplication algorithm to matrix multiplication, showing how this maintains the reduction in multiplication complexity of the original Karatsuba algorithm while reducing the complexity of the extra additions. Furthermore, we propose new matrix multiplication hardware architectures for efficiently exploiting this extension of the Karatsuba algorithm in custom hardware. We show that the proposed algorithm and hardware architectures can provide real area or execution time improvements for integer matrix multiplication compared to scalar Karatsuba or conventional matrix multiplication algorithms, while also supporting implementation through proven systolic array and conventional multiplier architectures at the core. We provide a complexity analysis of the algorithm and architectures and evaluate the proposed designs both in isolation and in an end-to-end accelerator system compared to baseline designs and prior state-of-the-art works implemented on the same type of compute platform, demonstrating their ability to increase the performance-per-area of matrix multiplication hardware.
Trevor E. Pogue, Nicola Nicolici
IEEE Trans. Computers1
2025 Strassen Multisystolic Array Hardware Architectures
abstract
While Strassen’s matrix multiplication algorithm reduces the complexity of naive matrix multiplication, general-purpose hardware is not suitable for achieving the algorithm’s promised theoretical speedups. This leaves the question of whether it could be better exploited in custom hardware architectures designed specifically for executing the algorithm. However, there is limited prior work on this and it is not immediately clear how to derive such architectures or whether they can ultimately lead to real improvements. We bridge this gap, presenting and evaluating new systolic array architectures that efficiently translate the theoretical complexity reductions of Strassen’s algorithm directly into hardware resource savings. Furthermore, the architectures are multisystolic array designs that can multiply smaller matrices with higher utilization than single-systolic array designs. The proposed designs implemented on FPGA reduce DSP requirements by a factor of$1.14^{r}$for r implemented Strassen recursion levels, and otherwise require overall similar soft logic resources when instantiated to support matrix sizes down to$32\times 32$and$24\times 24$at one to two levels of Strassen recursion, respectively. We evaluate the proposed designs in both isolation and an end-to-end machine learning accelerator compared with baseline designs and prior works, achieving state-of-the-art performance.
Trevor E. Pogue, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.1
2024 Fast Inner-Product Algorithms and Architectures for Deep Neural Network Accelerators
abstract
We introduce a new algorithm called the Free-pipeline Fast Inner Product (FFIP) and its hardware architecture that improve an under-explored fast inner-product algorithm (FIP) proposed by Winograd in 1968. Unlike the unrelated Winograd minimal filtering algorithms for convolutional layers, FIP is applicable to all machine learning (ML) model layers that can mainly decompose to matrix multiplication, including fully-connected, convolutional, recurrent, and attention/transformer layers. We implement FIP for the first time in an ML accelerator then present our FFIP algorithm and generalized architecture which inherently improve FIP's clock frequency and, as a consequence, throughput for a similar hardware cost. Finally, we contribute ML-specific optimizations for the FIP and FFIP algorithms and architectures. We show that FFIP can be seamlessly incorporated into traditional fixed-point systolic array ML accelerators to achieve the same throughput with half the number of multiply-accumulate (MAC) units, or it can double the maximum systolic array size that can fit onto devices with a fixed hardware budget. Our FFIP implementation for non-sparse ML models with 8 to 16-bit fixed-point inputs achieves higher throughput and compute efficiency than the best-in-class prior solutions on the same type of compute platform.
Trevor E. Pogue, Nicola Nicolici
IEEE Trans. Computers1
2020 Incremental Fault Analysis: Relaxing the Fault Model of Differential Fault Attacks
abstract
This article presents a new fault analysis technique against cryptographic devices called the incremental fault analysis (IFA), which can be adapted into fault attacks using more traditional differential fault analysis (DFA) techniques in order to increase their feasibility under more practical fault injection conditions. Many previous attack methods require precise fault injection techniques such as clock glitching. By contrast, IFA is compatible with a more practical overclocking fault injection technique in which a cryptosystem is stressed at a constant level throughout the entire encryption, and this constant stress level is then increased between consecutive encryptions. It is observed that as new faults occur incrementally between increased stress levels, they often become superimposed upon faults first appearing at lower stress levels. IFA exploits these incremental fault differentials to deduce the cipher key more rapidly. Attacks were tested using practical fault injection methods on the advanced encryption standard (AES) both with and without IFA applied. Using IFA, allowed cipher keys to be retrieved with a success rate of 100% from 10 times less faulty ciphertexts and 6.4 times less computational time, requiring 16, 86, and 43 ciphertexts on average for AES-128, AES-192, and AES-256, respectively.
Trevor E. Pogue, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.1