Pierre Abillama

dblp:268/1796 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-9523-6152ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accelerating Block Low-Rank Foundation Model Inference on Memory-Constrained GPUs
Pierre Abillama, Changwoo Lee 0001, Juechu Dong, David T. Blaauw, Dennis Sylvester, Hun-Seok Kim
HPDC1
2025 MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
abstract
Transformers have achieved state-of-the-art performance across various tasks, but suffer from a notable quadratic complexity in sequence length due to the attention mechanism. In this work, we propose MonarchAttention -- a novel approach to sub-quadratic attention approximation via Monarch matrices, an expressive class of structured matrices. Based on the variational form of softmax, we describe an efficient optimization-based algorithm to compute an approximate projection of softmax attention onto the class of Monarch matrices with $\Theta(N\sqrt{N} d)$ computational complexity and $\Theta(Nd)$ memory/IO complexity. Unlike previous approaches, MonarchAttention is both (1) transferable, yielding minimal performance loss with no additional training, even when replacing every attention layer of the transformer, and (2) hardware-efficient, utilizing the highest-throughput tensor core units on modern GPUs. With optimized kernels, MonarchAttention achieves substantial speed-ups in wall-time over FlashAttention-2: $1.4\times$ for shorter sequences $(N=256)$, $4.5\times$ for medium-length sequences $(N=4K)$, and $8.2\times$ for longer sequences $(N=16K)$. We demonstrate the quality of MonarchAttention on diverse tasks and architectures in vision and language problems, showing that it flexibly and accurately approximates softmax attention in a variety of contexts.
Can Yaras, Alec S. Xu, Pierre Abillama, Changwoo Lee 0001, Laura Balzano
NeurIPS3
2023 SONA: An Accelerator for Transform-Domain Neural Networks with Sparse-Orthogonal Weights
abstract
Recent advances in model pruning have enabled sparsity-aware deep neural network accelerators that improve the energy-efficiency and performance of inference tasks. We introduce SONA, a novel transform-domain neural network accelerator in which convolution operations are replaced by element-wise multiplications with sparse-orthogonal weights. SONA employs an output stationary dataflow coupled with an energy-efficient memory organization to reduce the overhead of sparse-orthogonal transform-domain kernels that are concurrently processed without any conflicts. Weights in SONA are non-uniformly quantized with bit-sparse canonical-signed-digit representations to reduce multiplications to simple additions. Moreover, for sparse fully-connected layers (FCLs), SONA introduces column-based-block structured pruning, which is integrated into the same architecture that maintains full multiply-and-accumulate (MAC) array utilization. Compared to prior dense and sparse neural networks accelerators, SONA can reduce inference energy by$5.1\times$and$2.4 \times$and increase performance by$5.2\times$and$2.1\times$, respectively, for convolution layers. For sparse FCLs, SONA can reduce inference energy by$2.4\times$and increase performance by$2\times$compared to prior work.
Pierre Abillama, Zichen Fan, Yu Chen 0070, Hyochan An, Qirui Zhang 0001, Seungkyu Choi, David T. Blaauw, Dennis Sylvester, Hun-Seok Kim
ASAP1
2023 TaskFusion: An Efficient Transfer Learning Architecture with Dual Delta Sparsity for Multi-Task Natural Language Processing
abstract
The combination of pre-trained models and task-specific fine-tuning schemes, such as BERT, has achieved great success in various natural language processing (NLP) tasks. However, the large memory and computation costs of such models make it challenging to deploy them in edge devices. Moreover, in real-world applications like chatbots, multiple NLP tasks need to be processed together to achieve higher response credibility. Running multiple NLP tasks with specialized models for each task increases the latency and memory cost latency linearly with the number of tasks. Though there have been recent works on parameter-shared tuning that aim to reduce the total parameter size by partially sharing weights among multiple tasks, computation remains intensive and redundant despite different tasks using the same input. In this work, we identify that a significant portion of activations and weights can be reused among different tasks, to reduce cost and latency for efficient multi-task NLP. Specifically, we propose TaskFusion, an efficient transfer learning software-hardware co-design that exploits delta sparsity in both weights and activations to boost data sharing among tasks. For training, TaskFusion uses ℓ1 regularization on delta activation to learn inter-task data redundancies. A novel hardware-aware sub-task inference algorithm is proposed to exploit the dual delta sparsity. We then designed a dedicated heterogeneous architecture to accelerate multi-task inference with an optimized scheduling to increase hardware utilization and reduce off-chip memory access. Extensive experiments demonstrate that TaskFusion can reduce the number of floating point operations (FLOPs) by over 73% in multi-task NLP with negligible accuracy loss, while adding a new task at the cost of only < 2% parameter size increase. With the proposed architecture and optimized scheduling, Task-Fusion can achieve 1.48--2.43× performance and 1.62--3.77× energy efficiency than those using state-of-the-art single-task accelerators for multi-task NLP applications.
Zichen Fan, Qirui Zhang 0001, Pierre Abillama, Sara Shoouri, Changwoo Lee 0001, David T. Blaauw, Hun-Seok Kim, Dennis Sylvester
ISCA3
2022 Approximate Constant-Coefficient Multiplication Using Hybrid Binary-Unary Computing for FPGAs
abstract
Multipliers are used in virtually all Digital Signal Processing (DSP) applications such as image and video processing. Multiplier efficiency has a direct impact on the overall performance of such applications, especially when real-time processing is needed, as in 4K video processing, or where hardware resources are limited, as in mobile and IoT devices. We propose a novel, low-cost, low energy, and high-speed approximate constant coefficient multiplier (CCM) using a hybrid binary-unary encoding method. The proposed method implements a CCM using simple routing networks with no logic gates in the unary domain, which results in more efficient multipliers compared to Xilinx LogiCORE IP CCMs and table-based KCM CCMs (Flopoco) on average. We evaluate the proposed multipliers on 2-D discrete cosine transform algorithm as a common DSP module. Post-routing FPGA results show that the proposed multipliers can improve the {area, area × delay, power consumption, and energy-delay product} of a 2-D discrete cosine transform on average by {30%, 33%, 30%, 31%}. Moreover, the throughput of the proposed 2-D discrete cosine transform is on average 5% more than that of the binary architecture implemented using table-based KCM CCMs. We will show that our method has fewer routability issues compared to binary implementations when implementing a DCT core.
S. Rasoul Faraji, Pierre Abillama, Kia Bazargan
ACM Trans. Reconfigurable Technol. Syst.2
2021 HTNN: Deep Learning in Heterogeneous Transform Domains with Sparse-Orthogonal Weights
abstract
Convolutional neural networks (CNNs) achieved great success on various tasks in recent years. Their applications to low power and low cost hardware platforms, however, have been often limited due to extensive complexity of convolution layers. We present a new class of transform domain deep neural networks (DNNs), where convolution operations are replaced by element-wise multiplications in heterogeneous transform domains. To further reduce the network complexity, we propose a framework to learn sparse-orthogonal weights in heterogeneous transform domains co-optimized with a hardware-efficient accelerator architecture to minimize the overhead of handling sparse weights. Furthermore, sparse-orthogonal weights are non-uniformly quantized with canonical-signed-digit (CSD) representations to substitute multiplications with simpler additions. The proposed approach reduces the complexity by a factor of 4.9– 6.8 $\times$ without compromising the DNN accuracy compared to equivalent CNNs that employ sparse (pruned) weights. The code is available at https://github.com/unchenyu/HTNN.
Yu Chen 0070, Bowen Liu 0001, Pierre Abillama, Hun-Seok Kim
ISLPED3
2020 Low-Cost Approximate Constant Coefficient Hybrid Binary-Unary Multiplier for DSP Applications
abstract
Multipliers are used in virtually all Digital Signal Processing (DSP) applications, such as image and video processing. Multiplier efficiency has a direct impact on the overall performance of such applications, especially when real-time processing is needed, as in 4K video processing, or where hardware resources are limited, as in mobile and IoT devices. We propose a novel, low-cost, low energy and high-speed approximate constant coefficient multiplier (CCM) using a hybrid binary-unary encoding method. The proposed method implements a CCM using simple routing networks with no logic gates in the unary domain, which results in more efficient multipliers compared to Xilinx LogiCORE IP CCMs and table-based KCM CCMs on average. We evaluate the proposed multipliers on 2-D discrete cosine transform and fast Fourier transform algorithms as two common DSP modules. Post-routing FPGA results show that the proposed multipliers can improve the {area × delay cost, and energy consumption per sample} of 8-bit fixed-point 2-D discrete cosine transform on average by {33%, 36%}. The improvement for 128-point 16-bit fixed-point fast Fourier transform on these metrics is {45%, 54%}. Moreover, the throughput of the proposed 2-D discrete cosine transform and 128-point fast Fourier transform architectures are on average 1.04× and 1.76× of the throughput of the binary architectures implemented using Xilinx LogiCORE IP CCMs, respectively.
S. Rasoul Faraji, Pierre Abillama, Kia Bazargan
FCCM2
2020 HBUCNNA: Hybrid Binary-Unary Convolutional Neural Network Accelerator
abstract
Convolutional layers account for 90% of the total computational power of Convolutional Neural Networks (CNNs). Field programmable gate arrays (FPGAs) have shown great potential for accelerating inference tasks in CNNs. However, it is harder for FPGA platforms to deliver the best performance due to high computational power and high memory bandwidth requirements of today's CNNs. In this paper, we propose a reconfigurable parallel-pipelined Hybrid Binary-Unary CNN Accelerator (HBUCNNA) to implement low-cost, high-performance convolutional layers of a ResNet-18 architecture. We use the hybrid binary-unary method to implement banks of constant-coefficient multipliers, which are used to implement convolutional kernels. Moreover, we propose hybrid binary-unary batch normalization units to further improve the total hardware costs. These two units reduce {area, area×delay} costs by {50%, 30%} and {44%, 65%} on average compared to their conventional binary counterparts, respectively. The proposed accelerator stores control signals for reconfigurability instead of the numeric value of weights, which in turn reduces the memory footprint on average by 20%. Overall, the proposed HBUCNNA architecture reduces the {area, latency, power, energy, area×delay} costs on average by {25.5%, 40%, 15%, 40%, 47%} and {53%, 61%, 47%, 62%, 67%} compared to the constant-coefficient multiplier-based and variable size multiplier-based binary architectures, respectively. Moreover, the proposed accelerator improves the throughput by about 1.4 × compared to both of the mentioned architectures.
S. Rasoul Faraji, Pierre Abillama, Gaurav Singh 0008, Kia Bazargan
ISCAS2