Esmerald Aliaj

dblp:227/8129 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2024
0009-0008-0625-4833ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Durin: CPU-FPGA Heterogeneous Platform for Scalable Low-Dimensional Data Clustering
abstract
Reconfigurable hardware accelerators, known for their high performance and power efficiency, have yet to be fully leveraged for clustering low-dimensional data at realistic scales. In this work, we identify and address two hurdles for accelerating this important class of applications: the overhead of implementing indexing data structures in hardware, and the PCIe bottleneck when the data capacity spills over from accelerator memory to host storage. We overcome these hurdles for the first time with Durin, a CPU-FPGA heterogeneous system with a hardware-software codesigned index structure, which minimizes the hardware resource overhead of high-performance neighbor search. It also minimizes PCIe overhead by facilitating asynchronous acceleration of distance calculation, and block floating-point compression. We show that a desktop computer with Durin implemented on a mid-range Alveo U50 FPGA can outperform a 32-thread Xeon server by almost 20× , with an order of magnitude power efficiency improvements. Furthermore, Durin outperforms even the best-case projection of a conventional standalone accelerator design, which implements the entirety of the clustering algorithm in the FPGA and its High Bandwidth Memory (HBM), by 2×.
Se-Min Lim, Esmerald Aliaj, Sang Woo Jun
IEEE Big Data2
2024 Morbius: Platform-Adaptive Hardware Accelerator for Scalable Sequence Motif Discovery
abstract
Motif finding is one of the fundamental tools of computational biology. Unfortunately, the benefits of high-performance, low-power acceleration have not been available to this application due to the limited parallelism available within efficient classes of algorithms, such as probabilistic Gibbs sampling. In this work, we present Morbius, which demonstrates the benefits of reconfigurable application-specific hardware acceleration using FPGAs as a solution to the execution time and cost overhead of motif finding. Morbius employs a novel, hardware-optimized data structure called base pair matrix to minimize off-chip data movement, and implements a small number of deep hardware pipelines to achieve high sequential performance. Furthermore, we develop performance and chip space prediction models based on microarchitectural parameters of the accelerator, to facilitate optimal performance on a wide range of accelerators potential users may already have. We compare Morbius on a wide range of FPGA platforms spanning low-profile $250 M.2 FPGA cards to Amazon F1, and demonstrate up to orders of magnitude performance improvements compared to costly server machines, and even higher cost and power efficiency.
Se-Min Lim, Esmerald Aliaj, Sang Woo Jun
e-Science2
2023 Barad-dur: Near-Storage Accelerator for Training Large Graph Neural Networks
abstract
Graph Neural Networks (GNNs) enable effective machine learning on graph-structured data, but their performance and scalability are often limited by the irregular structure and large size of real-world graphs. Conventional accelerators such as GPUs suffer a sharp performance loss when the target graph exceeds its fast random-access memory, because the overhead of partitioning graphs into memory-size chunks quickly dominates performance. We address this issue with Barad-dur, a near-storage GNN accelerator which processes the entire graph in cost-efficient solid-state storage (SSD) to remove the partitioning overhead. Barad-dur minimizes the performance impact of using relatively slow SSDs instead of DRAM, via storage-optimized graph encodings, hardware-accelerated access reordering, and the higher internal storage bandwidth available to near-storage accelerators. We demonstrate that Barad-dur can better maintain computational throughput on larger graphs compared to in-memory CPU and GPU systems. On even moderately large graphs such as the Twitter graph, Barad-dur achieves an order of magnitude higher performance compared to a highly optimized Pytorch Geometric implementation running on the NVIDIA V100 GPU, while requiring a fraction of capital cost and energy.
Jiyoung An, Esmerald Aliaj, Sang Woo Jun
PACT2
2023 FarSlayer: Turnkey Acceleration of Legacy Software on Commodity FPGA Cards
abstract
Application-specific hardware acceleration of computation-intensive kernels can often provide significant performance and power efficiency improvements over general-purpose software, but it is difficult and costly to incorporate them into existing software systems. Designing hardware accelerators and modifying legacy software to incorporate them are already complex tasks. Furthermore, identifying a suitable kernel for acceleration is complicated by the PCIe-attached architecture of commodity FPGA cards, meaning bandwidth and latency overhead must be considered for kernel selection. For example, small blocking kernels may not benefit from acceleration due to PCIe latency. As a remedy, we present FarSlayer, a high-level source-to-source compiler for end-to-end acceleration of legacy software. FarSlayer analyzes existing software code and emits an accelerated version of it, where the kernel is automatically selected considering data movement over PCIe. Specifically, FarSlayer identifies kernels which can be called asynchronously to hide the PCIe latency, while also having a high operational intensity for low bandwidth requirements. The entire process is automatic, meaning the programmer does not necessarily need to understand the existing code, or reason about hardware development. We demonstrate FarSlayer on multiple existing scientific computing software systems, and demonstrate it can automatically achieve significant performance improvements.
Esmerald Aliaj, Alberto Krone-Martins, Joshua Garcia, Sang Woo Jun
ASAP1
2023 PreCog: Near-Storage Accelerator for Heterogeneous CNN Inference
abstract
Computational Storage Devices (CSD) with near-storage acceleration is gaining popularity for data-intensive applications, by moving power-efficient hardware acceleration closer to data. However, because the power-constrained near-storage accelerator is often not powerful enough by itself to handle all computation requirements of an application, it must intelligently cooperate with other computation and acceleration units in the system. In this work, we explore how a near-storage accelerator can best fit into a larger computer system in the context of CNN inference. We demonstrate that an attractive configuration is using the near-storage accelerator to offload only the first convolution and pooling layers, where the accelerator can achieve almost an order of magnitude better performance compared to a general convolution accelerator. Targeting only the first layer allows some FPGA-specific floating-point computation optimizations such as pre-determine the range of output exponents, performing a costly floating point normalization task only once, as well as pack more input into the datapath to mitigate the performance impact of wide strides of the convolution filter. We package these optimizations into a flexible library we call Static Range Float (SRFloat), and construct a prototype system called PreCog. We evaluate PreCog implemented on the Samsung SmartSSD platform, and demonstrate over$\mathbf{3}\times$performance efficiency compared to conventional convolution accelerators on prominent CNN models without introducing a communications bottleneck.
Jiyoung An, Esmerald Aliaj, Sang Woo Jun
ASAP2
2022 GAROTA: Generalized Active Root-Of-Trust Architecture (for Tiny Embedded Devices)
Esmerald Aliaj, Ivan Oliveira Nunes, Gene Tsudik
USENIX Security Symposium1