Atiyehsadat Panahi

dblp:259/0375 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2023 Making BRAMs Compute: Creating Scalable Computational Memory Fabric Overlays
abstract
The increasing density of distributed BRAMs diffused throughout modern Field Programmable Gate Arrays (FP-GAs) is ideal for forming processor in/near memory architectures. This breaks the traditional von Neumann memory bottleneck limiting concurrency and degrading energy efficiency. Ideally, processing density should scale linearly with BRAM capacity, and clock frequencies should be set by the read/write access times of the BRAM. In this paper, we present a PIM overlay that achieves these goals. We observe an improvement of performance by 2.25 x, logic resource utilization by 2 x, and accumulation delay by 17 x compared to prior published work.
M. D. Arafat Kabir, Joshua Hollis, Atiyehsadat Panahi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FCCM3
2023 FPGA Processor In Memory Architectures (PIMs): Overlay or Overhaul ?
abstract
The dominance of machine learning and the ending of Moore's law have renewed interests in Processor in Memory (PIM) architectures. This interest has produced several recent proposals to modify an FPGA's BRAM architecture to form a next-generation PIM reconfigurable fabric [1], [2]. PIM architectures can also be realized within today's FPGAs as overlays without the need to modify the underlying FPGA architecture. To date, there has been no study to understand the comparative advantages of the two approaches. In this paper, we present a study that explores the comparative advantages between two proposed custom architectures and a PIM overlay running on a commodity FPGA. We created PiCaSO, a Processor in/near Memory Scalable and Fast Overlay architecture as a representative PIM overlay. The results of this study show that the PiCaSO overlay achieves up to 80% of the peak throughput of the custom designs with 2.56 x shorter latency and 25% - 43% better BRAM memory utilization efficiency. We then show how several key features of the PiCaSO overlay can be integrated into the custom PIM designs to further improve their throughput by 18%, latency by 19.5%, and memory efficiency by 6.2%.
M. D. Arafat Kabir, Ehsan Kabir, Joshua Hollis, Eli Levy-Mackay, Atiyehsadat Panahi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FPL5
2022 High-Rate Machine Learning for Forecasting Time-Series Signals
abstract
"Active structures" are physical structures that incorporate real-time monitoring and control. Examples include active vibration damping or blast mitigation systems. Evaluating physics-based models in real-time is generally not feasible for such systems having high-rate dynamics which require microsecond response times, but data-driven machine-learning-based models can potentially offer a solution. This paper compares the cost and performance of two FPGA-based implementations of real-time, continuously-trained models for forecasting time-series signals with non-stationarities, with one using High-Level Synthesis (HLS) and the other a programmable overlay architecture. The proposed model accepts a uni-variate vibration signal and seeks to forecast future samples to inform high-rate controllers. The proposed forecasting method performs two concurrent neural inference operations. One inference forecasts the state of the signal f samples into the future as a function of the most recent h samples, while the other forecasts the current sample given h samples starting from h+f−1 samples into the past. The first forecast produces the forecast while the second forecast allows the system to calculate the model’s loss and perform an immediate model update before the next sample period.
Atiyehsadat Panahi, Ehsan Kabir, Austin R. J. Downey, David Andrews 0001, Miaoqing Huang, Jason D. Bakos
FCCM1
2021 A Customizable Domain-Specific Memory-Centric FPGA Overlay for Machine Learning Applications
abstract
This paper presents an overview and performance analysis of a software-programmable domain-customizable System-on-Chip (SoC) overlay for low-latency inferencing of variable and low-precision Machine Learning (ML) networks targeting Internet-of-Things (IoT) edge devices. The SoC includes a 2-D processor array that can be customized at design time for FPGA logic families. The overlay resolves historic issues of poor designer productivity associated with traditional Field Programmable Gate Array (FPGA) design flows without the performance losses normally incurred by overlays. A standard Instruction Set Architecture (ISA) allows different ML networks to be quickly compiled and run on the overlay without the need to resynthesize. Performance results are presented that show the overlay achieves $1.3\times-8.0\times$ speedup over custom designs while still allowing rapid changes to ML algorithms on the FPGA through standard compilation.
Atiyehsadat Panahi, Suhail Basalama, Ange-Thierry Ishimwe, Joel Mandebi, David Andrews 0001
FPL1
2020 FPGA-Based Gesture Recognition with Capacitive Sensor Array using Recurrent Neural Networks
abstract
This work presents a prototype of an FPGA-based hand motion recognition system using a capacitive sensor array (CSA). The prototype system is being developed as a tool to evaluate upper-limb motor skills for assistive or rehabilitative applications. A light-weight gesture segmentation algorithm was developed that uses summation and moving average filtering of quantized capacitive sensing data to segment motions. The time-series hand motions are then recognized through a recurrent classifier based on long short-term memory (LSTM) neural networks. The classifier model is trained on uni-stroke hand written digit ('0'–'9') samples obtained from four volunteers. A total of 12,000 hand motion samples are collected. The accuracy of 10-fold and leave-one-user-out cross-validation accuracy is respectively 97.5% and 91.3% using a two-layer LSTM network. The LSTM classifier is implemented on a Zynq FPGA device. The experiment demonstrated that the FPGA implementation of the LSTM-based classifier can achieve real-time gesture classification with capacitive sensor data.
Haoyan Liu 0002, Atiyehsadat Panahi, David Andrews 0001, Alexander Nelson 0001
FCCM2