G. Abarajithan

dblp:308/2199 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0001-9768-5349ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Reconfigurable Computing Challenge: Transformer for Jet Tagging on Versal AI Engines
abstract
Transformer-based models achieve strong performance for jet tagging at the CERN LHC, but deploying them in low-latency, resource-constrained trigger systems is challenging. We present an initial implementation of a quantized, integer-only transformer for jet tagging on the AMD Versal AI Engine (AIE), mapping dense and multi-head attention (MHA) layers to AIE tiles. The main contribution is a reusable software framework that represents transformer layers as composable AIE building blocks and automatically generates the corresponding Vitis graph code from a high-level Python model description. This framework provides a foundation for future research and is released as open-source software at https://github.com/KastnerRG/particle_transformer_aie.
Gram Koski, Sean Lipps, Zhenghua Ma, G. Abarajithan, Ryan Kastner
FCCM4
2026 cgra4ml: A Hardware/Software Framework to Implement Neural Networks for Scientific Edge Computing
abstract
The scientific community increasingly relies on machine learning (ML) for near-sensor processing, leveraging its strengths in tasks such as pattern recognition, anomaly detection, and real-time decision-making. These deployments demand accelerators that combine extremely high performance with programmability, ease of integration, and straightforward verification. We present cgra4ml , an open-source, modular framework that generates parameterizable CGRA accelerators in synthesizable SystemVerilog RTL, tailored to common ML compute patterns found in scientific applications. The framework supports seamless system integration through AXI-compliant interfaces and open-source DMA components, and it includes automatic firmware generation for programming the accelerator. A comprehensive verification suite and a runtime firmware stack further support deployment across diverse SoC platforms. cgra4ml provides a modular, full-stack infrastructure, including a Python API, SystemVerilog hardware, TCL toolflows, and a C runtime, which facilitates easy integration and experimentation, allowing scientists to focus on innovation rather than dealing with the intricacies of hardware design and optimization. We demonstrate the effectiveness of cgra4ml to implement common scientific edge neural networks using ASIC and FPGA design flows.
G. Abarajithan, Zhenghua Ma, Ravidu Munasinghe, Francesco Restuccia 0002, Ryan Kastner
ACM Trans. Reconfigurable Technol. Syst.1
2024 Tailor: Altering Skip Connections for Resource-Efficient Inference
abstract
Deep neural networks use skip connections to improve training convergence. However, these skip connections are costly in hardware, requiring extra buffers and increasing on- and off-chip memory utilization and bandwidth requirements. In this article, we show that skip connections can be optimized for hardware when tackled with a hardware-software codesign approach. We argue that while a network’s skip connections are needed for the network to learn, they can later be removed or shortened to provide a more hardware-efficient implementation with minimal to no accuracy loss. We introduce Tailor , a codesign tool whose hardware-aware training algorithm gradually removes or shortens a fully trained network’s skip connections to lower the hardware cost. Tailor improves resource utilization by up to 34% for block random access memories (BRAMs), 13% for flip-flops (FFs), and 16% for look-up tables (LUTs) for on-chip, dataflow-style architectures. Tailor increases performance by 30% and reduces memory bandwidth by 45% for a two-dimensional processing element array architecture.
Olivia Weng, Gabriel Marcano, Vladimir Loncar, Alireza Khodamoradi, G. Abarajithan, Nojan Sheybani, Andres Meza 0001, Farinaz Koushanfar, Kristof Denolf, Javier M. Duarte, Ryan Kastner
ACM Trans. Reconfigurable Technol. Syst.5
2022 A Mostly-Online CAS Teaching Experience
abstract
Mostly-online teaching experiences in circuits and systems (CAS) during 2020-2021 COVID-19 pandemic are presented. Three case studies are shared summarizing course details, tools and platforms, best practices and limitations across three universities representing different continents. The presented approaches attempt to address several limitations of passive online delivery of CAS courses via interactive simulations, formative assessments via interactive web content and slide-embedded polls, at-home labs with student acquired hardware and instructor-lead and student-centered activity based learning.
Chamith Wijenayake, Kithmin Wickremasinghe, G. Abarajithan, Arjuna Madanayake, Chamira U. S. Edussooriya, K. Samarasinghe
ISCAS3