EDBT 2026 Demo / reviewers in the wild / expert
Jim Ignowski
dblp:37/10854
· DBLP profile ↗
9ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-5091-3674ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 7 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analog In-Memory Computing Enhanced FPGA for High-Throughput and Energy-Efficient AccelerationabstractThe ever-growing demand for AI computing, coupled with slowing performance gains in chip manufacturing, has heightened the role of FPGA-based accelerators. FPGAs enable the implementation of application-customized parallel dataflows due to their reconfigurability, achieving high energy efficiency. However, the bit-level routing fabric on FPGAs often results in high overheads because large amounts of data must be shuttled between compute blocks and memory blocks on the FPGA. We propose enhancing FPGAs with in-memory computing macros, specifically analog Dot Product Engines based on non-volatile RRAM devices. Using the Verilog to Routing (VTR) framework, we simulate a novel 40 nm, 26.2 mm × 26.2 mm architecture and employ a custom event-driven simulator to evaluate its performance. Our design achieves 25.5 ×103TOPS/W, an average ×31.4 throughput improvement and an average ×9,380 energy efficiency improvement when compared to state-of-the-art FPGA implementations of AI models. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Xia Sheng, Giacomo Pedretti, Aman Arora 0001, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FCCM | 9 |
| 2025 | Enhancing FPGAs with Analog In-Memory Computing MacrosabstractWhile the AI computing needs are ever-increasing and the innovation in models generates tens of new architectures yearly, the performance gain from improvements in chip manufacturing has slowed down. Within this context, FPGA-based accelerators play a fundamental role. FPGAs are the backbone of specialized architectures, their reconfigurability being the key differentiation that enables an effective design space exploration. At the same time, to overcome the limitations induced by the memory bottleneck, the computing architectures community has proposed the in-memory computing paradigm: storage and computations are both performed in non-volatile memory devices. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Rand Jean, Xia Sheng, Giacomo Pedretti, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FPGA | 9 |
| 2025 | RACE-IT: A Reconfigurable Analog Computing Engine for In-Memory Transformer AccelerationabstractTransformer models represent the cutting edge of Deep Neural Networks (DNNs) and excel in a wide range of machine learning tasks. However, processing these models demands significant computational resources and results in a substantial memory footprint. While In-memory Computing (IMC) offers promise for accelerating Vector-Matrix Multiplications (VMMs) with high computational parallelism and minimal data movement, employing it for other crucial DNN operators remains a formidable task. This challenge is exacerbated by the extensive use of complex activation functions, Softmax, and data-dependent matrix multiplications (DMMuls) within Transformer models. To address this challenge, we introduce a Reconfigurable Analog Computing Engine (RACE) by enhancing Analog Content Addressable Memories (ACAMs) to support broader operations. Based on the RACE, we propose the RACE-IT accelerator (meaning RACE for In-memory Transformers) to enable efficient analog-domain execution of all core operations of Transformer models. Given the flexibility of our proposed RACE in supporting arbitrary computations, RACE-IT is well-suited for adapting to emerging and non-traditional DNN architectures without requiring hardware modifications. We compare RACE-IT with various accelerators. Results show that RACE-IT increases performance by 453× and 15×, and reduces energy by 354× and 122× over the state-of-the-art GPUs and existing Transformer-specific IMC accelerators, respectively. Aishwarya Natarajan, Luca Buonanno, Archit Gajjar, Ron M. Roth, Sergey Serebryakov, John Moon, Omar Eldash, Jim Ignowski, Giacomo Pedretti |
ICCD | 9 |
| 2025 | Boosting Task Scheduling Data Locality with Low-latency, HW-accelerated Label PropagationabstractTask Scheduling is a popular technique for exploiting parallelism in modern computing systems.In particular, HW-accelerated Task Scheduling has been shown to be effective at improving the performance of fine-grained workloads by dynamically assigning tasks to cores based on their data dependencies with minimal overhead, allowing the handling of tasks with execution times in the order of thousands of cycles.However, the performance of applications assisted by accelerated Task Scheduling is limited by the fact that once a task has all its dependencies fulfilled, it is typically executed on the first available core, which might not be locality-optimal.We thus propose a novel approach to Task Scheduling that leverages HW-accelerated Label Propagation (LP), a graph clustering algorithm, to group tasks with intersecting data patterns such that they are executed on the same core.We show that our approach can significantly improve the performance of task-based applications, improving overall program execution times by up to 1.50× while simultaneously reducing average task sizes by up to 1.81×, augmenting both synthetic benchmarks and real-world applications running on a 24-core RISC-V processor mapped to the Alveo U55C FPGA.These gains rely heavily on the low-latency nature of our proposed label propagation accelerator, which will typically cluster dynamic task graphs in under 300 cycles, up to 581× faster than an equivalent software implementation.Furthermore, by ensuring that ideal placement predictions are used as a hint rather than a hard constraint, we allow the system to benefit from improved data locality for memory-intensive applications while also maintaining high core utilization in compute-bound scenarios.Our results hence demonstrate the potential of HW-accelerated label propagation to improve the performance of Task Scheduling systems with low-latency, dynamic data locality optimization. Lucas Morais, Juan Miguel De Haro Ruiz, Alfredo Goldman, Guido Araujo, Giacomo Pedretti, Jim Ignowski, Michael Frank 0008, Xavier Martorell, Daniel Jiménez-González, Carlos Álvarez 0001 |
MICRO | 6 |
| 2024 | CAMSHAP: Accelerating Machine Learning Model Explainability with Analog CAMabstractThe recent success of machine learning (ML) models has led to increasing demands for model explanations - why a result was given - along with model predictions. Tree-based ML models are considered more explainable than deep neural networks and higher performers in several domains. However, algorithms computing model explanations are irregular and scale poorly with model size. While many custom accelerators for training and inference have been proposed, little attention has been paid to accelerating model explanations. This lack of explanatory capability has limited the use of these models for real-time decision-making systems in critical fields such as healthcare, autonomous operation and cybersecurity. John Moon, Giacomo Pedretti, Pedro Bruel, Sergey Serebryakov, Omar Eldash, Luca Buonanno, Catherine Graves, Paolo Faraboschi, Jim Ignowski |
ICCAD | 9 |
| 2024 | Memristive Quaternary Content-Addressable Memories for Implementing Boolean FunctionsabstractIn-memory computing is, in current literature, the most common paradigm used to counteract the Von-Neumann bottleneck, proposing the use of memory elements to define complex input-output relations of the computing kernels. While in classical CMOS computing a similar paradigm can be implemented with look-up tables (LUT), this solution is power and area-hungry. This paper presents the use of Quaternary Content-Addressable Memories (QCAMs), a generalization of the Ternary Content-Addressable Memories (TCAMs), for implementing boolean functions. Content-Addressable Memories can be used as a building block for in-memory processing, using the states of the cells to define a ${\mathbb{B}^{\text{N}}} \to {\mathbb{B}^{\text{M}}}$ function which projects the search word into a new string of bits. The quaternary alphabet allows to represent a more complex function space with respect to the TCAMs while using the same number of cells, enhancing area, power consumption and latency performances achieved when representing arbitrary functions with the CAM hardware. For comparison, it can be demonstrated that QCAMs represent arbitrary Boolean functions with half the number of cells than that would be needed in a standard TCAM implementation, and a ×10 smaller area with respect to SRAM-based LUTs. Along with the table of states and a toy example where the QCAM states are used to define the product among two 2-bit precision real values, this paper presents multiple circuit schemes and encoding schemes for memristor-based QCAMs. Luca Buonanno, Giacomo Pedretti, Aishwarya Natarajan, Todd Richmond, John Moon, Rand Jean, Xia Sheng, Ron M. Roth, Jim Ignowski |
ISCAS | 10 |
| 2024 | Realizing In-Memory Baseband Processing for Ultrafast and Energy-Efficient 6GabstractTo support emerging applications ranging from holographic communications to extended reality, next-generation mobile wireless communication systems require ultrafast and energy-efficient baseband processors. Traditional complementary metal-oxide-semiconductor (CMOS)-based baseband processors face two challenges in transistor scaling and the von Neumann bottleneck. To address these challenges, in-memory computing-based baseband processors using resistive random-access memory (RRAM) present an attractive solution. In this article, we propose and demonstrate RRAM-implemented in-memory baseband processing for the widely adopted multiple-input–multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) air interface. Its key feature is to execute the key operations, including discrete Fourier transform (DFT) and MIMO detection, using linear minimum mean square error (L-MMSE) and zero forcing (ZF), in one-step. In addition, RRAM-based channel estimation module is proposed and discussed. By prototyping and simulations, we demonstrate the feasibility of RRAM-based full-fledged communication system in hardware, and reveal it can outperform state-of-the-art baseband processors with a gain of$91.2\times $in latency and$671\times $in energy efficiency by large-scale simulations. Our results pave a potential pathway for RRAM-based in-memory computing to be implemented in the era of the sixth generation (6G) mobile communications. Qunsong Zeng, Mingrui Jiang, Yi Gong 0001, Yida Li 0004, Can Li 0024, Jim Ignowski, Kaibin Huang |
IEEE Internet Things J. | 9 |
| 2024 | RD-FAXID: Ransomware Detection with FPGA-Accelerated XGBoostabstractOver the last decade, there has been a rise in cyberattacks, particularly ransomware, causing significant disruption and financial repercussions across public and private sectors. Tremendous efforts have been spent on developing techniques to detect ransomware to, ideally, protect data or have as minimum data loss as possible. Ransomware attacks are becoming more frequent and sophisticated as there is a constant tussle between attackers and cybersecurity defenders. Machine Learning (ML) approaches have proven more effective in detecting ransomware than classical signature-based detection. In particular, tree-based algorithms such as Decision Trees (DT), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) spike up interest among cybersecurity researchers. However, due to the nature of the problem, traditional CPUs and GPUs fail to keep up with the desired performance, especially for large data workloads. Thus, the problem demands a customized solution to detect the ransomware. Here, we propose an FPGA accelerated tree-based ML model for multi-dataset ransomware detection. We show the capability of the proposed prototype to address the problem from more than one set of features, reducing false positive and negative rates to have robust predictions by looking at Hardware Performance Counters (HPCs), Operating System (OS) calls, and network traffic information simultaneously. With 1,000 samples per batch, the FPGA prototype has 65.8 \({\times}\) and 4.1 \({\times}\) lower latency over the CPU and GPU, respectively. Moreover, the FPGA design is up to 11.3 \({\times}\) cost-effective and 643 \({\times}\) energy-efficient compared to the CPU and 3 \({\times}\) cost-effective and 16.8 \({\times}\) energy-efficient over the GPU. Archit Gajjar, Priyank Kashyap, Aydin Aysu, Paul D. Franzon, Chris Cheng, Giacomo Pedretti, Jim Ignowski |
ACM Trans. Reconfigurable Technol. Syst. | 8 |
| 2009 | Voltage transient detection and induction for debug and testabstractVoltage transients from circuit activity impact operation, testing and debug of complex designs. This paper describes a system which enables voltage transient detection and a capability to induce voltage transients in a controlled manner. Usage models and silicon results are described, along with limitations and future options for improvements. Rex Petersen, Pankaj Pant, Pablo Lopez, Aaron Barton, Jim Ignowski, Doug Josephson |
ITC | 5 |