Rasha Karakchi

dblp:184/5949 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0004-1391-0166ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Poster: Hybrid Monitoring for Side-Channel Security in Edge SoCs
abstract
Edge-class SoCs, widely used in IoT and embedded devices such as IoT gateways, smart meters, medical devices, surveillance cameras, and industrial controllers, face increasing exposure to side-channel and fault injection attacks due to their deployment in untrusted environments and strict power and resource constraints. Traditional defenses, such as constant-time coding or hardware redundancy, are often infeasible on these lightweight platforms. To address this challenge, we present a hybrid cryptographic monitoring system that combines statistical thresholding with machine learning (ML) to detect anomalies in real time. Using injected timing delays and ciphertext alteration, we evaluated two complementary detectors: a lightweight threshold-based monitor and a Random Forest classifier leveraging timing and cipher-text features. Implemented on the PYNQ-Z1, the framework achieves <5 ms inference latency with under 30% resource utilization. The results show that the hybrid approach improves accuracy and reduces false positives/negatives compared to threshold-only detection, offering a practical solution to strength cryptographic resilience in edge-class SoCs.
Nishant Chinnasami, Rye Stahle-Smith, Rasha Karakchi
SEC3
2024 Developing a Self-Explanatory Transformer
abstract
While IoT devices provide significant benefits, their rapid growth results in larger data volumes, increased complexity, and higher security risks. To manage these issues, techniques like encryption, compression, and mapping are used to process data efficiently and securely. General-purpose and AI platforms handle these tasks well, but mapping in natural language processing is often slowed by training times. This work explores a self-explanatory, training-free mapping transformer based on non-deterministic finite automata, designed for Field-Programmable Gate Arrays (FPGAs). Besides highlighting the advantages of this proposed approach in providing real-time, cost-effective processing and dataset-loading, we also address the challenges and considerations for enhancing the design in future iterations.
Rasha Karakchi, Ryan Karbowniczak
SEC1
2023 NAPOLY: A Non-deterministic Automata Processor OverLaY
abstract
Deterministic and Non-deterministic Finite Automata (DFA and NFA) comprise the core of many big data applications. Recent efforts to develop Domain-Specific Architectures (DSAs) for DFA/NFA have taken divergent approaches, but achieving consistent throughput for arbitrarily-large pattern sets, state activation rates, and pattern match rates remains a challenge. In this article, we present NAPOLY (Non-Deterministic Automata Processor OverLaY), an FPGA overlay and associated compiler. A common limitation of prior efforts is a limit on NFA size for achieving the advertised throughput. NAPOLY is optimized for fast re-programming to permit practical time-division multiplexing of the hardware and permit high asymptotic throughput for NFAs of unlimited size, unlimited state activation rate, and high pattern reporting rate. NAPOLY also allows for offline generation of configurations having tradeoffs between state capacity and transition capacity. In this article, we (1) evaluate NAPOLY using benchmarks packaged in the ANMLZoo benchmark suite, (2) evaluate the use of an SAT solver for allocating physical resources, and (3) compare NAPOLY’s performance against existing solutions. NAPOLY performs most favorably on larger benchmarks, benchmarks with higher state activation frequency, and benchmarks with higher reporting frequency. NAPOLY outperforms the fastest of the CPU and GPU implementations in 10 out of 12 benchmarks.
Rasha Karakchi, Jason D. Bakos
ACM Trans. Reconfigurable Technol. Syst.1
2019 An Overlay Architecture for Pattern Matching
abstract
Deterministic and Non-deterministic Finite Automata (DFA and NFA) comprise the fundamental unit of work for many emerging big-data applications, motivating recent efforts to develop Domain-Specific architectures (DSAs) to exploit fine-grain parallelism available in automata workloads. In this paper we present NAPOLY (Non-Deterministic Automata Processor OverLaY), an overlay architecture and associated software that attempts to maximally exploit on-chip memory parallelism for NFA evaluation. In order to avoid an upper bound on NFA size that commonly affects prior efforts, NAPOLY is optimized for runtime reconfiguration, allowing for full reconfiguration in 10s of microseconds. NAPOLY is also parameterizable, allowing for offline generation of a repertoire of overlay configurations with various trade-offs between state capacity and transition capacity. In this paper we evaluate NAPOLY using our proposed state mapping heuristic and the ANMLZoo benchmark suite, and we compare NAPOLY's performance against existing CPU and GPU implementations. To the best of the authors' knowledge this is the first example of a runtime-reprogrammable FPGA-based automata processor overlay.
Rasha Karakchi, Charles Daniels, Jason D. Bakos
ASAP1
2016 Two-Hit Filter Synthesis for Genomic Database Search
abstract
Advancements in genomic sequencing technology is causing genomic database growth to outpace Moore's Law. This continues to make genomic database search a difficult problem and a popular target for emerging processing technologies. The de facto software tool for genomic database search is NCBI BLAST, which operates by transforming each database query into a filter that is subsequently applied to the database. This requires a database scan for every query, fundamentally limiting its performance by I/O bandwidth. In this paper we present a functionally-equivalent variation on the NCBI BLAST algorithm that maps more suitably to an FPGA implementation. This variation of the algorithm attempts to reduce the I/O requirement by leveraging FPGA-specific capabilities, such as high pattern matching throughput and explicit on chip memory structure and allocation. Our algorithm transforms the database -- not the query -- into a filter that is stored as a hierarchical arrangement of three tables, the first two of which are stored on chip and the third off chip. Our results show that -- while performance is data dependent -- it is possible to achieve speedups of up to 8X based on the relative reduction in I/O of our approach versus that of NCBI BLAST. More importantly, the performance relative to NCBI BLAST improves with larger databases and query workload sizes.
Jordan A. Bradshaw, Rasha Karakchi, Jason D. Bakos
FCCM2