Arish Sateesan

dblp:276/2220 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-8197-0097ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 4 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ArchE-Q: A DSP-Free Dataflow Accelerator for Quantized Neural Networks in Sensor-Aided Millimeter-Wave Edge Connectivity
abstract
Sensor-aided wireless edge applications, such as LiDAR-based beam prediction for millimeter-wave communications, demand intelligent on-device processing of high-volume sensor data. However, the computational cost of machine learning models often exceeds the tight power and resource constraints of edge hardware. While quantized neural networks (QNNs) reduce resource requirements, typical FPGA accelerators still rely on power-hungry digital signal processing (DSP) slices and incur avoidable data-movement overheads. To bridge this gap, we propose ArchE-Q, a dataflow accelerator for QNNs combined with efficient data preprocessing. Our design is fundamentally multiplier-less, utilizing: (1) an application-specific first-layer kernel that exploits binarized sensor inputs to remove multipliers; (2) the eXtended Vector Activation Unit (XVAU), fusing convolution, activation, and pooling to reduce buffering and data transfers; and (3) memory-centric buffering for efficient data reuse. Implemented on a Xilinx ZCU104 FPGA, ArchE-Q achieves 13.3% lower latency and up to 28% lower dynamic power than the FINN-R baseline, while eliminating DSP usage.
Arish Sateesan, Ljiljana Simic, Marina Petrova
DATE1
2026 A Hardware-Aware Performance Analysis of Machine Learning Models for Sensor-Aided Millimeter-Wave Beam Prediction
abstract
Millimeter-wave (mm-wave) and sub-terahertz communication are cornerstones for 6G networks, but are hindered by their reliance on highly directional beams and the high latency of conventional beam alignment and beam management procedures. Sensor- and machine learning (ML)-based beam prediction offers a low-latency alternative, but remains constrained by hardware limitations at the wireless edge. This paper presents a systematic, hardware-aware analysis of ML models for sensor-aided beam prediction, specifically focusing on LiDAR, spanning deep learning, classical ML, and a novel class of proposed hybrid models that combine a CNN-based feature extractor with lightweight classical classifiers. Using the public DeepSense 6G dataset, models are benchmarked on accuracy, inference latency, memory footprint, and computational complexity. Our results show that classical ML models, particularly linear and tree-based methods, deliver strong accuracy-efficiency trade-offs, achieving up to 95× lower inference latency and 67× smaller memory footprint than CNNs with minimal accuracy loss. The proposed hybrid models further enhance this trade-off, maintaining CNN-level top-5 accuracy while drastically reducing re-training time by up to 20× and enabling fast, environment-specific adaptation. These findings establish classical and hybrid models, supported by efficient preprocessing, as promising candidates for real-time, resource-efficient beam prediction in future 6G edge systems.
Baris Sayitoglu, Arish Sateesan, Marina Petrova, Ljiljana Simic
ICC2
2026 Breaking the Scalability Barrier of Content Addressable Memories: A Probabilistic Alternative for Large-Key Associative Search
abstract
Content Addressable Memories (CAMs) offer high-speed, deterministic lookups but face significant scalability challenges with large input keys ( \( > \) 100 bits), leading to excessive power, silicon area, and memory costs. This article introduces Probabilistic CAM (P-CAM), a novel architecture designed to overcome these limitations by trading strict determinism for memory efficiency and scalability. P-CAM compresses high-dimensional inputs into fixed-size fingerprints using hashing, making memory requirements independent of key length. P-CAM preserves the constant-time lookup advantage of CAMs, while supporting applications with large keys, such as networking, bioinformatics, and machine learning, where conventional CAMs are impractical. FPGA implementation on Xilinx UltraScale+ devices shows that P-CAM maintains constant query latency and delivers 15 \(\times\) improvement in resource efficiency when handling 384-bit keys, compared to state-of-the-art deterministic CAMs designed for narrower inputs. Although P-CAM’s probabilistic nature introduces a small, controllable false-positive rate, it can be configured for fully deterministic operation under specific constraints. To the best of our knowledge, P-CAM is the first CAM architecture to employ a fingerprint-based probabilistic data structure as the primary storage mechanism for associative lookup, distinguishing it from prior probabilistic approaches that are limited to set membership checks, offering a robust and scalable alternative for modern data-intensive systems.
Arish Sateesan, Jo Vliegen, Nele Mentens
ACM Trans. Reconfigurable Technol. Syst.1
2025 ITERATOR: Interruptible Remote Attestation Through Cuckoo Filters
abstract
Remote attestation (RA) is emerging as a promising security mechanism that establishes trust in IoT devices by detecting the malware presence. Typically, RA consists of computing a hash over the device’s memory and is executed as anatomicprocedure to guarantee the reliability of the attestation evidence. However, in real-world situations, such as those involving real-time systems, energy-harvesting devices, or mission-critical operations, the IoT device may not be able to complete the attestation procedure due to various factors like task scheduling, limited battery life, or higher priority tasks. In such scenarios where flexibility, adaptability, and security are paramount, enablinginterruptibilityof RA is crucial. This paper presents a novel approach called ITERATOR which leverages hash-based storage to enable interruptible RA without any additional hardware requirements. Our proposal transforms the device attestation procedure from the traditional approach of memory hash computation to a lookup operation in a hash-based storage, namely, Cuckoo filter. The ITERATOR protocol divides the device’s memory into blocks associated with a Cuckoo filter bucket. This approach allows the device to perform RA in multiple rounds, ensuring secure interruptible attestation. We perform software simulations of ITERATOR, demonstrating its high effectiveness in detecting the malware presence. Due to its interruptible design, ITERATOR cannot guarantee 100% detection in a single attestation round; however, repeated rounds make long-term evasion by malware highly unlikely. In particular, the experiments showed that the probability of evading the detection ranges between 37% and less than 1%, depending on the protocol configuration. Moreover, we validate ITERATOR’s efficiency through two hardware proof-of-concept implementations that rely on ESP32 and FPGA platforms. The FPGA implementation shows the high efficiency of the protocol, with 34.3ns to attest a single memory block.
Nicoló Sponziello, Arish Sateesan, Md Masoom Rabbani, Nele Mentens, Nicola Dragoni, Edlira Dushku
IEEE Internet Things J.2
2024 SPArch: A Hardware-oriented Sketch-based Architecture for High-speed Network Flow Measurements
abstract
Network flow measurement is an integral part of modern high-speed applications for network security and data-stream processing. However, processing at line rate while maintaining the required data structure within the on-chip memory of the hardware platform is a challenging task for measurement algorithms, especially when accuracy is of primary importance, such as in network security applications. Most of the existing measurement algorithms are no exception to such issues when deployed in high-speed networking environments and are also not tailored for efficient hardware implementation. Sketch-based measurement algorithms minimize the memory requirement and are suitable for high-speed networks but possess a low memory-accuracy trade-off and lack the versatility of individual flow mapping. To address these challenges, we present a hardware-friendly data structure named Sketch-based Pseudo-associative array Architecture (SPArch). SPArch is highly accurate and extremely memory-efficient, making it suitable for network flow measurement and security applications. The parallelism in SPArch ensures minimal and constant memory access cycles. Unlike other sketch architectures, SPArch provides the functionality of individual flow mapping similar to associative arrays, and the optimized version of SPArch allows the organization of counters in multiple buckets based on the flow sizes. An in-depth analysis of SPArch is carried out in this article and implemented SPArch on the Alveo data center accelerator card, demonstrating its suitability for high-speed networks.
Arish Sateesan, Jo Vliegen, Simon Scherrer, Hsu-Chun Hsiao, Adrian Perrig, Nele Mentens
ACM Trans. Priv. Secur.1
2023 Evolving Non-cryptographic Hash Functions Using Genetic Programming for High-speed Lookups in Network Security Applications
Arish Sateesan, Jo Vliegen, Stjepan Picek, Nele Mentens
EvoApplications@EvoStar2
2023 ALBUS: a Probabilistic Monitoring Algorithm to Counter Burst-Flood Attacks
abstract
Modern DDoS defense systems rely on probabilistic monitoring algorithms to identify flows that exceed a volume threshold and should thus be penalized. Commonly, classic sketch algorithms are considered sufficiently accurate for usage in DDoS defense. However, as we show in this paper, these algorithms achieve poor detection accuracy under burst-flood attacks, i.e., volumetric DDoS attacks composed of a swarm of medium-rate sub-second traffic bursts. Under this challenging attack pattern, traditional sketch algorithms can only detect a high share of the attack bursts by incurring a large number of false positives. In this paper, we present ALBUS, a probabilistic monitoring algorithm that overcomes the inherent limitations of previous schemes: ALBUS is highly effective at detecting large bursts while reporting no legitimate flows, and therefore improves on prior work regarding both recall and precision. Besides improving accuracy, ALBUS scales to high traffic rates, which we demonstrate with an FPGA implementation, and is suitable for programmable switches, which we showcase with a P4 implementation.
Simon Scherrer, Jo Vliegen, Arish Sateesan, Hsu-Chun Hsiao, Nele Mentens, Adrian Perrig
SRDS3
2021 Novel Non-cryptographic Hash Functions for Networking and Security Applications on FPGA
abstract
This paper proposes the design and FPGA implementation of five novel non-cryptographic hash functions, that are suitable to be used in networking and security applications that require fast lookup and/or counting architectures. Our approach is inspired by the design of the existing non-cryptographic hash function Xoodoo-NC, which is constructed through the concatenation of several Xoodoo permutations. We similarly construct non-cryptographic hash functions based on the concatenation of several rounds of symmetric-key ciphers. The goal is to achieve high performance in combination with good avalanche properties, which are required in order to have a significant change in the output value as a result of a limited change in the input value. We simulate how many rounds are needed to achieve satisfactory avalanche scores and we implement the corresponding non-cryptographic hash functions on an FPGA to evaluate the occupied resources and the performance. One of the proposed non-cryptographic hash functions, namely GIFT-NC, outperforms all previously proposed non-cryptographic hash functions in terms of throughput and latency, in exchange for an acceptable increase in FPGA resources.
Thomas Claesen, Arish Sateesan, Jo Vliegen, Nele Mentens
DSD2
2021 Speed Records in Network Flow Measurement on FPGA
abstract
Network traffic measurement keeps track of the amount of traffic sent by each flow in the network. It is a core functionality in applications such as traffic engineering and network intrusion detection. In high-speed networks, it is impossible to keep an exact count of the flow traffic, due to limitations with respect to memory and computational speed. Therefore, probabilistic data structures, such as sketches, are used. This paper proposes Approximate Count-Min sketch or ACM sketch, a novel variant of the Count-Min sketch algorithm that uses less memory and has a higher throughput compared to other FPGA-based sketch implementations. A-CM sketch relies on optimizations at two levels: (1) it uses approximate counters and the newly proposed Hardware-oriented Simple Active Counter algorithm to efficiently implement these counters; (2) it uses a distribution of the embedded memory, optimized towards maximum operating frequency. To the best of our knowledge, A-CM sketch outperforms all other FPGA-based sketch implementations.
Arish Sateesan, Jo Vliegen, Simon Scherrer, Hsu-Chun Hsiao, Adrian Perrig, Nele Mentens
FPL1
2021 Low-Rate Overuse Flow Tracer (LOFT): An Efficient and Scalable Algorithm for Detecting Overuse Flows
abstract
Current probabilistic flow-size monitoring can only detect heavy hitters (e.g., flows utilizing 10 times their permitted bandwidth), but cannot detect smaller overuse (e.g., flows utilizing 50-100 % more than their permitted bandwidth). Thus, these systems lack accuracy in the challenging environment of high-throughput packet processing, where fast-memory resources are scarce. Nevertheless, many applications rely on accurate flow-size estimation, e.g., for network monitoring, anomaly detection and Quality of Service. We design, analyze, implement, and evaluate LOFT, a new approach for efficiently detecting overuse flows that achieves dramatically better properties than prior work. LOFT can detect 1.50x overuse flows in one second, whereas prior approaches can only reliably detect flows that overuse their allocation by at least 3x. We demonstrate LOFT's suitability for high-speed packet processing with implementations in the DPDK framework and on an FPGA.
Simon Scherrer, Che-Yu Wu, Yu-Hsi Chiang, Benjamin Rothenberger, Daniele Enrico Asoni, Arish Sateesan, Jo Vliegen, Nele Mentens, Hsu-Chun Hsiao, Adrian Perrig
SRDS6
2021 A Survey of Algorithmic and Hardware Optimization Techniques for Vision Convolutional Neural Networks on FPGAs
Arish Sateesan, Sharad Sinha, Kavallur Gopi Smitha, A. Prasad Vinod 0001
Neural Process. Lett.1
2020 Novel Bloom filter algorithms and architectures for ultra-high-speed network security applications
abstract
This paper proposes novel Bloom filter algorithms and FPGA architectures for high-speed searching applications. A Bloom filter is a memory structure that is used to test whether input search data are present in a table of stored data. Bloom filters are extensively used in network security solutions that apply traffic flow monitoring or deep packet inspection. Improving the speed of Bloom filters can therefore have a significant impact on the speed of many network applications. The most important components determining the speed of Bloom filters are hash functions. While hash functions in Bloom filters do not require strong cryptographic properties, they do need a minimized computational delay. We take on the challenge of developing ultra-high-speed Bloom filters on FPGAs by proposing a new noncryptographic hash function, called Xoodoo-NC, derived from the cryptographic permutation Xoodoo. Xoodoo-NC is a reducedround, reduced-state version of Xoodoo, inheriting Xoodoo's desired avalanche properties and low logical depth, resulting in an ultra-low-latency non-cryptographic hash function. We evaluate the performance of Bloom filter architectures based on Xoodoo-NC on a Xilinx UltraScale+FPGA and we compare the performance and resource occupation to existing Bloom filter implementations. We additionally compare our results to memories that use the built-in CAM cores in Xilinx UltraScale+ FPGAs. Our proposed algorithmic and architectural advances lead to Bloom filters that, to the best of our knowledge, outperform all other FPGA-based solutions.
Arish Sateesan, Jo Vliegen, Joan Daemen, Nele Mentens
DSD1