Sercan Aygün

dblp:202/7292 · also Sercan Aygun · DBLP profile ↗
← Back
39ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0002-4615-7914ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 7 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deterministic Hyperdimensional Learning with Rank Refinement (Student Abstract)
abstract
Hyperdimensional Computing (HDC) represents data as high-dimensional hypervectors that are robust and efficient for learning. Existing methods often rely on pseudo-random hypervector generation, which can suffer from poor orthogonality and high variance across runs, ultimately slowing convergence. These approaches typically require numerous iterations (20– 100) to achieve acceptable accuracy. We propose a method that utilizes deterministic Sobol-based linear projections and rank-based retraining to construct more stable and discriminative hypervectors, thereby reducing class confusion. Unlike pseudo-random initialization, our projections guarantee reproducibility and better coverage of the feature space. As a result, our approach achieves up to 97% accuracy in only 5 iterations. This makes our model up to 20× faster while simultaneously improving accuracy.
Abu Kaisar Mohammad Masum, Sercan Aygün
AAAI2
2026 Late Breaking Results: POSEiDON: Pose Estimation in Dynamic On-device Networks via Hyperdimensional Computing
abstract
On-device pose estimation is widely available through models, yet efficiently classifying the resulting landmarks on resource-constrained mobile platforms remains challenging. We propose POSEiDON, a hyperdimensional computing (HDC)-based pose classification pipeline that combines landmark-symbol and joint-angle encodings with low-discrepancy (Sobol) sequences. On a custom 4-class mobile dataset and a 5-class YOGA benchmark, POSEiDON achieves up to 97.79% and 88.44% accuracy, respectively, using a single training pass, while iterative baselines require 10–200 passes or estimators to reach comparable accuracy. An Artix-7 FPGA prototype and an Android implementation further show that the HDC stage adds only marginal power, latency, and memory overhead, indicating that HDC is a promising primitive for lightweight on-device pose recognition.
Colin Dupuis, Emilien Meyer, Abu Kaisar Mohammad Masum, Sercan Aygün
DATE4
2026 Late Breaking Results: SP-HD: Stochastic Projection-Based HyperDimensional Architecture for Near-Sensor Image Classification
abstract
This paper presents SP-HD, a near-sensor image classification architecture that combines stochastic computing (SC) and hyperdimensional computing (HDC) to enable energy-efficient and compact embedded intelligence. The proposed approach introduces a stochastic projection mechanism that converts input features into bitstreams, enabling bipolar multiplications to be performed with simple logic and in-memory accumulation, thereby eliminating costly multipliers and level hypervectors. A mixed-signal ReRAM-based implementation further reduces data movement by performing projection and accumulation directly within the memory fabric, while binary-weight classification minimizes circuit complexity. SP-HD achieves competitive accuracy across multiple image datasets and delivers 3μJ energy per inference with a compact 2.56mm2hardware footprint, significantly outperforming prior ReRAM compute-in-memory accelerators in both energy and area efficiency.
Ahmed Mamdouh, Sabrina Hassan Moon, Abu Kaisar Mohammad Masum, Emilien Meyer, Sercan Aygün, Dayane Reis
DATE5
2026 Late Breaking Results: ADC-FIST: ADC-Free In/Near-Sensor Stochastic Object Tracking
Mehran Shoushtari Moghadam, Sepehr Tabrizchi, Ali Shafiee Sarvestani, Sercan Aygün, Arman Roohi, M. Hassan Najafi
DATE4
2026 Late Breaking Results: DAIQUIRI: Dynamic Quantization with Layer-wise Sensitivity Ranking for Hardware-Efficient LLMs
Tanha Tasfia, Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, M. Hassan Najafi, Sercan Aygün
DATE5
2026 Efficient On-Device Estimation of Transcutaneous Oxygen Using Machine Learning for Photoluminescence-Based Wearables
Gokalp Cevik, Hakan Burak Karli, Vladimir Vakhter, Hohyeon Kim, Farnaz Niroui, Sercan Aygün, Bige D. Unluturk, Ulkuhan Guler 0001
ISLPED6
2026 XL-HD: Extended Learning in Hyperdimensional Computing via Deterministic Projections for In-Memory Accelerators
abstract
Hyperdimensional computing (HDC) is a promising approach for energy-efficient edge machine learning (ML), where low latency, low power, and tight memory budgets are essential. However, traditional HDC relies on symbolic binding and pseudo-random high-dimensional vectors, which require large dimensionality and heuristic updates to reach competitive accuracy, limiting deployment on edge hardware. We introduce XL-HD, a deterministic, projection-based, fully learnable HDC framework tailored for in-memory acceleration within edge computing systems. The method uses a fixed Sobol sequence to project binary inputs, extending learning beyond conventional HDC. During training, class prototypes are optimized in real-valued space and later binarized, enabling an entirely binary dot-product inference pipeline ideal for IMC hardware such as ReRAM crossbars. XL-HD achieves competitive accuracy on MNIST, UCIHAR, and ISOLET while maintaining a compact IMC-based inference engine with 0.395 mm2 area and only 0.40 μJ per single-cycle inference.
Sabrina Hassan Moon, Abu Kaisar Mohammad Masum, Sercan Aygün, Dayane Reis
ISLPED3
2026 MITRA: Reconfigurable, Low-Latency, and Power-Efficient In-Memory Stochastic Architecture for Transcendental Functions
Farzad Razi, Mehran Moghadam, M. Hassan Najafi, Sercan Aygün, Marc D. Riedel
ISLPED4
2026 Independent and Dynamic Vector Symbolic Architecture for Hardware-Efficient Edge AI
abstract
Hyperdimensional computing (HDC), also known as vector symbolic architecture (VSA), is a brain-inspired paradigm offering lightweight and hardware-efficient cognitive learning. By encoding data into high-dimensional hypervectors (HVs), HDC supports single-pass training and inherent robustness, making it highly attractive for edge AI. Yet, two challenges impede its deployment: efficient on-chip generation of orthogonal HVs and adaptation to dynamic data sizes without costly retraining. This work introduces the Independent and Dynamic VSA (ID-VSA), which advances HDC through five key innovations. First, we propose a compact single-source HV generator based on low-discrepancy (LD) sequences, enabling orthogonal symbol vectors with minimal hardware cost. Second, we presentGaussian Polygon, a multiscale learning mechanism that performs Gaussian-like interpolation directly in the HV domain. Third, we extend HV generation to quasi-normal distributions (QNDs), supporting both symbol and level vectors from the same randomness source. Fourth, we incorporate true-random number generation to exploit device-level noise for unbiased HV creation. Finally, we demonstrate flexible multi-assignment encoding for efficient$n$-gram processing. Evaluations demonstrate the proposed methods achieve accuracy improvements of up to 1.07% and 2.50% for image datasets MNIST and Pneumonia MNIST, and 18.36% on the language dataset over conventional HDC models. For larger-scale workloads, the proposedGaussian Polygon-based designs achieve up to 4.33% improvement on the EuroSAT remote sensing dataset and up to 0.71% improvement on FractureMNIST3D medical dataset. Hardware synthesis in 45 nm technology confirms efficiency, achieving up to$370\times $lower power and$109\times $smaller area, establishingID-VSAas a scalable solution for real-time, hardware-efficient edge AI.
Mehran Shoushtari Moghadam, Abu Kaisar Mohammad Masum, Sercan Aygün, M. Hassan Najafi
IEEE Trans. Very Large Scale Integr. Syst.3
2025 Comparison-Free Bit-Stream Generation for Cost-Efficient Unary Computing
abstract
Today, unconventional hardware design techniques based on simple data representations are receiving more and more attention. Unary computing is one of these techniques that processes data in the form of uniform bit-streams. The simplicity of implementing complex arithmetic operations and high tolerance to noise are the crucial advantages of unary systems. However, converting data from weighed binary radix to unary representation with existing comparator-based unary number generators is expensive regarding footprint area and power consumption. The problem aggregates when the number of inputs and data precision increase. This work proposes a low-cost, comparison-free, unary number generation mechanism for efficient data conversion from binary radix to unary representation. We introduce a serial and two parallel (an exact and an approximate) unary number generators. Synthesis results show that the proposed method reduces the hardware area, power consumption, and area-delay product for both serial and parallel designs compared to the state-of-the-art converter. We evaluate the efficiency of the proposed converter in four use cases.
Faeze S. Banitaba, Amir Hossein Jalilvand, M. Hassan Najafi, Sercan Aygün
DAC4
2025 Late Breaking Results: Automated Topology Generation for Power Amplifier Designs through BiLSTM-based DNN and Multi-objective Optimizations
abstract
This work presents an automated, intelligent methodology for optimizing power amplifier (PA) design by predicting the most suitable circuit topology-specifically, the input and output matching networks-for a given high electron mobility transistor (HEMT). A classification-based bidirectional long short-term memory (BiLSTM) deep neural network (DNN) is trained to determine the optimal PA topology, while multi-objective Pareto front-based optimization techniques refine the network’s hyperparameters, including the number of hidden layers and neurons. The proposed approach is adaptable to various HEMT models and is validated through the design and optimization of high-performance PAs using lumped elements and transmission lines, operating within the $1-2 \mathrm{GHz}$ frequency range. The method is demonstrated using the Cree CGH40010 GaN HEMT on a Rogers RO4350B substrate, achieving a power output of approximately 40 dBm, a power-added efficiency (PAE) of at least 50%, and a power gain exceeding 10dB.
Lida Kouhalvandi, Sercan Aygün, M. Hassan Najafi, Arman Roohi
DAC2
2025 All-in-Memory Stochastic Computing using ReRAM
abstract
As the demand for efficient, low-power computing in embedded and edge devices grows, traditional computing methods are becoming less effective for handling complex tasks. Stochastic computing (SC) offers a promising alternative by approximating complex arithmetic operations, such as addition and multiplication, using simple bitwise operations, like majority or AND, on random bit-streams. While SC operations are inherently fault-tolerant, their accuracy largely depends on the length and quality of the stochastic bit-streams (SBS). These bit-streams are typically generated by CMOS-based stochastic bit-stream generators that consume over 80% of the SC system’s power and area. Current SC solutions focus on optimizing the logic gates but often neglect the high cost of moving the bit-streams between memory and processor. This work leverages the physics of emerging ReRAM devices to implement the entire SC flow in place: ❶ generating low-cost true random numbers and SBSs, ❷ conducting SC operations, and ❸ converting SBSs back to binary. Considering the low reliability of ReRAM cells, we demonstrate how SC’s robustness to errors copes with ReRAM’s variability. Our evaluation shows significant improvements in throughput (1.39 ×, 2.16 ×) and energy consumption (1.15 ×, 2.8 ×) over state-of-the-art (CMOS- and ReRAM-based) solutions, respectively, with an average image quality drop of 5% across multiple SBS lengths and image processing tasks.
João Paulo C. de Lima, Mehran Shoushtari Moghadam, Sercan Aygün, Jerónimo Castrillón, M. Hassan Najafi, Asif Ali Khan
DAC3
2025 Late Breaking Results: On-the-Fly Hadamard Hypervector Processing for Efficient Hyperdimensional Computing
abstract
Inspired by the human brain, Hyperdimensional Computing (HDC) processes information efficiently by operating in high-dimensional space using hypervectors. While previous works focus on optimizing pregenerated hypervectors in software, this study introduces a novel on-the-fly vector generation method in hardware with $O(1)$ complexity, compared to the $O(N)$ iterative search used in conventional approaches to find the best orthogonal hypervectors. Our approach leverages Hadamard binary coefficients and unary computing to simplify encoding into addition-only operations after the generation stage in ASIC, implemented using inmemory computing. The proposed design significantly improves accuracy and computational efficiency across multiple benchmark datasets.
Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, Sabrina Hassan Moon, Ahmed Mamdouh Mohamed Ahmed, M. Hassan Najafi, Dayane Reis, Sercan Aygün
DAC7
2025 In-Memory Arithmetic: Enabling Division with Stochastic Logic
abstract
Designing an efficient arithmetic division circuit has long been a major challenge. Traditional binary computation methods rely on complex algorithms that require multiple cycles, complex control logic, and substantial hardware resources. Implementing division with emerging in-memory computing technologies is even more challenging due to susceptibility to noise, process variation, and the complexity of binary division. In this work, we propose an in-memory division architecture leveraging stochastic computing (SC), an emerging technology known for its high fault tolerance and low-cost design. Our approach utilizes a magnetic tunnel junction (MTJ)-based memory architecture to efficiently execute logic-in-memory operations. Experimental results across various process variation conditions demonstrate the robustness of our method against hardware variations. To assess its practical effectiveness, we apply our approach to the Retinex Algorithm for image enhancement, demonstrating its viability in real-world applications.
Farzad Razi, Mehran Shoushtari Moghadam, M. Hassan Najafi, Sercan Aygün, Marc D. Riedel
DAC4
2025 Breaking New Ground: Division Directly in Memory
abstract
In-memory computing (IMC) has emerged as a promising paradigm for overcoming the limitations of traditional von Neumann architectures by reducing data movement and enhancing computational efficiency. Despite significant advancements in this area, implementing complex arithmetic operations, such as division, directly within memory has remained an elusive challenge. This paper introduces a pioneering technique for performing division operations directly in memory, representing the first successful integration of such functionality into the IMC framework. Our approach leverages an innovative circuit based on an unconventional model of computing-stochastic computing. Our technique extends the computational capabilities of IMC systems and paves the way for lightweight division operations.
Farzad Razi, Mehran Shoushtari Moghadam, M. Hassan Najafi, Sercan Aygün, Marc D. Riedel
FCCM4
2025 Single-Pass Symbolic Learning for Real-Time Embedded Security
Alaaddin Goktug Ayar, Sercan Aygün, Martin Margala
ACM Great Lakes Symposium on VLSI2
2025 ParaHDC: Leveraging GPU Acceleration for Scalable Hyperdimensional Learning
Abu Kaisar Mohammad Masum, Sercan Aygün
ACM Great Lakes Symposium on VLSI2
2025 Quantum Image Processing: A Comparative Study of NEQR and FRQI Encoding Schemes with Hybrid Processing
Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, Lida Kouhalvandi, M. Hassan Najafi, Sercan Aygün
ACM Great Lakes Symposium on VLSI5
2025 Robust Data Processing for Vector Symbolic Computing
Mehran Shoushtari Moghadam, Abu Kaisar Mohammad Masum, Sercan Aygün, M. Hassan Najafi
ACM Great Lakes Symposium on VLSI3
2025 ReX-HD: A Deterministic ReRAM-Based Hyperdimensional Computing Framework for Edge Computing
Sabrina Hassan Moon, Ahmed Mamdouh, Abu Kaisar Mohammad Masum, Sercan Aygün, Dayane Reis
ACM Great Lakes Symposium on VLSI4
2025 GAN-BiLSTM-HDC: A Hybrid Framework for Robust and Hardware-Efficient Malware Detection
abstract
Hyperdimensional Computing (HDC) has emerged as a hardware-efficient paradigm for embedded malware detection, offering strong parallelism and low complexity. However, the accuracy and robustness of HDC classifiers remain highly dependent on the diversity and quality of training data, leaving them vulnerable to novel threats. To address this challenge, we introduce a generative adversarial network (GAN)-assisted augmentation framework for the Microprocessor without Interlocked Pipelined Stages-32 (MIPS32) malware generation. The GAN is trained on real-world MIPS32 malware binaries to produce previously unseen instruction sequences. The synthetic code stacks are filtered using a custom MIPS32 assembler for syntactic validation and a Bidirectional Long Short-Term Memory (BiLSTM)-based semantic critic to ensure logical coherence. Only validated samples are retained to expand the training set for the HDC classifier, thereby strengthening generalization and resilience against novel malware variants. Our preliminary results show an average generator loss ($\mathbf{G}$) of 2.68 over 200 epochs and a discriminator loss (D) converging to 0.63, indicating that the GAN is learning to generate realistic and diverse outputs. This hybrid GAN-BiLSTM-HDC framework shows strong potential for enhancing classification accuracy, resilience, and efficiency in resource-constrained, real-time malware detection systems.
Emilien Meyer, Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, Lida Kouhalvandi, Gourav Datta, Sercan Aygün, M. Hassan Najafi
ICCD6
2025 AMS-HD: Acute Mountain Sickness Detection with Hyperdimensional Computing
abstract
Acute mountain sickness (AMS) is a potentially life-threatening condition that affects many individuals traveling to high altitudes. Early diagnosis is crucial, especially for travelers who may not have immediate access to medical resources. While traditional machine learning (ML) methods have been used to detect AMS using biomedical data (e.g., heart rate, blood oxygen saturation, respiration rate, blood pressure, and body temperature), hyperdimensional computing (HDC) has yet to be explored for this purpose using the few of biomedical data. Previous classification methods fall short of balancing accuracy with low hardware complexity, but HDC offers a promising solution. HDC provides a hardware-efficient alternative solution, making it well-suited for resource-constrained environments, such as wearable devices. Its lightweight architecture and efficient memory management make it ideal for embedded systems, enabling real-time AMS detection with accuracy comparable to traditional ML models. We introduce AMS-HD, a novel framework that leverages custom feature engineering and quasi-random hyper-vector encoding to further enhance the efficiency and accuracy of HDC for AMS detection. The proposed framework demonstrates the potential for seamless integration into wearable biomedical devices for on-the-go health monitoring.
Abu Kaisar Mohammad Masum, Reeti Pradhananga, Jonas I. Schmidt, Mehran Shoushtari Moghadam, M. Hassan Najafi, Bige D. Unluturk, Ulkuhan Guler 0001, Sercan Aygün
ISCAS8
2025 TRUE-BSG: A True Random Bit-Stream Generator for Fast and Efficient Stochastic Computing
abstract
Stochastic computing (SC) leverages random bitstreams to perform arithmetic operations, offering ultra-low-cost, fault-tolerant, and highly parallelizable computations. The quality of these bit-streams is crucial for the accuracy and reliability of SC. This paper introduces TRUE-BSG, a novel true random bit-stream generator designed for fast and energy-efficient SC. Unlike state-of-the-art (SoTA) pseudo-random and quasi-random bit-stream generators, TRUE-BSG utilizes a high-quality true random number generator (TRNG), capable of producing random bits at a rate of 1 Gigabit per second. Our TRNG ensures high entropy and minimal correlation. TRUE-BSG shows comparable accuracy to software-based generators and better energy efficiency than SoTA bit-stream generators, making it an ideal solution for resource-constrained devices.
Mehran Shoushtari Moghadam, Shelby Williams, Abu Kaisar Mohammad Masum, M. Hassan Najafi, Sercan Aygün, Magdy A. Bayoumi
ISCAS5
2025 ID-VS A: Independent and Dynamic Vector Symbolic Architecture for Energy-Efficient Edge Al
abstract
Hyperdimensional computing (HDC), also known as Vector Symbolic Architecture, has gained significant attention for its hardware-efficient and accurate cognitive processing capabilities. By leveraging high-dimensional vector representations (hypervectors-HVs), HDC enables lightweight, single-pass learning. However, efficient and dynamic HV generation remains a key challenge, particularly for fully online learning in edge Al applications. Most existing approaches rely on pseudo-random, offline-generated HVs, which are neither adaptive nor software-independent, limiting their practicality in scenarios with varying data sizes, such as multi-resolution image processing. This work introduces three key innovations to advance online HDC for edge Al. First, we propose a lightweight, dynamic HV generator that operates entirely on-chip, eliminating the need for pre-generated vectors. Second, we introduce Gaussian Polygon, a novel multi-scale learning mechanism inspired by the Gaussian Pyramid, which performs Gaussian-like interpolation directly in binary HVs, achieving high efficiency without traditional upscaling techniques. Third, we show how Gaussian Polygon learning enables dynamic adaptation in HDC without conventional retraining mechanisms. Our hardware implementation in 45nm technology demonstrates up to 490 × reductions in power consumption and 676 × in hardware area, establishing the proposed framework as a practical and scalable solution for real-time edge learning.
Mehran Shoushtari Moghadam, Abu Kaisar Mohammad Masum, Sercan Aygün, M. Hassan Najafi
ISLPED3
2025 Always-On Sensing in Energy-Harvested Systems via Stochastic Intermittent Computing
abstract
This paper introduces Stochastic Intermittent Computing (STIC), a framework that integrates intermittent computing (ImC) and stochastic computing (SC) to enable always-on sensing in energy-harvested systems. STIC dynamically adjusts computational precision based on available energy, eliminating the need for non-volatile memory checkpointing traditionally used in ImC systems. By adapting precision in real-time, STIC ensures continuous operation even under severe power fluctuations, significantly improving energy efficiency and system resilience. Evaluation results demonstrate that STIC achieves substantial reductions in area, power, and energy consumption owing to the simplicity of SC and its tolerance to aggressive voltage scaling. Evaluations across multiple neural networks and charging traces confirm that STIC enables robust, low-power edge intelligence for resource-constrained environments.
Sepehr Tabrizchi, Mehran Moghadam, Ali Shafiee Sarvestani, Sercan Aygün, M. Hassan Najafi, Arman Roohi
ISLPED4
2025 Sobol Sequence Optimization for Hardware-Efficient Vector Symbolic Architectures
abstract
Hyperdimensional computing (HDC) is an emerging computing paradigm with significant promise for efficient and robust learning. In HDC, objects are encoded with high-dimensional vector symbolic sequences called hypervectors. The quality of hypervectors, defined by their distribution and independence, directly impacts the performance of HDC systems. Despite a large body of work on the processing parts of HDC systems, little to no attention has been paid to data encoding and the quality of hypervectors. Most prior studies have generated hypervectors using inherent random functions, such as MATLAB’s or Python’s random function. This work introduces an optimization technique for generating hypervectors by employing quasi-random sequences. These sequences have recently demonstrated their effectiveness in achieving accurate and low-discrepancy data encoding in stochastic computing systems. The study outlines the optimization steps for utilizing Sobol sequences to produce high-quality hypervectors in HDC systems. An optimization algorithm is proposed to select the most suitable Sobol sequences via indexes for generating minimally correlated hypervectors, particularly in applications related to symbol-oriented architectures. The performance of the proposed technique is evaluated in comparison to two traditional approaches of generating hypervectors based on linear-feedback shift registers and MATLAB random functions. The evaluation is conducted for three applications: 1) language; 2) headline; and 3) medical image classification. Our experimental results demonstrate accuracy improvements of up to 10.79%, depending on the vector size. Additionally, the proposed encoding hardware exhibits reduced energy consumption and a superior area-delay product.
Sercan Aygün, M. Hassan Najafi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Regional Weather Variable Predictions by Machine Learning With Near-Surface Observational and Atmospheric Numerical Data
abstract
Accurate and timely regional weather prediction is vital for sectors dependent on weather-related decisions. Traditional prediction methods, based on atmospheric equations, often struggle with coarse temporal resolutions and inaccuracies. This article presents a novel machine learning (ML) model, called Micro-Macro (MiMa), that integrates both near-surface observational data from Kentucky Mesonet stations (collected every 5 min, known as Micro data) and hourly atmospheric numerical outputs (termed as Macro data) for fine-resolution weather forecasting. The MiMa model employs an encoder-decoder transformer structure, with two encoders for processing multivariate data from both datasets and a decoder for forecasting weather variables over short time horizons. Each instance of the MiMa model, called a modelet, predicts the values of a specific weather parameter at an individual mesonet station. The approach is extended with Regional MiMa (Re-MiMa) modelets, which are designed to predict weather variables at ungauged locations by training on multivariate data from a few representative stations in a region, tagged with their elevations. Re-MiMa can provide highly accurate predictions across an entire region, even in areas without observational stations. Experimental results show that MiMa significantly outperforms current models, with Re-MiMa offering precise short-term forecasts for ungauged locations, marking a significant advancement in weather forecasting accuracy and applicability.
Yihe Zhang 0001, Bryce Turney, Purushottam Sigdel, Xu Yuan 0001, Eric Rappin, Adrian Lago, Sytske K. Kimball, Li Chen 0019, Paul J. Darby, Lu Peng 0001, Sercan Aygün, Yazhou Tu, M. Hassan Najafi, Nian-Feng Tzeng
IEEE Trans. Geosci. Remote. Sens.11
2025 Sorting it out in Hardware: A State-of-the-Art Survey
abstract
Sorting is a fundamental operation in various applications and a traditional research topic in computer science. Improving the performance of sorting operations can have a significant impact on many application domains. Much attention has been paid to hardware-based solutions for high-performance sorting. These are often realized with application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). Recently, in-memory sorting solutions have also been proposed to address the movement cost issue between memory and processing units, also known as the Von Neumann bottleneck. Due to the complexity of the sorting algorithms, achieving an efficient hardware implementation for sorting data is challenging. A large body of prior solutions is built on compare-and-swap (CAS) units. These are categorized as comparison-based sorting. Some recent solutions offer comparison-free sorting. In this survey, we review the latest works in the area of hardware-based sorting. We also discuss the recent hardware solutions for partial and stream sorting. Finally, we discuss some important concerns that need to be considered in the future designs of sorting systems.
Amir Hossein Jalilvand, Faeze S. Banitaba, Seyedeh Newsha Estiri, Sercan Aygün, M. Hassan Najafi
ACM Trans. Design Autom. Electr. Syst.4
2024 P2LSG: Powers-of-2 Low-Discrepancy Sequence Generator for Stochastic Computing
abstract
Stochastic Computing (SC) is an unconventional computing paradigm processing data in the form of random bit-streams. The accuracy and energy efficiency of SC systems highly depend on the stochastic number generator (SNG) unit that converts the data from conventional binary to stochastic bit-streams. Recent work has shown significant improvement in the efficiency of SC systems by employing low-discrepancy (LD) sequences such as Sobol and Halton sequences in the SNG unit. Still, the usage of many well-known random sequences for SC remains unexplored. This work studies some new random sequences for potential application in SC. Our design space exploration proposes a promising random number generator for accurate and energy-efficient SC. We propose P2LSG, a low-cost and energy-efficient Low-discrepancy Sequence Generator derived from Powers-of-2 Van der Corput (VDC) sequences. We evaluate the performance of our novel bit-stream generator for two SC image and video processing case studies: image scaling and scene merging. For the scene merging task, we propose a novel SC design for the first time. Our experimental results show higher accuracy and lower hardware cost and energy consumption compared to the state-of-the-art.
Mehran Shoushtari Moghadam, Sercan Aygün, Mohsen Riahi Alam, M. Hassan Najafi
ASPDAC2
2024 Late Breaking Results: TriSC: Low-Cost Design of Trigonometric Functions with Quasi Stochastic Computing
abstract
Low-cost and hardware-efficient design of trigonometric functions is challenging. Stochastic computing (SC), an emerging computing model processing random bit-streams, offers promising solutions for this problem. The existing implementations, however, often overlook the importance of the data converters necessary to generate the needed bit-streams. While recent advancements in SC bit-stream generators focus on basic arithmetic operations such as multiplication and addition, energy-efficient SC design of non-linear functions demands attention to both the computation circuit and the bit-stream generator. This work introduces TriSC, a novel approach for SC-based design of trigonometric functions enjoying state-of-the-art (SOTA) quasi-random bit-streams. Unlike SOTA SC designs of trigonometric functions that heavily rely on delay elements to decorrelate bit-streams, our approach avoids delay elements while improving the accuracy of the results. TriSC yields significant energy savings of up to 92% compared to SOTA. As two novel use cases studied for the first time in SC literature, we employ the proposed design for 2D image transformation and forward kinematics of a robotic arm, two computation-intensive applications demanding low-cost trigonometric designs.
Sercan Aygün, Mehran Shoushtari Moghadam, M. Hassan Najafi
DAC1
2024 uHD: Unary Processing for Lightweight and Dynamic Hyperdimensional Computing
abstract
Hyperdimensional computing (HDC) is a novel computational paradigm that operates on long-dimensional vectors known as hypervectors. The hypervectors are constructed as long bit-streams and form the basic building blocks of HDC systems. In HDC, hypervectors are generated from scalar values without considering bit significance. HDC is efficient and robust for various data processing applications, especially computer vision tasks. To construct HDC models for vision applications, the current state-of-the-art practice utilizes two parameters for data encoding: pixel intensity and pixel position. However, the intensity and position information embedded in high-dimensional vectors are generally not generated dynamically in the HDC models. Consequently, the optimal design of hypervectors with high model accuracy requires powerful computing platforms for training. A more efficient approach is to generate hypervectors dynamically during the training phase. To this aim, this work uses low-discrepancy sequences to generate intensity hypervectors, while avoiding position hypervectors. Doing so eliminates the multiplication step in vector encoding, resulting in a power-efficient HDC system. For the first time in the literature, our proposed approach employs lightweight vector generators utilizing unary bit-streams for efficient encoding of data instead of using conventional comparator-based generators.
Sercan Aygün, Mehran Shoushtari Moghadam, M. Hassan Najafi
DATE1
2024 Word2HyperVec: From Word Embeddings to Hypervectors for Hyperdimensional Computing
abstract
Word-aware sentiment analysis has posed a significant challenge over the past decade. Despite the considerable efforts of recent language models, achieving a lightweight representation suitable for deployment on resource-constrained edge devices remains a crucial concern. This study proposes a novel solution by merging two emerging paradigms, the Word2Vec language model and Hyperdimensional Computing, and introduces an innovative framework named Word2HyperVec. Our framework prioritizes model size and facilitates low-power processing during inference by incorporating embeddings into a binary space. Our solution demonstrates significant advantages, consuming only 2.2 W, up to 1.81 × more efficient than alternative learning models such as support vector machines, random forest, and multi-layer perceptron.
Alaaddin Goktug Ayar, Sercan Aygün, M. Hassan Najafi, Martin Margala
ACM Great Lakes Symposium on VLSI2
2024 All You Need is Unary: End-to-End Unary Bit-stream Processing in Hyperdimensional Computing
abstract
Hyperdimensional Computing (HDC) is a brain-inspired computing paradigm introduced to achieve energy efficiency with a lightweight and single-pass training model. Hypervectors (HVs) at the heart of the HDC systems play a fundamental role in elevating the accuracy and obtaining the desired performance. Image-based HV encoding requires two types of HVs: Position and Level HVs. State-of-the-art approaches utilize pseudo-random methods for generating these HVs, which might degrade system performance and cause higher power consumption due to poor randomness in HV generation. These conventional methods require iteratively calculating orthogonal Positional HVs for acceptable accuracy. This work proposes a fast, ultra-lightweight, and high-quality HV generator incorporating low-discrepancy random sequences and the emerging unary bit-stream processing. For the first time, we employ unary computing (UC) to generate Level HVs, demonstrating that there is no need for randomness in HDC systems. We generate Position HVs using a single-source quasi-random sequence with a recurrence property. Our proposed HV generation technique improves the overall HDC accuracy by up to 6.4% for the medical MNIST dataset while reducing the power consumption of HV generation by 98%.
Mehran Shoushtari Moghadam, Sercan Aygün, Faeze S. Banitaba, M. Hassan Najafi
ISLPED2
2023 A Linear-Time, Optimization-Free, and Edge Device-Compatible Hypervector Encoding
abstract
Hyperdimensional computing (HDC) offers a single-pass learning system by imitating the brain-like signal structure. HDC data structure is in random hypervector format for better orthogonality. Similarly, in bit-stream processing - aka stochastic computing- systems, low-discrepancy (LD) sequences are used for the efficient generation of uncorrelated bit-streams. However, LD-based hypervector generation has never been investigated before. This work studies the utilization of LD Sobol sequences as a promising alternative for encoding hypervectors. The new encoding technique achieves highly-accurate classification with a single-time training step without needing to iterate repeatedly over random rounds. The accuracy evaluations in an embedded environment exhibit a classification rate improvement of up to 9. 79% compared to the conventional random hypervector encoding.
Sercan Aygün, M. Hassan Najafi, Mohsen Imani
DATE1
2023 Reconvergent Path-aware Simulation of Bit-stream Processing
abstract
Few studies have explored the complex circuit simulation of stochastic and unary computing systems, which are referred to under the umbrella term of bit-stream processing. The computer simulation of multi-level cascaded circuits with reconvergent paths has not been largely examined in the context of bit-stream processing systems. This study addresses this gap and proposes a contingency table -based reconvergent path-aware simulation method for fast and efficient simulation of multi-level circuits. The proposed method exhibits significantly better runtime and accuracy.
Sercan Aygün, M. Hassan Najafi, Mohsen Imani, Ece Olcay Günes
ACM Great Lakes Symposium on VLSI1
2023 Bit-Stream Processing with No Bit-Stream: Efficient Software Simulation of Stochastic Vision Machines
abstract
Stochastic computing (SC) is an emerging paradigm that has come to the fore in computer vision applications in the last decade. Complex arithmetic circuitry is reduced to simple logic gates, fed with uniform random bit-streams. Due to the requirement of long bit-streams, the computer-aided simulation of SC systems is facing run-time and memory-use challenges. This work presents an efficient approach for emulating SC-based systems. The proposed simulation technique does not utilize actual bit-streams but produces similar results as if the traditional stochastic bit-streams were processed. The data are processed with the aid of a correlation-controlled contingency table (CT) construct. Our technique emulates three state-of-the-art stochastic bit-streams, namely, bit-streams with binomial distribution, pseudo-random, and low-discrepancy bit-streams. We validate the proposed technique by emulating three new SC image processing designs. We propose novel SC designs for (i) template matching, (ii) image compositing, and (iii) bilinear interpolation. Our experimental results show that our simulation technique provides comparable accuracy to processing actual bit-streams, but at a significantly lower run-time and memory usage.
Sercan Aygün, M. Hassan Najafi, Mohsen Imani, Ece Olcay Günes
ACM Great Lakes Symposium on VLSI1
2023 Optimizing Indoor Localization Accuracy with Neural Network Performance Metrics and Software-Defined IEEE 802.11az Wi-Fi Set-Up
abstract
Accurately classifying regions based on Wi-Fi signals can be a difficult task, especially when considering different frequency values. In this study, we aimed to improve the accuracy of indoor localization by developing a novel approach that does not rely on pre-trained models. To achieve this, fingerprints from the IEEE 802.11az standard were randomly selected, and the data samples were trained using parameterized station characteristics and neural network hyperparameters. The impact of each parameter on the localization accuracy was measured, and performance monitoring metrics such as F1-Measure and confusion matrix-based metrics were evaluated. Furthermore, the Thompson sampling (TS) algorithm was employed to determine the optimal parameters, which helped to achieve the best possible accuracy. The proposed approach demonstrated improved accuracy in region localization compared to conventional heuristic approaches which typically yield an accuracy range of 65% to 77%. The proposed approach achieved up to 80% accuracy in region localization and could be a promising solution for indoor localization in various settings.
Lida Kouhalvandi, Sercan Aygün, Ladislau Matekovits, Farshad Miramirkhani
WINCOM2
2023 Agile Simulation of Stochastic Computing Image Processing With Contingency Tables
abstract
The rapid computerized simulation of stochastic computing (SC) systems is a challenging problem. A method for agile simulation of SC image processing is proposed in this work. The input operands are processed with the aid of a correlation-controlled contingency table (CT) construct without using actual stochastic bit-streams. The proposed approach underlines the validity of CT simulation with 1) image compositing; 2) pattern detection; and 3) bilinear interpolation case studies. Using the corresponding error models, we emulate the state-of-the-art pseudo-random and quasi-random bit-streams. Experimental results show that the proposed approach achieves similar computation accuracy to the traditional SC simulation while performing runtime- and memory-efficient computations. The execution time reduces more than$200\times $for the image compositing task when emulating random bit-streams with CT. Pattern detection and bilinear interpolation further showed$76\times $and$22\times $lower memory usage, respectively, when employing CT.
Sercan Aygün, M. Hassan Najafi, Mohsen Imani, Ece Olcay Günes
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Sound Source Localization Using Stochastic Computing
abstract
Stochastic computing (SC) is an alternative computing paradigm that processes data in the form of long uniform bit-streams rather than conventional compact weighted binary numbers. SC is fault-tolerant and can compute on small, efficient circuits, promising advantages over conventional arithmetic for smaller computer chips. SC has been primarily used in scientific research, not in practical applications. Digital sound source localization (SSL) is a useful signal processing technique that locates speakers using multiple microphones in cell phones, laptops, and other voice-controlled devices. SC has not been integrated into SSL in practice or theory. In this work, for the first time to the best of our knowledge, we implement an SSL algorithm in the stochastic domain and develop a functional SC-based sound source localizer. The developed design can replace the conventional design of the algorithm. The practical part of this work shows that the proposed stochastic circuit does not rely on conventional analog-to-digital conversion and can process data in the form of pulse-width-modulated (PWM) signals. The proposed SC design consumes up to 39% less area than the conventional baseline design. The SC-based design can consume less power depending on the computational accuracy, for example, 6% less power consumption for 3-bit inputs. The presented stochastic circuit is not limited to SSL and is readily applicable to other practical applications such as radar ranging, wireless location, sonar direction finding, beamforming, and sensor calibration.
Peter Schober, Seyedeh Newsha Estiri, Sercan Aygün, Nima Taherinejad, M. Hassan Najafi
ICCAD3