VLDB 2026 Research / reviewers in the wild / expert
Mehran Shoushtari Moghadam
dblp:354/6216
· DBLP profile ↗
17ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-1325-1664ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 7 first-author · 17 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: ADC-FIST: ADC-Free In/Near-Sensor Stochastic Object Tracking
Mehran Shoushtari Moghadam, Sepehr Tabrizchi, Ali Shafiee Sarvestani, Sercan Aygün, Arman Roohi, M. Hassan Najafi |
DATE | 1 |
| 2026 | Late Breaking Results: DAIQUIRI: Dynamic Quantization with Layer-wise Sensitivity Ranking for Hardware-Efficient LLMs
Tanha Tasfia, Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, M. Hassan Najafi, Sercan Aygün |
DATE | 3 |
| 2026 | Independent and Dynamic Vector Symbolic Architecture for Hardware-Efficient Edge AIabstractHyperdimensional computing (HDC), also known as vector symbolic architecture (VSA), is a brain-inspired paradigm offering lightweight and hardware-efficient cognitive learning. By encoding data into high-dimensional hypervectors (HVs), HDC supports single-pass training and inherent robustness, making it highly attractive for edge AI. Yet, two challenges impede its deployment: efficient on-chip generation of orthogonal HVs and adaptation to dynamic data sizes without costly retraining. This work introduces the Independent and Dynamic VSA (ID-VSA), which advances HDC through five key innovations. First, we propose a compact single-source HV generator based on low-discrepancy (LD) sequences, enabling orthogonal symbol vectors with minimal hardware cost. Second, we presentGaussian Polygon, a multiscale learning mechanism that performs Gaussian-like interpolation directly in the HV domain. Third, we extend HV generation to quasi-normal distributions (QNDs), supporting both symbol and level vectors from the same randomness source. Fourth, we incorporate true-random number generation to exploit device-level noise for unbiased HV creation. Finally, we demonstrate flexible multi-assignment encoding for efficient$n$-gram processing. Evaluations demonstrate the proposed methods achieve accuracy improvements of up to 1.07% and 2.50% for image datasets MNIST and Pneumonia MNIST, and 18.36% on the language dataset over conventional HDC models. For larger-scale workloads, the proposedGaussian Polygon-based designs achieve up to 4.33% improvement on the EuroSAT remote sensing dataset and up to 0.71% improvement on FractureMNIST3D medical dataset. Hardware synthesis in 45 nm technology confirms efficiency, achieving up to$370\times $lower power and$109\times $smaller area, establishingID-VSAas a scalable solution for real-time, hardware-efficient edge AI. Mehran Shoushtari Moghadam, Abu Kaisar Mohammad Masum, Sercan Aygün, M. Hassan Najafi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | All-in-Memory Stochastic Computing using ReRAMabstractAs the demand for efficient, low-power computing in embedded and edge devices grows, traditional computing methods are becoming less effective for handling complex tasks. Stochastic computing (SC) offers a promising alternative by approximating complex arithmetic operations, such as addition and multiplication, using simple bitwise operations, like majority or AND, on random bit-streams. While SC operations are inherently fault-tolerant, their accuracy largely depends on the length and quality of the stochastic bit-streams (SBS). These bit-streams are typically generated by CMOS-based stochastic bit-stream generators that consume over 80% of the SC system’s power and area. Current SC solutions focus on optimizing the logic gates but often neglect the high cost of moving the bit-streams between memory and processor. This work leverages the physics of emerging ReRAM devices to implement the entire SC flow in place: ❶ generating low-cost true random numbers and SBSs, ❷ conducting SC operations, and ❸ converting SBSs back to binary. Considering the low reliability of ReRAM cells, we demonstrate how SC’s robustness to errors copes with ReRAM’s variability. Our evaluation shows significant improvements in throughput (1.39 ×, 2.16 ×) and energy consumption (1.15 ×, 2.8 ×) over state-of-the-art (CMOS- and ReRAM-based) solutions, respectively, with an average image quality drop of 5% across multiple SBS lengths and image processing tasks. João Paulo C. de Lima, Mehran Shoushtari Moghadam, Sercan Aygün, Jerónimo Castrillón, M. Hassan Najafi, Asif Ali Khan |
DAC | 2 |
| 2025 | Late Breaking Results: On-the-Fly Hadamard Hypervector Processing for Efficient Hyperdimensional ComputingabstractInspired by the human brain, Hyperdimensional Computing (HDC) processes information efficiently by operating in high-dimensional space using hypervectors. While previous works focus on optimizing pregenerated hypervectors in software, this study introduces a novel on-the-fly vector generation method in hardware with $O(1)$ complexity, compared to the $O(N)$ iterative search used in conventional approaches to find the best orthogonal hypervectors. Our approach leverages Hadamard binary coefficients and unary computing to simplify encoding into addition-only operations after the generation stage in ASIC, implemented using inmemory computing. The proposed design significantly improves accuracy and computational efficiency across multiple benchmark datasets. Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, Sabrina Hassan Moon, Ahmed Mamdouh Mohamed Ahmed, M. Hassan Najafi, Dayane Reis, Sercan Aygün |
DAC | 2 |
| 2025 | In-Memory Arithmetic: Enabling Division with Stochastic LogicabstractDesigning an efficient arithmetic division circuit has long been a major challenge. Traditional binary computation methods rely on complex algorithms that require multiple cycles, complex control logic, and substantial hardware resources. Implementing division with emerging in-memory computing technologies is even more challenging due to susceptibility to noise, process variation, and the complexity of binary division. In this work, we propose an in-memory division architecture leveraging stochastic computing (SC), an emerging technology known for its high fault tolerance and low-cost design. Our approach utilizes a magnetic tunnel junction (MTJ)-based memory architecture to efficiently execute logic-in-memory operations. Experimental results across various process variation conditions demonstrate the robustness of our method against hardware variations. To assess its practical effectiveness, we apply our approach to the Retinex Algorithm for image enhancement, demonstrating its viability in real-world applications. Farzad Razi, Mehran Shoushtari Moghadam, M. Hassan Najafi, Sercan Aygün, Marc D. Riedel |
DAC | 2 |
| 2025 | Breaking New Ground: Division Directly in MemoryabstractIn-memory computing (IMC) has emerged as a promising paradigm for overcoming the limitations of traditional von Neumann architectures by reducing data movement and enhancing computational efficiency. Despite significant advancements in this area, implementing complex arithmetic operations, such as division, directly within memory has remained an elusive challenge. This paper introduces a pioneering technique for performing division operations directly in memory, representing the first successful integration of such functionality into the IMC framework. Our approach leverages an innovative circuit based on an unconventional model of computing-stochastic computing. Our technique extends the computational capabilities of IMC systems and paves the way for lightweight division operations. Farzad Razi, Mehran Shoushtari Moghadam, M. Hassan Najafi, Sercan Aygün, Marc D. Riedel |
FCCM | 2 |
| 2025 | Quantum Image Processing: A Comparative Study of NEQR and FRQI Encoding Schemes with Hybrid Processing
Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, Lida Kouhalvandi, M. Hassan Najafi, Sercan Aygün |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Robust Data Processing for Vector Symbolic Computing
Mehran Shoushtari Moghadam, Abu Kaisar Mohammad Masum, Sercan Aygün, M. Hassan Najafi |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | GAN-BiLSTM-HDC: A Hybrid Framework for Robust and Hardware-Efficient Malware DetectionabstractHyperdimensional Computing (HDC) has emerged as a hardware-efficient paradigm for embedded malware detection, offering strong parallelism and low complexity. However, the accuracy and robustness of HDC classifiers remain highly dependent on the diversity and quality of training data, leaving them vulnerable to novel threats. To address this challenge, we introduce a generative adversarial network (GAN)-assisted augmentation framework for the Microprocessor without Interlocked Pipelined Stages-32 (MIPS32) malware generation. The GAN is trained on real-world MIPS32 malware binaries to produce previously unseen instruction sequences. The synthetic code stacks are filtered using a custom MIPS32 assembler for syntactic validation and a Bidirectional Long Short-Term Memory (BiLSTM)-based semantic critic to ensure logical coherence. Only validated samples are retained to expand the training set for the HDC classifier, thereby strengthening generalization and resilience against novel malware variants. Our preliminary results show an average generator loss ($\mathbf{G}$) of 2.68 over 200 epochs and a discriminator loss (D) converging to 0.63, indicating that the GAN is learning to generate realistic and diverse outputs. This hybrid GAN-BiLSTM-HDC framework shows strong potential for enhancing classification accuracy, resilience, and efficiency in resource-constrained, real-time malware detection systems. Emilien Meyer, Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, Lida Kouhalvandi, Gourav Datta, Sercan Aygün, M. Hassan Najafi |
ICCD | 3 |
| 2025 | AMS-HD: Acute Mountain Sickness Detection with Hyperdimensional ComputingabstractAcute mountain sickness (AMS) is a potentially life-threatening condition that affects many individuals traveling to high altitudes. Early diagnosis is crucial, especially for travelers who may not have immediate access to medical resources. While traditional machine learning (ML) methods have been used to detect AMS using biomedical data (e.g., heart rate, blood oxygen saturation, respiration rate, blood pressure, and body temperature), hyperdimensional computing (HDC) has yet to be explored for this purpose using the few of biomedical data. Previous classification methods fall short of balancing accuracy with low hardware complexity, but HDC offers a promising solution. HDC provides a hardware-efficient alternative solution, making it well-suited for resource-constrained environments, such as wearable devices. Its lightweight architecture and efficient memory management make it ideal for embedded systems, enabling real-time AMS detection with accuracy comparable to traditional ML models. We introduce AMS-HD, a novel framework that leverages custom feature engineering and quasi-random hyper-vector encoding to further enhance the efficiency and accuracy of HDC for AMS detection. The proposed framework demonstrates the potential for seamless integration into wearable biomedical devices for on-the-go health monitoring. Abu Kaisar Mohammad Masum, Reeti Pradhananga, Jonas I. Schmidt, Mehran Shoushtari Moghadam, M. Hassan Najafi, Bige D. Unluturk, Ulkuhan Guler 0001, Sercan Aygün |
ISCAS | 4 |
| 2025 | TRUE-BSG: A True Random Bit-Stream Generator for Fast and Efficient Stochastic ComputingabstractStochastic computing (SC) leverages random bitstreams to perform arithmetic operations, offering ultra-low-cost, fault-tolerant, and highly parallelizable computations. The quality of these bit-streams is crucial for the accuracy and reliability of SC. This paper introduces TRUE-BSG, a novel true random bit-stream generator designed for fast and energy-efficient SC. Unlike state-of-the-art (SoTA) pseudo-random and quasi-random bit-stream generators, TRUE-BSG utilizes a high-quality true random number generator (TRNG), capable of producing random bits at a rate of 1 Gigabit per second. Our TRNG ensures high entropy and minimal correlation. TRUE-BSG shows comparable accuracy to software-based generators and better energy efficiency than SoTA bit-stream generators, making it an ideal solution for resource-constrained devices. Mehran Shoushtari Moghadam, Shelby Williams, Abu Kaisar Mohammad Masum, M. Hassan Najafi, Sercan Aygün, Magdy A. Bayoumi |
ISCAS | 1 |
| 2025 | ID-VS A: Independent and Dynamic Vector Symbolic Architecture for Energy-Efficient Edge AlabstractHyperdimensional computing (HDC), also known as Vector Symbolic Architecture, has gained significant attention for its hardware-efficient and accurate cognitive processing capabilities. By leveraging high-dimensional vector representations (hypervectors-HVs), HDC enables lightweight, single-pass learning. However, efficient and dynamic HV generation remains a key challenge, particularly for fully online learning in edge Al applications. Most existing approaches rely on pseudo-random, offline-generated HVs, which are neither adaptive nor software-independent, limiting their practicality in scenarios with varying data sizes, such as multi-resolution image processing. This work introduces three key innovations to advance online HDC for edge Al. First, we propose a lightweight, dynamic HV generator that operates entirely on-chip, eliminating the need for pre-generated vectors. Second, we introduce Gaussian Polygon, a novel multi-scale learning mechanism inspired by the Gaussian Pyramid, which performs Gaussian-like interpolation directly in binary HVs, achieving high efficiency without traditional upscaling techniques. Third, we show how Gaussian Polygon learning enables dynamic adaptation in HDC without conventional retraining mechanisms. Our hardware implementation in 45nm technology demonstrates up to 490 × reductions in power consumption and 676 × in hardware area, establishing the proposed framework as a practical and scalable solution for real-time edge learning. Mehran Shoushtari Moghadam, Abu Kaisar Mohammad Masum, Sercan Aygün, M. Hassan Najafi |
ISLPED | 1 |
| 2024 | P2LSG: Powers-of-2 Low-Discrepancy Sequence Generator for Stochastic ComputingabstractStochastic Computing (SC) is an unconventional computing paradigm processing data in the form of random bit-streams. The accuracy and energy efficiency of SC systems highly depend on the stochastic number generator (SNG) unit that converts the data from conventional binary to stochastic bit-streams. Recent work has shown significant improvement in the efficiency of SC systems by employing low-discrepancy (LD) sequences such as Sobol and Halton sequences in the SNG unit. Still, the usage of many well-known random sequences for SC remains unexplored. This work studies some new random sequences for potential application in SC. Our design space exploration proposes a promising random number generator for accurate and energy-efficient SC. We propose P2LSG, a low-cost and energy-efficient Low-discrepancy Sequence Generator derived from Powers-of-2 Van der Corput (VDC) sequences. We evaluate the performance of our novel bit-stream generator for two SC image and video processing case studies: image scaling and scene merging. For the scene merging task, we propose a novel SC design for the first time. Our experimental results show higher accuracy and lower hardware cost and energy consumption compared to the state-of-the-art. Mehran Shoushtari Moghadam, Sercan Aygün, Mohsen Riahi Alam, M. Hassan Najafi |
ASPDAC | 1 |
| 2024 | Late Breaking Results: TriSC: Low-Cost Design of Trigonometric Functions with Quasi Stochastic ComputingabstractLow-cost and hardware-efficient design of trigonometric functions is challenging. Stochastic computing (SC), an emerging computing model processing random bit-streams, offers promising solutions for this problem. The existing implementations, however, often overlook the importance of the data converters necessary to generate the needed bit-streams. While recent advancements in SC bit-stream generators focus on basic arithmetic operations such as multiplication and addition, energy-efficient SC design of non-linear functions demands attention to both the computation circuit and the bit-stream generator. This work introduces TriSC, a novel approach for SC-based design of trigonometric functions enjoying state-of-the-art (SOTA) quasi-random bit-streams. Unlike SOTA SC designs of trigonometric functions that heavily rely on delay elements to decorrelate bit-streams, our approach avoids delay elements while improving the accuracy of the results. TriSC yields significant energy savings of up to 92% compared to SOTA. As two novel use cases studied for the first time in SC literature, we employ the proposed design for 2D image transformation and forward kinematics of a robotic arm, two computation-intensive applications demanding low-cost trigonometric designs. Sercan Aygün, Mehran Shoushtari Moghadam, M. Hassan Najafi |
DAC | 2 |
| 2024 | uHD: Unary Processing for Lightweight and Dynamic Hyperdimensional ComputingabstractHyperdimensional computing (HDC) is a novel computational paradigm that operates on long-dimensional vectors known as hypervectors. The hypervectors are constructed as long bit-streams and form the basic building blocks of HDC systems. In HDC, hypervectors are generated from scalar values without considering bit significance. HDC is efficient and robust for various data processing applications, especially computer vision tasks. To construct HDC models for vision applications, the current state-of-the-art practice utilizes two parameters for data encoding: pixel intensity and pixel position. However, the intensity and position information embedded in high-dimensional vectors are generally not generated dynamically in the HDC models. Consequently, the optimal design of hypervectors with high model accuracy requires powerful computing platforms for training. A more efficient approach is to generate hypervectors dynamically during the training phase. To this aim, this work uses low-discrepancy sequences to generate intensity hypervectors, while avoiding position hypervectors. Doing so eliminates the multiplication step in vector encoding, resulting in a power-efficient HDC system. For the first time in the literature, our proposed approach employs lightweight vector generators utilizing unary bit-streams for efficient encoding of data instead of using conventional comparator-based generators. Sercan Aygün, Mehran Shoushtari Moghadam, M. Hassan Najafi |
DATE | 2 |
| 2024 | All You Need is Unary: End-to-End Unary Bit-stream Processing in Hyperdimensional ComputingabstractHyperdimensional Computing (HDC) is a brain-inspired computing paradigm introduced to achieve energy efficiency with a lightweight and single-pass training model. Hypervectors (HVs) at the heart of the HDC systems play a fundamental role in elevating the accuracy and obtaining the desired performance. Image-based HV encoding requires two types of HVs: Position and Level HVs. State-of-the-art approaches utilize pseudo-random methods for generating these HVs, which might degrade system performance and cause higher power consumption due to poor randomness in HV generation. These conventional methods require iteratively calculating orthogonal Positional HVs for acceptable accuracy. This work proposes a fast, ultra-lightweight, and high-quality HV generator incorporating low-discrepancy random sequences and the emerging unary bit-stream processing. For the first time, we employ unary computing (UC) to generate Level HVs, demonstrating that there is no need for randomness in HDC systems. We generate Position HVs using a single-source quasi-random sequence with a recurrence property. Our proposed HV generation technique improves the overall HDC accuracy by up to 6.4% for the medical MNIST dataset while reducing the power consumption of HV generation by 98%. Mehran Shoushtari Moghadam, Sercan Aygün, Faeze S. Banitaba, M. Hassan Najafi |
ISLPED | 1 |