EDBT 2026 Demo / reviewers in the wild / expert
Deniz Najafi
dblp:348/4864
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0008-2734-8935ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | INSPIRE: In-Sensor Compressed Weight Retrieval for Enhancing ViT Efficiency at EdgeabstractDeploying Vision Transformer (ViT) models on edge devices poses significant challenges due to the high bandwidth, energy demands, and latency associated with transmitting large weight parameter sets to the sensing unit, along with limited on-chip memory resources, which are often insufficient for storing these parameters. To address these constraints, we present a software-hardware co-design framework that incorporates a novel in-sensor Compressed Weight Retrieval mechanism within an intelligent vision sensor. This framework offers two key contributions. First, we propose an innovative hardware-friendly weight compression algorithm that substantially reduces bandwidth and power consumption by optimizing on-chip memory usage for storing weight parameters. Second, we leverage the exceptional efficiency of Silicon Photonic (SiPh) devices and design a novel in-sensor accelerator called INSPIRE for the first time to perform in-sensor retrieval of the compressed weights and parallel fine-grained convolution operations next to the pixel array, enabling low-power adaptable ViT inference on resource-constrained edge platforms. Our extensive simulation results show that INSPIRE can remarkably reduce the memory footprint of ViT results with favorable accuracy. Besides, INSPIRE significantly reduces the bandwidth and power requirements associated with storing weight parameters in on-chip memory. INSPIRE achieves up to 245.4 Kilo FPS/W and reduces the data transfer energy by a factor of ∼11× on average compared with 4-bit quantized ViTs. Deniz Najafi, Mohaiminul Al Nahian, Navid Khoshavi, Abdullah Al Arafat, Mamshad Nayeem Rizve, Mahdi Nikdast, Adnan Siraj Rakin, Shaahin Angizi |
DATE | 2 |
| 2026 | Photonics-Enabled Edge Processing: A Vision for Near-Sensor Optical IntelligenceabstractEdge intelligence is rapidly shifting computation from centralized cloud infrastructure toward the point of data generation. This shift is especially important for visual sensing systems, where continuous streams of high-dimensional pixel data must be converted, stored, transmitted, and processed under strict energy and latency constraints. While processing-in-sensor and processing-near-sensor architectures have reduced data movement, they remain limited by analog-to-digital conversion, memory access, electronic bandwidth, and the difficulty of supporting increasingly complex models near the sensor. This invited paper argues that integrated photonics can provide a new substrate for edge processing by enabling high-bandwidth, low-latency, and naturally parallel analog computation close to the sensing interface. We review the basic principles of photonic computing and discuss how it can be used to realize near-sensor multiply-and-accumulate operations. We then use recent work from our group as representative case studies, including optical in-sensor acceleration, optical near-sensor acceleration with compressive acquisition, near-sensor neuro-symbolic photonic computing, and in-sensor compressed weight retrieval for vision transformers. These examples motivate a broader research agenda in which photonics is not only a fast accelerator for neural operations, but also a system-level enabler for data-centric, energy-aware, and real-time edge intelligence. Deniz Najafi, Shaahin Angizi, Mahdi Nikdast |
ACM Great Lakes Symposium on VLSI | 1 |
| 2026 | Shallow Enough? A Cross-Architecture Study of Ultra-Low-Depth Neural Networks for Edge InferenceabstractEdge deployment imposes strict latency, memory, and energy constraints that scale directly with network depth, yet the question of which architectural paradigm offers the best accuracy-efficiency tradeoff at ultra-low depth (four to six layers) remains open. We present the first systematic comparison of convolutional, pure transformer, and hybrid CNN-transformer models under fixed shallow depth budgets. We develop an analytical framework characterizing the representational cost of local convolutional versus global attention-based feature mixing as a function of depth, and validate it with experiments on ImageNet-1K across architectures spanning MobileNetV2, DeiT, Swin, MobileViT, and EfficientFormer. Beyond FLOPs and accuracy, we report real-world latency and energy measurements on an NVIDIA Jetson Nano and examine deployment feasibility on Cortex-M class microcontrollers. Our experiments reveal how accuracy, latency, and memory scale across all three architecture families as depth decreases, providing practitioners with direct guidance on which model class to choose for a given depth and hardware budget. Chengwei Zhou, Haotian Yu, Shoma Yukawa, Deniz Najafi, Shaahin Angizi, Gourav Datta |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | DeepCompress-ViT: Rethinking Model Compression to Enhance Efficiency of Vision Transformers at the EdgeabstractVision Transformers (ViTs) excel in tackling complex vision tasks, yet their substantial size poses significant challenges for applications on resource-constrained edge devices. The increased size of these models leads to higher overhead (e.g., energy, latency) when transmitting model weights between the edge device and the server. Hence, ViTs are not ideal for edge devices where the entire model may not fit on the device. Current model compression techniques often achieve high compression ratios at the expense of performance degradation, particularly for ViTs. To overcome the limitations of existing works, we rethink model compression strategy for ViTs from first principle approach and develop an orthogonal strategy called DeepCompress-ViT. The objective of the DeepCompress-ViT is to encode the model weights to a highly compressed encoded representation using a novel training method, denoted as Unified Compression Training (UCT). Proposed UCT is accompanied by a decoding mechanism during inference, which helps to gain any loss of accuracy due to high compression ratio. We further optimize this decoding step by reordering the decoding operation using associative property of matrix multiplication, ensuring that the compressed weights can be decoded during inference without incurring any computational overhead. Our extensive experiments across multiple ViT models on modern edge devices show that DeepCompress-ViT can successfully compress ViTs at high compression ratios (> 14×). DeepCompress-ViT enables the entire model to be stored on edge device, resulting in unprecedented reductions in energy consumption (> 1470×) and latency (> 68×) for edge ViT inference. Our code is available at https://github.com/ML-Security-Research-LAB/DeepCompress-ViT. Abdullah Al Arafat, Deniz Najafi, Akhlak Mahmood, Mamshad Nayeem Rizve, Mohaiminul Al Nahian, Ranyang Zhou, Shaahin Angizi, Adnan Siraj Rakin |
CVPR | 3 |
| 2025 | Maximizing Sub-Array Resource Utilization in Digital Processing-in-Memory: A Versatile Hardware-Aware Approach
Gamana Aragonda, Deniz Najafi, Deepak Vungarala, Sepehr Tabrizchi, Arman Roohi, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Event-Driven Spatiotemporal Processing-In-Sensor with Phase Change Memory-based Optical Acceleration
Mehrdad Morsali, Deniz Najafi, Amin Shafiee, Sepehr Tabrizchi, Pietro Mercati, Mohsen Imani, Arman Roohi, Navid Khoshavi, Mahdi Nikdast, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Magnetic In/Near-Sensor Architectures: From Raw Sensing to Smart Processing
Sepehr Tabrizchi, Ali Shafiee Sarvestani, Md Hasibul Amin, Deniz Najafi, Shaahin Angizi, Ramtin Zand, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon PhotonicsabstractVision Transformers (ViTs) have emerged as a powerful architecture for computer vision tasks due to their ability to model long-range dependencies and global contextual relationships. However, their substantial compute and memory demands hinder efficient deployment in scenarios with strict energy and bandwidth limitations. In this work, we propose Opto-ViT, the first near-sensor, region-aware ViT accelerator leveraging silicon photonics (SiPh) for real-time and energy-efficient vision processing. Opto-ViT features a hybrid electronic-photonic architecture, where the optical core handles compute-intensive matrix multiplications using Vertical-Cavity Surface-Emitting Lasers (VCSELs) and Microring Resonators (MRs), while nonlinear functions and normalization are executed electronically. To reduce redundant computation and patch processing, we introduce a lightweight Mask Generation Network (MGNet) that identifies regions of interest in the current frame and prunes irrelevant patches before ViT encoding. We further co-optimize the ViT backbone using quantization-aware training and matrix decomposition tailored for photonic constraints. Experiments across device fabrication, circuit and architecture co-design, to classification, detection, and video tasks demonstrate that Opto-ViT achieves 100.4 KFPS/W with up to 84% energy savings with less than 1.6% accuracy loss, while enabling scalable and efficient ViT deployment at the edge. Mehrdad Morsali, Chengwei Zhou, Deniz Najafi, Sreetama Sarkar, Pietro Mercati, Navid Khoshavi, Peter A. Beerel, Mahdi Nikdast, Gourav Datta, Shaahin Angizi |
ICCAD | 3 |
| 2024 | Lightator: An Optical Near-Sensor Accelerator with Compressive Acquisition Enabling Versatile Image ProcessingabstractThis paper proposes a high-performance and energy-efficient optical near-sensor accelerator for vision applications, called Lightator. Harnessing the promising efficiency offered by photonic devices, Lightator features innovative compressive acquisition of input frames and fine-grained convolution operations for low-power and versatile image processing at the edge for the first time. This will substantially diminish the energy consumption and latency of conversion, transmission, and processing within the established cloud-centric architecture as well as recently designed edge accelerators. Our device-to-architecture simulation results show that with favorable accuracy, Lightator achieves 84.4 Kilo FPS/W and reduces power consumption by a factor of ~24× and 73× on average compared with existing photonic accelerators and GPU baseline. Mehrdad Morsali, Brendan Reidy, Deniz Najafi, Sepehr Tabrizchi, Mohsen Imani, Mahdi Nikdast, Arman Roohi, Ramtin Zand, Shaahin Angizi |
DAC | 3 |
| 2024 | OISA: Architecting an Optical In-Sensor Accelerator for Efficient Visual ComputingabstractTargeting vision applications at the edge, in this work, we systematically explore and propose a high-performance and energy-efficient Optical In-Sensor Accelerator architecture called OISA for the first time. Taking advantage of the promising efficiency of photonic devices, the OISA intrinsically implements a coarse-grained convolution operation on the input frames in an innovative minimum-conversion fashion in low-bit-width neural networks. Such a design remarkably reduces the power consumption of data conversion, transmission, and processing in the conventional cloud-centric architecture as well as recently-presented edge accelerators. Our device-to-architecture simulation results on various image data-sets demonstrate acceptable accuracy while OISA achieves 6.68 TOp/s/W efficiency. OISA reduces power consumption by a factor of 7.9 and 18.4 on average compared with existing electronic in-/near-sensor and ASIC accelerators. Mehrdad Morsali, Sepehr Tabrizchi, Deniz Najafi, Mohsen Imani, Mahdi Nikdast, Arman Roohi, Shaahin Angizi |
DATE | 3 |
| 2024 | Hybrid Magneto-electric FET-CMOS Integrated Memory Design for Instant-on ComputingabstractThe surge in the number of normally-off power-constraint Internet of Things (IoT) devices in recent years has amplified the demand for high-performance and energy-efficient in-memory computing architectures built on top of various non-volatile memories. Magneto-Electric Field Effect Transistors (MEFETs) have presented compelling design features suitable for logic and memory integration as an emerging post-CMOS FET. These include high-speed switching, minimal power usage, and non-volatility. This work introduces a new in-memory computing architecture designed for edge applications, leveraging emerging MEFETs. The proposed architecture enables the execution of both Boolean logic operations and Binary Content Addressable Memory (BCAM) operations within a single cycle. Furthermore, the energy consumption during the write operation of the proposed cell is optimized by introducing a new write circuitry. The outcomes of our device-to-architecture evaluation reveal approximately 43.5% and 96.9% reduction in read and write energy consumption, respectively, compared to the counterpart non-volatile memories. At the application level, the proposed architecture is applied to implement Binary Neural Networks (BNNs) based on AlexNet and VGG16. Our results showcase a decrease of approximately 54% in the overall energy consumption when implementing these networks using the proposed design compared to non-volatile in-memory computing designs. Deniz Najafi, Sepehr Tabrizchi, Ranyang Zhou, Mohammadreza Amel Solouki, Andrew Marshall, Arman Roohi, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | Accelerating Low Bit-width Neural Networks at the Edge, PIM or FPGA: A Comparative StudyabstractDeep Neural Network (DNN) acceleration with digital Processing-in-Memory (PIM) platforms at the edge is an actively-explored domain with great potential to not only address memory-wall bottlenecks but to offer orders of performance improvement in comparison to the von-Neumann architecture. On the other side, FPGA-based edge computing has been followed as a potential solution to accelerate compute-intensive workloads. In this work, adopting low-bit-width neural networks, we perform a solid and comparative inference performance analysis of a recent processing-in-SRAM tape-out with a low-resource FPGA board and a high-performance GPU to provide a guideline for the research community. We explore and highlight the key architectural constraints of these edge candidates that impact their overall performance. Our experimental data demonstrate that the processing-in-SRAM can obtain up to ~160x speed-up and up to 228x higher efficiency (img/s/W) compared to the under-test FPGA on the CIFAR-10 dataset. Nakul Kochar, Lucas Ekiert, Deniz Najafi, Deliang Fan, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 3 |