VLDB 2026 Research / reviewers in the wild / expert
Brendan Reidy
dblp:266/8219 · also Brendan C. Reidy
· DBLP profile ↗
8ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0004-4320-6890ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpikeViT: A Memory-Efficient Mobile Spiking Vision TransformabstractSpiking Transformers constitute an emerging class of neural architectures that seek to unify the representational power of Transformer-based models with the computational efficiency of spiking neural networks (SNNs). By leveraging discrete spike-based communication and event-driven processing, Spiking Transformers enable temporally sparse computation while maintaining the global context modeling and scalability inherent to self-attention mechanisms. This integration facilitates energy-efficient sequence modeling and opens new avenues for deploying large-scale attention-based models on neuromorphic hardware. However, existing Spiking Transformer architectures often incur substantial memory overhead, limiting their suitability for deployment in resource-constrained environments such as edge devices. To address this limitation, we propose SpikeViT, an efficient Spiking Transformer architecture designed to minimize memory consumption while preserving representational capacity. The architecture adopts a parallel design, combining a convolutional SNN with a lightweight transformer, connected via bidirectional cross-modal bridges that enable efficient tokenization and integration of spike-based features. Experimental results on the CIFAR10-DVS dataset show that SpikeViT achieves competitive accuracy while reducing memory footprint by up to 50% compared to state-of-the-art models, making it well-suited for deployment in energy- and memory-constrained neuromorphic systems. James Seekings, Hasti Zanganeh, Brendan Reidy, Jason Kamran Eshraghian, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | PixelPrune: Optimizing AIoT Vision Systems via In-Sensor Segmentation and Adaptive Data Transfer
Mohammadreza Mohammadi, Mehrdad Morsali, Sepehr Tabrizchi, Brendan Reidy, Arman Roohi, Shaahin Angizi, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | Lightator: An Optical Near-Sensor Accelerator with Compressive Acquisition Enabling Versatile Image ProcessingabstractThis paper proposes a high-performance and energy-efficient optical near-sensor accelerator for vision applications, called Lightator. Harnessing the promising efficiency offered by photonic devices, Lightator features innovative compressive acquisition of input frames and fine-grained convolution operations for low-power and versatile image processing at the edge for the first time. This will substantially diminish the energy consumption and latency of conversion, transmission, and processing within the established cloud-centric architecture as well as recently designed edge accelerators. Our device-to-architecture simulation results show that with favorable accuracy, Lightator achieves 84.4 Kilo FPS/W and reduces power consumption by a factor of ~24× and 73× on average compared with existing photonic accelerators and GPU baseline. Mehrdad Morsali, Brendan Reidy, Deniz Najafi, Sepehr Tabrizchi, Mohsen Imani, Mahdi Nikdast, Arman Roohi, Ramtin Zand, Shaahin Angizi |
DAC | 2 |
| 2024 | HiRISE: High-Resolution Image Scaling for Edge ML via In-Sensor Compression and Selective ROIabstractWith the rise of tiny IoT devices powered by machine learning (ML), many researchers have directed their focus toward compressing models to fit on tiny edge devices. Recent works have achieved remarkable success in compressing ML models for object detection and image classification on microcontrollers with small memory, e.g., 512kB SRAM. However, there remain many challenges prohibiting the deployment of ML systems that require high-resolution images. Due to fundamental limits in memory capacity for tiny IoT devices, it may be physically impossible to store large images without external hardware. To this end, we propose a high-resolution image scaling system for edge ML, called HiRISE, which is equipped with selective region-of-interest (ROI) capability leveraging analog in-sensor image scaling. Our methodology not only significantly reduces the peak memory requirements, but also achieves up to 17.7× reduction in data transfer and energy consumption. Brendan Reidy, Sepehr Tabrizchi, Mohammadreza Mohammadi, Shaahin Angizi, Arman Roohi, Ramtin Zand |
DAC | 1 |
| 2023 | Heterogeneous Integration of In-Memory Analog Computing Architectures with Tensor Processing UnitsabstractTensor processing units (TPUs), specialized hardware accelerators for machine learning tasks, have shown significant performance improvements when executing convolutional layers in convolutional neural networks (CNNs). However, they struggle to maintain the same efficiency in fully connected (FC) layers, leading to suboptimal hardware utilization. In-memory analog computing (IMAC) architectures, on the other hand, have demonstrated notable speedup in executing FC layers. This paper introduces a novel, heterogeneous, mixed-signal, and mixed-precision architecture that integrates an IMAC unit with an edge TPU to enhance mobile CNN performance. To leverage the strengths of TPUs for convolutional layers and IMAC circuits for dense layers, we propose a unified learning algorithm that incorporates mixed-precision training techniques to mitigate potential accuracy drops when deploying models on the TPU-IMAC architecture. The simulations demonstrate that the TPU-IMAC configuration achieves up to 2.59× performance improvements, and 88% memory reductions compared to conventional TPU architectures for various CNN models while maintaining comparable accuracy. The TPU-IMAC architecture shows potential for various applications where energy efficiency and high performance are essential, such as edge computing and real-time processing in mobile devices. The unified training algorithm and the integration of IMAC and TPU architectures contribute to the potential impact of this research on the broader machine learning landscape. Mohammed E. Elbtity, Brendan Reidy, Md Hasibul Amin, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | Work in Progress: Real-time Transformer Inference on Edge AI AcceleratorsabstractTransformer models have become a dominant architecture in the world of machine learning. From natural language processing to more recent computer vision applications, Transformers have shown remarkable results and established a new state-of-the-art in many domains. However, this increase in performance has come at the cost of ever-increasing model sizes requiring more resources to deploy. Machine learning (ML) models are used in many real-world systems, such as robotics, mobile devices, and internet of things (IoT) devices, that require fast inference with low energy consumption. For batterypowered devices, lower energy consumption directly translates into longer battery life. To address these issues, several edge AI accelerators have been developed. Among these, the Coral Edge TPU has shown promising results for image classification while maintaining very low energy consumption. Many of these devices, including the Coral TPU, were originally designed to accelerate convolutional neural networks, making deployment of Transformers challenging. Here, we propose a methodology to deploy Transformers on Edge TPU. We provide extensive latency, power, and energy comparisons among the leading edge devices and show that our methodology allows for real-time inference of Transformers while maintaining the lowest power and energy consumption of other edge devices on the market. Brendan Reidy, Mohammadreza Mohammadi, Mohammed E. Elbtity, Heath Smith, Ramtin Zand |
RTAS | 1 |
| 2022 | APTPU: Approximate Computing Based Tensor Processing UnitabstractWe propose an approximate tensor processing unit (APTPU), which includes two main components: (1) approximate processing elements (APEs) consisting of a low-precision multiplier and an approximate adder, and (2) pre-approximate units (PAUs) which are shared among the APEs in the APTPU’s systolic array, functioning as the steering logic to pre-process the operands and feed them to the APEs. We conduct extensive experiments to evaluate the performance of the APTPU across various configurations and various workloads. The results show that the APTPU’s systolic array achieves up to$5.2\times \textit {TOPS}/mm^{2}$and$4.4\times \textit {TOPS}/W$improvements compared to that of a conventional systolic array design. The comparison between the proposed APTPU and in-house TPU designs shows that we can achieve approximately$2.5\times $and$1.2\times $area and power reduction, respectively, while realizing comparable accuracy. Finally, a comparison with the state-of-the-art approximate systolic arrays shows that the APTPU can realize up to$1.58\times $,$2\times $, and$1.78\times $, reduction in delay, power, and area, respectively, while using similar design specifications and synthesis constraints. Mohammed E. Elbtity, Peyton Chandarana, Brendan Reidy, Jason Kamran Eshraghian, Ramtin Zand |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | TSV Extrusion Morphology Classification Using Deep Convolutional Neural NetworksabstractIn this paper, we utilize deep convolutional neural networks (CNNs) to classify the morphology of through-silicon via (TSV) extrusion in three dimensional (3D) integrated circuits (ICs). TSV extrusion is a crucial reliability concern which can deform and crack interconnect layers in 3D ICs and cause device failures. Herein, the white light interferometry (WLI) technique is used to obtain the surface profile of the extruded TSVs. We have developed a program that uses raw data obtained from WLI to create a TSV extrusion morphology dataset, including TSV images with 54 × 54 pixels that are labeled and categorized into three morphology classes. Four CNN architectures with different network complexities are implemented and trained for TSV extrusion morphology classification application. Data augmentation and dropout approaches are utilized to realize a balance between overfitting and underfitting in the CNN models. Results obtained show that the CNN model with optimized complexity, dropout, and data augmentation can achieve a classification accuracy comparable to that of a human expert. Brendan Reidy, Golareh Jalilvand, Tengfei Jiang, Ramtin Zand |
ICMLA | 1 |