Tianrui Ma

dblp:326/4027 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-0894-6653ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 4 first-author · 17 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CommILP: Synthesizing Communication Infrastructure for Domain Computing Platforms
abstract
Domain-specific accelerator-rich platforms are promising for high-throughput, low-power edge systems. Prior domain design-space exploration (Domain-DSE) methods [1] efficiently allocate processing elements (PEs) and bind applications, but largely assume simplified, centralized communication within the hardware tile, which limits scalability and wastes resources. Novel research work is needed that takes advantage of designtime known communication patterns to synthesize custom intratile interconnects that minimize area and energy while preserving application throughput. This paper introduces CommILP for synthesizing intra-tile communication infrastructure given a platform allocation and PE-to-PE communication graphs for multiple applications. CommILP instantiates Communication Elements (CEs), determines a minimal interconnect topology, and assigns routing paths for each app and connection. It guarantees that application throughput is preserved while minimizing area and energy. CommILP introduces a hierarchical ILP formulation across platform, application, and hop-stack levels, and supports both App-Level and Domain-Aggregated modeling modes to balance solution quality and runtime. Evaluated on ASIC and FPGA backends using both real (OpenVX-40) and synthetic (RNDComm-100) domains, CommILP achieves up to 69.4% area and 73% energy savings over centralized baselines, while remaining scalable to over 100 applications. It complements existing platform DSE tools by bridging the gap between computation binding and hardware-efficient communication synthesis.
Qucheng Jiang, Jacob Ginesin, Oscar Kellner, Tianrui Ma, Gunar Schirner
ASP-DAC4
2026 Hardwired-Neuron Language Processing Units as General-Purpose Cognitive Substrates
abstract
The rapid advancement of Large Language Models (LLMs) has established language as a core general-purpose cognitive substrate, driving the demand for specialized Language Processing Units (LPUs) tailored for LLM inference. To overcome the growing energy consumption of LLM inference systems, this paper proposes a Hardwired-Neurons Language Processing Unit (HNLPU), which physically hardwires LLM weight parameters into the computational fabric, achieving several orders of magnitude computational efficiency improvement by extreme specialization. However, a significant challenge still lies in the scale of modern LLMs. A straightforward hardwiring of GPT-OSS-120B would require fabricating photomask sets valued at over 6 billion dollars, rendering this straightforward solution economically impractical.
Yang Liu 0466, Yongwei Zhao 0001, Yifan Hao 0001, Zifu Zheng, Weihao Kong, Zhangmai Li, Dongchen Jiang, Ruiyang Xia, Zhihong Ma, Zisheng Liu, Zhaoyong Wan, Yunqi Lu, Hongrui Guo, Zhe Wang 0017, Tianrui Ma, Mo Zou, Rui Zhang 0040, Ling Li 0001, Xing Hu 0001, Zidong Du, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002, Yunji Chen
ASPLOS (2)18
2026 Cambricon-CIM: Enabling Energy-Efficient and Error-Resilient Analog CIM Acceleration via Reformation of Coding Bases
abstract
Recently, multi-bit slicing has emerged as a promising technique to improve the energy efficiency of charge-domain Compute-In-Memory (CIM) accelerators by reducing the number of Analog-to-Digital (A/D) conversions. However, multi-bit slicing requires shift-and-add operations to reconstruct outputs, which exponentially amplify errors and cause significant accuracy degradation. Existing works mainly rely on hardware-aware retraining or noise-suppression techniques, incurring considerable design or power overhead. Thus, multi-bit CIM designs often face the dilemma of trading off energy efficiency for error resilience. In this paper, we propose Cambricon-CIM, a charge-domain multi-bit CIM accelerator that achieves both high energy efficiency and strong error resilience, without requiring retraining. The core insight is that the error amplification is proportional to digit weights; and by redefining these digit weights with smaller non-binary coding bases, it is possible to reduce the total error amplification. Leveraging this principle, CambriconCIM dynamically selects the minimal coding bases for every analog dot-product. With novel circuit and architectural support, Cambricon-CIM enables fast, low-overhead reconfiguration of coding bases at runtime. Experimental results show that Cambricon-CIM achieves 2.27× energy efficiency and 3.06× performance over RAELLA, a state-of-the-art error-resilient multi-bit slicing CIM architecture.
Hongrui Guo, Tianrui Ma, Zidong Du, Mo Zou, Yifan Hao 0001, Yongwei Zhao 0001, Rui Zhang 0040, Wei Li 0008, Xing Hu 0001, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002
HPCA2
2026 SPLATONIC: Architectural Support for 3D Gaussian Splatting SLAM via Sparse Processing
abstract
3D Gaussian splatting (3DGS) has emerged as a promising direction for SLAM due to its high-fidelity reconstruction and rapid convergence. However, 3DGS-SLAM algorithms remain impractical for mobile platforms due to their high computational cost, especially for their tracking process. This work introduces Splatonic, a sparse and efficient realtime 3DGS-SLAM algorithm-hardware co-design for resourceconstrained devices. Inspired by classical SLAMs, we propose an adaptive sparse pixel sampling algorithm that reduces the number of rendered pixels by up to$256 \times$while retaining accuracy. To unlock this performance potential on mobile GPUs, we design a novel pixel-based rendering pipeline that improves hardware utilization via Gaussian-parallel rendering and preemptive$\alpha$-checking. Together, these optimizations yield up to$121.7 \times$speedup on the bottleneck stages and$14.6 \times$end-toend speedup on off-the-shelf GPUs. To further address new bottlenecks introduced by our rendering pipeline, we propose a pipelined architecture that simplifies the overall design while addressing newly emerged bottlenecks in projection and aggregation. Evaluated across four 3DGS-SLAM algorithms, Splatonic achieves up to$274.9 \times$speedup and$4738.5 \times$energy savings over mobile GPUs and up to$25.2 \times$speedup and$241.1 \times$energy savings over state-of-the-art accelerators, all with comparable accuracy.
Xiaotong Huang, Tianrui Ma, Yuxiang Xiong, Fangxin Liu, Zhezhi He, Yiming Gan, Zihan Liu 0002, Jingwen Leng, Yu Feng 0007, Minyi Guo
HPCA3
2026 Systematic Methodology of Modeling and Design Space Exploration for CMOS Image Sensors
abstract
CMOS Image Sensors (CIS) are integral to both human and computer vision tasks, necessitating continuous improvements in key performance metrics such as latency, power, and noise. Despite experienced designers being able to make informed design decisions, novice designers and system architects face challenges due to the complex and expansive design space of CIS. This paper introduces a systematic methodology that elucidates the trade-offs among CIS performance metrics and enables efficient design space exploration. Specifically, we propose a first-principle-based CIS modeling method. By exposing low-level circuit parameters, our modeling method explicitly reveals the impacts of design changes on high-level metrics. Based on the modeling method, we propose a design space exploration process that swiftly evaluates and identifies the optimal CIS design, capable of exploring over 109 designs in under a minute without the need for time-consuming SPICE simulations. Our approach is validated through a case study and comparisons with real-world designs, demonstrating its practical utility in guiding early-stage CIS design.
Tianrui Ma, Ramakrishna Kakarala, Charles Shan, Weidong Cao 0001, Xuan Zhang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 SNAPPIX: Efficient-Coding-Inspired In-Sensor Compression for Edge Vision
abstract
Energy-efficient image acquisition on the edge is crucial for enabling remote sensing applications where the sensor node has weak compute capabilities and must transmit data to a remote server/cloud for processing. To reduce the edge energy consumption, this paper proposes a sensor-algorithm co-designed system called SNAPPIX, which compresses raw pixels in the analog domain inside the sensor. We use coded exposure (CE) as the in-sensor compression strategy as it offers the flexibility to sample, i.e., selectively expose pixels, both spatially and temporally. SnapPix has three contributions. First, we propose a task-agnostic strategy to learn the sampling/exposure pattern based on the classic theory of efficient coding. Second, we codesign the downstream vision model with the exposure pattern to address the pixel-level non-uniformity unique to CE-compressed images. Finally, we propose lightweight augmentations to the image sensor hardware to support our in-sensor CE compression. Evaluating on action recognition and video reconstruction, SnapPix outperforms state-of-the-art video-based methods at the same speed while reducing the energy by up to $15.4 \times$. We have open-sourced the code at: https://github.com/horizonresearch/SnapPix.
Weikai Lin, Tianrui Ma, Adith Boloor, Yu Feng 0007, Ruofan Xing, Xuan Zhang 0001, Yuhao Zhu 0001
DAC2
2025 DenSparSA: A Balanced Systolic Array Approach for Dense and Sparse Matrix Multiplication
abstract
Numerous studies have proposed hardware architectures to accelerate sparse matrix multiplication, but these approaches often incur substantial area and power overhead, significantly compromising their usage in dense scenarios. On the other hand, systolic arrays deliver high efficiency for dense matrix operations, but their application to sparse matrices remains challenging. An ideal design should process both dense and sparse matrices with high efficiency to satisfy performance and versatility requirements.In this paper, we introduce DenSparSA, a balanced systolic array centralized architecture that can execute sparse matrix computations with minimal overhead to original dense matrix computations. DenSparSA supports both single-side and dual-side unstructured sparse matrix multiplications with high efficiency. At the same time, the additional hardware required for managing sparsity is compact and decoupled from the conventional systolic array, allowing for minimal power overhead when switched back to dense matrix operations via circuit gating. The proposed design is implemented with Nangate 45 nm. Implementation results show that DenSparSA achieves a speedup ranging from $1.9 \times$ to $22 \times$ compared to the classic systolic array for sparse workloads, while maintaining relatively low area and power overhead. For dense workloads, the power overhead can be reduced to $\mathbf{1 2 \%}$ for BF16 and 5% for FP32. Compared with existing solutions for sparse acceleration, DenSparSA delivers competitive ($0.82 \times-1.32 \times$) efficiency in sparse scenarios and $1.17 \times-2.28 \times$ better efficiency for dense scenarios, indicating a better balance between both situations.
Tianrui Ma, An Zou
DAC4
2025 SEAL: A Single-Event Architecture for In-Sensor Visual Localization
abstract
Image sensors have low costs and broad applications, but the large data volume they generate can result in significant energy and latency overheads during data transfer, storage, and processing.This paper explores how shifting from traditional binary encoding to delay-based codes early on can address these inefficiencies, enabling keypoint detection and tracking within digital pixel sensors.The result is SEAL, a Single-Event Architecture for In-Sensor Localization.SEAL optimizes the entire pipeline between the pixel array and the sensor-processor interface by introducing a temporal processor co-designed with analog-to-time converters, followed by a custom heavily quantized frontend processor.Its implementation is fully digital, relying on off-the-shelf CMOS cells and EDA tools, and adheres to race logic's single-wire-per-variable and single-eventper-wire policies to maximize energy and area efficiency wherever possible.Our evaluation-combining analog and digital simulations, FPGA prototyping, and an end-to-end system analysis incorporating a host processor for visual inertial odometry (VIO) backend tasks-demonstrates a 16-61× reduction in the latency of keypoint detection and tracking compared to software baselines running on the host processor, and a 7× reduction in energy consumption compared to a standard digital pixel sensor without processing capabilities.Meanwhile, SEAL preserves robust tracking accuracy: on the EuRoC dataset, the average root mean square absolute trajectory error decreases by 1.0 cm for HybVIO and increases by just 0.3 cm for VINS-Mono compared to their original implementations.
Ryan Hou, Thomas Twomey, Vasileios Milionis, Evangelos Dikopoulos, Tianrui Ma, Yuhao Zhu 0001, Georgios Tzimpragos
ISCA5
2025 LoRASensE: Learnable Low-Rank Acquisition in Sensors for Efficient Edge Machine Vision
abstract
Integrating deep learning with ubiquitous image sensors has empowered various edge vision applications such as classification, segmentation, and detection. Deploying these data-driven applications requires holistic optimizations, from front-end sensing to back-end processing, within the limited resources of edge devices. While significant advances have been made in the efficient processing of sensory data in the back end with optimizations of learning algorithms (e.g., compression) and development of computing hardware (e.g., accelerators), the energy efficiency of front-end sensors remains significantly limited due to conventional high-fidelity image acquisition and the resulting massive off-chip data transfer.This paper proposes a domain-specific visual acquisition method, LoRASensE, learnable low-rank acquisition in sensors tailored for efficient data-driven edge vision applications. LoRASensE is an algorithm-hardware co-design framework that integrates a learned low-rank compressor into image sensors to acquire compressed features. Specifically, this compressor is optimized alongside downstream vision tasks to ensure end-to-end accuracy and is implemented with efficient analog processing hardware. Our extensive evaluations on real-world datasets across various vision applications demonstrate that LoRASensE can achieve a 12.5× compression ratio with a just 1-b compressor, minimal accuracy loss, and 86.9% energy saving compared to the conventional high-fidelity acquisition. Multi-dimensional comparisons further show that LoRASensE also significantly outperforms existing in-sensor compression methods.
Zhiqiang Yi, Tianrui Ma, Weidong Cao 0001
ISLPED3
2025 HARD: Hardening Real-Time Scheduling and Analysis for Accelerator Enabled Computing
abstract
Despite the advancements in supporting artificial intelligence, accelerator-enabled computing architectures still struggle to meet strict timing constraints due to the complex interactions between CPU cores and accelerators. Although various scheduling and response-time analysis techniques have been developed, a significant gap remains between the conservative hard real-time schedulability (i.e., worst-case response times) and the average measured schedulability on real systems. This pessimism significantly limits the deployment of hard real-time tasks on accelerator-enabled computing platforms. To address this, we propose HARD, a real-time scheduling approach that integrates scheduling strategies, response time analysis, and practical scheduler designs for general accelerator-enabled computing platforms. Benefiting the subtask level segmented characteristics that are ignored by classic schedulers, the proposed HARD can significantly improve the theoretically guaranteed hard real-time schedulability. Extensive experiments on off-the-shelf Intel CPUs and NVIDIA GPUs show that HARD outperforms state-of-the-art scheduling and analysis approaches, delivering a 11.3% improvement in hard real-time schedulability and a remarkable 45.1 % reduction in pessimism.
Yinchen Ni, Tianrui Ma, Jintao Chen 0001, Chongye Yang, Siwei Ye, Yuankai Xu, Yier Jin, An Zou
RTAS2
2025 PrivateEye: In-Sensor Privacy Preservation Through Optical Feature Separation
abstract
We address privacy issues in applications where images captured by an edge device (camera) are sent to the cloud for inference on utility tasks such as classification. Sending raw images to the cloud exposes them to data sniffing attacks and misuse by untrusted third-party service providers beyond the user's intended tasks. We propose an encoding scheme that not only evades direct visual inspection to the images or image reconstruction, but also prevents sensitive information from being ascertained. Unlike commonly used adversarial learning approaches, the proposed method is two-fold: first, it uses a diffractive optical neural network to spatially separate features corresponding to different tasks on the sensor plane in the optical domain. Then only the pixels corresponding to the utility task region are read. This encoding ensures that private features are never digitally stored on the edge device, thereby preventing privacy leakage. The proposed method successfully reduces the privacy retrieval in binary tasks with minimal accuracy loss (~ 2%) of the utility task, while reducing private task accuracy by ~ 35% and defending against reconstruction attacks with SSIM score of 0.43.
Adith Boloor, Weikai Lin, Tianrui Ma, Yu Feng 0007, Yuhao Zhu 0001, Xuan Zhang 0001
WACV3
2025 RoSE-Opt: Robust and Efficient Analog Circuit Parameter Optimization With Knowledge-Infused Reinforcement Learning
abstract
Design automation of analog circuits has long been sought. However, achieving robust and efficient analog design automation remains challenging. This article proposes a learning framework, RoSE-Opt, to achieve robust and efficient analog circuit parameter optimization. RoSE-Opt has two important features. First, it incorporates key domain knowledge of analog circuit design, such as circuit topology, couplings between circuit specifications, and variations of process, supply voltage, and temperature, into the learning loop. This strategy facilitates the training of an artificial agent capable of achieving design goals by identifying device parameters that are optimal and robust. Second, it exploits a two-level optimization method, that is, integrating Bayesian optimization (BO) with reinforcement learning (RL) to improve sample efficiency. In particular, BO is used for a coarse yet quick search of an initial starting point for optimization. This sets a solid foundation to efficiently train the RL agent with fewer samples. Experimental evaluations on benchmarking circuits show promising sample efficiency, extraordinary figure-of-merit in terms of design efficiency and design success rate, and Pareto optimality in circuit performance of our framework, compared to previous methods. Furthermore, this work thoroughly studies the performance of different RL optimization algorithms, such as deep deterministic policy gradients (DDPGs) with an off-policy learning mechanism and proximal policy optimization (PPO) with an on-policy learning mechanism. This investigation provides users with guidance on choosing the appropriate RL algorithms to optimize the device parameters of analog circuits. Finally, our study also demonstrates RoSE-Opt’s promise in parasitic-aware device optimization for analog circuits. In summary, our work reports a knowledge-infused BO-RL design automation framework for reliable and efficient optimization of analog circuits’ device parameters. Code implementation of our method can be found athttps://github.com/xz-group/RoSE.
Weidong Cao 0001, Tianrui Ma, Mouhacine Benosman, Xuan Zhang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 BlissCam: Boosting Eye Tracking Efficiency with Learned In-Sensor Sparse Sampling
abstract
Eye tracking is becoming an increasingly important task domain in emerging computing platforms such as Augmented/Virtual Reality (AR/VR). Today’s eye tracking system suffers from long end-to-end tracking latency and can easily eat up half of the power budget of a mobile VR device. Most existing optimization efforts exclusively focus on the computation pipeline by optimizing the algorithm and/or designing dedicated accelerators while largely ignoring the front-end of any eye tracking pipeline: the image sensor. This paper makes a case for co-designing the imaging system with the computing system. In particular, we propose the notion of “in-sensor sparse sampling”, whereby the pixels are drastically downsampled (by $20 \times$) within the sensor. Such in-sensor sampling enhances the overall tracking efficiency by significantly reducing 1) the power consumption of the sensor readout chain and sensor-host communication interfaces, two major power contributors, and 2) the work done on the host, which receives and operates on far fewer pixels. With careful reuse of existing pixel circuitry, our proposed BlissCam requires little hardware augmentation to support the in-sensor operations. Our synthesis results show up to $8.2 \times$ energy reduction and $1.4 \times$ latency reduction over existing eye tracking pipelines.
Yu Feng 0007, Tianrui Ma, Yuhao Zhu 0001, Xuan Zhang 0001
ISCA2
2024 Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN Inference
abstract
Charge-domain SRAM-based Computing-in-memory (CIM) proves to be a promising method for DNN inference, and benefits from avoiding data movement between computing units and memory. However, the high resolution Analog-to-Digital Converters (ADCs) dominates the energy consumption (up to 64%), limiting the energy efficiency of SRAM-CIM architectures. The main reason is the wide range of input analog values, requiring high resolution ADCs to convert the high precision averaged analog voltages into high bitwidth digital data. In this paper, to reduce the ADC overhead, we propose Cambricon-M, a novel Fibonacci-coded SRAM-based charge-domain CIM accelerator for DNN inference. Cambricon-M features the Fibonacci coding, which guarantees low density of ‘1’ in operands (i.e., the adjacent two bits of each ‘1’ are both ‘0’), narrowing the output voltage range and enabling low resolution ADCs. Further, Cambricon-M exploits the high bit-level sparsity to address the extra energy and area overhead caused by the larger bitwidth in Fibonacci coding. Specifically, Cambricon-M proposes zero-skipping methods to reduce ineffectual input/output, and the bit-slice based compression method to reduce memory capacity/bandwidth pressure. Experimental results show that Cambricon-M reduces ADC energy by 68.7%, and improves the energy efficiency 3.48× and 1.62× compared to TPUv4 and an ISAAC-based charge-domain SRAM-CIM accelerator.
Hongrui Guo, Mo Zou, Yifan Hao 0001, Zidong Du, Erxiang Ren, Yang Liu 0466, Yongwei Zhao 0001, Tianrui Ma, Rui Zhang 0040, Xing Hu 0001, Fei Qiao, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002
MICRO8
2023 Invited Paper: Learned In-Sensor Visual Computing: From Compression to Eventification
abstract
Visual computing is vital for numerous applications. In conventional visual computing systems, CMOS image sensors (CIS) act as pure imaging devices for capturing images, however, recent CIS designs increasingly integrate processing capabilities such as Deep Neural Networks (DNN), which give rise to a notion of in-sensor computing. In this paper, we propose a new concept, learned in-sensor visual computing, which exploits end-to-end optimization of in-sensor processing and downstream vision tasks to achieve better overall algorithm accuracy and adopts hardware/algorithm co-design to achieve ultra-low sensor energy consumption. Two examples of the learned in-sensor visual computing, Leca and EDGAzE, are demonstrated.
Yu Feng 0007, Tianrui Ma, Adith Boloor, Yuhao Zhu 0001, Xuan Zhang 0001
ICCAD2
2023 CAMJ: Enabling System-Level Energy Modeling and Architectural Exploration for In-Sensor Visual Computing
abstract
CMOS Image Sensors (CIS) are fundamental to emerging visual computing applications. While conventional CIS are purely imaging devices for capturing images, increasingly CIS integrate processing capabilities such as Deep Neural Network (DNN). Computational CIS expand the architecture design space, but to date no comprehensive energy model exists. This paper proposes CamJ, a detailed energy modeling framework that provides a component-level energy breakdown for computational CIS and is validated against nine recent CIS chips. We use CamJ to demonstrate three use-cases that explore architectural trade-offs including computing in vs. off CIS, 2D vs. 3D-stacked CIS design, and analog vs. digital processing inside CIS. The code of CamJ is available at: https://github.com/horizon-research/CamJ.
Tianrui Ma, Yu Feng 0007, Xuan Zhang 0001, Yuhao Zhu 0001
ISCA1
2023 LeCA: In-Sensor Learned Compressive Acquisition for Efficient Machine Vision on the Edge
abstract
With the rapid advances of deep learning-based computer vision (CV) technology, digital images are increasingly consumed, not by humans, but by downstream CV algorithms. However, capturing high-fidelity and high-resolution images is energy-intensive. It not only dominates the energy consumption of the sensor itself (i.e. in low-power edge devices), but also contributes to significant memory burdens and performance bottlenecks in the later storage, processing, and communication stages. In this paper, we systematically explore a new paradigm of in-sensor processing, termed "learned compressive acquisition" (LeCA). Targeting machine vision applications on the edge, the LeCA framework exploits the joint learning of a sensor autoencoder structure with the downstream CV algorithms to effectively compress the original image into low-dimensional features with adaptive bit depth. We employ column-parallel analog-domain processing directly inside the image sensor to perform the compressive encoding of the raw image, resulting in meaningful hardware savings, and energy efficiency improvements. Evaluated within a modern machine vision processing pipeline, LeCA achieves 4×, 6×, and 8× compression ratios prior to any digital compression, with minimal accuracy loss of 0.97%, 0.98%, and 2.01% on ImageNet, outperforming existing methods. Compared with the conventional full-resolution image sensor and the state-of-the-art compressive sensing sensor, our LeCA sensor is 6.3× and 2.2× more energy-efficient while reaching a 2× higher compression ratio.
Tianrui Ma, Adith Boloor, Xiangxing Yang, Weidong Cao 0001, Patrick Williams, Nan Sun 0001, Ayan Chakrabarti, Xuan Zhang 0001
ISCA1
2022 HOGEye: Neural Approximation of HOG Feature Extraction in RRAM-Based 3D-Stacked Image Sensors
abstract
Many computer vision tasks, ranging from recognition to multi-view registration, operate on feature representation of images rather than raw pixel intensities. However, conventional pipelines for obtaining these representations incur significant energy consumption due to pixel-wise analog-to-digital (A/D) conversions and costly storage and computations. In this paper, we propose HOGEye, an efficient near-pixel implementation for a widely-used feature extraction algorithm—Histograms of Oriented Gradients (HOG). HOGEye moves the key but computation-intensive derivative extraction (DE) and histogram generation (HG) steps into the analog domain by applying a novel neural approximation method in a resistive random-access memory (RRAM)-based 3D-stacked image sensor. The co-location of perception (sensor) and computation (DE and HG) and the alleviation of A/D conversions allow HOGEye design to achieve significant energy saving. With negligible detection rate degradation, the entire HOGEye sensor system consumes less than 48μ[email protected] for an image resolution of 256 × 256 (equivalent to 24.3pJ/pixel) while the processing part only consumes 14.1pJ/pixel, achieving more than 2.5 × energy efficiency improvement than the state-of-the-art designs.
Tianrui Ma, Weidong Cao 0001, Fei Qiao, Ayan Chakrabarti, Xuan Zhang 0001
ISLPED1