Ningyuan Cao

dblp:180/5504 · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-5323-1051ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Towards Generalizable and Efficient Circuit Topology Design: A Graph-Transformer-based Surrogate Model with Curriculum Learning
abstract
Unlike circuit parameter and sizing optimizations, the automated design of analog circuit topologies poses significant challenges for learning-based approaches. One challenge arises from the combinatorial growth of the topology space with circuit size, which limits the topology optimization efficiency. Moreover, traditional circuit evaluation methods are time-consuming, while the presence of data discontinuity in the topology space makes the accurate prediction of circuit performance exceptionally difficult for unseen topologies. To tackle these challenges, we design a novel Graph-Transformer-based Network (GTN) as the surrogate model for circuit evaluation, offering a substantial acceleration in the speed of circuit topology optimization without sacrificing performance. Our GTN model architecture is designed to embed voltage changes in circuit loops and current flows in connected devices, enabling accurate performance predictions for circuits with unseen topologies. To address the cold start problem when scaling GTN to large-scale circuits, we further introduce a curriculum learning strategy that progressively trains GTN from small-scale to large-scale circuits. This approach enables the model to first learn fundamental physical principles from simpler topologies and gradually adapt to complex configurations, effectively bridging the circuit complexity gap and improving prediction accuracy. Taking the power converter circuit design as an experimental task, our GTN model significantly outperforms an analytical approach and baseline methods directly utilizing graph neural networks. Furthermore, GTN achieves less than 5% relative error and 196× speed-up compared with high-fidelity simulation. Notably, our GTN surrogate model empowers an automatic circuit design framework to discover circuits of comparable quality to those identified through high-fidelity simulation while reducing the time required by up to 98.2%. With curriculum learning, the enhanced GTN achieves a 51% improvement for performance prediction of large-scale circuits compared to the GTN model without this strategy. These advancements establish GTN as a scalable framework for automated analog circuit design across varying circuit complexity levels.
Haoshu Lu, Shaoze Fan, Ningyuan Cao, Xin Zhang 0025, Jing Li 0025
ACM Trans. Design Autom. Electr. Syst.4
2025 Towards Uncertainty-aware Robotic Perception via Mixed-signal BNN Engine Leveraging Probabilistic Quantum Tunneling
abstract
Integrating deep learning with environmental perception enhances robotic adaptability to complex tasks. However, its “black-box” nature, such as the lack of uncertainty quantification, poses challenges for safety-critical applications, particularly in unstructured and noisy environments. Bayesian neural networks (BNNs) offer uncertainty quantification but are limited by high hardware overhead, restricting real-time implementation on resource-constrained robots. This paper presents a mixedsignal hardware accelerator for BNNs, utilizing probabilistic quantum tunneling in fully depleted silicon-on-insulator (FDSOI) transistors to enable efficient, real-time uncertainty quantification. Device measurements indicate high-quality Gaussian random variable generation, validated through quantile-quantile plot analysis, with a high correlation coefficient ($r=0.997$) at $200 \mathrm{fJ} /$ sample. Leveraging such compact randomness, the parallel architecture achieved $10^{3}-10^{4} \times$ latency reduction at less than $2 \times$ area cost. Finally, in uncertainty-aware visual localization application of autonomous underwater vehicles, the BNN model effectively distinguishes data noise from model uncertainty, yielding significant information gain and enhancing the resampling efficiency by $4.5 \times$ at same accuracy.
Likai Pei, Xingtian Wang, Xueji Zhao, Wanxin Huang, Boyang Cheng, Halid Mulaosmanovic, Stefan Dünkel, Dominik Kleimaier, Sven Beyer, Kai Ni 0004, Mengxue Hou, Michael T. Niemier, Ningyuan Cao
DAC14
2025 A Physically Unclonable Bio-Signal Encoder for Privacy-Preserving IoMT Applications
abstract
Next-generation Internet of Medical Things (IoMT) must balance efficient local decision-making with strong privacy protection for remote monitoring and diagnosis. However, the resource constraints of IoMT devices make this difficult. This paper introduces a novel bio-signal encoder within the hyperdimensional computing (HDC) framework. It leverages inherent transistor variations for physically unclonable encoding. This technique is called variation-based analog entropy (VAE). VAE reduces memory footprint and power consumption while enhancing security. It offers a scalable, energy-efficient solution that addresses IoT’s resource limitations while ensuring secure, intelligent healthcare applications.The VAE cell achieves high entropy robustness (30.23-57.76 dB signal-to-noise ratio) with only a 10-transistor footprint. It reduces HDC vector dimensions by 14.3× and improves accuracy by 2%. Compared to an SRAM baseline, it shrinks encoder area by 1.3-4.4× and cuts leakage power by 327×. Custom analog circuits for entropy management eliminate data conversion, boosting energy efficiency to 48.5 nJ per query. Evaluations on experimentally collected bio-fluid data demonstrate classification accuracy of 94.9 % and 97.6 % in 5-class and 3-class viscosity sensing, respectively. This highlights the potential of the system for IoMT applications. Furthermore, the VAE significantly enhances security, lowering attacker-restored data peak SNR by 16 dB, and making unauthorized recovery indistinguishable.
Boyang Cheng, Xueji Zhao, Steven Davis, Xiaoguang Dong 0001, Ningyuan Cao
IEEE Internet Things J.7
2024 VAE-HDC: Efficient and Secure Hyper-dimensional Encoder Leveraging Variation Analog Entropy
abstract
Hyperdimensional computing (HDC) is a bio-inspired machine learning paradigm utilizing hyperdimensional spaces for data representation. HDC significantly improves the ability to learn from sparse data and enhances noise robustness, and also enables parallel computation. Despite these advantages, HDC's reliance on high dimensionality and operational simplicity can lead to increased hardware costs and potential security vulnerabilities. This paper introduces a novel HDC encoding strategy using variation-based analog entropy (VAE), aiming to reduce memory footprint, lower power/energy consumption, and enhance security with physically-unclonable entropy generation. The VAE cell, with high entropy robustness (30.23 -- 57.76 dB SNR) and a small footprint (10 transistors), allows HDC to achieve a 14.3× reduction in vector dimensions, a 4.4× decrease in unit entropy cell area, and a 2% increase in accuracy compared to binary/multi-bit HDC. These benefits lead to a 1.3 -- 4.4× area and a 327× leakage power reduction when compared to an SRAM baseline. We have designed custom low-power circuits that enable end-to-end analog entropy storage, distribution management, binding, permutation, and bundling. This analog implementation prevents data conversion during feature vector encoding, thereby significantly enhancing energy efficiency (48.5nJ per query). Furthermore, with hardware-secured basis vectors, data security is significantly improved, as evidenced by the markedly degraded visual distinguish-ability of retrieved image data and maximum of 11 dB lower PSNR.
Boyang Cheng, Steven Davis, Zephan M. Enciso, Yiyang Zhang 0006, Ningyuan Cao
DAC6
2024 Graph-Transformer-based Surrogate Model for Accelerated Converter Circuit Topology Design
abstract
Unlike circuit parameter and sizing optimizations, the automated design of analog circuit topologies poses significant challenges for learning-based approaches. One challenge arises from the combinatorial growth of the topology space with circuit size, which limits the topology optimization efficiency. Moreover, traditional circuit evaluation methods are time-consuming, while the presence of data discontinuity in the topology space makes the accurate prediction of circuit performance exceptionally difficult for unseen topologies. To tackle these challenges, we design a novel Graph-Transformer-based Network (GTN) as the surrogate model for circuit evaluation, offering a substantial acceleration in the speed of circuit topology optimization without sacrificing performance. Our GTN model architecture is designed to embed voltage changes in circuit loops and current flows in connected devices, enabling accurate performance predictions for circuits with unseen topologies. Taking the power converter circuit design as an experimental task, our GTN model significantly outperforms an analytical approach and baseline methods directly utilizing graph neural networks. Furthermore, GTN achieves less than 5% relative error and 196× speed-up compared with high-fidelity simulation. Notably, our GTN surrogate model empowers an automatic circuit design framework to discover circuits of comparable quality to those identified through high-fidelity simulation while reducing the time required by up to 98.2%.
Shaoze Fan, Haoshu Lu, Ningyuan Cao, Xin Zhang 0025, Jing Li 0025
DAC4
2024 Towards Uncertainty-Quantifiable Biomedical Intelligence: Mixed-signal Compute-in-Entropy for Bayesian Neural Networks
abstract
To enhance AI robustness of mission-critical biomedical applications, Bayesian Neural Networks (BNNs) are instrumental for their structured approach to AI uncertainty estimation. However, implementing BNNs on edge devices is challenging due to significant resource demands for dynamic model updates and extensive inference sampling. Addressing this, we introduce a novel mixed-signal Compute-in-Memory with Entropy (CIE) hardware architecture that segregates dynamically-generated weights into analog entropy and digital parameters within a compute-in-memory framework, greatly reducing hardware overhead. We conducted thorough evaluations of the CIE architecture, assessing its performance against varying hardware imperfections, such as digital quantization errors, analog distribution imperfections, and device process variations, with a focus on both general and specialized tasks like Ventricular Arrhythmia (VA) detection. Our contributions include (1) a generic BNN acceleration strategy suitable for various CIM techniques and emerging devices, (2) a custom circuit design that improves hardware efficiency by 19.2×-440× compared to existing BNN accelerators, (3) a CIE-based BNN for VA detection enhancing accuracy, reducing uncertainty estimation time and energy/latency to 1.29μJ/1.55ms, and (4) identification of tolerable quantization error and device variation limits for BNNs in uncertainty estimation.
Likai Pei, Zephan M. Enciso, Boyang Cheng, Steven Davis, Zhenge Jia, Michael T. Niemier, Yiyu Shi 0001, Xiaobo Sharon Hu, Ningyuan Cao
ICCAD11
2024 Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures
abstract
Large Language Models (LLMs) deployed on edge devices learn through fine-tuning and updating a certain portion of their parameters. Although such learning methods can be optimized to reduce resource utilization, the overall required resources remain a heavy burden on edge devices. Instead, Retrieval-Augmented Generation (RAG), a resource-efficient LLM learning method, can improve the quality of the LLM-generated content without updating model parameters. However, the RAG-based LLM may involve repetitive searches on the profile data in every user-LLM interaction. This search can lead to significant latency along with the accumulation of user data. Conventional efforts to decrease latency result in restricting the size of saved user data, thus reducing the scalability of RAG as user data continuously grows. It remains an open question: how to free RAG from the constraints of latency and scalability on edge devices? In this paper, we propose a novel framework to accelerate RAG via Computing-in-Memory (CiM) architectures. It accelerates matrix multiplications by performing in-situ computation inside the memory while avoiding the expensive data transfer between the computing unit and memory. Our framework, Robust CiM-backed RAG (RoCR), utilizing a novel contrastive learning-based training method and noise-aware training, can enable RAG to efficiently search profile data with CiM. To the best of our knowledge, this is the first work utilizing CiM to accelerate RAG.
Ruiyang Qin, Zheyu Yan, Dewen Zeng, Zhenge Jia, Dancheng Liu, Ahmed Abbasi, Zhi Zheng 0002, Ningyuan Cao, Kai Ni 0004, Jinjun Xiong, Yiyu Shi 0001
ICCAD9
2024 LaMAGIC: Language-Model-based Topology Generation for Analog Integrated Circuits
abstract
In the realm of electronic and electrical engineering, automation of analog circuit is increasingly vital given the complexity and customized requirements of modern applications. However, existing methods only develop search-based algorithms that require many simulation iterations to design a custom circuit topology, which is usually a time-consuming process. To this end, we introduce LaMAGIC, a pioneering language model-based topology generation model that leverages supervised finetuning for automated analog circuit design. LaMAGIC can efficiently generate an optimized circuit design from the custom specification in a single pass. Our approach involves a meticulous development and analysis of various input and output formulations for circuit. These formulations can ensure canonical representations of circuits and align with the autoregressive nature of LMs to effectively addressing the challenges of representing analog circuits as graphs. The experimental results show that LaMAGIC achieves a success rate of up to 96% under a strict tolerance of 0.01. We also examine the scalability and adaptability of LaMAGIC, specifically testing its performance on more complex circuits. Our findings reveal the enhanced effectiveness of our adjacency matrix-based circuit formulation with floating-point input, suggesting its suitability for handling intricate circuit designs. This research not only demonstrates the potential of language models in graph generation, but also builds a foundational framework for future explorations in automated analog circuit design.
Chen-Chia Chang, Yikang Shen, Shaoze Fan, Jing Li 0025, Ningyuan Cao, Yiran Chen 0001, Xin Zhang 0025
ICML6
2024 CIPUF: Towards On-chip Learnable Anomaly Detection with Compute-In-PUF Architecture
abstract
With the rising threats of side-channel-attacks (SCA) and complexities of both on-chip and ambient environment, it is demanding to incorporate on-chip learnability into SCA anomaly detection. This will enable offline-trained models to adapt to the new power profiles of emerging SCA schemes, workloads, and varying environments. Existing SCA detection techniques often fall short in in-situ learning or pose excessive on-chip integration challenges due to resource and data demands. This paper presents a novel neuromorphic "compute-in-PUF" (CIPUF) architecture designed for SCA detection with on-chip learning capability and optimized area/energy/data overheads. We harness the PUF-based key generator as a hyperdimensional encoder, fostering few-shot learning capabilities. It showcases a state-of-the-art accuracy of 96% with offline training. While deployed on-chip, our architecture can adeptly re-calibrate its model at the introduction of unseen power profiles, and regain model accuracy by 45% with as few as 254 power trace samples during 0.45ms time frame. Meanwhile, compared with baseline design using separate PUF and learning modules, it achieves a area savings of 4.15X and energy savings of 12.8X. Nevertheless, it introduces a unique scalability advantages for both hardware key repository and learning accuracy for future technology.
Boyang Cheng, Zephan M. Enciso, Steven Davis, Ningyuan Cao
ISLPED5
2024 In-Situ Privacy via Mixed-Signal Perturbation and Hardware-Secure Data Reversibility
abstract
The swift proliferation of edge intelligence and ubiquitous data generation have heightened privacy into a pressing societal need. State-of-the-art reversible privacy protection requires significant hardware resources at the edge with distinct architecture for sensors and security, leading to a rise in hardware overhead and expanded attack surfaces. To address these challenges, we propose a time-domain mixed-signal (TD-MS) circuit architecture facilitating in-situ privacy (ISP) with hardware-secured data reversibility. The proposed TD-MS ISP unites data acquisition, data conversion, key generation, and protection while providing authorized device-specific unclonable data recovery for forensic purposes. At the system level, we demonstrate the attack resilience and privacy-preserving computation performance by implementing a custom embedded system applied to real-world surveillance scenarios. At the circuit level, we showcase custom TD-MS circuits, evaluating their energy and area efficiency against a digital baseline implemented in 65nm technology. With full-stack SPICE simulations for both the baseline digital and proposed TD-MS circuits, we measured a$670\times$energy/frame savings against the embedded system,$3\times$area reduction and$3.2\times$energy TD-MS gains over digital.
Steven Davis, Boyang Cheng, Muya Chang, Ningyuan Cao
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Privacy-by-Sensing with Time-domain Differentially-Private Compressed Sensing
abstract
With the ubiquitous IoT sensors and enormous real-time data generation, data privacy is becoming a critical societal concern. State-of-the-art privacy protection methods all demand significant hardware overhead due to computation-insensitive algorithms and divided sensor/security architecture. In this paper, we propose a generic time-domain circuit architecture that protects raw data by enabling a differentially-private compressed sensing (DP-CS) algorithm secured by physical unclonable functions (PUF). To address privacy concerns and hardware overhead at the same time, a robust unified PUF and time-domain mixed-signal (TD-MS) module are designed, where PUF enables private and secure entropy generation. To evaluate the proposed design against a digital baseline, we performed experiments based on synthesized circuits and SPICE simulation and measured a 2.9x area reduction and 3.2x energy gains. We also measured high-quality PUF generation with TD-MS circuit with a inter-die Hamming distance of 52% and a low intra-die Hamming distance of 2.8%. Furthermore, we performed attack and algorithm performance measurements demonstrating the proposed design preserves data privacy even under attack, and the machine learning performance has minimal degradation (within 2%) compared to the digital baseline.
Boyang Cheng, Pengyu Zeng, Steven Davis, Muya Chang, Ningyuan Cao
DATE6
2023 Memory-Based Computing for Energy-Efficient AI: Grand Challenges
abstract
The remarkable progress in artificial intelligence (AI) has ushered in a new era characterized by models with billions of parameters, enabling extraordinary capabilities across diverse domains. However, these achievements come at a significant cost in terms of memory and energy consumption. The growing demand for computational resources raises grand challenges for the sustainable development of energy-efficient AI systems. This paper delves into the paradigm of memory-based computing as a promising avenue to address these challenges. By capitalizing on the inherent characteristics of memory and its efficient utilization, memory-based computing offers a novel approach to enhance AI performance while reducing the associated energy costs. Our paper systematically analyzes the multifaceted aspects of this paradigm, highlighting its potential benefits and outlining the challenges it poses. Through an exploration of various methodologies, architectures, and algorithms, we elucidate the intricate interplay between memory utilization, computational efficiency, and AI model complexity. Furthermore, we review the evolving area of hardware and software solutions for memory-based computing, underscoring their implications for achieving energy-efficient AI systems. As AI continues its rapid evolution, identifying the key challenges and insights presented in this paper serve as a foundational guide for researchers striving to navigate the complex field of memory-based computing and its pivotal role in shaping the future of energy-efficient AI.
Foroozan Karimzadeh, Mohsen Imani, Bahar Asgari, Ningyuan Cao, Yingyan (Celine) Lin, Yan Fang 0002
VLSI-SoC4
2023 Power Converter Circuit Design Automation Using Parallel Monte Carlo Tree Search
abstract
The tidal waves of modern electronic/electrical devices have led to increasing demands for ubiquitous application-specific power converters. A conventional manual design procedure of such power converters is computation- and labor-intensive, which involves selecting and connecting component devices, tuning component-wise parameters and control schemes, and iteratively evaluating and optimizing the design. To automate and speed up this design process, we propose an automatic framework that designs custom power converters from design specifications using Monte Carlo Tree Search. Specifically, the framework embraces the upper-confidence-bound-tree (UCT), a variant of Monte Carlo Tree Search, to automate topology space exploration with circuit design specification-encoded reward signals. Moreover, our UCT-based approach can exploit small offline data via the specially designed default policy and can run in parallel to accelerate topology space exploration. Further, it utilizes a hybrid circuit evaluation strategy to substantially reduce design evaluation costs. Empirically, we demonstrated that our framework could generate energy-efficient circuit topologies for various target voltage conversion ratios. Compared to existing automatic topology optimization strategies, the proposed method is much more computationally efficient—the sequential version can generate topologies with the same quality while being up to 67% faster. The parallelization schemes can further achieve high speedups compared to the sequential version.
Shaoze Fan, Ningyuan Cao, Jing Li 0025, Xin Zhang 0025
ACM Trans. Design Autom. Electr. Syst.4
2022 Stochastic Mixed-Signal Circuit Design for In-Sensor Privacy
abstract
The ubiquitous data acquisition and extensive data exchange of sensors pose severe security and privacy concerns for the end-users and the public. To enable real-time protection of raw data, it is demanding to facilitate privacy-preserving algorithms at data generation, or in-sensory privacy. However, due to the severe sensor resource constraints and intensive computation/security cost, it remains an open question of how to enable data protection algorithms with efficient circuit techniques. To answer this question, this paper discusses the potential of a stochastic mixed-signal (SMS) circuit for ultra-low-power, small-foot-print data security. In particular, this paper discusses digitally-controlled-oscillators (DCO) and their advantages in (1) seamless analog interface, (2) stochastic computation efficiency, and (3) unified entropy generation over conventional digital circuit baselines. With DCO as an illustrative case, we target (1) SMS privacy-preserving architecture definition and systematic SMS analysis on its performance gains across various hardware/software configurations, and (2) revisit analog/mixed-signal voltage/transistor scaling in the context of entropy-based data protection.
Ningyuan Cao, Boyang Cheng, Muya Chang
ICCAD1
2021 From Specification to Topology: Automatic Power Converter Design via Reinforcement Learning
abstract
The tidal waves of modern electronic/electrical devices have led to increasing demands for ubiquitous application-specific power converters. A conventional manual design procedure of such power converters is computation- and labor-intensive, which involves selecting and connecting component devices, tuning component-wise parameters and control schemes, and iteratively evaluating and optimizing the design. To automate and speed up this design process, we propose an automatic framework that designs custom power converters from design specifications using reinforcement learning. Specifically, the framework embraces upper-confidence-bound-tree-based (UCT-based) reinforcement learning to automate topology space exploration with circuit design specification-encoded reward signals. Moreover, our UCT-based approach can exploit small offline data via the specially designed default policy to accelerate topology space exploration. Further, it utilizes a hybrid circuit evaluation strategy to substantially reduces design evaluation costs. Empirically, we demonstrated that our framework could generate energy-efficient circuit topologies for various target voltage conversion ratios. Compared to existing automatic topology optimization strategies, the proposed method is much more computationally efficient - it can generate topologies with the same quality while being up to 67% faster. Additionally, we discussed some interesting circuits discovered by our framework.
Shaoze Fan, Ningyuan Cao, Jing Li 0025, Xin Zhang 0025
ICCAD2
2021 A Hardware-Friendly Approach Towards Sparse Neural Networks Based on LFSR-Generated Pseudo-Random Sequences
abstract
The increase in the number of edge devices has led to the emergence of edge computing where the computations are performed on the device. In recent years, deep neural networks (DNNs) have become the state-of-the-art method in a broad range of applications, from image recognition, to cognitive tasks to control. However, neural network models are typically large and computationally expensive and therefore not deployable on power and memory constrained edge devices. Sparsification techniques have been proposed to reduce the memory foot-print of neural network models. However, they typically lead to substantial hardware and memory overhead. In this article, we propose a hardware-aware pruning method using linear feedback shift register (LFSRs) to generate the locations of non-zero weights in real-time during inference. We call this LFSR-generated pseudorandom sequence based sparsity (LGPS) technique. We explore two different architectures for our hardware-friendly LGPS technique, based on (1) row/column indexing with LFSRs and (2) column-wise indexing with nested LFSRs, respectively. Using the proposed method, we present a total saving of energy and area up to 37.47% and 49.93% respectively and speed up of 1.53× w.r.t the baseline pruning method, for the VGG-16 network on down-sampled ImageNet.
Foroozan Karimzadeh, Ningyuan Cao, Brian Crafton, Justin K. Romberg, Arijit Raychowdhury
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 Hardware-Aware Pruning of DNNs using LFSR-Generated Pseudo-Random Indices
abstract
Deep neural networks (DNNs) have been emerged as the state-of-the-art algorithms in broad range of applications. To reduce the memory foot-print of DNNs, in particular for embedded applications, sparsification techniques have been proposed. Unfortunately, these techniques come with a large hardware overhead. In this paper, we present a hardware-aware pruning method where the locations of non-zero weights are derived in real-time from a Linear Feedback Shift Registers (LFSRs). Using the proposed method, we demonstrate a total saving of energy and area up to 63.96% and 64.23% for VGG-16 network on down-sampled ImageNet, respectively for iso-compression-rate and iso-accuracy.
Foroozan Karimzadeh, Ningyuan Cao, Brian Crafton, Justin K. Romberg, Arijit Raychowdhury
ISCAS2
2018 A Light-Powered Smart Camera With Compressed Domain Gesture Detection
abstract
This paper presents an ultralow power smart camera with gesture detection. Low power is achieved by directly extracting gesture features from the compressed measurements, which are the block averages and the linear combinations of the image sensor's pixel values. We present two classifier techniques to allow low computational and storage requirements. The system has been implemented on an analog devices BlackFin ULP vision processor. By enabling ultralow energy consumption, we demonstrate that the system is powered by ambient light harvested through photovoltaic cells whose output is regulated by TI's dc-dc buck converter with maximum power point tracking. Measured data reveals that with only 400 compressed measurements (768× compression ratio) per frame, the system is able to recognize key wake-up gestures with greater than 80% accuracy and only 95mJ of energy per frame. Owing to its fully self-powered operation, the proposed system can find wide applications in “always-on” vision systems, such as in surveillance, robotics, and consumer electronics with touch-less operation.
Amaravati Anvesha, Shaojie Xu, Ningyuan Cao, Justin K. Romberg, Arijit Raychowdhury
IEEE Trans. Circuits Syst. Video Technol.3
2016 A Light-powered, "Always-On", Smart Camera with Compressed Domain Gesture Detection
abstract
In this paper we propose an energy-efficient camera-based gesture recognition system powered by light energy for "always on" applications. Low energy consumption is achieved by directly extracting gesture features from the compressed measurements, which are the block averages and the linear combinations of the image sensor's pixel values. The gestures are recognized using a nearest-neighbour (NN) classifier followed by Dynamic Time Warping (DTW). The system has been implemented on an Analog Devices Black Fin ULP vision processor and powered by PV cells whose output is regulated by TI's DC-DC buck converter with Maximum Power Point Tracking (MPPT). Measured data reveals that with only 400 compressed measurements (768x compression ratio) per frame, the system is able to recognize key wake-up gestures with greater than 80% accuracy and only 95mJ of energy per frame. Owing to its fully self-powered operation, the proposed system can find wide applications in "always-on" vision systems such as in surveillance, robotics and consumer electronics with touch-less operation.
Amaravati Anvesha, Shaojie Xu, Ningyuan Cao, Justin K. Romberg, Arijit Raychowdhury
ISLPED3