VLDB 2026 Research / reviewers in the wild / expert
Weidong Cao 0001
dblp:61/5251-1
· DBLP profile ↗
27ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0001-7539-8250ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 10 first-author · 21 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASPA: Reassigning DDR5 Parity BandwidthabstractRecent memory advancements, such as DDR5, HBM3, and emerging memory-centric accelerators, primarily focus on increasing bandwidth capacity, yet often omit the significance of effective bandwidth utilization, i.e., bandwidth efficiency. Motivated by the suboptimal channel allocation in DDR5, where parity accounts for 25% bandwidth overhead, we propose ASPA, an efficiency-oriented solution that reallocates parity bandwidth to boost regular data transfer. The objective of ASPA is to enhance bandwidth efficiency without compromising reliability while maintaining low hardware overhead. In particular, we leverage existing CRC (Cyclic Redundancy Check) units in DRAM chips to opportunistically generate second-tier parity for the existing ECC (Error Correction Code) parity. For bulk-sized memory accesses, only the small second-tier parity is transmitted. This reduces the parity bandwidth consumption, allowing data chips to reuse the freed parity bandwidth for data transfer. Our observation indicates that the protection capability is sufficient when the second-tier parity (with a 64-bit size) is used exclusively for error detection. Furthermore, by exploiting underutilized resources in high-performance memory systems, ASPA is implemented with negligible hardware overhead. Qiufeng Li, Yanan Guo 0002, Weidong Cao 0001, Xin Xin 0008 |
HPCA | 4 |
| 2026 | HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECC
Ruizhi Zhu, Yanan Guo 0002, Huize Li, Weidong Cao 0001, Qian Lou, Xin Xin 0008 |
ISCA | 4 |
| 2026 | Systematic Methodology of Modeling and Design Space Exploration for CMOS Image SensorsabstractCMOS Image Sensors (CIS) are integral to both human and computer vision tasks, necessitating continuous improvements in key performance metrics such as latency, power, and noise. Despite experienced designers being able to make informed design decisions, novice designers and system architects face challenges due to the complex and expansive design space of CIS. This paper introduces a systematic methodology that elucidates the trade-offs among CIS performance metrics and enables efficient design space exploration. Specifically, we propose a first-principle-based CIS modeling method. By exposing low-level circuit parameters, our modeling method explicitly reveals the impacts of design changes on high-level metrics. Based on the modeling method, we propose a design space exploration process that swiftly evaluates and identifies the optimal CIS design, capable of exploring over 109 designs in under a minute without the need for time-consuming SPICE simulations. Our approach is validated through a case study and comparisons with real-world designs, demonstrating its practical utility in guiding early-stage CIS design. Tianrui Ma, Ramakrishna Kakarala, Charles Shan, Weidong Cao 0001, Xuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | LEAP: Lightweight Neural Network Inference Through Proactive Early-Exiting PredictionabstractIn recent years, the incorporation of early exit layers into deep neural networks has allowed inference to terminate earlier while maintaining accuracy. However, the passive decision-making involved in the these static exit placement creates a dilemma: fine-grained placement may cause high performance and energy overhead due to frequent exit layer execution, while coarse-grained placement may miss early exit opportunities. Moreover, common energy-saving techniques like adjusting processor configurations are not applicable once inference begins. To overcome these challenges and improve computation and energy efficiency, we propose LEAP, a software-hardware co-design approach. On the software side, LEAP proactively predicts exit points at runtime, reducing computation by enabling early exits without requiring every pre-placed exit layer to be executed. On the hardware side, LEAP adjusts processor settings—such as frequency and voltage—based on single or multiple predicted exits to optimize energy consumption while adhering to latency requirements. Extensive experimental results show that LEAP significantly improves efficiency. Compared to standard inference, LEAP reduces computation by up to 76.4% and saves up to 83.2% in energy. Compared to state-of-the-art early exit methods, LEAP achieves up to 27.9% less computation and 57.1% more energy savings, while maintaining similar accuracy and latency. Yingtao Shen, Xiangjie Li, Yehan Ma, Weidong Cao 0001, An Zou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | EVA: An Efficient and Versatile Generative Engine for Targeted Discovery of Novel Analog CircuitsabstractAnalog circuit design has traditionally depended on manual expertise, slowing the discovery of novel topologies essential for advanced technologies like AI, $5 \mathrm{G} / 6 \mathrm{G}$, and quantum computing. While AI-driven methods have accelerated hardware design workflows, most of them focus on topology synthesis, often reusing known structures to achieve specific goals. The challenge of discovering entirely new, high-performance topologies remains largely underexplored due to its abstract nature. In this work, we introduce EVA, an efficient and versatile generative engine for discovering novel analog circuit topologies. EVA employs a bottom-up generation framework, using a decoder-only transformer to sequentially predict device pin connections and create diverse circuits from scratch. Pretraining on unlabeled circuit topologies builds foundational knowledge about circuit connectivity, achieving baseline discovery efficiency by generating valid circuits and reducing performance-labeled samples needed in fine-tuning. For targeted discovery of highperformance designs, EVA leverages two fine-tuning strate-gies-proximal policy optimization (PPO) and direct preference optimization (DPO)-to further enhance discovery efficiency for relevant, high-performing topologies. Experimental results across various circuit types highlight EVA’s strengths in validity, novelty, versatility, and both training sample and discovery efficiency. Weimin Fu, Xiaolong Guo 0001, Weidong Cao 0001, Xuan Zhang 0001 |
DAC | 4 |
| 2025 | Late Breaking Results: Opera: An Open and Efficient Platform for Data-driven Synthesis of Analog CircuitsabstractThe front-end synthesis of analog circuits has been a long-standing challenge since the advent of integrated circuits. Many methods, ranging from conventional optimization-based techniques to emerging learning-based approaches, have been extensively explored to address this challenge. Yet, these methods are data-driven and often suffer from low design efficiency, due to their heavy reliance on time-consuming circuit simulators, which are frequently used in the synthesis loop for real-time evaluation of the evolving circuit design. In addition, benchmarking these methods is also largely unachievable due to their exclusive use of commercial semiconductor technology for evaluation. This “Late Breaking Results” introduces Opera, an open and efficient platform for the data-driven synthesis of analog circuits. Specifically, Opera develops efficient surrogate models for various circuits and integrates them into open-source OpenAI Gym-like environments to enable efficient synthesis. Case studies on exemplary circuits show that this platform can accelerate the conventional data-driven synthesis flow by up to $40 \times$. It also enables the benchmarking of various synthesis methods with standardized environments built upon an open-source semiconductor process. Shikai Wang, Yaolong Hu, Zhiqiang Yi, Taiyun Chi, Weidong Cao 0001 |
DAC | 5 |
| 2025 | AnalogGenie: A Generative Engine for Automatic Discovery of Analog Circuit TopologiesabstractThe massive and large-scale design of foundational semiconductor integrated circuits (ICs) is crucial to sustaining the advancement of many emerging and future technologies, such as generative AI, 5G/6G, and quantum computing.
Excitingly, recent studies have shown the great capabilities of foundational models in expediting the design of digital ICs.
Yet, applying generative AI techniques to accelerate the design of analog ICs remains a significant challenge due to critical domain-specific issues, such as the lack of a comprehensive dataset and effective representation methods for analog circuits.
This paper proposes, $\textbf{AnalogGenie}$, a $\underline{\textbf{Gen}}$erat$\underline{\textbf{i}}$ve $\underline{\textbf{e}}$ngine for automatic design/discovery of $\underline{\textbf{Analog}}$ circuit topologies--the most challenging and creative task in the conventional manual design flow of analog ICs.
AnalogGenie addresses two key gaps in the field: building a foundational comprehensive dataset of analog circuit topology and developing a scalable sequence-based graph representation universal to analog circuits.
Experimental results show the remarkable generation performance of AnalogGenie in broadening the variety of analog ICs, increasing the number of devices within a single design, and discovering unseen circuit topologies far beyond any prior arts.
Our work paves the way to transform the longstanding time-consuming manual design flow of analog ICs to an automatic and massive manner powered by generative AI.
Our source code is available at https://github.com/xz-group/AnalogGenie. Weidong Cao 0001, Xuan Zhang 0001 |
ICLR | 2 |
| 2025 | AnalogGenie-Lite: Enhancing Scalability and Precision in Circuit Topology Discovery through Lightweight Graph ModelingabstractThe sustainable performance improvements of integrated circuits (ICs) drive the continuous advancement of nearly all transformative technologies. Since its invention, IC performance enhancements have been dominated by scaling the semiconductor technology. Yet, as Moore's law tapers off, a crucial question arises: ***How can we sustain IC performance in the post-Moore era?*** Creating new circuit topologies has emerged as a promising pathway to address this fundamental need. This work proposes AnalogGenie-Lite, a decoder-only transformer that discovers novel analog IC topologies with significantly enhanced scalability and precision via lightweight graph modeling.
AnalogGenie-Lite makes several unique contributions, including concise device-pin representations (i.e., advancing the best prior art from $O\left(n^2\right)$ to $O\left(n\right)$), frequent sub-graph mining, and optimal sequence modeling. Compared to state-of-the-art circuit topology discovery methods, it achieves $5.15\times$ to $71.11\times$ gains in scalability and 23.5\% to 33.6\% improvements in validity. Case studies on other domains' graphs are also provided to show the broader applicability of the proposed graph modeling approach. Source code: https://github.com/xz-group/AnalogGenie-Lite. Weidong Cao 0001, Xuan Zhang 0001 |
ICML | 2 |
| 2025 | LoRASensE: Learnable Low-Rank Acquisition in Sensors for Efficient Edge Machine VisionabstractIntegrating deep learning with ubiquitous image sensors has empowered various edge vision applications such as classification, segmentation, and detection. Deploying these data-driven applications requires holistic optimizations, from front-end sensing to back-end processing, within the limited resources of edge devices. While significant advances have been made in the efficient processing of sensory data in the back end with optimizations of learning algorithms (e.g., compression) and development of computing hardware (e.g., accelerators), the energy efficiency of front-end sensors remains significantly limited due to conventional high-fidelity image acquisition and the resulting massive off-chip data transfer.This paper proposes a domain-specific visual acquisition method, LoRASensE, learnable low-rank acquisition in sensors tailored for efficient data-driven edge vision applications. LoRASensE is an algorithm-hardware co-design framework that integrates a learned low-rank compressor into image sensors to acquire compressed features. Specifically, this compressor is optimized alongside downstream vision tasks to ensure end-to-end accuracy and is implemented with efficient analog processing hardware. Our extensive evaluations on real-world datasets across various vision applications demonstrate that LoRASensE can achieve a 12.5× compression ratio with a just 1-b compressor, minimal accuracy loss, and 86.9% energy saving compared to the conventional high-fidelity acquisition. Multi-dimensional comparisons further show that LoRASensE also significantly outperforms existing in-sensor compression methods. Zhiqiang Yi, Tianrui Ma, Weidong Cao 0001 |
ISLPED | 4 |
| 2025 | RoSE-Opt: Robust and Efficient Analog Circuit Parameter Optimization With Knowledge-Infused Reinforcement LearningabstractDesign automation of analog circuits has long been sought. However, achieving robust and efficient analog design automation remains challenging. This article proposes a learning framework, RoSE-Opt, to achieve robust and efficient analog circuit parameter optimization. RoSE-Opt has two important features. First, it incorporates key domain knowledge of analog circuit design, such as circuit topology, couplings between circuit specifications, and variations of process, supply voltage, and temperature, into the learning loop. This strategy facilitates the training of an artificial agent capable of achieving design goals by identifying device parameters that are optimal and robust. Second, it exploits a two-level optimization method, that is, integrating Bayesian optimization (BO) with reinforcement learning (RL) to improve sample efficiency. In particular, BO is used for a coarse yet quick search of an initial starting point for optimization. This sets a solid foundation to efficiently train the RL agent with fewer samples. Experimental evaluations on benchmarking circuits show promising sample efficiency, extraordinary figure-of-merit in terms of design efficiency and design success rate, and Pareto optimality in circuit performance of our framework, compared to previous methods. Furthermore, this work thoroughly studies the performance of different RL optimization algorithms, such as deep deterministic policy gradients (DDPGs) with an off-policy learning mechanism and proximal policy optimization (PPO) with an on-policy learning mechanism. This investigation provides users with guidance on choosing the appropriate RL algorithms to optimize the device parameters of analog circuits. Finally, our study also demonstrates RoSE-Opt’s promise in parasitic-aware device optimization for analog circuits. In summary, our work reports a knowledge-infused BO-RL design automation framework for reliable and efficient optimization of analog circuits’ device parameters. Code implementation of our method can be found athttps://github.com/xz-group/RoSE. Weidong Cao 0001, Tianrui Ma, Mouhacine Benosman, Xuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Addition is Most You Need: Efficient Floating-Point SRAM Compute-in-Memory by Harnessing Mantissa AdditionabstractThe compute-in-memory (CIM) paradigm holds great promise to efficiently accelerate machine learning workloads. Among memory devices, static random-access memory (SRAM) stands out as a practical choice for its exceptional reliability in the digital domain and excellent scalability. Recently, there has been a growing interest in accelerating floating-point (FP) deep neural networks (DNNs) with SRAM CIM due to their critical importance in DNN training and high-accurate inference. This paper proposes an energy-efficient SRAM CIM macro for FP DNNs. To achieve the design, we identify a lightweight approach that decomposes conventional FP mantissa multiplication into two parts: mantissa sub-addition (sub-ADD) and mantissa sub-multiplication (sub-MUL). Our study shows that while mantissa sub-MUL is compute-intensive, it only contributes to the minority of FP products, whereas mantissa sub-ADD, although compute-light, accounts for the majority of FP products. Recognizing "Addition is Most You Need", we develop a novel hybrid-domain SRAM CIM macro to accurately handle mantissa sub-ADD in the digital domain while improving the energy efficiency of mantissa sub-MUL using analog computing. Experiments with the MLPerf benchmark show its remarkable improvement in energy efficiency on average by 3×~ 3.6× (2.5×~3.1×) in inference (training) compared to a fully digital baseline without any accuracy loss, showcasing its great potential for FP DNN acceleration. Weidong Cao 0001, Xin Xin 0008, Xuan Zhang 0001 |
DAC | 1 |
| 2024 | Efficient Processing of MLPerf Mobile Workloads Using Digital Compute-In-Memory MacrosabstractCompute-in-memory (CIM) has recently emerged as a promising design paradigm to accelerate deep neural network (DNN) processing. Continuously better energy and area efficiency at the macrolevel had been reported through many testchips over the last few years. However, in those macro design-oriented studies, accelerator-level considerations, such as memory accesses and processing of entire DNN workloads have not been investigated in-depth. In this article, we aim to fill this gap starting with the characteristics of our latest CIM macro fabricated with cutting-edge FinFET CMOS technology at 4-nm node. We then study, through an accelerator simulator developed in-house, three key items that would determine the efficiency of our CIM macro in the accelerator context while running MLPerf Mobile suite: 1) dataflow optimization; 2) optimal selection of CIM macro dimensions to further improve macro utilization; and 3) optimal combination of multiple CIM macros. Although there is typically a stark contrast between macro-level peak and accelerator-level average throughput and energy efficiency, the aforementioned optimizations are shown to improve the macro utilization by$3.04\times $and reduce the energy-delay product (EDP) to$0.34\times $compared to the original macro on MLPerf Mobile inference workloads. While we exploit a digital CIM macro in this study, the findings and proposed methods remain valid for other types of CIM (such as analog CIM and analog–digital–hybrid CIM) as well. Xiaoyu Sun 0001, Weidong Cao 0001, Brian Crafton, Kerem Akarvardar, Haruki Mori, Hidehiro Fujiwara, Hiroki Noguchi, Yu-Der Chih, Meng-Fan Chang, Yih Wang, Tsung-Yung Jonathan Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical FrameworkabstractAugmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration. Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2023 | RoSE: Robust Analog Circuit Parameter Optimization with Sampling-Efficient Reinforcement LearningabstractDesign automation of analog circuits has been a long-standing challenge in the integrated circuit field. Recently, multiple methods based on learning or optimization have demonstrated great promise in automating device sizing for analog circuits. However, they often ignore the strong susceptibility of analog circuits to process, voltage, and temperature (PVT) variations or suffer from low sampling efficiency to train algorithms. To address these critical limitations, this paper proposes RoSE, the first Robust analog circuit parameter optimization framework with high Sampling Efficience by synergistically combining Bayesian Optimization (BO) and reinforcement learning (RL). Its core is to use the fast convergence of BO to find an optimized starting point for the backbone RL agent to notably improve its sampling efficiency during the learning process. With this pre-optimization, we further leverage the RL’s superior optimization ability to achieve robust device sizing by incorporating sufficient features of PVT variations into the representation learning loop. Experimental results of our proposed method on exemplary circuits show 3.25×∼16× improvement of sampling efficiency and 6.8× ∼ 24× improvement of figure-of-merit (FoM, defined with design efficiency and design accuracy) as compared to prior methods. Weidong Cao 0001, Xuan Zhang 0001 |
DAC | 2 |
| 2023 | PDNSig: Identifying Multi-Tenant Cloud FPGAs with Power Distribution Network-Based SignaturesabstractThe increasing use of Field Programmable Gate Arrays (FPGAs) in modern cloud data centers, such as Amazon's EC2 F1 instances, has led to a rising concern regarding remote side-channel attacks, which necessitates thorough investigation and analysis. Existing threat models crucially depend on a critical assumption: attackers could uniquely identify the target FPGA chip (or the specific die in a chip). However, this assumption is impractical in real cloud scenarios since security measures routinely implemented by cloud FPGA providers can anonymize the devices. To address this critical limitation, we propose PDNSig-a power distribution network (PDN)-based signature generation framework. By recognizing the complexity and irregularity of the PDN network and its susceptibility to process variation, we reveal that the impedance profile of an FPGA's PDN can uniquely distinguish different FPGAs. Particularly, we inject pseudo-random noises into the PDN by turning on or off power-hungry circuits (e.g., ring oscillators). The corresponding response of PDN is subsequently captured by on-chip sensors (e.g., time-to-digital converter), followed by a statistical analysis to obtain the PDN impedance at different frequencies. This proposed novel random process-based PDN measurement methodology can be directly applied to prior attack infrastructures with low hardware overhead. We perform thorough characterizations and demonstrate the effectiveness of PDNSig by conducting multiple real-world experiments on 40 Amazon cloud FPGA chips (including 120 dies). Experimental results show that the extracted PDN-based signatures can distinguish all 40 chips reliably. Additionally, a 99% true positive rate and 0.4% false positive rate are also achieved when identifying the 120 distinctive dies associated with these FPGA chips. Huifeng Zhu, Weidong Cao 0001, Xuan Zhang 0001 |
ICCAD | 2 |
| 2023 | CktGNN: Circuit Graph Neural Network for Electronic Design Automation
Zehao Dong, Weidong Cao 0001, Muhan Zhang, Dacheng Tao, Yixin Chen 0001, Xuan Zhang 0001 |
ICLR | 2 |
| 2023 | LeCA: In-Sensor Learned Compressive Acquisition for Efficient Machine Vision on the EdgeabstractWith the rapid advances of deep learning-based computer vision (CV) technology, digital images are increasingly consumed, not by humans, but by downstream CV algorithms. However, capturing high-fidelity and high-resolution images is energy-intensive. It not only dominates the energy consumption of the sensor itself (i.e. in low-power edge devices), but also contributes to significant memory burdens and performance bottlenecks in the later storage, processing, and communication stages. In this paper, we systematically explore a new paradigm of in-sensor processing, termed "learned compressive acquisition" (LeCA). Targeting machine vision applications on the edge, the LeCA framework exploits the joint learning of a sensor autoencoder structure with the downstream CV algorithms to effectively compress the original image into low-dimensional features with adaptive bit depth. We employ column-parallel analog-domain processing directly inside the image sensor to perform the compressive encoding of the raw image, resulting in meaningful hardware savings, and energy efficiency improvements. Evaluated within a modern machine vision processing pipeline, LeCA achieves 4×, 6×, and 8× compression ratios prior to any digital compression, with minimal accuracy loss of 0.97%, 0.98%, and 2.01% on ImageNet, outperforming existing methods. Compared with the conventional full-resolution image sensor and the state-of-the-art compressive sensing sensor, our LeCA sensor is 6.3× and 2.2× more energy-efficient while reaching a 2× higher compression ratio. Tianrui Ma, Adith Boloor, Xiangxing Yang, Weidong Cao 0001, Patrick Williams, Nan Sun 0001, Ayan Chakrabarti, Xuan Zhang 0001 |
ISCA | 4 |
| 2023 | Non-Hermitian Physics-Inspired Voltage-Controlled Oscillators with Resistive TuningabstractThis paper presents a non-Hermitian physics-inspired voltage-controlled oscillator (VCO) topology, which is termed parity-time-symmetric topology. The VCO consists of two coupled inductor-capacitor (LC) cores with a balanced gain and loss profile. Due to the interplay between the gain/loss and their coupling, an extra degree of freedom is enabled via resistive tuning, which can enhance the frequency tuning range (FTR) beyond the bounds of conventional capacitive or inductive tuning. A silicon prototype is implemented in a standard 130 nm bulk CMOS process with a core area of$0.15\mathbf{mm}^{2}$. Experimental results show that it achieves a$3.1\times$FTR improvement and 30% phase noise reduction of the baseline VCO with the same amount of capacitive tuning ability. Weidong Cao 0001, Hua Wang 0006, Xuan Zhang 0001 |
ISCAS | 1 |
| 2023 | A/D Alleviator: Reducing Analog-to-Digital Conversions in Compute-In-Memory with Augmented Analog AccumulationabstractCompute-in-memory (CIM) has shown great promise in accelerating numerous deep-learning tasks. However, existing analog CIM (ACIM) accelerators often suffer from frequent and energy-intensive analog-to-digital (A/D) conversions, severely limiting their energy efficiency. This paper proposes A/D Alleviator, an energy-efficient augmented analog accumulation data flow to reduce A/D conversions in ACIM accelerators. To make it, switched-capacitor-based multiplication and accumulation circuits are used to connect the bitlines (BLs) of memory crossbar arrays and the final A/D conversion stage. In this way, analog partial sums can be accumulated both spatially across all adjacent BLs that store high-precision weights and temporarily across all input cycles before the final quantization, thereby minimizing the need for explicit A/D conversions. Evaluations demonstrate that A/D Alleviator can improve energy efficiency by 4.9× and 1.9× with a high signal-to-noise ratio, as compared to state-of-the-art ACIM accelerators. Weidong Cao 0001, Xuan Zhang 0001 |
ISCAS | 1 |
| 2022 | Domain knowledge-infused deep learning for automated analog/radio-frequency circuit parameter optimizationabstractThe design automation of analog circuits is a longstanding challenge. This paper presents a reinforcement learning method enhanced by graph learning to automate the analog circuit parameter optimization at the pre-layout stage, i.e., finding device parameters to fulfill desired circuit specifications. Unlike all prior methods, our approach is inspired by human experts who rely on domain knowledge of analog circuit design (e.g., circuit topology and couplings between circuit specifications) to tackle the problem. By originally incorporating such key domain knowledge into policy training with a multimodal network, the method best learns the complex relations between circuit parameters and design targets, enabling optimal decisions in the optimization process. Experimental results on exemplary circuits show it achieves human-level design accuracy (~99%) with 1.5× efficiency of existing best-performing methods. Our method also shows better generalization ability to unseen specifications and optimality in circuit performance optimization. Moreover, it applies to design radio-frequency circuits on emerging semiconductor technologies, breaking the limitations of prior learning methods in designing conventional analog circuits. Weidong Cao 0001, Mouhacine Benosman, Xuan Zhang 0001 |
DAC | 1 |
| 2022 | PowerTouch: A Security Objective-Guided Automation Framework for Generating Wired Ghost Touch Attacks on TouchscreensabstractThe wired ghost touch attacks are the emerging and severe threats against modern touchscreens. The attackers can make touchscreens falsely report nonexistent touches (i.e., ghost touches) by injecting common-mode noise (CMN) into the target devices via power cables. Existing attacks rely on reverse-engineering the touchscreens, then manually crafting the CMN waveforms to control the types and locations of ghost touches. Although successful, they are limited in practicality and attack capability due to the touchscreens' black-box nature and the immense search space of attack parameters. To overcome the above limitations, this paper presents PowerTouch, a framework that can automatically generate wired ghost touch attacks. We adopt a software-hardware co-design approach and propose a domain-specific genetic algorithm-based method that is tailored to account for the characteristics of the CMN waveform. Based on the security objectives, our framework automatically optimizes the CMN waveform towards injecting the desired type of ghost touches into regions specified by attackers. The effectiveness of PowerTouch is demonstrated by successfully launching attacks on touchscreen devices from two different brands given nine different objectives. Compared with the state-of-the-art attack, we seminally achieve controlling taps on an extra dimension and injecting swipes on both dimensions. We can place an average of 84.2% taps on the targeted side of the screen, with the location error in the other dimension no more than 1.53mm. An average of 94.5% of injected swipes with correct directions is also achieved. The quantitative comparison with the state-of-the-art method shows that a better attack performance can be achieved by PowerTouch. Huifeng Zhu, Zhiyuan Yu 0001, Weidong Cao 0001, Ning Zhang 0017, Xuan Zhang 0001 |
ICCAD | 3 |
| 2022 | HOGEye: Neural Approximation of HOG Feature Extraction in RRAM-Based 3D-Stacked Image SensorsabstractMany computer vision tasks, ranging from recognition to multi-view registration, operate on feature representation of images rather than raw pixel intensities. However, conventional pipelines for obtaining these representations incur significant energy consumption due to pixel-wise analog-to-digital (A/D) conversions and costly storage and computations. In this paper, we propose HOGEye, an efficient near-pixel implementation for a widely-used feature extraction algorithm—Histograms of Oriented Gradients (HOG). HOGEye moves the key but computation-intensive derivative extraction (DE) and histogram generation (HG) steps into the analog domain by applying a novel neural approximation method in a resistive random-access memory (RRAM)-based 3D-stacked image sensor. The co-location of perception (sensor) and computation (DE and HG) and the alleviation of A/D conversions allow HOGEye design to achieve significant energy saving. With negligible detection rate degradation, the entire HOGEye sensor system consumes less than 48μ[email protected] for an image resolution of 256 × 256 (equivalent to 24.3pJ/pixel) while the processing part only consumes 14.1pJ/pixel, achieving more than 2.5 × energy efficiency improvement than the state-of-the-art designs. Tianrui Ma, Weidong Cao 0001, Fei Qiao, Ayan Chakrabarti, Xuan Zhang 0001 |
ISLPED | 2 |
| 2022 | Neural-PIM: Efficient Processing-In-Memory With Neural Approximation of PeripheralsabstractProcessing-in-memory (PIM) architecture has demonstrated great potentials in accelerating numerous deep learning tasks. In particular, resistive random-access memory (RRAM) technology provides a promising hardware substrate for PIM accelerators, because it can support efficient in-situ vector-matrix multiplications (VMMs) with high-density RRAM crossbar arrays. However, such accelerators suffer from frequent and energy-intensive analog-to-digital (A/D) conversions, severely limiting their performance. This paper proposes a new PIM architecture to efficiently accelerate deep learning tasks by minimizing the required A/D conversions with neural approximated peripheral circuits. By characterizing the existing dataflows of state-of-the-art PIM architectures, we first propose a new dataflow by extending shift and add (S+A) operations into the analog domain before the final A/D conversion, which can remarkably reduce the required A/D conversions for a dot-product. We then elaborate on a neural approximation method to design both accumulation circuits (S+A) and quantization circuits (ADC) using RRAM crossbar arrays. Finally, we apply them to build a RRAM-based PIM accelerator--\textbf{Neural-PIM} based on the proposed analog dataflow and evaluate its system-level performances. Evaluations on different DNN benchmarks demonstrate that Neural-PIM can improve energy efficiency by 5.36x (1.73x) and speed up throughput by 3.43x (1.59x) without losing accuracy, compared to state-of-the-art RRAM-based PIM accelerators, i.e., ISAAC} (CASCADE) Weidong Cao 0001, Yilong Zhao 0004, Adith Boloor, Yinhe Han 0001, Xuan Zhang 0001, Li Jiang 0002 |
IEEE Trans. Computers | 1 |
| 2021 | Evaluating Neural Network-Inspired Analog-to-Digital Conversion With Low-Precision RRAMabstractRecent work has demonstrated great potentials of neural network-inspired analog-to-digital converters (NNADCs) in many emerging applications. These NNADCs often rely on resistive random-access memory (RRAM) devices to realize basic NN operations, and usually need high-precision RRAM (6-12 b) to achieve moderate quantization resolutions (4-8 b). Such an optimistic assumption of RRAM precision, however, is not well supported by practical RRAM arrays in the large-scale production process. In this article, we evaluate two new designs of NNADC with low-precision RRAM devices. They take advantage of traditional two-stage/pipelined hardware architecture and a custom deep-learning-based building block design methodology. Results obtained from SPICE simulations demonstrate a robust design of an 8-b subranging NNADC using 4-b RRAM devices, as well as a 14-b pipelined NNADC using 3-b RRAM devices. The evaluations on the two NNADCs suggest that pipelined architecture is better to achieve higher-resolution using lower precision RRAM. We also perform design space exploration on the building blocks of NNADCs to achieve a balanced performance tradeoff. Comprehensive comparisons reveal improved power, speed performance, and competitive figure of merits (FoMs) of the pipelined NNADC, compared with state-of-the-art NNADCs and traditional ADCs. In addition, the proposed pipelined NNADC can support reconfigurable high-resolution nonlinear quantization with high conversion speed and low conversion energy, enabling intelligent analog-to-information interfaces for near-sensor processing. Weidong Cao 0001, Liu Ke 0001, Ayan Chakrabarti, Xuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | NeuADC: Neural Network-Inspired Synthesizable Analog-to-Digital ConversionabstractTraditional analog-to-digital converters (ADCs) employ dedicated analog and mixed-signal (AMS) circuits, requiring time-consuming manual design process. They also exhibit limited configurability to support diverse quantization schemes on the same circuitry. In this paper, we propose NeuADC-an automated design approach to synthesizing an analog-to-digital (A/D) interface that can approximate the desirable quantization function using a neural network (NN) with a single hidden layer. We leverage the mixed-signal resistive random-access memory (RRAM) crossbar architecture to design a novel dual-path configuration for the implementation of the basic NN operations at the circuit level. We exploit alternative bits encoding scheme to the conventional binary encoding to improve the training accuracy. Our method incorporates nonidealities at the device and circuit level into the training process to ensure NeuADC's robustness against variations of process, supply voltage, and temperature (PVT). Results obtained from SPICE simulation based on RRAM and standard 130-nm CMOS technology suggest that not only can NeuADC deliver promising performance compared to the state-of-the-art ADCs and other emerging converter designs across comprehensive design metrics, but it can also intrinsically support multiple configurable quantization schemes using the same hardware substrate, paving ways for future adaptable application-driven signal conversion. Our systematic evaluations on the proposed NeuADC framework also quantify the impacts on the ADC quantization quality from hidden neuron sizes, RRAM resistance imprecision, and PVT variations, and reveal the design tradeoff between speed, power, and area in a NeuADC circuit. Weidong Cao 0001, Xin He 0011, Ayan Chakrabarti, Xuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | NeuADC: Neural Network-Inspired RRAM-Based Synthesizable Analog-to-Digital Conversion with Reconfigurable Quantization SupportabstractTraditional analog-to-digital converters (ADCs) employ dedicated analog and mixed-signal (AMS) circuits and require time-consuming manual design process. They also exhibit limited reconfigurability and are unable to support diverse quantization schemes using the same circuitry. In this paper, we propose NeuADC — an automated design approach to synthesizing an analog-to-digital (A/D) interface that can approximate the desired quantization function using a neural network (NN) with a single hidden layer. Our design leverages the mixed-signal resistive random-access memory (RRAM) crossbar architecture in a novel dual-path configuration to realize basic NN operations at the circuit level and exploits smooth bit-encoding scheme to improve the training accuracy. Results obtained from SPICE simulations based on 130nm technology suggest that not only can NeuADC deliver promising performance compared to the state-of-art ADC designs across comprehensive design metrics, but also it can intrinsically support multiple reconfigurable quantization schemes using the same hardware substrate, paving the ways for future adaptable application-driven signal conversion. The robustness of NeuADC’s quantization quality under moderate RRAM resistance precision is also evaluated using SPICE simulations. Weidong Cao 0001, Xin He 0011, Ayan Chakrabarti, Xuan Zhang 0001 |
DATE | 1 |
| 2019 | Neural Network-Inspired Analog-to-Digital Conversion to Achieve Super-Resolution with Low-Precision RRAM DevicesabstractRecent works propose neural network- (NN-) inspired analog-to-digital converters (NNADCs) and demonstrate their great potentials in many emerging applications. These NNADCs often rely on resistive random-access memory (RRAM) devices to realize the NN operations and require high-precision RRAM cells (6~12-bit) to achieve a moderate quantization resolution (4~8-bit). Such optimistic assumption of RRAM resolution, however, is not supported by fabrication data of RRAM arrays in large-scale production process. In this paper, we propose an NN-inspired super-resolution ADC based on low-precision RRAM devices by taking the advantage of a co-design methodology that combines a pipelined hardware architecture with a custom NN training framework. Results obtained from SPICE simulations demonstrate that our method leads to robust design of a 14-bit super-resolution ADC using 3-bit RRAM devices with improved power and speed performance and competitive figure-of-merits (FoMs). In addition to the linear uniform quantization, the proposed ADC can also support configurable high-resolution nonlinear quantization with high conversion speed and low conversion energy, enabling future intelligent analog-to-information interfaces for near-sensor analytics and processing. Weidong Cao 0001, Liu Ke 0001, Ayan Chakrabarti, Xuan Zhang 0001 |
ICCAD | 1 |