EDBT 2026 Demo / reviewers in the wild / expert
Yu Gong 0002
dblp:76/3005-2
· DBLP profile ↗
15ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-2736-4635ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A High-Accuracy MRAM-Based Computing-in-Memory Macro for Secure Edge AI InferenceabstractComputing-in-memory (CIM) represents a pivotal technology for overcoming the speed and power bottlenecks posed by the “memory wall” and “power wall” existing in von Neumann architectures. With the fast development of Internet of Things, data security has become one of the most attractive research topics for CIM in edge applications, as well as computing accuracy and energy efficiency. This paper proposes a highly accurate and secure MRAM-based CIM macro that is designed to reduce the multiply accumulate (MAC) computation errors and protect weight bits in untrusted environments. The architecture employs a series-connected structure to enhance computational linearity and introduces a dynamic reference column to increase reliability. Meanwhile, a lightweight encryption mechanism based on physical unclonable function (PUF) is implemented to protect weight bits. The results demonstrate that, after the obfuscation, the prediction accuracy of machine learning-based attacks on the PUF is reduced to approximately 50%. The CIM macro achieves impressive inference accuracy of 93.73% on the CIFAR-10 dataset with energy efficiency of 38.3 TOPS/W. Furthermore, security verification performed on the CNN model indicates that weight encryption degrades inference accuracy to 10%, thereby providing robust protection against potential attacks. You Wang 0002, Jiaao Dai, Shuo Fan, Yijun Cui, Yu Gong 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | PreDAC: An Efficient Framework of Pre-Refining Enhanced Design Space Exploration for Approximate ComputingabstractApproximate computing has emerged as a promising solution in energy-efficiency applications. Recently, attention has shifted from approximate components to Design Space Exploration (DSE) algorithms. However, traditional DSE algorithms face challenges in efficiently obtaining optimal solutions within large and complex design spaces. This paper introduces a prerefining enhanced design space exploration framework that provides customized design space and cost-performance formula for applications. Experimental results demonstrate that integrating this pre-refining step into various DSE algorithms leads to substantial performance gains, including up to $87 \times$ speedup and a 23% improvement in hardware overhead. Moreover, the innovative cost-performance-based DSE algorithm attains a $7.7 \times$ acceleration and further optimizes hardware metrics by an additional 8.8% compared to advanced frameworks employing the same pre-refinement. Ziying Cui, Ke Chen 0018, Bi Wu 0002, Yu Gong 0002, Chenggang Yan 0002, Weiqiang Liu 0001 |
DAC | 4 |
| 2025 | Rank-based Multi-objective Approximate Logic Synthesis via Monte Carlo Tree SearchabstractApproximate Logic Synthesis (ALS) is an automated technique designed for error-tolerant applications, optimizing delay, area, and power under specified error constraints. However, existing methods typically focus on either delay reduction or area minimization, often leading to local optima in multi-objective optimization. This paper proposes a rankbased multi-objective ALS framework using Monte Carlo Tree Search (MCTS). It develops non-dominated circuit ranking, to guide MCTS in exploring local approximate changes (LACs) across the entire circuit and generate approximate circuit sets with great optimization potential. Additionally, a Rank-Transformer model is introduced to predict pathdomain ranks, enhancing the application of high-quality LACs within circuit paths. Experimental results show that our framework achieves faster and more efficient optimization in delay and area simultaneously compared to state-of-the-art methods. Yuyang Ye 0001, Xiangfei Hu, Peng Xu 0052, Yu Gong 0002, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi |
DAC | 5 |
| 2025 | Exploring Teaching Methods for Courses on Radiation Hardening Technology in ICsabstractWith the rapid advancement of space exploration technology, the use of intelligent equipment and systems is increasing at an accelerated pace. As the core component of intelligent systems, integrated circuits (ICs) have become a key area of research in space applications. However, the complex space environment significantly degrades the reliability of ICs due to radiation effects. As a result, radiation hardening technology is critical for ICs used in space applications. Unlike general consumer electronics, students majoring in ICs are often unfamiliar with radiation hardening technologies, which is a disadvantage for those who may work in industries such as aerospace, nuclear, or medical electronics after graduation. This paper explores teaching methods for a course on radiation hardening technology in ICs. Through interdisciplinary collaboration and joint university-enterprise teaching, as well as classroom interaction and project-based learning, students will gain an in-depth understanding of radiation sources, radiation effects, hardening techniques, and irradiation testing. You Wang 0002, Erya Deng, Yu Gong 0002, Zhongkun Shen, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2025 | A High-Performance In-Memory Multi-Bit Adder Based on TST-MRAMabstractIn traditional von-Neumann architecture, the memory and the arithmetic logic unit (ALU) are separated. The extra overhead caused by data transfer limits the performance of ALUs in data-intensive scenarios. Compute-in-memory (CiM) architecture based on emerging non-volatile memories (NVMs) has been proven to be effective in addressing the "memory wall" issue. However, current NV-CiM schemes primarily focus on the Boolean logic paradigm, limiting the parallelism of multi-bit computations. In this paper, we propose a high-performance in-memory multi-bit adder based on the toggle spin torque MRAM crossbar array. We use the time-based sensing amplifier to implement majority logic in the crossbar array. Based on majority logic gate, we design a multi-bit parallel-prefix adder. Compared to the multi-bit adders implemented based on Boolean logic gate, our work reduces the number of memory read/write operations and improves the computational parallelism. Moreover, the proposed multi-bit addition scheme exhibits O(log2(n)) latency and requires 6n cells for n-bit adder. Erya Deng, Zhongkun Shen, Yu Gong 0002, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2025 | Dynamic Challenge Cross-Selection Physical Unclonable Function Based on MRAMabstractThe rapid development of Internet of Things (IoT) devices has triggered massive data transmission. Meanwhile, advances in artificial intelligence (AI) introduce new security vulnerabilities in device interactions. These challenges demand lightweight yet robust security solutions. In this context, physical unclonable functions (PUFs) serve as critical hardware security primitives, enabling reliable authentication for edge devices. Nevertheless, PUF is increasingly susceptible to novel threats, notably machine learning attacks. To address this security vulnerability to attacks, we propose a novel double-layer dynamic challenge cross-selection magnetoresistive random access memory PUF (MPUF). This design leverages the inherent process variation in spin-transfer torque magnetoresistive random access memory (STT-MRAM) as an entropy source. The proposed structure incorporates an obfuscation decode circuit (ODC) that combinesxorgates and shift registers. It dynamically obfuscates interlayer relationships between two PUF arrays to enhance circuit nonlinearity. The simulation results demonstrate uniformity of 50.16%, uniqueness of 49.94%, a worst bit error rate (BER) of 2.34% for$- 25~^{\circ } $C to$125~^{\circ } $C and 1.56% for$0.5\sim 1.1$V. In addition, four common machine learning models are used to attack this PUF, achieving accuracies of 50.49%, 50.49%, 50.48%, and 58.41%, which are close to a random guess. Compared with traditional PUF implementations, this work exhibits higher reliability and enhanced security while maintaining low power consumption of approximately 9.975 fJ/bit. Siying Wu, Yu Gong 0002, Jiaao Dai, Shouzhong Peng, Yue Zhang 0010, You Wang 0002, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | A Time Efficient Comprehensive Model of Approximate Multipliers for Design Space ExplorationabstractMultipliers play an essential role in various data processing applications and have garnered significant attention in approximate computing (AxC) for their energy-efficient features. However, formulating a precise error model for approximate data processing algorithms in conjunction with hardware metrics presents a challenge, leading to substantial time consumption in the design space exploration. This paper introduces an analytical model for approximate multipliers while considering input patterns. This model furnishes accurate error metrics, along with high-precision hardware metrics for various approximate multiplier configurations, impervious to variations in input data distribution. The proposed error model reduces the runtime by an average factor of 120.85 and, in some instances, by as much as 2,500 times, when contrasted with simulation-based methods. The design space exploration is performed on a 3×3 convolution circuit, revealing a comparable Pareto-optimal set and substantial reductions of up to 79.46% in the Power-Delay-Product (PDP) and 71.98% in area compared to the accurate counterpart. Additionally, the result of the Gaussian Blur application experiment demonstrates a 68.59% reduction in PDP and a 56.21% reduction in area, all while maintaining a PSNR of 30 dB. Ziying Cui, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Yu Gong 0002, Weiqiang Liu 0001 |
ARITH | 5 |
| 2022 | An Efficient BCNN Deployment Method Using Quality-Aware Approximate ComputingabstractAs the artificial intelligence and Internet of Things (AIoT) develop rapidly, the deployment of artificial neural networks in edge computing is becoming significant with great challenge. The binarized convolutional neural network (BCNN) is one of the most widely adopted light-weight ANNs in AIoT, which can achieve the balance of system accuracy and hardware resource consumption, compared to others. To achieve high power and area efficiency in BCNN deployment, many approximate computing (AxC) techniques are integrated to make full use of the resilience of BCNN. As the research focused on the integration of AxC in circuit design, the design of AxC itself is not fully considered when applied to specific applications or domains. Based on circuit-architecture-system co-design, this article proposes an efficient BCNN deployment method, including a quality-circuit co-design method for approximate adder generation, a quality-aware intercompensation approach for addition tree, and a computing quality involved retraining approach for BCNN deployment. Experimental results show that the proposed quality model can achieve 86.43% in average accuracy while evaluating nine types of typical approximate adders. The proposed method is conducted on the applications of keyword spotting of GSCD, MNIST, and CIFAR-10, and we can further rise the approximation degree by 50%–75%, while reducing the accuracy by less than 1%. Bo Liu 0019, Xuetao Wang, Anfeng Xue, Qiao Shen 0001, Na Xie, Yu Gong 0002, Zhen Wang 0019, Jun Yang 0006, Hao Cai 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2022 | Quality Driven Systematic Approximation for Binary-Weight Neural Network DeploymentabstractNeural networks (NNs) with large scales of artificial neurons are increasingly used in recognition and classification tasks. In power-constrained scenarios, the tradeoff between performance and hardware consumptions must be carefully evaluated before silicon tape-out. In this paper, we proposed a systematic approach to design ultra-low power NN system. This work is motivated by the facts that NNs are resilient to approximation in many of the computations and NNs are outputting statistical tensors which are acceptable to less-than-perfect results. We resort to the front-back end approach with a twofold aim: (1) a fast and accurate design approach is proposed by estimating the computing quality of low-power approximate adder arrays, and it is adopted to evaluate the neural network system; (2) a quality configurable engine with different approximation degrees while processing NNs is implemented. The proposed work is demonstrated with a comprehensive keyword spotting (KWS) system as an ultra-low power NN engine. The experimental environment is setup with ten keywords from the google speech command dataset (GSCD) using an industrial 22-nm ultra-low-leakage (ULL) process. Comparing to the state-of-the-art KWS processors, the proposed approximate NN engine can demonstrate over 60% improvement in power efficiency and$1.1\times $area efficiency while achieving similar recognition accuracy. Yu Gong 0002, Hao Cai 0001, Haige Wu, Hao Yan 0002, Zhen Wang 0019, Longxing Shi, Bo Liu 0019 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | More is Less: Domain-Specific Speech Recognition Microprocessor Using One-Dimensional Convolutional Recurrent Neural NetworkabstractLow-power keywords recognition has been a focus of acoustic signal processing for several decades. This work investigates the domain-specific speech recognition microprocessor based on optimized one-dimensional convolutional recurrent neural network (1D-CRNN). Compared to previous DNN based frameworks, the proposed 1D-CRNN can process both the feature extraction and keywords classification, and achieve high recognition accuracy with reduced computation operations under wide range background noise SNRs. An energy-efficient 1D-CRNN accelerator is implemented to dynamically reconfigure and process the different layers. This accelerator has the characteristics of “More is Less” in three aspects: 1) the hybrid network with more complex layers is much more compact and requires less computation; 2) although the weight width quantized to 8 bits requires more memory size and multiplication energy cost, the required network neurons can be reduced and hardware utilization can be improved; 3) an energy-aware self-compensation tensor multiplication unit with dual power supply based on approximation design method can be utilized for 1D-CRNN computing. Compared to the state-of-the-art architectures, the novel more-is-less architecture can achieve a much lower power consumption of$1.4~\mu \text{W}\sim 2.1~\mu \text{W}$(over 80% reduced) under an industry 22nm technology, while maintaining higher system adaptability (support SNRs: −5dB~Clean) for 1~5 real-time keywords recognition. Bo Liu 0019, Hao Cai 0001, Xiaoling Ding, Yu Gong 0002, Weiqiang Liu 0001, Jinjiang Yang, Zhen Wang 0019, Jun Yang 0006 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | A 1D-CRNN Inspired Reconfigurable Processor for Noise-robust Low-power Keywords RecognitionabstractA low-power high-accuracy reconfigurable processor is proposed for noise-robust keywords recognition and evaluated in 22nm technology, which is based on an optimized one-dimensional convolutional recurrent neural network (1D-CRNN). In traditional DNN-based keywords recognition system, the speech feature extraction based on traditional algorithms and the DNN based keywords classification are two independent modules. Compared to the traditional architecture, both the feature extraction and keywords classification are processed by the proposed 1D-CRNN with weight/data bit width quantized to 8/8 bits. Therefore unified training and optimization framework can be performed for various application scenarios and input loads. The proposed 1D-CRNN based keywords recognition system can achieve a higher recognition accuracy with reduced computation operations. Based on system-architecture co-design, an energy-efficient DNN accelerator which can be dynamically reconfigured to process the 1D-CRNN with different configurations is proposed. The processing circuits of the accelerator are optimized to further improve the energy efficiency using a fine-grained precision reconfigurable approximate multiplier. Compared to the state-of-the-art architectures, this work can support 1~5 real-time keywords recognition with lower power consumption, while maintaining higher system capability and adaptability. Bo Liu 0019, Zeyu Shen 0003, Lepeng Huang, Yu Gong 0002, Hao Cai 0001 |
DATE | 4 |
| 2021 | Semi-Analytical Path Delay Variation Model With Adjacent Gates Decorrelation for Subthreshold CircuitsabstractThe subthreshold circuit is a practical design style for the ultralow-power applications, but its timing estimation is a challenge due to the increasing local variation effects. The delay variation of adjacent gates is not independent because of input slew variation caused by the precedent gate, so their correlation effects are difficult to model and estimate. This article proposes a semi-analytical statistical delay model considering local variation for the subthreshold region, that is, the combination of analytical and simulation-based method. First, it decorrelates the slew influence between adjacent stages by dividing delay and output slew model into fast/slow input cases and dividing delay variation model into process variation and input slew variation. Then, it can be applied into multi-PVT conditions with a one-time SPICE nominal simulation by analyzing the independence of variability and relative variability of step input gate delay variance with output load capacitance and process, voltage, and temperature (PVT). Finally, experiments are carried out for different benchmarks, processes, voltages, and temperatures (BPVTs). The average errors of variance on different BPVTs are 4.8%, 3.1%, 4.0%, and 4.7%. Compared with other analytical works, the accuracies' improvements of three metrics (variance, variability, and max delay) are 8.3×, 9.6×, and 2.7× by the mean error at all test benchmarks. Compared with industrial method LVF, it has a comparable error in max path delay and runtime, and less three orders of magnitudes than LVF in characterization time and stored data (from TB to GB) at all test benchmarks. Peng Cao 0002, Mengxiao Li, Yu Gong 0002, Zhiyuan Liu 0011, Geng Bai, Jun Yang 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | An Ultra-low Power Keyword-Spotting Accelerator Using Circuit-Architecture-System Co-design and Self-adaptive Approximate Computing Based BWNabstractThis paper proposed an ultra-low power keyword-spotting (KWS) accelerator using circuit-architecture-system co-design and precision self-adaptive approximate computing based binarized weight network (BWN). To reduce the power consumption while maintaining the system recognition accuracy for different background noise, we first proposed a bit-by-bit layer-by-layer quantization method to quantize the deep neural network (DNN) to BWN. Then, we proposed a precision self-adaptive approximate addition unit to further reduce the BWN energy consumption. Evaluated under TSMC22nm ULL process technology, this work can support up to 10 keywords real time recognition under different background noise types and SNRs (from 5dB to near microphone) with power consumption of 13.6uW. Bo Liu 0019, Hao Cai 0001, Zeyu Shen 0003, Yu Gong 0002, Lepeng Huang, Zhen Wang 0019 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2020 | A Background Noise Self-adaptive VAD Using SNR Prediction Based Precision Dynamic Reconfigurable Approximate ComputingabstractThis paper proposed a background-noise self-adaptive voice activity detection (VAD) accelerator using SNR prediction based precision dynamic reconfigurable approximate computing. To improve the energy efficiency while maintaining high recognition accuracy for different background noises, two optimization techniques are proposed. Firstly, we proposed a SNR prediction module to analyze and pre-classify the back-ground noise into different levels, and a binarized weight network (BWN) accelerator with reconfigurable data bit width to implement the feature classification of VAD. Then, we proposed an approximate computing architecture with precision self-adaptive approximate addition unit to further reduce the energy consumption of BWN accelerator. Evaluated under 28nm process technology, this work can achieve high recognition accuracy (speech/none-speech hit rate: 95%/92% @10dB, 90%/87% @5dB, and 85%/80% @-5dB) under different background noise (SNR-5dB) with a low power consumption of 2 ~ 8uW. Bo Liu 0019, Yan Li 0056, Lepeng Huang, Hao Cai 0001, Shisheng Guo, Yu Gong 0002, Zhen Wang 0019 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2020 | Binarized Weight Neural-Network Inspired Ultra-Low Power Speech Recognition Processor with Time-Domain Based Digital-Analog Mixed Approximate ComputingabstractIn this paper, an ultra-low power speech recognition processor is implemented based on an optimized binarized weight neural-network (BWN). To accelerate the BWN and make it energy efficient, we proposed an approximate computing architecture for the quantized BWN based on time-domain digital-analog mixed addition unit and precision optimization with fault-tolerant training method. Experimental results show that the proposed digital-analog mixed approximate computing architecture can significantly reduce the power consumption while maintaining the recognition accuracy. Implemented under TSMC 28nm, the proposed processor can support 10 keywords real time recognition under different noise types and SNRs, while the power consumption is 56μW. Bo Liu 0019, Hao Cai 0001, Yu Gong 0002, Yan Li 0056, Zhen Wang 0019 |
ISCAS | 3 |