Yang Liu 0062

dblp:51/3710-62 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-0615-7036ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 I2E: Real-Time Image-to-Event Conversion for High-Performance Spiking Neural Networks
abstract
Spiking neural networks (SNNs) promise highly energy-efficient computing, but their adoption is hindered by a critical scarcity of event-stream data. This work introduces I2E, an algorithmic framework that resolves this bottleneck by converting static images into high-fidelity event streams. By simulating microsaccadic eye movements with a highly parallelized convolution, I2E achieves a conversion speed over 300x faster than prior methods, uniquely enabling on-the-fly data augmentation for SNN training. The framework's effectiveness is demonstrated on large-scale benchmarks. An SNN trained on the generated I2E-ImageNet dataset achieves a state-of-the-art accuracy of 60.50%. Critically, this work establishes a powerful sim-to-real paradigm where pre-training on synthetic I2E data and fine-tuning on the real-world CIFAR10-DVS dataset yields an unprecedented accuracy of 92.5%. This result validates that synthetic event data can serve as a high-fidelity proxy for real sensor data, bridging a long-standing gap in neuromorphic engineering. By providing a scalable solution to the data problem, I2E offers a foundational toolkit for developing high-performance neuromorphic systems. The open-source algorithm and all generated datasets are provided to accelerate research in the field.
Liwei Meng, Guanchao Qiao, Ning Ning 0002, Yang Liu 0062, Shaogang Hu
AAAI5
2026 Reconfigurable In-Memory Computing With Compute-Fusion SAR ADC for Area-Efficient Pipelined MAC Operations
abstract
This brief proposes a reconfigurable compute-in-memory (CIM) macro using a compute-fusion successive approximation register ADC (CF-SAR) that integrates charge-domain multiply-and-accumulate (MAC), SAR quantization, and partial-sum accumulation in a unified data path. By reusing the capacitor array for both analog MAC and ADC conversion, the design reduces area overhead. A reconfigurable in-memory-computing storage and logic (RSL) element merges SAR result storage with in-memory serial addition (IMSA), enabling a pipelined ping-pong quantization/serial accumulation scheme without external digital adders. The macro supports 1-/2-/4-bit signed and unsigned weights viaand/nandselection and bank swapping without extra hardware. Implemented in 40 nm CMOS, the macro reduces transistor count by 44%, achieving up to 250.73 TOPS/W and 255.5 GOPS, with normalized efficiencies of 273.92 TOPS/W and 197.61 TOPS/mm2, and 98.54%/86.62% accuracy on MNIST/CIFAR-10.
Shiqin Yan, Junjie Wang 0008, Shuang Liu 0019, Y. T. Liu, H. D. Mao, Xiangzhan Wang, Tupei Chen, Yang Liu 0062
IEEE Trans. Very Large Scale Integr. Syst.10
2026 Neuromorphic Hybrid Information Processing Architecture for High-Performance Information Compression
abstract
The rapid development of artificial intelligence-of-things (AIOT) has led to a significant increase in intelligent nodes, presenting substantial challenges to internode communication. Semantic communication systems based on artificial neural networks (ANNs) have demonstrated greater robustness than traditional Shannon systems. However, the significant resource overhead makes hardware implementation challenging, and the low efficiency of information compression places considerable strain on communication bandwidth. This work proposed a hybrid semantic system that employs ANNs for high-precision semantic extraction at the server and spiking neural networks (SNNs) for high-performance information compression in the spatiotemporal dimension and semantic comprehending with low-hardware cost at the edge. A hardware-friendly, event-based neuromorphic core has been developed with low-hardware resource consumption and a high sampling rate for SNN deployment. The evaluation results indicate that the hybrid semantic system reduces the bandwidth by 95.8%, while maintaining a high sampling rate of 1000 Sa/s, outperforming traditional systems. Meanwhile, it enables a low entropy of the receiving end after four samples. Considering both bandwidth overhead and the entropy of the receiving end, the hybrid system achieves a high information compression ratio (ICR), surpassing the ANN-based system over 20 times. This work is expected to reveal the significance of neuromorphic computing in enabling high-performance information compression and transmission (Tx) in semantic communication.
Pujun Zhou, Qi Yu 0002, Liwei Meng, C. X. Xiong, G. L. Yang, Guanchao Qiao, Ning Ning 0002, Yang Liu 0062, Shaogang Hu
IEEE Trans. Very Large Scale Integr. Syst.9
2025 A Neuromorphic Transformer Architecture Enabling Hardware-Friendly Edge Computing
abstract
The transformer model has demonstrated significant capabilities in various intelligent tasks, attracting widespread attention in recent years. However, it involves numerous complex operations, including large-bit-width multiplication, division, matrix transposition, and exponentiation. These require substantial storage and computational resources, making it challenging to deploy on edge devices. This work introduces a neuromorphic transformer architecture with low hardware cost for AI edge computing (AI-EC). At the structural level, it absorbs scaling factors within the self-attention mechanism into weight matrixes, thereby eliminating the division caused by the scaling operation. Additionally, a transposition calculation method is proposed to perform matrix transposition using dedicated memory access strategies and optimized data flow designs, which reduces logic resource overhead and avoids memory access discontinuities. At the computing paradigm level, the architecture employs spike-driven computing, substituting multi-bit multipliers with AND logic for synaptic operations. The paradigm introduces high sparsity to computational data, which is effectively exploited to reduce the computational workload of the architecture. The results indicate that the architecture successfully eliminates high-cost operators and significantly reduces computational expenses. Eventually, this architecture is verified as a prototype using a 28 nm CMOS process library, demonstrating a compact logic area of sub-0.2 mm2and a high energy efficiency of 0.34 pJ/SOP @ 50MHz. This work is expected to promote the application of transformers in edge computing and the development of intelligent edge applications.
Pujun Zhou, R. C. Ma, Z. T. Liu, Liwei Meng, Guanchao Qiao, Yang Liu 0062, Qi Yu 0002, Shaogang Hu
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 A 0.96 pJ/SOP Heterogeneous Neuromorphic Chip Toward Energy-Efficient Edge Visual Applications
abstract
Edge devices require low power consumption and compact area, which poses challenges for visual signal processing. This work introduces an energy-efficient heterogeneous neuromorphic system-on-chip (SoC) for edge visual computing. The neuromorphic core design incorporates advanced technologies, such as sparse-aware synaptic calculation, partial membrane potential update, non-uniform weight quantization, and partial parallel computing, achieving excellent energy efficiency, computing performance, and area utilization. Twenty neuromorphic cores and twelve multi-mode connected-matrix-based routers form a network-on-chip (NoC) with fullerene-like topology. Its average degree of communication nodes exceeds traditional topologies by 32 % and maintains a minimum degree variance of 0.93, thereby enabling advanced decentralized on-chip communication. Moreover, the NoC can be scaled up through extended off-chip high-level router nodes. At the top layer of the SoC, a RISC-V CPU and a 20-core neuromorphic processor are tightly coupled to form a heterogeneous architecture. Eventually, the chip is fabricated within a 3.41 mm2die area under 55 nm CMOS technology, achieving a low power density of 0.52 mW/mm2and a high neuron density of 30.23 K/mm2. Its effectiveness is verified across different visual tasks, with a best energy efficiency of 0.96 pJ/SOP. This work is expected to promote the development of neuromorphic computing in edge visual applications.
Pujun Zhou, Guanchao Qiao, Qi Yu 0002, Junjie Wang 0008, Ning Ning 0002, Yang Liu 0062, Shaogang Hu
IEEE Trans. Circuits Syst. Video Technol.9
2024 A&B BNN: Add&Bit-Operation-Only Hardware-Friendly Binary Neural Network
abstract
Binary neural networks utilize 1-bit quantized weights and activations to reduce both the model's storage demands and computational burden. However, advanced binary architectures still incorporate millions of inefficient and non-hardware-friendly full-precision multiplication operations. A&B BNN is proposed to directly remove part of the multiplication operations in a traditional BNN and replace the rest with an equal number of bit operations, introducing the mask layer and the quantized RPReLU structure based on the normalizer-free network architecture. The mask layer can be removed during inference by leveraging the intrinsic characteristics of BNN with straightforward mathematical transformations to avoid the associated multiplication operations. The quantized RPReLU structure enables more efficient bit operations by constraining its slope to be integer powers of 2. Experimental results achieved 92.30%, 69.35%, and 66.89% on the CIFAR-10, CIFAR-100, and ImageNet datasets, respectively, which are competitive with the state-of-the-art. Ablation studies have verified the efficacy of the quantized RPReLU structure, leading to a 1.14% enhancement on the ImageNet compared to using a fixed slope RLeakyReLU. The proposed add&bit-operation-only BNN offers an innovative approach for hardware-friendly network architecture.
Guanchao Qiao, Yian Liu 0001, Liwei Meng, Ning Ning 0002, Yang Liu 0062, Shaogang Hu
CVPR6
2024 Live Demonstration for Input-Sparsity-Aware RRAM Processing-in-Memory Chip
abstract
This paper presents a live demonstration of an RRAM processing-in-memory (PIM) chip in which the input sparsity is exploited to reduce power consumption and increase the throughput of the PIM chip. An offline quantization-aware training (QAT) is employed to fine-tune models to be suitable for the 4-bit PIM chip. Post-QAT, the model exhibited accuracy of 90.08% on the test dataset. Interestingly, we found that the input sparsity of input activation is always over 90%. This high level of sparsity proves advantageous, contributing substantially to both throughput and energy efficiency of the PIM chip. This design yields a throughput of 410 Gops, which is 9 times higher than the design without input sparsity awareness.
Junjie Wang 0008, Shuang Liu 0019, Ruicheng Pan, Shiqin Yan, Yang Liu 0062
ISCAS6
2023 An efficient pruning and fine-tuning method for deep spiking neural network
Liwei Meng, Guanchao Qiao, Yue Zuo, Pujun Zhou, Yang Liu 0062, Shaogang Hu
Appl. Intell.7
2019 A Low Voltage 10-Bit Non-Binary 2B/Cycle Time and Voltage Based SAR ADC
abstract
This paper proposes a low power 10-bit 2b/cycle time and voltage based-successive approximation register analog-to-digital converter (ADC). At low supply voltage, there will be a significant difference in comparator decision time for different input voltages. By taking advantage of the fact, this ADC converts the reference voltage to the corresponding comparator decision time, achieving 2b/cycle quantization to improve the conversion speed. In addition, by obtaining reference delays with duplicated circuits and using non-binary capacitor arrays, the ADC can tolerate process, voltage and temperature (PVT) variations and decision errors. To validate these concepts, a 10-bit 2 MS/s SAR ADC is designed using 130nm CMOS process with 0.5 V power supply voltage. Simulation results shows the ADC achieve signal-to-noise distortion ratio (SNDR) of 59.63 dB, corresponding to an effective number of bits (ENOB) of 9.61 bits and consumes 3.2 µW, resulting in a figure of merit (FOM) of 2.06 fJ/c-s.
Jian Luo 0004, Jing Li 0022, Ning Ning 0002, Kejun Wu, Zhen Liu 0013, Yang Liu 0062, Qi Yu 0002
ISCAS6
2018 A 0.9-V 12-bit 100-MS/s 14.6-fJ/Conversion-Step SAR ADC in 40-nm CMOS
Jian Luo 0004, Jing Li 0022, Ning Ning 0002, Yang Liu 0062, Qi Yu 0002
IEEE Trans. Very Large Scale Integr. Syst.4
2015 A Supply Voltage and Temperature Variation-Tolerant Relaxation Oscillator for Biomedical Systems Based on Dynamic Threshold and Switched Resistors
abstract
A fully integrated supply voltage and temperature variation-tolerant relaxation oscillator for biomedical systems has been presented. Concepts of dynamic threshold and switched resistors are proposed to improve the frequency stability against power supply and temperature variations, respectively. This design was verified in a 0.35-$\mu{\rm m}$standard CMOS process with a 3 V supply. Measurement results show the frequency drift of 0.6% from 2.4 to 4.0 V and temperature stability of 53.9 ppm/$^{\circ}{\rm C}$as temperature varied from${-}{30}{}^{\circ}{\rm C}$to 120$^{\circ}{\rm C}$at a typical working frequency of 4 MHz. With the consideration of resistor and transistor matching, the oscillator was implemented in a core area of 0.05${\rm mm}^{2}$.
Zhentao Xu, Wei Wang 0152, Ning Ning 0002, Wei Meng Lim, Yang Liu 0062, Qi Yu 0002
IEEE Trans. Very Large Scale Integr. Syst.5
2014 A 10-bit 100MS/s subrange SAR ADC with time-domain quantization
abstract
This paper presents a 10-bit subrange successive approximation register analog-to-digital converter (SAR ADC). A 3.5-bit time-domain coarse ADC converts the analog input to the time delay of two pulse signals and a time-to-digital converter (TDC) is used to quantize the delay. The coarse ADC controls the switching of the higher 3-bit capacitors in the digital-to-analog converter (DAC). A 7-bit SAR controls the remaining capacitors. The 1-bit redundancy corrects the linearity and mismatch error of the coarse ADC. The proposed 10-bit 100MS/s ADC is designed in a 65nm CMOS technology with 1.2V power supply. Simulation results show that this design achieves 59.7dB SNDR and consumes 2.69mW. The figure-of-merit (FOM) is 34.2fJ/conversion-step.
Shuangyi Wu, Ning Ning 0002, Qi Yu 0002, Yang Liu 0062
ISCAS6