VLDB 2026 Research / reviewers in the wild / expert
Qi Liu 0010
dblp:95/2446-10
· DBLP profile ↗
36ranked-venue papers
0as first author
34since 2021 · last 2026
0000-0001-7062-831XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 31 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NS-FPS: Accelerating Farthest Point Sampling via Neighbor Search in Large-Scale Point Clouds
Jiapei Zheng, Shuan Yang, Siqi He, Qi Liu 0010, Chixiao Chen |
ISCA | 4 |
| 2026 | An RRAM-based Neuromorphic Sleep Monitoring System for Energy-efficient Edge Healthcare Applications
Fangduo Zhu, Jingsong Zhang, Jinhao Liang, Xumeng Zhang, Qi Liu 0010 |
ISCAS | 9 |
| 2026 | An RRAM-based Multi-Timescale Spiking Processor with Reconfigurable Neurons
Jinhao Liang, Fangduo Zhu, Siyuan Ouyang, Jingsong Zhang, Xumeng Zhang, Qi Liu 0010, Ming Liu 0022 |
ISCAS | 8 |
| 2026 | A Temperature-Adaptive Bias Generator with Threshold-Voltage Compensation from 3.6 K to 410 K
Yingzhe Sha, Jiaxuan Weng, Zhangcheng Huang 0001, Qi Liu 0010 |
ISCAS | 5 |
| 2026 | PSVS: A Parallel-Series Voltage-Sensing 1T4R RRAM Macro Achieving 15-Level Storage and 17.88 Mb/mm2 Density for LLM Inference
Kaijun Zhang, Wencong Wu, Jiapei Zheng, Jinru Lai, Xiaoxin Xu, Qi Liu 0010, Chixiao Chen |
ISCAS | 7 |
| 2026 | A Pipelined NoC-Based Membrane Shortcut SNN Architecture for Low-Latency Spike Sorting
Jingsong Zhang, Fangduo Zhu, Siyuan Ouyang, Jinhao Liang, Xumeng Zhang, Qi Liu 0010 |
ISCAS | 9 |
| 2026 | A 13-GS/s 9-bit Time-Interleaved Pipelined-SAR ADC With Common-Mode Regulated Floating-Inverter-Amplifier and Rapid-Tracking Bootstrapped SwitchabstractThis article presents a 13GS/s 9-bit 8-channel time-interleaved (TI) Pipelined-SAR (Pipe-SAR) ADC. A common-mode regulated floating-inverter-amplifier (CMR-FIA) is proposed to overcome the common-mode voltage variation due to the charge leakage through the parasitic capacitor, thereby eliminating the need for common-mode feedback (CMFB) circuitry embedded in the body of FIA. By combining an adaptively biased (AB) technique, the proposed FIA facilitates the high-speed and robust Pipe-SAR ADCs. In addition, a rapid-tracking bootstrapped signal generation is introduced to achieve high-linearity with short sampling time in an ultra-high speed sampling network. The ADC prototype is fabricated is a 28nm-CMOS process, the achieved spurious-free dynamic range (SFDR) and signal-to-noise and distortion ratio (SNDR) at the Nyquist input are 56.4dB and 41.75dB, respectively. Consuming 97mW at 13GS/s, it yields a Schreier figure of merit ($\text{FoM}_{\mathrm {S}}$) of 150dB. With the proposed CMR-FIA, the ADC’s SNDR variation is within 2.48dB across the input common-mode range of 0.3V to 0.8V. Ji Guo, Danfeng Zhai, Wenning Jiang, Qi Liu 0010, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | A 1024-Ch 583-nW/Ch Spike-Sorting SoC With Sparsity-Aware Spike Detection Scratchpad and Ultra-Low-Leakage Dual-Voltage 5T-SRAM for 16K-Template ClusteringabstractThis paper presents an energy-efficient spike-sorting system-on-chip (SoC) designed for closed-loop brain-computer interfaces of massive probing channels. The design first incorporates a sparsity/similarity-aware spike detection scratchpad, leveraging a bit-wise differential encoder and zero-friendly read-out circuits, reducing the dynamic power consumption of spike detection by 77.7%. To mitigate static power dissipation, it also introduces an ultra-low-leakage dual-voltage 5T-SRAM array with level-shifter embedded sense amplifiers, achieving an 82.2% leakage power reduction of neural signal buffering by applying half$V_{DD}$on SRAM cells. Additionally, a memory hierarchy architecture combining on-chip SRAM and off-chip FeRAM, along with a firing-rate-based Osort for cluster template management, minimizes off-chip memory access to only 9.7% with a latency of$11.7\mu $s for 1024-channel spike sorting. A silicon prototype is fabricated in 28-nm CMOS technology, which achieves a power consumption of 583nW/channel and an area consumption of 0.0012mm2/channel. The chip supports real-time spike sorting with up to 16K templates,$21.3\times $greater than the state-of-the-art spike-sorting processor. Hao Jiang 0024, Zexing Chen, Jiajun Lu, Siqi He, Liangjian Lyu, Jiamin Xu, Shiwei Liu 0002, Yingping Chen, Chixiao Chen, Qi Liu 0010, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2026 | COPIC: A Codesigned Accelerator for Efficient Octree-Based Learned Point Cloud Compression at EdgeabstractWith the growing adoption of light detection and ranging (LiDAR) technology and the increasing resolution of captured data, efficient point cloud compression (PCC) has become a critical challenge. While PCC is essential for reducing bandwidth and storage demands, existing octree-based learned PCC methods face two fundamental bottlenecks: 1) octree construction and serialization that involve irregular, pointwise memory accesses that are difficult to parallelize and 2) autoregressive entropy model inference that requires heavy computation. These issues lead to high latency and energy consumption, making real-time deployment on edge devices impractical. This article introduces COPIC, a hardware accelerator designed for real-time PCC on edge devices. COPIC integrates two key components. The first is OctCAM, a content-addressable-memory (CAM)-based module for rapid octree construction and serialization. By enabling hardware-level parallel search, OctCAM significantly reduces the latency and energy consumption caused by irregular memory access patterns in octree processing. The second is PCC-Engine, a dedicated neural network inference module. Combined with a lightweight PCC-TinyAttention network and a window-based key-value caching approach, it further minimizes execution latency and power usage. Experimental results show that COPIC achieves a$25.9\times $speedup compared to GPU implementations, with a total processing latency of 40.1 ms and an average energy consumption of 0.23 J per frame, while delivering a compression rate 12.4% higher than standard proposals. Jiapei Zheng, Yutong Su, Siqi He, Haozhe Zhu, Qi Liu 0010, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | Hydra: Harnessing Expert Popularity for Efficient Mixture-of-Expert Inference on Chiplet SystemabstractThe rapid growth of model sizes in advanced artificial intelligence algorithms, particularly in Transformerbased large language models (LLMs), has led to significant computational overhead. Mixture-of-Expert (MoE) models offer a solution through their sparsely gating mechanism but introduce new challenges of extensive all-to-all communication and model computational inefficiencies. This paper presents Hydra, a software/hardware co-design aimed at accelerating MoE inference on chiplet-based architectures. In software, Hydra employs a popularity-aware expert mapping strategy to optimize interchiplet communication. In hardware, it incorporates Content Addressable Memory (CAM) to eliminate expensive explicit token (un)-permutation based on sparse matrix multiplications and a redundant-calculation-skipping softmax engine to bypass unnecessary division and exponential operations. Evaluated in 22 nm technology, Hydra achieves latency reductions of $14.2 \times$ and $3.5 \times$ and power reductions of $169.1 \times$ and $18.9 \times$ over GPU and state-of-the-art MoE accelerator, respectively, thereby offering a scalable and efficient solution for MoE model deployment. Siqi He, Haozhe Zhu, Jiapei Zheng, Lizhou Wu, Bo Jiao 0003, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen |
DAC | 6 |
| 2025 | DRAFT: Decoupling Backpropagation from Pre-trained Backbone for Efficient Transformer Fine-Tuning on EdgeabstractTransformers have demonstrated outstanding performance across diverse applications recently, necessitating finetuning to optimize their performance for downstream tasks. However, fine-tuning remains challenging due to the substantial computational costs and storage overhead of backpropagation (BP). The existing fine-tuning techniques require the BP through the massive pre-trained backbone weights for computing the input gradient, resulting in significant computing overhead and memory footprint for resource-constrained edge devices. To address the challenge, this work proposes an algorithm-hardware co-design framework, DRAFT, for efficient Transformer finetuning by decoupling the BP from the backbone weights, thereby efficiently reducing the BP overhead. The framework employs Feedback Decoupling Approximation (FDA), an efficient finetuning algorithm that decouples BP into two low-complexity pathways: trainable adapter pathway and sparse ternary Bypass Network (BPN) pathway. The two pathways work collaboratively to approximate the conventional BP process. Further, a DRAFT accelerator is proposed, featuring a reconfigurable design with lightweight sparse gather networks and dynamic workflows to fully harness the sparsity and data parallelism inherent to the FDA. Experimental results demonstrate that DRAFT achieves a speedup of $4.9 \times$ and an energy efficiency improvement of $4.2 \times$ on average compared to baseline fine-tuning methods across multiple fine-tuning tasks with negligible accuracy loss. Zhirui Huang, Shiwei Liu 0002, Haozhe Zhu, Qi Liu 0010, Chixiao Chen |
DAC | 4 |
| 2025 | McPAL: Scaling Unstructured Sparse Inference with Multi-Chiplet HBM-PIM Architecture for LLMsabstractLarge language models (LLMs) have gained significant attention recently. However, executing LLM is memory-bound due to the extensive memory accesses. Process-in-memory (PIM) emerges as an energy-efficient solution for LLMs, delivering high memory bandwidth and compute parallelism. Nevertheless, the trend towards larger LLMs introduces escalating memory footprint challenges for monolithic PIM chips. This paper proposes McPAL, which tackles this challenge by emphasizing unstructured sparse compute within PIM and hierarchical multi-chiplet scaling. McPAL decomposes arbitrary sparse weight matrix into multiple irregular sparse vectors. The non-skipped computations in each vector are then routed via an in-memory butterfly network to the standard PIM array, enhancing the PIM utilization. In addition, we scale McPAL vertically by strategically organizing the 3D-HBM hierarchy to minimize the internal long-distance data travel. Meanwhile, a 2.5D IO chiplet scales McPAL horizontally, reducing die-to-die (D2D) data transfer and ensuring sparse workload balance. We conducted extensive experiments from Llama-7B to Llama-70B. The results show that McPAL achieves $1.57 \times$ to $3.12 \times$ speedup and $10.43 \times$ to $35.66 \times$ energy efficiency over Nvidia A100 GPU. Compared to SOTAs, McPAL also achieves $1.08 \times$ to $2.15 \times$ speedup and $1.65 \times$ to $5.14 \times$ energy efficiency. Shiwei Liu 0002, Zhirui Huang, Jiangnan Yu, Qi Liu 0010, Chixiao Chen |
DAC | 4 |
| 2025 | SDISC: A Spike-Driven Human-Machine Interface with In-Situ Computing for Real-Time Low-Power InteractionabstractFeature extraction and classification of bio-signals are crucial in human-machine interface (HMI), yet suffer from high delay and limited energy efficiency using conventional hardware. To mitigate this challenge, we propose an SDISC architecture, a neuromorphic HMI with the innovation from signal encoding, computing-in-memory (CIM) hardware, to algorithm-hardware co-optimization. The following strategies are implemented: (1) A spike-driven feature extractor, achieving > $10 \times$ sparser dataflow than frame-based method; (2) In-situ computing based on resistive random-access memory (RRAM), enabling energy-efficient (4.09 TOPS/W) spiking neural network (SNN) classifier; (3) A Spike-Activity-Distillation algorithm and an Aid-Loser-Only recovery scheme to alleviate the non-ideality of RRAM devices, ensuring SDISC maintains high accuracy ($\sim \mathbf{9 8. 0 \%}$) in long time inference ($\boldsymbol{\gt} \mathbf{1 5}$ days). We further develop an end-to-end SDISC system for real-time EMG-based robot control, achieving a low latency ($34 \mu \mathrm{~s}$) and low power ($39.72 \mu \mathrm{~W} /$ sample) interaction on edge. Fangduo Zhu, Jingsong Zhang, Xumeng Zhang, Siyuan Ouyang, Chenyang, Hao Jiang 0024, Qi Liu 0010 |
DAC | 9 |
| 2025 | EIGEN: Enabling Efficient 3DIC Interconnect with Heterogeneous Dual-Layer Network-on-Active-InterposerabstractChiplet-based 3DICs have emerged as modular solutions for large-scale, high-performance computing systems. However, unlike monolithic Network-on-Chips (NoCs), 3DICs using active interposers encounter difficulties in managing heterogeneous and high-load network traffic patterns. The emerging field of Networks-on-Active-Interposer (NOAI), independently designed from top dies’ interconnects, is desired to address these challenges by supporting flexible topologies, low-latency memory access, and reduced traffic congestion. To satisfy the requirements, we propose a heterogeneous dual-layer interconnect architecture, EIGEN, for chiplet-interposer systems, along with a reinforcement learning (RL)-based routing framework, which can provide efficient and flexible communication for chiplet-based 3DICs. EIGEN features an application-aware switch-programable interconnection layer (AspLayer) and a dynamic packet-routing interconnection layer (DynLayer) on the interposer, respectively. To optimize inter-chiplet data communication, we also develop an RL-based path routing framework tailored to this dual-layer architecture. The RL framework’s state space includes topology, network, and memory metrics, while the RL reward is set as the product of average packet latency and link utilization. The effectiveness of EIGEN is demonstrated through different applications where CPU, GPU, and AI accelerator chiplets are integrated via NOAI. Simulation results show that the proposed EIGEN achieves up to $\mathbf{6 7. 1 6 \%}$ latency reduction, $\mathbf{5 3. 8 9 \%}$ hops reduction and $\mathbf{1 1. 2 1 \%}$ runtime reduction compared to the state-of-the-art (SOTA) chiplets interconnect architectures. Furthermore, sensitivity analysis of network scaling shows that the EIGEN architecture and framework exhibit strong scalability, with latency reduction of $\mathbf{2 0. 9 6 \%}$ to $\mathbf{5 2. 7 7 \%}$ and hop reduction of $\mathbf{1 9. 7 2 \%}$ to $\mathbf{4 3. 4 1 \%}$ as the system scales from $4 \times 4$ to $32 \times 32$ chiplets. Siyao Jia, Bo Jiao 0003, Haozhe Zhu, Chixiao Chen, Qi Liu 0010, Ming Liu 0022 |
HPCA | 5 |
| 2025 | A 0.22 pJ/bit Processing-in-Controller GEMV Macro with Weight Prefetch for Efficient Near-Memory Computingabstract3D-stacked DRAM is a key technology enabling the development of large language models (LLMs). However, the intensive computational demands lead to substantial energy consumption resulting from large-scale data movement. Integrating processing-in-memory (PIM) within DRAM has been employed to reduce data movement, but it elevates manufacturing costs and incurs additional area and power consumption. On the other hand, implementing processing-near-memory (PNM) outside the DRAM controller fails to eliminate the high energy consumption associated with interconnects during data transfer. To address these challenges, we propose a processing-in-controller (PIC) architecture aimed at 3D-stacked DRAM for efficient data movement. A size-scalable General Matrix-Vector Multiplication (GEMV) structure is proposed, supporting configurations ranging from 16 × 16 to 128 × 128, thereby enhancing its adaptability for a variety of applications. To address the low throughput bottleneck caused by DDR read latency, a ping-pong buffer with weight prefetching is introduced in the PNM. Based on 28 nm CMOS technology, experimental results demonstrate that the energy consumption of this PIC is 0.22 pJ/bit, while data movement efficiency improves by nearly 92% compared to traditional DRAM operations. Jiapei Zheng, Siqi He, Lizhou Wu, Chen Mu, Haozhe Zhu, Liyu Lin, Qi Liu 0010, Chixiao Chen |
ISCAS | 8 |
| 2025 | RT-FLOW: FPGA Implementation of Real-Time Optical-Flow-Based SLAM for High-Speed Tracking and High-Quality MappingabstractSimultaneous Localization and Mapping (SLAM) is pivotal for autonomous robotics, yet feature-based SLAM systems struggle with sparse environmental representations and robustness under dynamic conditions. Optical-flow-based SLAM (OpF-SLAM) addresses these limitations by leveraging pixel-level motion data for dense mapping; however, its computational intensity hinders real-time deployment. This paper presents RT-FLOW, an FPGA-based accelerator for OpF-SLAM that achieves real-time performance through three key innovations: 1) A feature-context encoding engine that exploits inter-frame similarity to resolve data dependency in correlation construction, reducing latency by 77.5%. 2) A heterogeneous mixed-precision flow update engine guided by correlation sparsity, enabling 3.7× faster optical flow computation with negligible accuracy loss. 3) A pivoting-free linear solver using Householder transformations for stable pose optimization. Implemented on Xilinx XCZU7EV FPGA, RT-FLOW processes full-image pixels per frame at 65 fps with an energy efficiency of 0.358 μJ/point, outperforming previous FPGA designs. Evaluated on benchmark datasets, RT-FLOW demonstrates robustness in diverse environments while maintaining sub-110mJ/frame energy consumption. This work bridges the gap between algorithmic potential and hardware feasibility for high-density SLAM, empowering next-generation mobile robots with real-time scene understanding capabilities. Siqi He, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen, Haozhe Zhu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A Novel Neuromorphic Hardware Using Area-Efficient Chain RRAM-Based Synapses and Compact Neurons With (Anti-) Integration SchemeabstractThe neuromorphic system aims to implement large-scale spiking neural networks (SNN) through hardware, ultimately achieving human-level intelligence. At this stage, it is difficult for a single silicon-based chip to reach the density of the human brain, so it is important to improve the area efficiency of neuromorphic chips. We present a novel neuromorphic hardware that employs high area-efficiency chain RRAM synaptic array, and compact RRAM-based neurons. The proposed chain RRAM structure achieves a small cell size of 41.5F2 at 28 nm logic process. Compared to the conventional 1T1R structure, the chain structure has a 22.2% area reduction and a 58.8% parasitic capacitance reduction. As for the neuron design, we use RRAM instead of the capacitor to integrate membrane voltage. To alleviate the endurance issue of RRAM, we propose an (anti-) integration scheme that removes the operation of membrane voltage reset. The simulation results demonstrate an average energy consumption of 1.15pJ per spike and an area of$15.9\mu $m2. With the (anti-)integration scheme, the RRAM’s lifetime is demonstrated to extend by$2\times $and the energy of programming RRAM is demonstrated to be reduced by$2\times $. Qiqiao Wu, Honghu Yang, Yongkang Han, Haijun Jiang, Keji Zhou, Hailan Yi, Qi Liu 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | CEDAR: Computing-in-pixel Edge-aware Detection and Reconstruction Architecture for High-resolution 3D ImagingabstractLarge-format single-photon avalanche diode (SPAD)-based direct time of flight (dToF) sensors are expected to be widely applied in future L5 full driving automation. However, the high-power in-pixel TDCs and the huge amount of data generated by multiframe histogram sampling impose limitations on the pixel format of SPAD-based dToF sensors. To tackle this challenge, we proposed the Computing-in-pixel Edge-aware Detection and Reconstruction (CEDAR) architecture. In this architecture, edge pixels are recognized by charge-domain convolution (CDC) computing, and noise pixels are eliminated by in-memory denoising (IMD). Only few TDCs in these edge pixels are activated, resulting in significant power and data savings. Afterward, the full-format image is reconstructed by a U-Net using the obtained depth information from these edge pixels. For the first time, we proposed a high-resolution 512 × 512 SPAD-based dToF sensor with a low power of 83.3 mW, a distance accuracy of 0.9 cm, and a frame rate of 60 fps. The high-resolution 3D image can be reconstructed by only 3.5% sparse edge pixels, achieving a PSNR of 35.2 dB. The CEDAR architecture can achieve 16× pixel format and image resolution improvement under the same constraint of power dissipation. Bu Chen, Zhangcheng Huang 0001, Qi Zheng 0004, Weiyi Tang, Hankun Lv, Chixiao Chen, Jianlu Wang, Qi Liu 0010 |
DAC | 9 |
| 2024 | CAMPER: Exploring the Potential of Content Addressable Memory for 3D Point Cloud Efficient Range SearchabstractThe use of Light Detection and Ranging (LiDAR) for sensing has continuously improved the precision and performance of autonomous driving. At the same time, the large number of high-precision point clouds generated by LiDAR require real-time processing, and the range search is the key part of the processing pipeline. Content-Addressable Memory (CAM) has proven its efficiency for search tasks on switches and routers, but so far, there is still a lack of exploration on its application in point cloud range search. In this work, we propose CAMPER, a CAM-centered accelerator, aiming to explore the potential of CAM for point cloud range search. We developed a ripple comparison 13T (RC-13T) CAM cell for distance comparison, designed a spatial approximation search algorithm based on Chebyshev distance, and discussed the flexibility and scalability of the architecture. The results show that in the range search task of 64k@64k points, CAMPER achieves a latency of 0.83ms and a power consumption of 114.6mW. Compared with GPU, the throughput is increased by 10.4×; compared with SOTA accelerator, the energy efficiency is increased by about 228×. Jiapei Zheng, Lizhou Wu, Yutong Su, Zhangcheng Huang 0001, Chixiao Chen, Qi Liu 0010 |
DAC | 7 |
| 2024 | A 128×128 CMOS SPAD Receiver for 500Mbps Free Space Optical Communication with Column-wise Decoding and Fast Spot TrackingabstractThis work presents a 128×128 pixel array receiver based on single-photon avalanche diode (SPAD) for free space optical communication (FSOC). Each pixel incorporates an active quenching circuit and a delay-time-adjusting circuit to reduce the afterpulsing effect and the dead time. To address the challenge of high-speed transmission of massive data in a large-format SPAD array, a column-wise decoding circuit with reduced bus parasitic capacitance and voltage-sensitive discrimination is proposed, which significantly reduces the latency of data transmission. Additionally, the receiver includes cluster engines with highly parallelized computation capabilities for tracking the central addresses of a laser spot. The chip has been designed using a 130nm CMOS technology. Simulation results of the receiver indicate that a bit error rate (BER) of 3 × 10−4can be achieved at 500Mbps with a sensitivity of -41dBm, under random NRZOOK bitstreams. Furthermore, the chip demonstrates 100% accuracy in tracking the laser spot at a rate of 100kHz during 1000 transceiver simulations. Bu Chen, Zhangcheng Huang 0001, Qi Liu 0010 |
ISCAS | 3 |
| 2024 | A 6.4-Gbps 0.41-pJ/b fully-digital die-to-die interconnect PHY for silicon interposer based 2.5D integration
Yinglin Yang, Yunzhengmao Wang, Tengyue Yi, Chixiao Chen, Qi Liu 0010 |
Integr. | 5 |
| 2024 | FPIA: Communication-Aware Multi-Chiplet Integration With Field-Programmable Interconnect Fabric on Reusable Silicon InterposerabstractSilicon interposer re-usage is drawing attention for cost-effective multi-chiplet integrated systems. To address the communication awareness of inter/off-chiplet interconnect, the paper proposes a field-programmable interconnect fabric and develops its corresponding automatic physical integration tool. The tile-based fabric consists of turnout, cross-over boxes and parallel tracks. It features micro-bump-wise connecting flexibility and hardware efficiency. The automation flow performs chiplet location optimization and efficient bump-to-bump routing, supporting multi-lane bus interconnect and miscellaneous external ports. The methodology is validated by 9 different integration scenarios, where the routability is guaranteed when the local resource utilization ratio approaches 94.5%. The data’s maximum interconnect latency is 2.2 ns and the energy consumption is 1.18 pJ/bit at a bitrate of 1 Gbps. The latency consumes$16.5\times \sim ~53.4\times $fewer clock cycles than the state-of-the-art network-on-package-based reusable interposer architectures. Bo Jiao 0003, Haozhe Zhu, Jundong Zhu, Dexin Wen, Lingli Wang, Jun Tao 0001, Chixiao Chen, Yinhe Han 0001, Qi Liu 0010, Ninghui Sun, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 13 |
| 2024 | Write-Verify-Free MLC RRAM Using Nonbinary Encoding for AI Weight Storage at the EdgeabstractHigh-density and reliable multilevel-cell (MLC) resistive random access memory (RRAM) is expected to meet the ever-increasing demand for on-chip weight storages in the intelligent edge devices. However, due to the device variations, many write-and-verify (WAV) iterations are usually required to program the RRAM cell, which causes high power consumption, long latency, and degradation on the memory lifetime. To address this issue, we propose a write–verify-free MLC RRAM macro for weight storage with 1) a cascode-current-mirror multibit write (CCM-MW) driver and 2) a nonbinary programming scheme (NB-PS) with a radix not greater than 2. A 180-nm 400-Kb RRAM test chip is demonstrated in silicon. For 2-bit-per-cell MLC storage, the value error rates can be reduced by 24.13% after introducing two redundant bits (RBDs). In addition, compared to the single-level cell (SLC) storage scheme, a 37.50% reduction in the number of cells can be achieved to store the ResNet-8 model with a 0.79% loss in inference accuracy without the need for WAV iterations. Junjie An, Zhidao Zhou, Linfang Wang, Wang Ye, Weizeng Li, Hanghang Gao, Zhi Li 0062, Jinghui Tian, Hongyang Hu, Jinshan Yue, Lingyan Fan, Shibing Long, Qi Liu 0010, Chunmeng Dou |
IEEE Trans. Very Large Scale Integr. Syst. | 14 |
| 2024 | HARDSEA: Hybrid Analog-ReRAM Clustering and Digital-SRAM In-Memory Computing Accelerator for Dynamic Sparse Self-Attention in TransformerabstractSelf-attention-based transformers have outperformed recurrent and convolutional neural networks (RNN/ CNNs) in many applications. Despite the effectiveness, calculating self-attention is prohibitively costly due to quadratic computation and memory requirements. To solve this challenge, this article proposes a hybrid analog-ReRAM and digital-SRAM in-memory computing accelerator (HARDSEA), a computing-in-memory (CIM) accelerator supporting self-attention in transformer applications. To trade off between energy efficiency and algorithm accuracy, HARDSEA features an algorithm-architecture-circuit codesign. A product-quantization-based scheme dynamically facilitates self-attention sparsity by predicting lightweight token relevance. A hybrid in-memory computing architecture employs both high-efficiency analog ReRAM-CIM and high-precision digital SRAM-CIM to implement the proposed new scheme. The ReRAM-CIM, whose precision is sensitive to circuit nonidealities, takes charge of token relevance prediction where only computing monotonicity is demanded. The SRAM-CIM, utilized for exact sparse attention computing, is reorganized as an on-memory-boundary computing scheme, thus adapting to irregular sparsity patterns. In addition, we propose a time-domain winner-take-all (WTA) circuit to replace the expensive ADCs in ReRAM-CIM macros. Experimental results show that HARDSEA prunes BERT and GPT-2 models to 12%–33% sparsity without accuracy loss, achieving$13.5\times $–$28.5\times $speedup and$291.6\times $–$1894.3\times $energy efficiency over GPU. Compared to state-of-the-art transformer accelerators, HARDSEA has$1.2\times $–$14.9\times $better energy efficiency at the same level of throughput. Shiwei Liu 0002, Chen Mu, Hao Jiang 0024, Yunzhengmao Wang, Jinshan Zhang 0006, Keji Zhou, Qi Liu 0010, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2024 | Hi-NeRF: A Multicore NeRF Accelerator With Hierarchical Empty Space Skipping for Edge 3-D RenderingabstractNeural radiance field (NeRF) has proved to be promising in augmented/virtual-reality applications. However, the deployment of NeRF on edge devices suffers from inadequate throughput due to redundant ray sampling and congested memory access. To address these challenges, this article proposes Hi-NeRF, a multirendering-core accelerator for efficient edge NeRF rendering. On the architecture level, a hierarchical empty space skipping (HESS) scheme is adopted, which efficiently locates the effective samples with fewer skipping steps and thus accelerates the ray marching process. Furthermore, to alleviate the memory access bottleneck, a vertex-interleaved mapping (VIM) method that eliminates memory bank conflicts is also proposed. On the hardware level, ineffective sample filters (ISFs) and voxel access filters (VCFs) are introduced to further exploit spatial sparsity and data locality at run-time. The experimental results show that our work achieves$2.67\times $rendering throughput and$11.2\times $energy efficiency compared to a SOTA NeRF rendering accelerator. The energy efficiency can be improved by$561\times $compared to a commercial GPU. Lizhou Wu, Haozhe Zhu, Jiapei Zheng, Yinuo Cheng, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | TiPU: A Spatial-Locality-Aware Near-Memory Tile Processing Unit for 3D Point Cloud Neural NetworkabstractEnergy-efficient 3D point cloud neural network accelerators are desired for autonomous driving and AR/VR applications. This paper proposes TiPU, a spatial-locality-aware near-memory tile processing unit where the point clouds are partitioned into tiles to process spatial features locally. Intra-tile farthest point sampling and cross-tile neighbor search are employed to avoid unnecessary distance computing. To efficiently facilitate the tile operations, TiPU architecture consists of a tile-based unified distance computing unit, a near-CAM feature extractor, and a near-SRAM-computing MLP engine. The experimental results show that, compared to GPU implementation, TiPU achieves 15.7× processing speed and reduces 7308× energy consumption. Jiapei Zheng, Hao Jiang 0024, Xinkai Nie, Zhangcheng Huang 0001, Chixiao Chen, Qi Liu 0010 |
DAC | 6 |
| 2023 | A $2.53 \mu \mathrm{W}/\text{channel}$ Event-Driven Neural Spike Sorting Processor with Sparsity-Aware Computing-In-Memory MacrosabstractSpike sorting processors with high energy efficiency are widely used in large-scale neural signal processing tasks to monitor the activity of neurons in brains. This paper presents a low-power processor for high-accuracy spike sorting and on-chip incremental learning using an algorithm-hardware co-design approach. The processor introduces an event-driven mechanism with adaptive-threshold detection to conditionally activate the system in order to reduce power consumption. Sparsity-aware computing-in-memory (CIM) macros are also developed in our design to store templates and perform complicated computations efficiently. The prototype is designed using 28nm technology with an area of 0.018 mm2/channel and an overall power efficiency of$\mathbf{2.53} \mu \mathbf{W}/\mathbf{channel}$and 84nW/(channel.cluster) at the voltage of 0.72V. Moreover, the accuracy of the whole design can reach 94.5% in a 32-channel scenario. Hao Jiang 0024, Jiapei Zheng, Yunzhengmao Wang, Jinshan Zhang 0006, Haozhe Zhu, Liangjian Lyu, Yingping Chen, Chixiao Chen, Qi Liu 0010 |
ISCAS | 9 |
| 2023 | A Scalable Die-to-Die Interconnect with Replay and Repair Schemes for 2.5D/3D IntegrationabstractChiplet is a critical technology in the post-Moore era, and the die-to-die (D2D) interconnect is essential for communication between chiplets. Meanwhile, several edge-computing devices based on 2.5D/3D chiplet have recently emerged. However, a lightweight D2D interconnect for 2.5D/3D edge-computing systems is lacking. Given the differences between 2.5D/3D integration, a scalable D2D interconnect with replay and repair schemes is presented in this paper. A credit-based flow control scheme and a custom replay scheme are presented for high efficiency. An effective detection and repair scheme is proposed to enhance fault tolerance for the D2D interconnect. Compared with a previous D2D interconnect design, the proposed D2D interconnect delivers 1.07/1.09Gbps throughput ($\sim 2.4\times/\sim 3.9\times \text{for}\ \text{write}/\text{read}$) and significantly reduced energy/bit with only ∼1.7× increased hardware cost. Additionally, compared with a previous chip-to-chip interconnect design, the proposed D2D interconnect can be configured down to power consumption as low as 0.55pJ/bit and 38.40Gbps throughput, achieving ∼2.5× throughput and significantly reduced latency with a negligible increase in hardware cost. Bo Jiao 0003, Jinshan Zhang 0006, Shiwei Liu 0002, Hao Jiang 0024, Jun Tao 0001, Wenning Jiang, Qi Liu 0010, Lihua Zhang 0002, Haozhe Zhu, Chixiao Chen |
ISCAS | 8 |
| 2022 | A 11.6μ W Computing-on-Memory-Boundary Keyword Spotting Processor with Joint MFCC-CNN Ternary QuantizationabstractThis paper presents an ultra-low-power keyword spotting processor using an algorithm-architecture co-design approach. Joint MFCC-CNN ternary weight quantization is proposed to reduce power consumption. The Mel filter and the DCT module are merged into one matrix multiplication. The merged coefficients and the weights of the rest NN classifier are ternary-quantized, causing less than 3% accuracy loss but 39 × energy efficiency improvement. Moreover, a Computing-on-Memory-Boundary macro is adopted to store the quantized coefficients and weights, and perform matrix multiplications. Compared to the existing computing-in-memory technology, the proposed technique can reduce power consumption due to the higher utilization ratio. To verify the proposed techniques, a keyword spotting processor prototype is designed with 28nm CMOS technology. Simulation results show that the prototype achieves power consumption of 11.6μ W under a power supply of 0.72V and a clock frequency of 250KHz. Xinru Jia, Haozhe Zhu, Yunzheng Wang, Jinshan Zhang 0006, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Qi Liu 0010 |
ISCAS | 9 |
| 2022 | A neuromorphic core based on threshold switching memristor with asynchronous address event representation circuits
Jinsong Wei, Xumeng Zhang, Zuheng Wu, Mansun Chan, Qi Liu 0010, Hong Chen 0002 |
Sci. China Inf. Sci. | 9 |
| 2022 | High-Speed and Time-Interleaved ADCs Using Additive-Neural-Network-Based Calibration for Nonlinear Amplitude and Phase DistortionabstractThis paper presents a neural network-based digital calibration algorithm for high-speed and time-interleaved (TI) ADCs. In contrast with prior methods, the proposed work features joint amplitude-dependent and phase-dependent nonlinear distortion correction without prior-knowledge of ADC architecture feature. A dynamic calibration is first used to compensate for phase-dependent distortion. Two training optimizations, including a sub-range-sample-based batch schemes and a recursive foreground co-calibration flow are proposed to reduce the error and overfitting and further save hardware resources. A practical calibration engine is also investigated for interleaved ADCs with distributed weight and shared weight methods. To demonstrate the effectiveness of the method, the calibration engine is verified by two fabricated ADC prototypes, a 5 GS/s 16-way interleaved ADC and a 625 MS/s interleaving-SAR assisted pipeline ADC. Measurement results show that SFDR is improved between 16.9dB and 36.4dB before and after calibration for different frequency inputs. To trade-off between accuracy and power consumption, a quantized and pruned engine is implemented on both FPGA and 28nm CMOS technology. Experimental results show that the dedicated calibration on silicon consumes 8.64mW with 0.9V power supply at 333MHz clock rate. Measurement results show that the quantized hardware implementation has only 0.4-4 dB loss in SFDR. Danfeng Zhai, Wenning Jiang, Xinru Jia, Jingchao Lan, Mingqiang Guo, Sai-Weng Sin, Fan Ye 0001, Qi Liu 0010, Junyan Ren, Chixiao Chen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2022 | Conditional Uncorrelation and Efficient Subset Selection in Sparse RegressionabstractGiven$m~d$-dimensional responsors and$n~d$-dimensional predictors, sparse regression finds at most$k$predictors for each responsor for linear approximation,$1\leq k \leq d-1$. The key problem in sparse regression is subset selection, which usually suffers from high computational cost. In recent years, many improved approximate methods of subset selection have been published. However, less attention has been paid to the nonapproximate method of subset selection, which is very necessary for many questions in data analysis. Here, we consider sparse regression from the view of correlation and propose the formula of conditional uncorrelation. Then, an efficient nonapproximate method of subset selection is proposed in which we do not need to calculate any coefficients in the regression equation for candidate predictors. By the proposed method, the computational complexity is reduced from$O([{1}/{6}]{k^{3}}\!+(m+1)k^{2}\!+\!mkd)$to$O([{1}/{6}]{k^{3}}\!+[{1}/{2}](m+1)k^{2})$for each candidate subset in sparse regression. Because the dimension$d$is generally the number of observations or experiments and large enough, the proposed method can greatly improve the efficiency of nonapproximate subset selection. We also apply the proposed method in real scenarios of dental age assessment and sparse coding to validate the efficiency of the proposed method. Jianji Wang 0001, Qi Liu 0010, Shaoyi Du, Yu-Cheng Guo, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Cybern. | 3 |
| 2021 | Sparsity-Aware Clamping Readout Scheme for High Parallelism and Low Power Nonvolatile Computing-in-Memory Based on Resistive MemoryabstractThe input parallelism of resistive memory (RRAM) based nonvolatile computing-in-memory (nvCIM) structure is limited by the signal margin as well as the readout precision. In this work, we propose a sparsity-aware clamping (SAC) scheme and its circuit implementation for nvCIM by co-design of circuit and algorithm. It can adaptively tune the quantized range and resolution of the readout circuit according to the degree of sparsity in neural network models. As a result, the SAC scheme can effectively increase the input parallelism of nvCIMs without incurring degradation on the signal margin or increasing the hardware cost for analogue readout. A case study on processing a multi-layer perceptron (MLP) model with the proposed nvCIM structure shows that the SAC scheme can improve the throughput by 2 times and increase the energy efficiency by 25.35% with negligible inference accuracy loss. Linfang Wang, Wang Ye, Junjie An, Chunmeng Dou, Qi Liu 0010, Meng-Fan Chang, Ming Liu 0022 |
ISCAS | 5 |
| 2021 | Recent progress of integrated circuits and optoelectronic chips
Yue Hao 0001, Genquan Han, Jincheng Zhang 0001, Xiaohua Ma 0001, Zhangming Zhu, Yanan Han, Ling Yang 0003, Jiangyi Shi, Wei Zhang 0343, Biao Pan, Yangqi Huang, Qi Liu 0010, Yimao Cai, Xin Ou, Tiangui You, Huaqiang Wu, Bin Gao 0006, Guoping Guo, Yonghua Chen, Xiangfei Chen, Chunlai Xue, Lixia Zhao, Xihua Zou, Lianshan Yan |
Sci. China Inf. Sci. | 20 |
| 2018 | Flexible cation-based threshold selector for resistive switching memory integration
Xiangheng Xiao, Congyan Lu, Facai Wu, Rongrong Cao, Changzhong Jiang, Qi Liu 0010 |
Sci. China Inf. Sci. | 8 |
| 2010 | Formation and annihilation of Cu conductive filament in the nonpolar resistive switching Cu/ZrO2: Cu/Pt ReRAMabstractWe report a ZrO2-based resistive memory composed of a thin Cu doped ZrO2layer sandwiched between Pt bottom and Cu top electrode. The Cu/ZrO2:Cu/Pt shows excellent nonpolar resistive switching behaviors, such as free-electroforming, high ON/OFF resistance ratio (106), fast Set/Reset speed (50 ns/100 ns), and reliable data retention (>10 years). The temperature-dependent switching characteristics show that a metallic filamentary channel is responsible for the low resistance state. Further analysis reveals that the physical origin of this metallic filament is the nanoscale Cu conductive filament. On this basis, we propose that the set process and the reset process stem from the electrochemical reactions in the filament, in which a thermal effect is greatly involved. Ming Liu 0022, Qi Liu 0010, Shibing Long, Weihua Guan |
ISCAS | 2 |