Siqi He

dblp:280/3748 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NS-FPS: Accelerating Farthest Point Sampling via Neighbor Search in Large-Scale Point Clouds
Jiapei Zheng, Shuan Yang, Siqi He, Qi Liu 0010, Chixiao Chen
ISCA3
2026 CIM-Pruner: A Dual-Mode Compute-In-Memory Macro for Efficient VLMs with Intra-Chunk Token Pruning and Merging
Zhuojun Han, Siqi He, Chixiao Chen, Haozhe Zhu
ISCAS2
2026 A 1024-Ch 583-nW/Ch Spike-Sorting SoC With Sparsity-Aware Spike Detection Scratchpad and Ultra-Low-Leakage Dual-Voltage 5T-SRAM for 16K-Template Clustering
abstract
This paper presents an energy-efficient spike-sorting system-on-chip (SoC) designed for closed-loop brain-computer interfaces of massive probing channels. The design first incorporates a sparsity/similarity-aware spike detection scratchpad, leveraging a bit-wise differential encoder and zero-friendly read-out circuits, reducing the dynamic power consumption of spike detection by 77.7%. To mitigate static power dissipation, it also introduces an ultra-low-leakage dual-voltage 5T-SRAM array with level-shifter embedded sense amplifiers, achieving an 82.2% leakage power reduction of neural signal buffering by applying half$V_{DD}$on SRAM cells. Additionally, a memory hierarchy architecture combining on-chip SRAM and off-chip FeRAM, along with a firing-rate-based Osort for cluster template management, minimizes off-chip memory access to only 9.7% with a latency of$11.7\mu $s for 1024-channel spike sorting. A silicon prototype is fabricated in 28-nm CMOS technology, which achieves a power consumption of 583nW/channel and an area consumption of 0.0012mm2/channel. The chip supports real-time spike sorting with up to 16K templates,$21.3\times $greater than the state-of-the-art spike-sorting processor.
Hao Jiang 0024, Zexing Chen, Jiajun Lu, Siqi He, Liangjian Lyu, Jiamin Xu, Shiwei Liu 0002, Yingping Chen, Chixiao Chen, Qi Liu 0010, Ming Liu 0022
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 COPIC: A Codesigned Accelerator for Efficient Octree-Based Learned Point Cloud Compression at Edge
abstract
With the growing adoption of light detection and ranging (LiDAR) technology and the increasing resolution of captured data, efficient point cloud compression (PCC) has become a critical challenge. While PCC is essential for reducing bandwidth and storage demands, existing octree-based learned PCC methods face two fundamental bottlenecks: 1) octree construction and serialization that involve irregular, pointwise memory accesses that are difficult to parallelize and 2) autoregressive entropy model inference that requires heavy computation. These issues lead to high latency and energy consumption, making real-time deployment on edge devices impractical. This article introduces COPIC, a hardware accelerator designed for real-time PCC on edge devices. COPIC integrates two key components. The first is OctCAM, a content-addressable-memory (CAM)-based module for rapid octree construction and serialization. By enabling hardware-level parallel search, OctCAM significantly reduces the latency and energy consumption caused by irregular memory access patterns in octree processing. The second is PCC-Engine, a dedicated neural network inference module. Combined with a lightweight PCC-TinyAttention network and a window-based key-value caching approach, it further minimizes execution latency and power usage. Experimental results show that COPIC achieves a$25.9\times $speedup compared to GPU implementations, with a total processing latency of 40.1 ms and an average energy consumption of 0.23 J per frame, while delivering a compression rate 12.4% higher than standard proposals.
Jiapei Zheng, Yutong Su, Siqi He, Haozhe Zhu, Qi Liu 0010, Chixiao Chen
IEEE Trans. Very Large Scale Integr. Syst.3
2025 Lexical Diversity-aware Relevance Assessment for Retrieval-Augmented Generation
abstract
Zhange Zhang, Yuqing Ma, Yulong Wang, Shan He, Tianbo Wang, Siqi He, Jiakai Wang, Xianglong Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhange Zhang, Yuqing Ma, Siqi He, Jiakai Wang, Xianglong Liu 0001
ACL (1)6
2025 Hydra: Harnessing Expert Popularity for Efficient Mixture-of-Expert Inference on Chiplet System
abstract
The rapid growth of model sizes in advanced artificial intelligence algorithms, particularly in Transformerbased large language models (LLMs), has led to significant computational overhead. Mixture-of-Expert (MoE) models offer a solution through their sparsely gating mechanism but introduce new challenges of extensive all-to-all communication and model computational inefficiencies. This paper presents Hydra, a software/hardware co-design aimed at accelerating MoE inference on chiplet-based architectures. In software, Hydra employs a popularity-aware expert mapping strategy to optimize interchiplet communication. In hardware, it incorporates Content Addressable Memory (CAM) to eliminate expensive explicit token (un)-permutation based on sparse matrix multiplications and a redundant-calculation-skipping softmax engine to bypass unnecessary division and exponential operations. Evaluated in 22 nm technology, Hydra achieves latency reductions of $14.2 \times$ and $3.5 \times$ and power reductions of $169.1 \times$ and $18.9 \times$ over GPU and state-of-the-art MoE accelerator, respectively, thereby offering a scalable and efficient solution for MoE model deployment.
Siqi He, Haozhe Zhu, Jiapei Zheng, Lizhou Wu, Bo Jiao 0003, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen
DAC1
2025 PIMoE: Towards Efficient MoE Transformer Deployment on NPU-PIM System through Throttle-Aware Task Offloading
abstract
Mixture-of-experts (MoE) technique holds significant promise for scaling up Transformer models. However, the data transfer overhead and imbalanced workload hinder efficient deployment. This work presents PIMoE, a heterogeneous system combining processing-in-memory (PIM) and neural-processing-unit (NPU) to facilitate efficient MoE Transformer inference. We propose a throttle-aware task offloading method that addresses workload imbalance between NPU and PIM, achieving optimal task distribution. Furthermore, we design a near-memory-controller data condenser to address the mismatch of sparse data layout between NPU and PIM, enhancing data transfer efficiency. Experimental results demonstrate that PIMoE achieves $4.5 \times$ speedup and $13.7 \times$ greater energy efficiency compared to the A 100, and $1.4 \times$ speedup over a state-of-the-art MoE platform.
Lizhou Wu, Haozhe Zhu, Siqi He, Xuanda Lin, Xiaoyang Zeng, Chixiao Chen
DAC3
2025 A 0.22 pJ/bit Processing-in-Controller GEMV Macro with Weight Prefetch for Efficient Near-Memory Computing
abstract
3D-stacked DRAM is a key technology enabling the development of large language models (LLMs). However, the intensive computational demands lead to substantial energy consumption resulting from large-scale data movement. Integrating processing-in-memory (PIM) within DRAM has been employed to reduce data movement, but it elevates manufacturing costs and incurs additional area and power consumption. On the other hand, implementing processing-near-memory (PNM) outside the DRAM controller fails to eliminate the high energy consumption associated with interconnects during data transfer. To address these challenges, we propose a processing-in-controller (PIC) architecture aimed at 3D-stacked DRAM for efficient data movement. A size-scalable General Matrix-Vector Multiplication (GEMV) structure is proposed, supporting configurations ranging from 16 × 16 to 128 × 128, thereby enhancing its adaptability for a variety of applications. To address the low throughput bottleneck caused by DDR read latency, a ping-pong buffer with weight prefetching is introduced in the PNM. Based on 28 nm CMOS technology, experimental results demonstrate that the energy consumption of this PIC is 0.22 pJ/bit, while data movement efficiency improves by nearly 92% compared to traditional DRAM operations.
Jiapei Zheng, Siqi He, Lizhou Wu, Chen Mu, Haozhe Zhu, Liyu Lin, Qi Liu 0010, Chixiao Chen
ISCAS3
2025 RT-FLOW: FPGA Implementation of Real-Time Optical-Flow-Based SLAM for High-Speed Tracking and High-Quality Mapping
abstract
Simultaneous Localization and Mapping (SLAM) is pivotal for autonomous robotics, yet feature-based SLAM systems struggle with sparse environmental representations and robustness under dynamic conditions. Optical-flow-based SLAM (OpF-SLAM) addresses these limitations by leveraging pixel-level motion data for dense mapping; however, its computational intensity hinders real-time deployment. This paper presents RT-FLOW, an FPGA-based accelerator for OpF-SLAM that achieves real-time performance through three key innovations: 1) A feature-context encoding engine that exploits inter-frame similarity to resolve data dependency in correlation construction, reducing latency by 77.5%. 2) A heterogeneous mixed-precision flow update engine guided by correlation sparsity, enabling 3.7× faster optical flow computation with negligible accuracy loss. 3) A pivoting-free linear solver using Householder transformations for stable pose optimization. Implemented on Xilinx XCZU7EV FPGA, RT-FLOW processes full-image pixels per frame at 65 fps with an energy efficiency of 0.358 μJ/point, outperforming previous FPGA designs. Evaluated on benchmark datasets, RT-FLOW demonstrates robustness in diverse environments while maintaining sub-110mJ/frame energy consumption. This work bridges the gap between algorithmic potential and hardware feasibility for high-density SLAM, empowering next-generation mobile robots with real-time scene understanding capabilities.
Siqi He, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen, Haozhe Zhu
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 ARCTIC: Agile and Robust Compute-In-Memory Compiler with Parameterized INT/FP Precision and Built-In Self Test
abstract
Digital Compute-in-Memory (DCIM) architectures are playing an increasingly vital role in artificial intelligence (AI) applications due to their significant energy efficiency enhancement. Coupling memory and computing logic in DCIM requires extensive customization of custom cells and layouts, thus increasing design complexity and implementation effort. To adapt to the swiftly evolving AI algorithms, DCIM compiler for agile customization is required. Previous DCIM compilers accelerate the customization process but only focus on integer computation. Moreover, with technology node scaling down, design-for-test circuits are critical for robust chip design, while previous built-in-self-test (BIST) schemes for traditional memory fail to offer support for DCIM. This paper presents ARCTIC, an agile and robust DCIM compiler supporting parameterized integer/floating-point formats with corresponding BIST circuits. To support variable precision formats (including integer and floating-point), ARCTIC applies adaptive topology and layout optimization schemes for optimal performance. The compiler is also equipped with DCIM-friendly MarchCIM BIST circuits for efficient post-silicon tests with negligible area overhead. The energy efficiency of the generated DCIM macros remains competent with the state-of-the-art counterparts.
Haozhe Zhu, Siqi He, Chengchen Wang, Xiankui Xiong, Haidong Tian, Xiaoyang Zeng, Chixiao Chen
DATE3
2024 A 19.7 TFLOPS/W Multiply-less Logarithmic Floating-Point CIM Architecture with Error-Reduced Compensated Approximate Adder
abstract
The growing demand for high-precision neural network training and inference has driven the necessity for floating-point (FP) compute-in-memory (CIM) architectures. However, compared to the extensively studied INT-CIM, the energy efficiency of FP-CIM still requires further optimization and enhancement. This work presents an energy-efficient multiply-less digital SRAM-based FP-CIM architecture. Specifically, to improve the energy efficiency and minimize the area requirement, we propose to employ logarithmic approximate FP multiplication (LAM) within the FP-CIM architecture. The LAM approximates FP multiplication by converting it into a straightforward addition operation, thereby reducing the power consumption and area. Additionally, we propose an approximate adder with error-reduced compensation to address critical path delay issues associated with carry propagation, further minimizing power consumption and area overhead. A 24Kb SRAM CIM macro with the proposed techniques is designed in a 28nm CMOS technology and occupies an area of 0.033 mm2. The simulation results show that our work achieves an energy efficiency of 19.7 TFLOPS/W with bfloat16 representation at 0.9V and 200MHz.
Siqi He, Haozhe Zhu, Jinglei Liu, Zhenping Hu, Xiaoyang Zeng, Chixiao Chen
ISCAS3
2024 GauSPU: 3D Gaussian Splatting Processor for Real-Time SLAM Systems
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising technique in the realms of 3D vision and robotics. Its capacity for rapid rendering and high-fidelity reconstruction makes it an attractive candidate for integration into Simultaneous Localization and Mapping (SLAM) systems. However, existing 3DGS-based SLAM systems still suffer from inadequate tracking throughput due to tremendous recursion in volume rendering and irregular memory access for gradient backpropagation. To address these challenges, this paper proposes GauSPU, an algorithm-hardware co-designed accelerator for supporting real-time 3DGS-based SLAM. On the algorithm side, we present a sparse-tile-sampling (STS) method for efficient pose tracking. The STS focuses on informative image regions, discarding the rest to alleviate computational workload while maintaining accuracy. At the hardware level, we make twofold efforts. Firstly, we design a sparsity-adaptive ray recursion unit (SA-RRU) to accelerate volume rendering by leveraging irregular spatial sparsity. The SA-RRU introduces a sub-tile-wise execution pattern and a Morton-based thread allocation scheme to optimize sparsity utilization. Additionally, a sparsity-aware task dispatcher ensures efficient fine-grained task scheduling. Secondly, we propose a memory-access-relaxed backpropagation engine (MAR-BE) for efficient gradient aggregation. It comprises a gradient buffer unit (GBU) for coalescing partial gradients and a pose backward unit (PBU) for pipeline-fused backpropagation, collaboratively eliminating the costly atomic operations. Sufficient experiments demonstrate that, through the integration of GauSPU and GPU, the system achieves a throughput of 33.6 FPS for real-time pose tracking in 3DGS-SLAM, presenting a significant$63.9\times$improvement in energy efficiency compared to the RTX3090 baseline.
Lizhou Wu, Haozhe Zhu, Siqi He, Jiapei Zheng, Chixiao Chen, Xiaoyang Zeng
MICRO3
2024 TDS-NA: Blockchain-based trusted data sharing scheme with PKI authentication
Zhenshen Ou, Xiaofei Xing, Siqi He, Guojun Wang 0001
Comput. Commun.3
2023 Flower Image Identification and Feature Extracting Method Based on Transfer Learning
abstract
Artificial intelligence (AI) technology is booming in information society, and its role has become one of the hot research topics in the field of image identification. Flowers play essential role in daily life, but the identification and know well of some flower species is not clear. Therefore, the system is of great significance to help people obtain the detailed information of plant flowers. The proposed method is to use machine learning algorithm to process and identify flower images, mainly using transfer learning model, based on the trained model parameters, and taking 25 reptiles as the training base. After the optimal training and parameters adjustment, the extracted features were transferred to the flower model. The experimental results show that the effective recognition accuracy of different flowers is up to 98.20%. The proposed algorithm can be adopted to the small-scale training task, and can reduces the amount of machine operation. In addition, the algorithm also can obtain satisfactory recognition accuracy within a short time. The flower information is presented in the form of dynamic web pages, which improved readability of idenfication results.
Xiaofei Xing, Yunxuan Zeng, Siqi He
CSCWD5
2022 DIV-SC: A Data Integrity Verification Scheme for Centralized Database Using Smart Contract
abstract
When faced with massive data volumes, many current companies and institutions usually choose centralized databases or distributed databases to meet their data storage needs. But untrusted centralized third-party auditor can pose serious security problems. Malicious database service providers may tamper with or delete users data to achieve certain benefits, while returning false data integrity verification results to users. The traditional solution is to introduce a third-party auditor to ensure the reliability of the data verification results, but this third-party auditor may also be untrustworthy and partner with the database service provider to forge false data verification results. The centralization of the database system makes the verification of data integrity a difficult but necessary task. Therefore, we propose a data integrity verification scheme using smart contract (DIV-SC) to ensure the reliability of data integrity verification results in a centralized database environment. We introduce the blockchain as a decentralized third-party auditor. The immutability of the blockchain can ensure that the information stored on the blockchain will not be maliciously tampered with. Meanwhile, the smart contract deployed on the blockchain can ensure that the procedure of storing verification information and the verification procedure are correct and will not be affected by any malicious parties.
Siqi He, Xiaofei Xing
TrustCom1