Anqin Xiao

dblp:380/5546 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0000-5390-4335ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Self-Supervised Neuromorphic Processor Using High-Dimensional Representations for Cognitive Map Navigation
abstract
This work proposes a self-supervised neuromorphic processor using high-dimensional representations for cognitive map navigation. By employing the Cognitive Map Learner (CML), it enables agents to explore and understand diverse environments online through random walks. To enhance path planning, the agent’s actions and observations are embedded into high-dimensional state spaces. This embedding creates a sense of direction, simplifying navigation into a retrieval process within an Associative Memory (AM). We design an energy-efficient processor that features a scalable multi-core hardware architecture with precision flexibility, combined with an on-chip random walk training engine. To balance the precision of the model with hardware overhead, two hardware-software co-design strategies are proposed. The first is a Content-Addressable Memory (CAM)-based approach for AM access, which reduces the number of memory access by up to 25%. The second involves high-dimensional matrix sparsity optimizations, reducing computation operations to less than 8%. We simulate this processor by a 40-nm CMOS technology, which has 2.88 mm2core area with 15.8 mW power at a frequency of 140 MHz. Compared to previous processors, our experiments show that the proposed processor achieves outstanding success rates of 99.9%, 96%, and 98.7% on 100 2D nodes, 125 3D nodes, and 25 abstract map nodes with obstacles, respectively. In terms of energy efficiency, it delivers a path planning result of 28 nJ/node and 35 nJ/node in 2D and 3D maps, offering a 1.2x to 2.9x improvement over the state-of-the-art.
Anqin Xiao, Luyu Yang, Yuhan He, Hengtan Zhang, Ziyi Yang 0014, Lirong Zheng 0001, Zhuo Zou
DATE1
2026 CAMPRO: A CAM-Based Processing-in-Memory Processor for Hyperdimensional Computing
abstract
This work introduces CAMPRO, a Content Addressable Memory (CAM)-based Processing-In-Memory (PIM) processor customized for Hyperdimensional Computing (HDC). CAMPRO leverages a 6T Split Word Lines (SWL) cell structure for its CAM, enabling efficient column-wise search for ultra-wide Hypervector (HV) storage and an optimized associative PIM architecture tailored to HDC operations, significantly enhancing energy efficiency. The four key operators of HDC, binding, bundling, permutation, and similarity, are mapped to the proposed architecture. CAMPRO enhances operational parallelism via approximate bundling and employs a hierarchical permutation method to mitigate the gap in flexible shift support within CAM-based PIM architecture. The fine grained pipelined operations boost processing efficiency and dynamically reclaim memory space to support larger models. The Two-Phase Bit Pruning (TPBP) strategy prunes redundant bits in class HVs across two computing stages to eliminate unnecessary computations, reducing operation counts by 73.6% and energy consumption by 65.2% while maintaining query precision. Simulated in a 22 nm CMOS process, CAMPRO occupies 1.13 mm2and consumes 0.99 mW at 200 MHz. CAMPRO demonstrates robust versatility and scalability across five datasets, including MUTAG, CIFAR10, MNIST, language classification, and EMG gesture recognition, using diverse encoding schemes. It achieves excellent energy efficiency and low latency from small to large-scale datasets. In language classification, it reduces inference energy by 99.2% compared to the similiar work. For EMG gesture recognition, it improves training and inference energy efficiency by 2.6x and 6.7x, and reduce inference latency by 73x compared to related works. On MNIST, it enhances energy efficiency by 11.3x and latency by 1.6x to the prior work, making it an efficient Artificial Intelligence of Things (AIoT) solution.
Yuhan He, Tianxi Hu, Anqin Xiao, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.3
2026 An Always-On Event-Triggered DVS Fall Detection Processor With Precision-Adaptive Inference in 40-nm CMOS
Ziyi Yang 0014, Jinqiao Yang, Quanshu Yan, Anqin Xiao, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 CAP-HDC: A CAM-Based Processor for Hyperdimensional Computing
abstract
This paper presents CAP-HDC, a Content Addressable Memory (CAM)-based processor designed for Hyperdimensional Computing (HDC). CAP-HDC integrates the binding, bundling, permutation, and similarity operators of HDC into the in-memory associative processing framework that features high parallelism, thereby achieving low power consumption and latency. The CAM utilized in CAP-HDC is designed using the Split Word Lines (SWL) 6T bit cell, enabling column-wise searching across all rows simultaneously. The approximate bundling method is proposed to implement bundling in CAM with multiple Hypervectors (HVs) without sacrificing accuracy on the 5-Class Gesture dataset. The hierarchical permutation method is proposed to implement permutation with lower power consumption, achieving a reduction of 90.66% in power consumption compared to the direct circular shift method. CAP-HDC is simulated using the 22 nm CMOS process, occupying an area of 1.06 mm2 and consuming 1.08 mW at a clock frequency of 200 MHz with a 0.9 V power supply. Compared to previous works, CAP-HDC improves energy efficiency by 2.9x and latency by 2.4x on the MNIST dataset. For hand gesture prediction based on EMG signals, CAP-HDC achieves improvements of 3.1x in inference energy efficiency and 2.6x in encoding energy efficiency.
Yuhan He, Anqin Xiao, Tianxi Hu, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou
ISCAS2
2024 Spiking-HDC: A Spiking Neural Network Processor with HDC Classifier Enabling Transfer Learning
abstract
This work proposes Spiking-HDC, a spiking neural network (SNN) processing system with hyperdimensional computing (HDC) and its hardware design for domain transfer scenarios. The input data is firstly fed into a two-layer SNN, serving as a feature extractor. It is followed by a HDC classifier to process feature vectors using hypervectors in binary representation. Such a system leverages HDC’s capability of single-pass learning, which can be adopted to rapidly updating but highly similar tasks by fine-tuning the HDC classifier with limited labeled data. By our experiments, the proposed system demonstrates transfer learning accuracy of 94.76%, 87.12% and 94.37% with few-shot samples on N-MNIST, DVS-Gesture and MNIST datasets, respectively. To apply Spiking-HDC model to extreme edge inference tasks, a dedicated processor is designed and implemented. The simulated results in 40 nm CMOS process illustrate that it has 0.88 mm2core area and 1.8 mW power at 100 MHz frequency. In comparison to similar works, it achieves 3.8×-36× inference energy efficiency enhancement.
Anqin Xiao, Jinqiao Yang, Lirong Zheng 0001, Zhuo Zou
ISCAS1
2024 CorTile: A Scalable Neuromorphic Processing Core for Cortical Simulation With Hybrid-Mode Router and TCAM
abstract
In neuromorphic processors, simulating large-scale Spiking Neural Networks (SNNs) for cortical models necessitates a significant increase in communication traffic and memory capacity, due to the lack of exploiting the sparsity of connections. Therefore, this paper proposes CorTile, a scalable neuromorphic processing core designed for cortical simulation. We propose a hybrid-mode router that supports Remote Unicast and Local Broadcast (RULB) routing method, leveraging the high local connectivity and low distal connectivity observed in cortical models. This approach achieves reductions of 36.7% in average router load, 40.7% in peak load, 51.2% in average link traffic, 41.7% in peak traffic, respectively, compared to conventional routing methods. Additionally, the proposed Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture leads to 87.1% reduction in area and a 62.7% reduction in power consumption. These approaches effectively decrease communication traffic and mitigate the quadratic increase in memory requirements, achieving linear growth instead, thus achieving scalability. The proposed CorTile is simulated using UMC 40-nm CMOS process, occupying an area of 5.15 mm2, supporting a maximum of 8k neurons and 64M synapses. Evaluated using a typical macaque cortex model, it consumes 8.25 mW, with the router operating at 200 MHz and the other modules at 100 MHz. This design achieves an average router load of 12.33 Mpackets/s and peak link traffic of 21.16 MB/s. Thanks to the scalability of the proposed processing core that can be tiled into many-core processors, it paves the way for chiplets and multiple chip integration towards a brain-scale neuromorphic computing system.
Fanxi Yang, Yuhan He, Jinqiao Yang, Anqin Xiao, Lufei Fan, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.4