EDBT 2026 Demo / reviewers in the wild / expert
Fanxi Yang
dblp:295/6714
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance-Aware Design Space Exploration for Chiplet-Based Cortical Simulation Processors Considering Area Constraint
Fanxi Yang, Lufei Fan, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 1 |
| 2026 | CAMPRO: A CAM-Based Processing-in-Memory Processor for Hyperdimensional ComputingabstractThis work introduces CAMPRO, a Content Addressable Memory (CAM)-based Processing-In-Memory (PIM) processor customized for Hyperdimensional Computing (HDC). CAMPRO leverages a 6T Split Word Lines (SWL) cell structure for its CAM, enabling efficient column-wise search for ultra-wide Hypervector (HV) storage and an optimized associative PIM architecture tailored to HDC operations, significantly enhancing energy efficiency. The four key operators of HDC, binding, bundling, permutation, and similarity, are mapped to the proposed architecture. CAMPRO enhances operational parallelism via approximate bundling and employs a hierarchical permutation method to mitigate the gap in flexible shift support within CAM-based PIM architecture. The fine grained pipelined operations boost processing efficiency and dynamically reclaim memory space to support larger models. The Two-Phase Bit Pruning (TPBP) strategy prunes redundant bits in class HVs across two computing stages to eliminate unnecessary computations, reducing operation counts by 73.6% and energy consumption by 65.2% while maintaining query precision. Simulated in a 22 nm CMOS process, CAMPRO occupies 1.13 mm2and consumes 0.99 mW at 200 MHz. CAMPRO demonstrates robust versatility and scalability across five datasets, including MUTAG, CIFAR10, MNIST, language classification, and EMG gesture recognition, using diverse encoding schemes. It achieves excellent energy efficiency and low latency from small to large-scale datasets. In language classification, it reduces inference energy by 99.2% compared to the similiar work. For EMG gesture recognition, it improves training and inference energy efficiency by 2.6x and 6.7x, and reduce inference latency by 73x compared to related works. On MNIST, it enhances energy efficiency by 11.3x and latency by 1.6x to the prior work, making it an efficient Artificial Intelligence of Things (AIoT) solution. Yuhan He, Tianxi Hu, Anqin Xiao, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | PHENICS: A Scalable Neuromorphic FPGA Architecture for Million-Neuron Cortical Simulation With 4.6× Real-Time AccelerationabstractTraditional brain simulations using CPU/GPU architectures face critical limitations in both speed and scalability, particularly when modeling large-scale biological neural networks. Cortical models exhibit distinctive features, recurrent, random and sparse connectivity, and ultra-high synaptic density that fundamentally differ from the structured and dense architectures of deep neural networks (DNNs). As a result, conventional accelerators suffer from inefficiencies: communication overhead scales linearly with neuron count, and memory architectures exhibit poor utilization efficiency under complex connectivity, severely constraining scalability. To address these challenges, we presentPHENICS(Pyramidal Hierarchical Event-driven Neuromorphic Infrastructure for Cortical Simulation), a scalable FPGA-based architecture tailored for large-scale spiking neural network simulations. At the communication level, PHENICS introduces a pyramidal multi-tier on-chip network combined with aBusy-Aware Threshold Adaptation (BATA)routing strategy and a lightweight router design to mitigate network congestion and improve spike transmission efficiency. At the storage level, we employ a multi-level addressing scheme adapted to sparse, irregular synaptic connectivity and leverage high-bandwidth memory (HBM) for efficient large-scale synaptic access. PHENICS is implemented on the Xilinx Alveo U50 system, successfully simulating a one-million-neuron Leaky Integrate-and-Fire (LIF) cortical model with 4 billion synapses, achieving a$4.6\times $real-time acceleration. It achieves this with only 25% of the memory bandwidth available on GPU platforms, yet delivers a$5\times $speedup compared to GPU-based simulators. Furthermore, our architecture reduces synaptic storage overhead by$3\times $through compressed encoding and HBM-backed sparse addressing. Moreover, it exhibits sub-linear communication time growth with increasing neuron counts. These advancements pave the way for real-time simulation of billion-neuron brain-scale models, opening new frontiers in neuroscience and neuromorphic computing. Fanxi Yang, Lufei Fan, Yuhan He, Hanwen Ou, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | CAP-HDC: A CAM-Based Processor for Hyperdimensional ComputingabstractThis paper presents CAP-HDC, a Content Addressable Memory (CAM)-based processor designed for Hyperdimensional Computing (HDC). CAP-HDC integrates the binding, bundling, permutation, and similarity operators of HDC into the in-memory associative processing framework that features high parallelism, thereby achieving low power consumption and latency. The CAM utilized in CAP-HDC is designed using the Split Word Lines (SWL) 6T bit cell, enabling column-wise searching across all rows simultaneously. The approximate bundling method is proposed to implement bundling in CAM with multiple Hypervectors (HVs) without sacrificing accuracy on the 5-Class Gesture dataset. The hierarchical permutation method is proposed to implement permutation with lower power consumption, achieving a reduction of 90.66% in power consumption compared to the direct circular shift method. CAP-HDC is simulated using the 22 nm CMOS process, occupying an area of 1.06 mm2 and consuming 1.08 mW at a clock frequency of 200 MHz with a 0.9 V power supply. Compared to previous works, CAP-HDC improves energy efficiency by 2.9x and latency by 2.4x on the MNIST dataset. For hand gesture prediction based on EMG signals, CAP-HDC achieves improvements of 3.1x in inference energy efficiency and 2.6x in encoding energy efficiency. Yuhan He, Anqin Xiao, Tianxi Hu, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 4 |
| 2025 | A Neuromorphic Controller with On-Chip Learning for Robot Motion ControlabstractMotion control is one of the most fundamental issues in robotics, with kinematics and dynamics serving as its core components. While most existing control systems rely on general-purpose processors with large areas and high power consumption. This paper proposes a neuromorphic controller with on-chip learning, satisfying the requirements of high control performance and low cost for robot motion control. The proposed controller consists of an Operational Space Control (OSC) unit and a Spiking Neural Networks (SNNs) processing unit, offering kinematic and dynamic motion control across different (4, 6, 7, and 9) Degrees of Freedom (DoF). Under external disturbances, its control precision and the convergence speed are enhanced by 2.83× and 1.78×, respectively, compared to standard proportional integrated-error derivative (PID) OSC controller. The controller is simulated under 40 nm CMOS technology, occupying a core area of 0.755 mm2and consuming 2.4 mW of power at a frequency of 100 MHz. Compared with other chips used for robot motion control, the proposed controller achieves 2.15× and 65× enhancements in core area and power consumption. Hengtan Zhang, Jinqiao Yang, Yuhan He, Fanxi Yang, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 5 |
| 2024 | TSCM: A TCAM-Based Sparse Connection Memory Architecture in Neuromorphic Computing System for Cortical SimulationabstractThe connection matrix requires significant memory capacity in large-scale Spiking Neural Networks (SNNs). However, the sparsity of connections in cortical models leads to memory capacity and energy inefficiencies. This paper proposes Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture in neuromorphic computing systems for cortical simulation. The architecture consists of a TCAM for searching existing synaptic connections, a Static Random Access Memory (SRAM) for storing synapse addresses, and a three-stage circuit for memory access control. By leveraging the sparsity, the proposed memory architecture demonstrates improved area and energy efficiency compared to the conventional Direct Mapped Full-Address Memory (DMFAM) architecture. A TSCM macro is designed, simulated using UMC 40-nm CMOS technology, and evaluated across various scales of classical cortical models. Experimental results demonstrate that the TSCM architecture reduces area by 27.6% to 75.6% and energy consumption by 15.8% to 96.0% in cortical simulation with neurons ranging from 100k to 10M compared to DMFAM architecture. Fanxi Yang, Yuhan He, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 1 |
| 2024 | 360° video quality assessment based on saliency-guided viewport extraction
Fanxi Yang, Chao Yang 0021, Ping An 0001, Xinpeng Huang |
Multim. Syst. | 1 |
| 2024 | CorTile: A Scalable Neuromorphic Processing Core for Cortical Simulation With Hybrid-Mode Router and TCAMabstractIn neuromorphic processors, simulating large-scale Spiking Neural Networks (SNNs) for cortical models necessitates a significant increase in communication traffic and memory capacity, due to the lack of exploiting the sparsity of connections. Therefore, this paper proposes CorTile, a scalable neuromorphic processing core designed for cortical simulation. We propose a hybrid-mode router that supports Remote Unicast and Local Broadcast (RULB) routing method, leveraging the high local connectivity and low distal connectivity observed in cortical models. This approach achieves reductions of 36.7% in average router load, 40.7% in peak load, 51.2% in average link traffic, 41.7% in peak traffic, respectively, compared to conventional routing methods. Additionally, the proposed Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture leads to 87.1% reduction in area and a 62.7% reduction in power consumption. These approaches effectively decrease communication traffic and mitigate the quadratic increase in memory requirements, achieving linear growth instead, thus achieving scalability. The proposed CorTile is simulated using UMC 40-nm CMOS process, occupying an area of 5.15 mm2, supporting a maximum of 8k neurons and 64M synapses. Evaluated using a typical macaque cortex model, it consumes 8.25 mW, with the router operating at 200 MHz and the other modules at 100 MHz. This design achieves an average router load of 12.33 Mpackets/s and peak link traffic of 21.16 MB/s. Thanks to the scalability of the proposed processing core that can be tiled into many-core processors, it paves the way for chiplets and multiple chip integration towards a brain-scale neuromorphic computing system. Fanxi Yang, Yuhan He, Jinqiao Yang, Anqin Xiao, Lufei Fan, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Unsupervised blind image quality assessment based on joint structure and natural scene statistics features
Qinglin He, Chao Yang 0021, Fanxi Yang, Ping An 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | A Hybrid-Mode On-Chip Router for the Large-Scale FPGA-Based Neuromorphic PlatformabstractLarge-scale neuromorphic computing requires the multi-chip network to provide high computing power. Efficient routing schemes and on-chip router design are necessary for handling various inter-chip transmission patterns. In this paper, we propose a hybrid-mode on-chip router that supports both multicast and unicast routing for the large-scale neuromorphic simulation. Two routing schemes, namely Cache-like Spike Weight Indexing and General Unicast Flow Control, are proposed to accommodate the chip-to-chip transmission of spike and non-spike data. This work is evaluated on a neuromorphic platform built with an$8\times 8$FPGA chips array. Running a simulation of 1M neurons at 200MHz, the proposed router achieves a processing latency of 25ns and a chip-to-chip latency of 287ns. Working in the unicast mode, the router can synchronize status flags of all chips within$5 ~\mu \text{s}$. Moreover, it reduces the peak spike traffic by 25.65% with the help of Load-aware Multicast Routing, compared with other multicast routing strategies. Chen Ding 0010, Yuxiang Huan, Yulong Yan, Fanxi Yang, Lizheng Liu, Meigen Shen, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |