VLDB 2026 Research / reviewers in the wild / expert
Yuhan He
dblp:94/7562
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Self-Supervised Neuromorphic Processor Using High-Dimensional Representations for Cognitive Map NavigationabstractThis work proposes a self-supervised neuromorphic processor using high-dimensional representations for cognitive map navigation. By employing the Cognitive Map Learner (CML), it enables agents to explore and understand diverse environments online through random walks. To enhance path planning, the agent’s actions and observations are embedded into high-dimensional state spaces. This embedding creates a sense of direction, simplifying navigation into a retrieval process within an Associative Memory (AM). We design an energy-efficient processor that features a scalable multi-core hardware architecture with precision flexibility, combined with an on-chip random walk training engine. To balance the precision of the model with hardware overhead, two hardware-software co-design strategies are proposed. The first is a Content-Addressable Memory (CAM)-based approach for AM access, which reduces the number of memory access by up to 25%. The second involves high-dimensional matrix sparsity optimizations, reducing computation operations to less than 8%. We simulate this processor by a 40-nm CMOS technology, which has 2.88 mm2core area with 15.8 mW power at a frequency of 140 MHz. Compared to previous processors, our experiments show that the proposed processor achieves outstanding success rates of 99.9%, 96%, and 98.7% on 100 2D nodes, 125 3D nodes, and 25 abstract map nodes with obstacles, respectively. In terms of energy efficiency, it delivers a path planning result of 28 nJ/node and 35 nJ/node in 2D and 3D maps, offering a 1.2x to 2.9x improvement over the state-of-the-art. Anqin Xiao, Luyu Yang, Yuhan He, Hengtan Zhang, Ziyi Yang 0014, Lirong Zheng 0001, Zhuo Zou |
DATE | 3 |
| 2026 | CIMS: A CAM-based In-Memory Sorting Architecture for Efficient Top-K Ranking
Yuhan He, Siheng Lei, Tianxi Hu, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 1 |
| 2026 | CAMPRO: A CAM-Based Processing-in-Memory Processor for Hyperdimensional ComputingabstractThis work introduces CAMPRO, a Content Addressable Memory (CAM)-based Processing-In-Memory (PIM) processor customized for Hyperdimensional Computing (HDC). CAMPRO leverages a 6T Split Word Lines (SWL) cell structure for its CAM, enabling efficient column-wise search for ultra-wide Hypervector (HV) storage and an optimized associative PIM architecture tailored to HDC operations, significantly enhancing energy efficiency. The four key operators of HDC, binding, bundling, permutation, and similarity, are mapped to the proposed architecture. CAMPRO enhances operational parallelism via approximate bundling and employs a hierarchical permutation method to mitigate the gap in flexible shift support within CAM-based PIM architecture. The fine grained pipelined operations boost processing efficiency and dynamically reclaim memory space to support larger models. The Two-Phase Bit Pruning (TPBP) strategy prunes redundant bits in class HVs across two computing stages to eliminate unnecessary computations, reducing operation counts by 73.6% and energy consumption by 65.2% while maintaining query precision. Simulated in a 22 nm CMOS process, CAMPRO occupies 1.13 mm2and consumes 0.99 mW at 200 MHz. CAMPRO demonstrates robust versatility and scalability across five datasets, including MUTAG, CIFAR10, MNIST, language classification, and EMG gesture recognition, using diverse encoding schemes. It achieves excellent energy efficiency and low latency from small to large-scale datasets. In language classification, it reduces inference energy by 99.2% compared to the similiar work. For EMG gesture recognition, it improves training and inference energy efficiency by 2.6x and 6.7x, and reduce inference latency by 73x compared to related works. On MNIST, it enhances energy efficiency by 11.3x and latency by 1.6x to the prior work, making it an efficient Artificial Intelligence of Things (AIoT) solution. Yuhan He, Tianxi Hu, Anqin Xiao, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2026 | PHENICS: A Scalable Neuromorphic FPGA Architecture for Million-Neuron Cortical Simulation With 4.6× Real-Time AccelerationabstractTraditional brain simulations using CPU/GPU architectures face critical limitations in both speed and scalability, particularly when modeling large-scale biological neural networks. Cortical models exhibit distinctive features, recurrent, random and sparse connectivity, and ultra-high synaptic density that fundamentally differ from the structured and dense architectures of deep neural networks (DNNs). As a result, conventional accelerators suffer from inefficiencies: communication overhead scales linearly with neuron count, and memory architectures exhibit poor utilization efficiency under complex connectivity, severely constraining scalability. To address these challenges, we presentPHENICS(Pyramidal Hierarchical Event-driven Neuromorphic Infrastructure for Cortical Simulation), a scalable FPGA-based architecture tailored for large-scale spiking neural network simulations. At the communication level, PHENICS introduces a pyramidal multi-tier on-chip network combined with aBusy-Aware Threshold Adaptation (BATA)routing strategy and a lightweight router design to mitigate network congestion and improve spike transmission efficiency. At the storage level, we employ a multi-level addressing scheme adapted to sparse, irregular synaptic connectivity and leverage high-bandwidth memory (HBM) for efficient large-scale synaptic access. PHENICS is implemented on the Xilinx Alveo U50 system, successfully simulating a one-million-neuron Leaky Integrate-and-Fire (LIF) cortical model with 4 billion synapses, achieving a$4.6\times $real-time acceleration. It achieves this with only 25% of the memory bandwidth available on GPU platforms, yet delivers a$5\times $speedup compared to GPU-based simulators. Furthermore, our architecture reduces synaptic storage overhead by$3\times $through compressed encoding and HBM-backed sparse addressing. Moreover, it exhibits sub-linear communication time growth with increasing neuron counts. These advancements pave the way for real-time simulation of billion-neuron brain-scale models, opening new frontiers in neuroscience and neuromorphic computing. Fanxi Yang, Lufei Fan, Yuhan He, Hanwen Ou, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | BMC-Net: A Framework for IDH Genotyping of Gliomas Based on Bi-Directional Mamba Sequences
Shuaidan Wang, Shun Zou, Yuhan He, Zhuo Zhang 0020 |
ICIC (28) | 3 |
| 2025 | CAP-HDC: A CAM-Based Processor for Hyperdimensional ComputingabstractThis paper presents CAP-HDC, a Content Addressable Memory (CAM)-based processor designed for Hyperdimensional Computing (HDC). CAP-HDC integrates the binding, bundling, permutation, and similarity operators of HDC into the in-memory associative processing framework that features high parallelism, thereby achieving low power consumption and latency. The CAM utilized in CAP-HDC is designed using the Split Word Lines (SWL) 6T bit cell, enabling column-wise searching across all rows simultaneously. The approximate bundling method is proposed to implement bundling in CAM with multiple Hypervectors (HVs) without sacrificing accuracy on the 5-Class Gesture dataset. The hierarchical permutation method is proposed to implement permutation with lower power consumption, achieving a reduction of 90.66% in power consumption compared to the direct circular shift method. CAP-HDC is simulated using the 22 nm CMOS process, occupying an area of 1.06 mm2 and consuming 1.08 mW at a clock frequency of 200 MHz with a 0.9 V power supply. Compared to previous works, CAP-HDC improves energy efficiency by 2.9x and latency by 2.4x on the MNIST dataset. For hand gesture prediction based on EMG signals, CAP-HDC achieves improvements of 3.1x in inference energy efficiency and 2.6x in encoding energy efficiency. Yuhan He, Anqin Xiao, Tianxi Hu, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 1 |
| 2025 | A Neuromorphic Controller with On-Chip Learning for Robot Motion ControlabstractMotion control is one of the most fundamental issues in robotics, with kinematics and dynamics serving as its core components. While most existing control systems rely on general-purpose processors with large areas and high power consumption. This paper proposes a neuromorphic controller with on-chip learning, satisfying the requirements of high control performance and low cost for robot motion control. The proposed controller consists of an Operational Space Control (OSC) unit and a Spiking Neural Networks (SNNs) processing unit, offering kinematic and dynamic motion control across different (4, 6, 7, and 9) Degrees of Freedom (DoF). Under external disturbances, its control precision and the convergence speed are enhanced by 2.83× and 1.78×, respectively, compared to standard proportional integrated-error derivative (PID) OSC controller. The controller is simulated under 40 nm CMOS technology, occupying a core area of 0.755 mm2and consuming 2.4 mW of power at a frequency of 100 MHz. Compared with other chips used for robot motion control, the proposed controller achieves 2.15× and 65× enhancements in core area and power consumption. Hengtan Zhang, Jinqiao Yang, Yuhan He, Fanxi Yang, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 4 |
| 2025 | Modeling detail feature connections for infrared image enhancement
Ziang Meng, Kaiqi Han, Yuchun He, Yuhan He, Xingyuan Li 0005, Yang Zou 0004 |
Neurocomputing | 4 |
| 2025 | Deep Learning Super-Resolution-Based Channel Completion for Massive MISO SystemsabstractWith the deployment of large-scale antenna arrays, the already limited time-frequency resources are becoming increasingly scarce. In this study, we propose a novel Laplacian Pyramid Channel Completion Network (LPCCNet) designed for channel completion, thereby reducing the demand for time-frequency resources in massive MIMO systems. Compared with existing network models, the proposed LPCCNet, by employing a progressive upsampling architecture, effectively mitigates aliasing effects, suppresses error propagation, and achieves a substantial reduction in computational complexity. The simulation results show that LPCCNet achieves a superior channel completion quality compared to existing methods, particularly in rapidly time-varying scenarios. Keke Zu, Yuhan He, Hongyang Chen 0001, Yu Zheng 0004, Martin Haardt |
IEEE Signal Process. Lett. | 2 |
| 2024 | TSCM: A TCAM-Based Sparse Connection Memory Architecture in Neuromorphic Computing System for Cortical SimulationabstractThe connection matrix requires significant memory capacity in large-scale Spiking Neural Networks (SNNs). However, the sparsity of connections in cortical models leads to memory capacity and energy inefficiencies. This paper proposes Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture in neuromorphic computing systems for cortical simulation. The architecture consists of a TCAM for searching existing synaptic connections, a Static Random Access Memory (SRAM) for storing synapse addresses, and a three-stage circuit for memory access control. By leveraging the sparsity, the proposed memory architecture demonstrates improved area and energy efficiency compared to the conventional Direct Mapped Full-Address Memory (DMFAM) architecture. A TSCM macro is designed, simulated using UMC 40-nm CMOS technology, and evaluated across various scales of classical cortical models. Experimental results demonstrate that the TSCM architecture reduces area by 27.6% to 75.6% and energy consumption by 15.8% to 96.0% in cortical simulation with neurons ranging from 100k to 10M compared to DMFAM architecture. Fanxi Yang, Yuhan He, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 2 |
| 2024 | CorTile: A Scalable Neuromorphic Processing Core for Cortical Simulation With Hybrid-Mode Router and TCAMabstractIn neuromorphic processors, simulating large-scale Spiking Neural Networks (SNNs) for cortical models necessitates a significant increase in communication traffic and memory capacity, due to the lack of exploiting the sparsity of connections. Therefore, this paper proposes CorTile, a scalable neuromorphic processing core designed for cortical simulation. We propose a hybrid-mode router that supports Remote Unicast and Local Broadcast (RULB) routing method, leveraging the high local connectivity and low distal connectivity observed in cortical models. This approach achieves reductions of 36.7% in average router load, 40.7% in peak load, 51.2% in average link traffic, 41.7% in peak traffic, respectively, compared to conventional routing methods. Additionally, the proposed Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture leads to 87.1% reduction in area and a 62.7% reduction in power consumption. These approaches effectively decrease communication traffic and mitigate the quadratic increase in memory requirements, achieving linear growth instead, thus achieving scalability. The proposed CorTile is simulated using UMC 40-nm CMOS process, occupying an area of 5.15 mm2, supporting a maximum of 8k neurons and 64M synapses. Evaluated using a typical macaque cortex model, it consumes 8.25 mW, with the router operating at 200 MHz and the other modules at 100 MHz. This design achieves an average router load of 12.33 Mpackets/s and peak link traffic of 21.16 MB/s. Thanks to the scalability of the proposed processing core that can be tiled into many-core processors, it paves the way for chiplets and multiple chip integration towards a brain-scale neuromorphic computing system. Fanxi Yang, Yuhan He, Jinqiao Yang, Anqin Xiao, Lufei Fan, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |