EDBT 2026 Demo / reviewers in the wild / expert
Youming Yang 0002
dblp:250/5955-2
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0009-0593-4842ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DyNeuro: A Hybrid Neuromorphic Accelerator with Dynamic Spatio-Temporal Variation Adaptations
Youming Yang 0002, Yi Zhong 0002, Li Lun, Tao Zhang 0140, Xiaoxin Cui, Yuan Wang 0001 |
ISCAS | 1 |
| 2025 | CROSSCUT: A Multi-Core Neuromorphic Accelerator Improving Resource-UtilizationabstractNeuromorphic computing is attracting significant attention due to its bio-mimetic characteristics. Consequently, neuromorphic hardware platforms have emerged as innovative computing architectures for acceleration. However, the fixed nature of data flow and resources leads to considerable inefficiencies in storage and computation, thereby limiting both utilization efficiency and overall performance. This severely hinders the deployment of edge artificial intelligence (AI) models. To address these issues, we present a multi-core neuromorphic accelerator named CROSSCUT. This crossbar-based system supports both spiking neural network (SNN) and artificial neural network (ANN) paradigms and has a capacity of 256K neurons and 288M synapses. By leveraging the Neuron Package Mechanism (NPM) and Synapse Compress Mechanism (SCM), CROSSCUT can increase input data scale by 64 times and reduce wasted resources and computations by 46.7%, ensuring high compatibility with diverse network structures in machine learning models. Additionally, a Tree-Mesh hybrid network on chip (NoC) is constructed for inter-core communication. Implemented on Xilinx XCVU9P FPGA, CROSSCUT can achieve a peak performance of 431.9 GSOPS/s and 121.13 GSOPS/W energy efficiency. The inference accuracy on MNIST is 98.2%. Youming Yang 0002, Yi Zhong 0002, Zilin Wang 0001, Tao Zhang 0140, Li Lun, Yingying Cui, Xiaoxin Cui, Song Jia, Yuan Wang 0001 |
ISCAS | 1 |
| 2025 | A Bit-Partitioned Floating-Point 6T SRAM Computing-in-Memory Macro Based on Dual-Edge Time-Domain StructureabstractIn the computing-in-memory (CIM) field, floating-point (FP) CIM is afflicted with high computing latency and energy consumption due to the intricate procedures involved in exponent computation and processing. In this work, an 8Kb FP time-domain (TD) static-random-access-memory (SRAM) CIM macro is presented. Fabricated with a 180nm process, this macro exhibits low computational latency and high energy efficiency. A novel FP computing architecture is proposed, which is capable of concurrently executing exponent summation, maximum value finding, difference generation, and mantissa shifting. This architecture effectively reduces the overall delay in exponent computation and processing, thereby enhancing the throughput. Furthermore, a bit-partitioned computing concept and an exponent sparsity scheme are introduced. In this scheme, sparsity judgment is made solely by processing the high 4 bits of the exponent, which significantly reduces power consumption in the remaining exponent computation and processing steps. Additionally, based on the bit-partitioned concept, a dual-edge TD exponent summation and mantissa multiplication-and-accumulation (MAC) circuit is devised. This circuit not only suppresses nonlinear errors during multi-bit computation but also exploits both the rising and falling edges of pulses for computation, thus accelerating the macro’s operation speed. Compared to previous approaches, an extra 24% power reduction is achieved. At a sparsity level of 90%, a normalized energy efficiency of 14.418 TFLOPS/W and a normalized area efficiency of 0.041 TFLOPS/mm2are attained. When this work is applied to the ResNet-18 model with BF16 format for input, weight, and output, the accuracy loss on the CIFAR-100 dataset is merely −0.16%. Chang Xue, Youming Yang 0002, Gang Du, Yuan Wang 0001, Yandong He |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | NeuroREC: A 28-nm Efficient Neuromorphic Processor for Radar Emitter ClassificationabstractRadar emitter classification (REC) plays an important role in modern warfare. Traditional REC methods have difficulty identifying complex radar signals in the present day. Inspired by biology, spiking neural networks (SNNs) have gradually gained widespread attention due to their low power characteristics. Compared with convolutional neural networks (CNNs), SNNs are more suitable for application in the field of REC. The reason is that SNN can not only maintain higher accuracy in the presence of noise interference, but also reduce the power consumption of mobile devices. However, it is challenging to make full use of the input sparsity of radar emitter signals and the weight sparsity of pruned SNN models. In this paper, a 28-nm neuromorphic processor for REC named NeuroREC is proposed. It uses matrix compression algorithms to store sparse weights on chip, and designs corresponding spike detection circuits for this purpose. As a single-core design, we propose a ping-pong running mechanism to alleviate the imbalance between IO throughput and peak performance. Two SNN models for classifying RadioML2016.b and RadioML2018.a datasets are deployed on the chip, achieving competitive accuracy with only 8 timesteps, and demonstrating better robustness than CNN. Fabricated in 28-nm CMOS process, NeuroREC runs at frequencies ranging from 22.5MHz to 744MHz. Under specific sparsity conditions, it can reach an energy efficiency of 7.22TSOP/W for 8-bit weight. Zilin Wang 0001, Zehong Ou, Yi Zhong 0002, Youming Yang 0002, Li Lun, Hufei Li, Jian Cao 0002, Xiaoxin Cui, Song Jia, Yuan Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Marmotini: A Weight Density Adaptation Architecture With Hybrid Compression Method for Spiking Neural NetworkabstractBrain-inspired spiking neural network (SNN) has recently attracted widespread interest owing to its event-driven nature and relatively low-power hardware for transmitting highly sparse binary spikes. To further improve energy efficiency, some matrix compression algorithms are used for weight storage. However, the weight sparsity of different layers varies greatly. For a multicore neuromorphic system, it is difficult for the same compression algorithm to adapt to all the layers of SNN model. In this work, we propose a weight density adaptation architecture with hybrid compression method for SNN, named Marmotini. It is a multicore heterogeneous design, including three types of cores to complete computation of different weight sparsity. Benefiting from the hybrid compression method, Marmotini minimizes the waste of neurons and weights as much as possible. Besides, for better flexibility, a reconfigurable core that can be configured to compute convolutional layer or fully connected layer is proposed. Implemented on Xilinx Kintex UltraScale XCKU115 field-programmable gate array (FPGA) board, Marmotini can operate at 150-MHz frequency, achieving 244.6-GSOP/s peak performance and 54.1-GSOP/W energy efficiency at 0% spike sparsity. Zilin Wang 0001, Yi Zhong 0002, Zehong Ou, Youming Yang 0002, Xiaoxin Cui, Song Jia, Yuan Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |