EDBT 2026 Demo / reviewers in the wild / expert
Guanchao Qiao
dblp:340/8338
· DBLP profile ↗
14ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-4982-5938ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I2E: Real-Time Image-to-Event Conversion for High-Performance Spiking Neural NetworksabstractSpiking neural networks (SNNs) promise highly energy-efficient computing, but their adoption is hindered by a critical scarcity of event-stream data. This work introduces I2E, an algorithmic framework that resolves this bottleneck by converting static images into high-fidelity event streams. By simulating microsaccadic eye movements with a highly parallelized convolution, I2E achieves a conversion speed over 300x faster than prior methods, uniquely enabling on-the-fly data augmentation for SNN training. The framework's effectiveness is demonstrated on large-scale benchmarks. An SNN trained on the generated I2E-ImageNet dataset achieves a state-of-the-art accuracy of 60.50%. Critically, this work establishes a powerful sim-to-real paradigm where pre-training on synthetic I2E data and fine-tuning on the real-world CIFAR10-DVS dataset yields an unprecedented accuracy of 92.5%. This result validates that synthetic event data can serve as a high-fidelity proxy for real sensor data, bridging a long-standing gap in neuromorphic engineering. By providing a scalable solution to the data problem, I2E offers a foundational toolkit for developing high-performance neuromorphic systems. The open-source algorithm and all generated datasets are provided to accelerate research in the field. Liwei Meng, Guanchao Qiao, Ning Ning 0002, Yang Liu 0062, Shaogang Hu |
AAAI | 3 |
| 2026 | A Membrane Potential Communication-based Network-on-chip for Merge-core-free Spiking Neural Network Deployment
Weilai Chu, Zhantao Liu, Youming Liu, Pujun Zhou, Guanchao Qiao, Shaogang Hu |
ISCAS | 7 |
| 2026 | An Energy-Efficient Neuromorphic Self-Attention Core Exploiting Dual Sparsity in Neurons and Spikes
Pujun Zhou, R. C. Ma, Guanchao Qiao, Ning Ning 0002, Qi Yu 0002, Shaogang Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2026 | Neuromorphic Hybrid Information Processing Architecture for High-Performance Information CompressionabstractThe rapid development of artificial intelligence-of-things (AIOT) has led to a significant increase in intelligent nodes, presenting substantial challenges to internode communication. Semantic communication systems based on artificial neural networks (ANNs) have demonstrated greater robustness than traditional Shannon systems. However, the significant resource overhead makes hardware implementation challenging, and the low efficiency of information compression places considerable strain on communication bandwidth. This work proposed a hybrid semantic system that employs ANNs for high-precision semantic extraction at the server and spiking neural networks (SNNs) for high-performance information compression in the spatiotemporal dimension and semantic comprehending with low-hardware cost at the edge. A hardware-friendly, event-based neuromorphic core has been developed with low-hardware resource consumption and a high sampling rate for SNN deployment. The evaluation results indicate that the hybrid semantic system reduces the bandwidth by 95.8%, while maintaining a high sampling rate of 1000 Sa/s, outperforming traditional systems. Meanwhile, it enables a low entropy of the receiving end after four samples. Considering both bandwidth overhead and the entropy of the receiving end, the hybrid system achieves a high information compression ratio (ICR), surpassing the ANN-based system over 20 times. This work is expected to reveal the significance of neuromorphic computing in enabling high-performance information compression and transmission (Tx) in semantic communication. Pujun Zhou, Qi Yu 0002, Liwei Meng, C. X. Xiong, G. L. Yang, Guanchao Qiao, Ning Ning 0002, Yang Liu 0062, Shaogang Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | YOLO-fall: a YOLO-based fall detection model with high precision, shrunk size, and low latencyabstractAbstract According to recent research statistics, falling has become an important factor affecting the health and safety of the elderly. To reduce the computational cost of hardware and meet the demand for real-time fall detection, we propose a lightweight fall detection network called YOLO-fall oriented for mobile and small edge computing devices. We have made the following improvements based on you only look once (YOLO). First, the backbone network is designed to be lightweight. Then, the convolution module is reparameterized and the C3 structure is improved to ensure the balance between speed and accuracy. Finally, a 5 × 5 depth convolution is added to the detection head to improve the detection ability of large targets. The proposed YOLO-fall is trained and validated on the E-FPDS public dataset and achieves a 78.4% mean average precision (mAP) with 2.45 M parameters and 12.2 GFLOPs. Compared with YOLOv5s, YOLO-fall has a 6.1% improvement in mAP and a 65.1% reduction in parameters. Although Yolov9s has a higher mAP of 82.9%, YOLO-fall reduces the parameters and calculation quantities by 74.8 and 69.2%, respectively. Therefore, the proposed YOLO-fall has the potential to accurately perform real-time fall detection on mobile and small edge computing devices. Guanchao Qiao, Liwei Meng, Shaogang Hu |
Comput. J. | 3 |
| 2025 | A Neuromorphic Transformer Architecture Enabling Hardware-Friendly Edge ComputingabstractThe transformer model has demonstrated significant capabilities in various intelligent tasks, attracting widespread attention in recent years. However, it involves numerous complex operations, including large-bit-width multiplication, division, matrix transposition, and exponentiation. These require substantial storage and computational resources, making it challenging to deploy on edge devices. This work introduces a neuromorphic transformer architecture with low hardware cost for AI edge computing (AI-EC). At the structural level, it absorbs scaling factors within the self-attention mechanism into weight matrixes, thereby eliminating the division caused by the scaling operation. Additionally, a transposition calculation method is proposed to perform matrix transposition using dedicated memory access strategies and optimized data flow designs, which reduces logic resource overhead and avoids memory access discontinuities. At the computing paradigm level, the architecture employs spike-driven computing, substituting multi-bit multipliers with AND logic for synaptic operations. The paradigm introduces high sparsity to computational data, which is effectively exploited to reduce the computational workload of the architecture. The results indicate that the architecture successfully eliminates high-cost operators and significantly reduces computational expenses. Eventually, this architecture is verified as a prototype using a 28 nm CMOS process library, demonstrating a compact logic area of sub-0.2 mm2and a high energy efficiency of 0.34 pJ/SOP @ 50MHz. This work is expected to promote the application of transformers in edge computing and the development of intelligent edge applications. Pujun Zhou, R. C. Ma, Z. T. Liu, Liwei Meng, Guanchao Qiao, Yang Liu 0062, Qi Yu 0002, Shaogang Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | A 0.96 pJ/SOP Heterogeneous Neuromorphic Chip Toward Energy-Efficient Edge Visual ApplicationsabstractEdge devices require low power consumption and compact area, which poses challenges for visual signal processing. This work introduces an energy-efficient heterogeneous neuromorphic system-on-chip (SoC) for edge visual computing. The neuromorphic core design incorporates advanced technologies, such as sparse-aware synaptic calculation, partial membrane potential update, non-uniform weight quantization, and partial parallel computing, achieving excellent energy efficiency, computing performance, and area utilization. Twenty neuromorphic cores and twelve multi-mode connected-matrix-based routers form a network-on-chip (NoC) with fullerene-like topology. Its average degree of communication nodes exceeds traditional topologies by 32 % and maintains a minimum degree variance of 0.93, thereby enabling advanced decentralized on-chip communication. Moreover, the NoC can be scaled up through extended off-chip high-level router nodes. At the top layer of the SoC, a RISC-V CPU and a 20-core neuromorphic processor are tightly coupled to form a heterogeneous architecture. Eventually, the chip is fabricated within a 3.41 mm2die area under 55 nm CMOS technology, achieving a low power density of 0.52 mW/mm2and a high neuron density of 30.23 K/mm2. Its effectiveness is verified across different visual tasks, with a best energy efficiency of 0.96 pJ/SOP. This work is expected to promote the development of neuromorphic computing in edge visual applications. Pujun Zhou, Guanchao Qiao, Qi Yu 0002, Junjie Wang 0008, Ning Ning 0002, Yang Liu 0062, Shaogang Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | A&B BNN: Add&Bit-Operation-Only Hardware-Friendly Binary Neural NetworkabstractBinary neural networks utilize 1-bit quantized weights and activations to reduce both the model's storage demands and computational burden. However, advanced binary architectures still incorporate millions of inefficient and non-hardware-friendly full-precision multiplication operations. A&B BNN is proposed to directly remove part of the multiplication operations in a traditional BNN and replace the rest with an equal number of bit operations, introducing the mask layer and the quantized RPReLU structure based on the normalizer-free network architecture. The mask layer can be removed during inference by leveraging the intrinsic characteristics of BNN with straightforward mathematical transformations to avoid the associated multiplication operations. The quantized RPReLU structure enables more efficient bit operations by constraining its slope to be integer powers of 2. Experimental results achieved 92.30%, 69.35%, and 66.89% on the CIFAR-10, CIFAR-100, and ImageNet datasets, respectively, which are competitive with the state-of-the-art. Ablation studies have verified the efficacy of the quantized RPReLU structure, leading to a 1.14% enhancement on the ImageNet compared to using a fixed slope RLeakyReLU. The proposed add&bit-operation-only BNN offers an innovative approach for hardware-friendly network architecture. Guanchao Qiao, Yian Liu 0001, Liwei Meng, Ning Ning 0002, Yang Liu 0062, Shaogang Hu |
CVPR | 2 |
| 2024 | A 0.96pJ/SOP, 30.23K-neuron/mm2 Heterogeneous Neuromorphic Chip With Fullerene-like Interconnection Topology for Edge-AI ComputingabstractEdge-AI computing requires high energy efficiency, low power consumption, and relatively high flexibility and compact area, challenging the AI-chip design. This work presents a 0.96 pJ/SOP heterogeneous neuromorphic system-on-chip (SoC) with fullerene-like interconnection topology for edge-AI computing. The neuromorphic core integrates different technologies to augment computing energy efficiency, including sparse computing, partial membrane potential updates, and non-uniform weight quantization. Multiple neuromorphic cores and multi-mode routers form a fullerene-like network-on-chip (NoC). The average degree of communication nodes exceeds traditional topologies by 32%, with a minimal degree variance of 0.93, allowing advanced decentralized on-chip communication. Additionally, the NoC can be scaled up through extended off-chip high-level router nodes. A RISC-V CPU and a neuromorphic processor are tightly coupled and fabricated within a 5.42 mm2die area under 55 nm CMOS technology. The chip has a low power density of 0.52 mW/mm2, reducing 67.5% compared to related works, and achieves a high neuron density of 30.23 K/mm2. Eventually, the chip is demonstrated to be effective on different datasets and achieves 0.96 pJ/SOP energy efficiency. Pujun Zhou, Qi Yu 0002, Liwei Meng, Yue Zuo, Ning Ning 0002, Shaogang Hu, Guanchao Qiao |
ISCAS | 10 |
| 2023 | An efficient pruning and fine-tuning method for deep spiking neural network
Liwei Meng, Guanchao Qiao, Yue Zuo, Pujun Zhou, Yang Liu 0062, Shaogang Hu |
Appl. Intell. | 2 |
| 2023 | Batch normalization-free weight-binarized SNN based on hardware-saving IF neuron
Guanchao Qiao, Nanning Zheng 0001, Yue Zuo, Pujun Zhou, M. L. Sun, Shaogang Hu, Qi Yu 0002 |
Neurocomputing | 1 |
| 2021 | Direct training of hardware-friendly weight binarized spiking neural network with surrogate gradient learning towards spatio-temporal event-based dynamic data recognition
Guanchao Qiao, Ning Ning 0002, Yue Zuo, Shaogang Hu, Qi Yu 0002 |
Neurocomputing | 1 |
| 2021 | Quantized STDP-based online-learning spiking neural network
Shaogang Hu, Guanchao Qiao, Tupei Chen, Qi Yu 0002, L. M. Rong |
Neural Comput. Appl. | 2 |
| 2020 | STBNN: Hardware-friendly spatio-temporal binary neural network with high pattern recognition accuracy
Guanchao Qiao, Shaogang Hu, Tupei Chen, L. M. Rong, Ning Ning 0002, Qi Yu 0002 |
Neurocomputing | 1 |