EDBT 2026 Demo / reviewers in the wild / expert
Faquan Chen
dblp:343/5578
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeuroUNI: A Unified Event-Driven Multi-Core Architecture Optimizing Neuromorphic Primitives for Brain-Inspired ComputingabstractThe hardware convergence of Artificial Neural Networks (ANNs) and Spiking Neural Networks (SNNs) is hindered by conflicting computational paradigms: dense tensor parallelism versus asynchronous sparse dynamics. Existing unifications typically rely on inefficient spatial partitioning or mode-reconfigurable datapaths, limiting the flexibility needed by heterogeneous ANN-SNN hybrid models requiring frequent cross-domain interaction. To resolve this, we present NeuroUNI, a unified event-driven multi-core architecture. Unlike partitioned designs, NeuroUNI unifies computation at the primitive level using a novel Five-Tuple Event Model, abstracting both continuous activations and discrete spikes to naturally leverage dynamic sparsity. The architecture features a co-optimized hierarchical Macro-Micro-$\mu $OP ISA, a superscalar SIMD-based microarchitecture, and a decentralized multi-core synchronization protocol. Validated in TSMC 28nm technology via post-synthesis simulation and on a Xilinx VCU129 FPGA prototype, NeuroUNI demonstrates competitive cross-paradigm efficiency. It achieves$35.0\times $and$1.21\times $the ANN energy efficiency of the NVIDIA V100 and EyerissV2, respectively, while delivering$5.7\times $the SNN throughput of TrueNorth. In a unified mapless navigation workload, NeuroUNI attains 422.6 GOPS/W (ANN) and 190.8GSOPS/W (SNN), outperforming Loihi 1 with$56.7\times $the throughput and$3.85\times $the energy efficiency, proving the viability of a primitive/ISA-level unified silicon substrate. Faquan Chen, Qingyang Tian, Ziren Wu, Xiangcheng Shi, Rendong Ying, Fei Wen 0005 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | Low-cost Deployment and Acceleration of Event-based Spiking Convolutional Neural NetworksabstractBy simulating the neurodynamics of biological brains, Spiking Neural Networks (SNNs) leverage sparse spike signal, eliminating the continuous multiply-accumulate operations of traditional Artificial Neural Networks (ANNs). Event-driven SNN processing offers significant advantages in energy efficiency and latency, making it ideal to be deployed on low-end processors. Spiking Convolutional Neural Networks (SC-NNs), which incorporate event-based processing, are increasingly employed for their power efficiency and ability to process spatio-temporal information. Unlike fully-connected networks, which rely on regular vector calculations, convolution operations present challenges for event-driven computation due to their sliding window nature.In this work, we deployed a compact 7-layer SCNN on the ARM Cortex-A9 core of PYNQ Z2 development board, using event-based acceleration. By optimizing data storage and processing sequences, we achieved efficient low-cost deployment. Offline training on the DVS128-Gesture dataset for an object recognition task yielded an accuracy of 93.40%. Following low-precision quantization and deployment, the model maintained a considerable accuracy of 92.36%. Compared to traditional periodic computation, event-based convolution processing achieved a 12.87× speedup. Furthermore, by exploiting the parallelism of feature map data storage along the channel dimension and ARM NEON instruction set, we gained an additional 2.97× speedup. Qingyang Tian, Faquan Chen, Lisheng Xie, Ziren Wu, Liangshun Wu, Rendong Ying |
ISCAS | 2 |
| 2024 | SPRCPl: An Efficient Tool for SNN Models Deployment on Multi-Core Neuromorphic Chips via Pilot RunningabstractThis paper introduce SPRCpl, an efficient compiler/toolkit for deploying Spiking Neural Network (SNN) models on multi-core neuromorphic chips. It uses "pilot running" to optimize the deployment process. It includes a front-end compiler, synapse pruning and regeneration optimizer, and a mapping tool. SPRCpl proposes a synapse pruning scheme based on spike firing statistics obtained through pilot running, dynamically reducing model size. It also presents a mapping scheme that minimizes strikes within and between clusters using spike firing statistics and multi-objective optimization. Experimental results demonstrate SPRCpl’s effectiveness in maintaining model accuracy during pruning and outperforming SpiNeMap in terms of communication count, execution time, and memory usage. It achieves lower latency, reduced power consumption, and higher throughput, making it a promising tool for SNN model deployment on multi-core neuromorphic chips. Liangshun Wu, Lisheng Xie, Jianwei Xue, Faquan Chen, Qingyang Tian, Ziren Wu, Rendong Ying |
ISCAS | 4 |
| 2024 | Combining contrastive learning and shape awareness for semi-supervised medical image segmentationabstractFor computer-aided diagnosis(CAD) to be successful, automatic segmentation needs to be reliable and efficient. Semi-supervised segmentation (SSL) techniques make extensive use of unlabeled data to address the issue of the high acquisition cost of medically labeled data. However, different anatomical regions and boundaries in medical images may exhibit similar gray-level features. The discrimination of similar regions and the geometrical limitations on boundaries are disregarded by current semi-supervised algorithms for segmenting medical images. In this work, we propose a framework for multi-task pixel-level representation learning that is led by certainty pixels. Specifically, we concentrate on the task of segmentation prediction as the primary task and shape-aware level set representation as a collaborative task to enforce local boundary constraints on unlabeled data. We construct dual decoders to obtain predictions and uncertainty maps from different perspectives, which can enhance the capacity to distinguish similar regions. In addition, we introduce certainty pixels to guide the computation of pixel-level contrastive loss to strengthen the correlation between pixels. Finally, experiments on two open datasets demonstrate that our strategy outperforms current approaches. The code will be released at https://github.com/yqimou/SAMT-PCL. Faquan Chen, Chenxi Huang 0001 |
Expert Syst. Appl. | 2 |
| 2024 | NHD-YOLO: Improved YOLOv8 using optimized neck and head for product surface defect detection with data augmentationabstractAbstract Surface defect detection is an essential task for ensuring the quality of products. Many excellent object detectors have been employed to detect surface defects in resent years, which has achieved outstanding success. To further improve the detection performance, a defect detector based on state‐of‐the‐art YOLOv8, named improved YOLOv8 by neck, head and data (NHD‐YOLO), is proposed. Specifically, YOLOv8 from three crucial aspects including neck, head and data is improved. First, a shortcut feature pyramid network is designed to effectively fuse features from backbone by improving the information transmission. Then, an adaptive decoupled head is proposed to alleviate the feature spatial misalignment between the classification and regression tasks. Finally, to enhance the training on small objects, a data augmentation method named selective small object copy and paste is proposed. Extensive experiments are conducted on three real‐world datasets: detection dataset from Northeastern University (NEU‐DET), printed circuit boards from Peking University (PKU‐Market‐PCB) and common objects in context (COCO). According to the results, NHD‐YOLO achieves the highest detection accuracy and exhibits outstanding inference speed and generalisation performance. Faquan Chen, Miaolei Deng, Xiaoya Yang, Dexian Zhang |
IET Image Process. | 1 |
| 2023 | SpikeNC: An Accurate and Scalable Simulator for Spiking Neural Network on Multi-Core Neuromorphic HardwareabstractMulti-core neuromorphic hardware for spiking neural networks (SNNs) has garnered considerable attention due to its biological plausibility and energy efficiency. However, the performance of SNN applications on such hardware is constrained by the rigid architecture and interconnection among neuron cores. To enable early-stage evaluation of SNN performance on multi-core neuromorphic hardware, we introduce an accurate and scalable simulator, SpikeNC. We present the entire workflow, ranging from SNN model training to simulation, providing comprehensive insights into both model and Network-on-Chip (NoC) related statistics. Moreover, we identify a considerable amount of time wastage in the widely adopted tick-based synchronous scheme. A three-stage agent-based asynchronous scheme is proposed for fast simulation. We evaluate the performance of deep spiking neural networks (DSNNs) with various scales trained on spike-converted datasets using SpikeNC. The results demonstrate that SpikeNC achieves precise and scalable simulation for SNNs on multi-core neuromorphic hardware. Additionally, the proposed asynchronous scheme significantly reduces the simulation cycles and absolute simulation time by approximately 63 % and 56 % respectively, compared to the synchronous scheme. We also delve into the trade-offs between different design parameters and explore the influence of mapping schemes utilizing SpikeNC. Lisheng Xie, Jianwei Xue, Liangshun Wu, Faquan Chen, Qingyang Tian, Rendong Ying |
HiPC | 4 |
| 2023 | ParallelNN: A Parallel Octree-based Nearest Neighbor Search Accelerator for 3D Point CloudsabstractAs Light Detection And Ranging (LiDAR) increasingly becomes an essential component in robotic navigation and autonomous driving, the processing of high throughput 3D point clouds in real time is widely required. This work considers the point cloud k-Nearest Neighbor (kNN) search, which is an important 3D processing kernel. Although applying fine-grained parallelism optimization on internal processing, e.g., using multiple workers, has demonstrated high efficiency, previous accelerators with DDR external memory are fundamentally limited by the external bandwidth bottleneck. To break this bottleneck, this work proposes a highly parallel architecture, namely ParallelNN, for highly efficient kNN search processing of high throughput point clouds. First, we optimize the multichannel cache based on High Bandwidth Memory (HBM) and on-chip memory to provide large external bandwidth. Then, a novel parallel depth-first octree construction algorithm is proposed and mapped onto multiple construction branches with trace-coded construction queues, which can regularize random accesses and perform multi-branch octree construction efficiently. Furthermore, in the search stage, we present algorithm-architecture co-optimization, including parallel keyframe-based scheduling and multi-branch flexible search engines, to provide conflict-free access and maximum reuse opportunities for reference points, which achieves more than 27.0× speedup compared with baseline architectures. We prototype ParallelNN on Virtex HBM FPGA and perform extensive benchmarking on the KITTI dataset. The results demonstrate that ParallelNN achieves up to 107.7× and 12.1× speedup over CPU and GPU implementations, while being more energy efficient, e.g., outperforming CPU and GPU implementations by 73.6× and 31.1×, respectively. Besides, with the proposed algorithm-architecture co-optimization, ParallelNN achieves 11.4× speedup over state-of-the-art architecture. Moreover, ParallelNN is configurable and can be easily generalized to similar octree-based applications. Faquan Chen, Rendong Ying, Jianwei Xue, Fei Wen 0005 |
HPCA | 1 |
| 2023 | Decoupled Consistency for Semi-supervised Medical Image Segmentation
Faquan Chen, Jingjing Fei |
MICCAI (1) | 1 |
| 2023 | SFANC: Scalable and Flexible Architecture for Neuromorphic ComputingabstractSpiking neural networks (SNNs) are recognized as the third generation of neural networks, boasting remarkable computational capabilities, which also require massive computational resources and flexibility to simulate biological neural functions. In this work, we present SFANC, a scalable and flexible neuromorphic architecture for SNNs based on optimized router architecture of network-on-chip (NoC) and highly programmable neuromorphic cores (NCs). SFANC includes 16 NCs, supporting 8000 neurons and four million synapses. The NCs are based on RISC-V with SNN-specific instructions, providing remarkable flexibility and computational speed. The NoC’s spiking routers facilitate multicast routing, flow estimation, and buffer sharing, contributing to increased scalability. Additionally, we propose a spectral cluster mapping approach for efficiently deploying SNNs onto SFANC, ensuring high flexibility and parallelism. We analyze the processing speedup for typical SNN topologies using the leaky-integrate-and-fire (LIF) neuron model within the NCs, achieving up to an$8.6\times $speedup over the RISC-V core on average under different SNN applications. Our enhanced spiking router demonstrates a spike latency reduction of up to 47.4% under various spiking coding schemes. Moreover, when applied to typical SNN topologies, our method exhibits an average spike latency decrease of up to 32.5% compared to the sequential mapping utilized by SpiNNaker. In summary, this work demonstrates high performance, flexibility, and scalability for simulating and accelerating SNNs, showcasing its potential as a promising solution for neuromorphic architectures. Jianwei Xue, Rendong Ying, Faquan Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |