EDBT 2026 Demo / reviewers in the wild / expert
Jianwei Xue
dblp:246/5844
· DBLP profile ↗
6ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0003-1992-218XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SPRCPl: An Efficient Tool for SNN Models Deployment on Multi-Core Neuromorphic Chips via Pilot RunningabstractThis paper introduce SPRCpl, an efficient compiler/toolkit for deploying Spiking Neural Network (SNN) models on multi-core neuromorphic chips. It uses "pilot running" to optimize the deployment process. It includes a front-end compiler, synapse pruning and regeneration optimizer, and a mapping tool. SPRCpl proposes a synapse pruning scheme based on spike firing statistics obtained through pilot running, dynamically reducing model size. It also presents a mapping scheme that minimizes strikes within and between clusters using spike firing statistics and multi-objective optimization. Experimental results demonstrate SPRCpl’s effectiveness in maintaining model accuracy during pruning and outperforming SpiNeMap in terms of communication count, execution time, and memory usage. It achieves lower latency, reduced power consumption, and higher throughput, making it a promising tool for SNN model deployment on multi-core neuromorphic chips. Liangshun Wu, Lisheng Xie, Jianwei Xue, Faquan Chen, Qingyang Tian, Ziren Wu, Rendong Ying |
ISCAS | 3 |
| 2023 | SpikeNC: An Accurate and Scalable Simulator for Spiking Neural Network on Multi-Core Neuromorphic HardwareabstractMulti-core neuromorphic hardware for spiking neural networks (SNNs) has garnered considerable attention due to its biological plausibility and energy efficiency. However, the performance of SNN applications on such hardware is constrained by the rigid architecture and interconnection among neuron cores. To enable early-stage evaluation of SNN performance on multi-core neuromorphic hardware, we introduce an accurate and scalable simulator, SpikeNC. We present the entire workflow, ranging from SNN model training to simulation, providing comprehensive insights into both model and Network-on-Chip (NoC) related statistics. Moreover, we identify a considerable amount of time wastage in the widely adopted tick-based synchronous scheme. A three-stage agent-based asynchronous scheme is proposed for fast simulation. We evaluate the performance of deep spiking neural networks (DSNNs) with various scales trained on spike-converted datasets using SpikeNC. The results demonstrate that SpikeNC achieves precise and scalable simulation for SNNs on multi-core neuromorphic hardware. Additionally, the proposed asynchronous scheme significantly reduces the simulation cycles and absolute simulation time by approximately 63 % and 56 % respectively, compared to the synchronous scheme. We also delve into the trade-offs between different design parameters and explore the influence of mapping schemes utilizing SpikeNC. Lisheng Xie, Jianwei Xue, Liangshun Wu, Faquan Chen, Qingyang Tian, Rendong Ying |
HiPC | 2 |
| 2023 | ParallelNN: A Parallel Octree-based Nearest Neighbor Search Accelerator for 3D Point CloudsabstractAs Light Detection And Ranging (LiDAR) increasingly becomes an essential component in robotic navigation and autonomous driving, the processing of high throughput 3D point clouds in real time is widely required. This work considers the point cloud k-Nearest Neighbor (kNN) search, which is an important 3D processing kernel. Although applying fine-grained parallelism optimization on internal processing, e.g., using multiple workers, has demonstrated high efficiency, previous accelerators with DDR external memory are fundamentally limited by the external bandwidth bottleneck. To break this bottleneck, this work proposes a highly parallel architecture, namely ParallelNN, for highly efficient kNN search processing of high throughput point clouds. First, we optimize the multichannel cache based on High Bandwidth Memory (HBM) and on-chip memory to provide large external bandwidth. Then, a novel parallel depth-first octree construction algorithm is proposed and mapped onto multiple construction branches with trace-coded construction queues, which can regularize random accesses and perform multi-branch octree construction efficiently. Furthermore, in the search stage, we present algorithm-architecture co-optimization, including parallel keyframe-based scheduling and multi-branch flexible search engines, to provide conflict-free access and maximum reuse opportunities for reference points, which achieves more than 27.0× speedup compared with baseline architectures. We prototype ParallelNN on Virtex HBM FPGA and perform extensive benchmarking on the KITTI dataset. The results demonstrate that ParallelNN achieves up to 107.7× and 12.1× speedup over CPU and GPU implementations, while being more energy efficient, e.g., outperforming CPU and GPU implementations by 73.6× and 31.1×, respectively. Besides, with the proposed algorithm-architecture co-optimization, ParallelNN achieves 11.4× speedup over state-of-the-art architecture. Moreover, ParallelNN is configurable and can be easily generalized to similar octree-based applications. Faquan Chen, Rendong Ying, Jianwei Xue, Fei Wen 0005 |
HPCA | 3 |
| 2023 | SFANC: Scalable and Flexible Architecture for Neuromorphic ComputingabstractSpiking neural networks (SNNs) are recognized as the third generation of neural networks, boasting remarkable computational capabilities, which also require massive computational resources and flexibility to simulate biological neural functions. In this work, we present SFANC, a scalable and flexible neuromorphic architecture for SNNs based on optimized router architecture of network-on-chip (NoC) and highly programmable neuromorphic cores (NCs). SFANC includes 16 NCs, supporting 8000 neurons and four million synapses. The NCs are based on RISC-V with SNN-specific instructions, providing remarkable flexibility and computational speed. The NoC’s spiking routers facilitate multicast routing, flow estimation, and buffer sharing, contributing to increased scalability. Additionally, we propose a spectral cluster mapping approach for efficiently deploying SNNs onto SFANC, ensuring high flexibility and parallelism. We analyze the processing speedup for typical SNN topologies using the leaky-integrate-and-fire (LIF) neuron model within the NCs, achieving up to an$8.6\times $speedup over the RISC-V core on average under different SNN applications. Our enhanced spiking router demonstrates a spike latency reduction of up to 47.4% under various spiking coding schemes. Moreover, when applied to typical SNN topologies, our method exhibits an average spike latency decrease of up to 32.5% compared to the sequential mapping utilized by SpiNNaker. In summary, this work demonstrates high performance, flexibility, and scalability for simulating and accelerating SNNs, showcasing its potential as a promising solution for neuromorphic architectures. Jianwei Xue, Rendong Ying, Faquan Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | No-Reference Stereoscopic Image Quality Assessment Based on Local to Global Feature RegressionabstractIn this paper, we propose a two-channel deep convolutional neural network (DCNN) through local to global regression for no-reference (NR) stereoscopic images quality assessment (SIQA). Firstly, in most deep learning based methods, they use the given subjective mean opinion score (MOS) or the differential MOS (DMOS) value to adjust the network parameters. But it is unreasonable, especially for the asymmetrical distortion stereoscopic image. To alleviate the problem, we propose to use feature similarity index (FSIM) to provide pseudo labels for the left and right view respectively, named local regression, so that the left and right channel are trained better. Then, we use the given DMOS to finetune the locally trained model parameters, named global regression. So we achieve an end-to-end network to measure the stereoscopic image quality. The experimental results show that the proposed method is superior to other existing SIQA methods. Sumei Li, Jianwei Xue, Yongtian Han |
ICME | 2 |
| 2019 | Stereoscopic Video Quality Assessment Based on The Two-step-training Binocular Fusion NetworkabstractIn this paper, we propose a novel binocular fusion network for stereoscopic video quality assessment (SVQA). In this network, we construct a long-term fusion, competition, and processing process by simulating the long-term complex process of the whole visual pathway. And we employ a two-step-training strategy for this network, which solves the problem that the network is difficult to fit caused by using the same value to label the different quality regions and views of the same stereoscopic video. In the first step, we use the computed quality scores of different patches to train the local network, namely local regression. And then the global regression is performed by using MOS value based on the first step trained model. Besides, considering temporal information, we take spatiotemporal saliency feature flows as the inputs of the proposed network. The proposed method is tested on public stereoscopic video databases, and results show that our method outperforms any other methods. Sumei Li, Jianwei Xue, Yixiu Ding, Guanghui Yue 0001 |
VCIP | 3 |