EDBT 2026 Demo / reviewers in the wild / expert
Yuehai Chen
dblp:286/7033
· DBLP profile ↗
20ranked-venue papers
8as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention DataflowabstractSpiking neural networks (SNNs) have emerged as a promising alternative to artificial neural networks (ANNs), offering improved energy efficiency by leveraging sparse and event-driven computation. However, existing hardware implementations of SNNs still suffer from the inherent spike sparsity and multi-timestep execution, which significantly increase latency and reduce energy efficiency. This study presents NEURAL, a novel neuromorphic architecture based on a hybrid dataevent execution paradigm by decoupling sparsity-aware processing from neuron computation and using elastic first-in-firstout (FIFO). NEURAL supports on-the-fly execution of spiking QKFormer by embedding its operations within the baseline computing flow without requiring dedicated hardware units. It also integrates a novel window-to-time-to-first-spike (W2TTFS) mechanism to replace average pooling and enable full-spike execution. Furthermore, we introduce a knowledge distillation (KD)-based training framework to construct single-timestep SNN models with competitive accuracy. NEURAL is implemented on a Xilinx Virtex-7 FPGA and evaluated using ResNet-11, QKFResNet-11, and VGG-11. Experimental results demonstrate that, at the algorithm level, the VGG-11 model trained with KD improves accuracy by $3.20 \%$ on CIFAR-10 and $5.13 \%$ on CIFAR-100. At the architecture level, compared to existing SNN accelerators, NEURAL achieves a $50 \%$ reduction in resource utilization and a $1.97 \times$ improvement in energy efficiency. Yuehai Chen, Farhad Merchant |
ASP-DAC | 1 |
| 2026 | PenPy-DETR: A Penta-Pyramid Framework for Small Target Perception in Low-Light
Yanheng Wen, Yuehai Chen, Jing Yang 0014 |
IV | 2 |
| 2026 | SODBoost: Density-Guided Zoom-In Mosaic for Small Object Detection
Jinlun Yu, Yuehai Chen, Jing Yang 0014 |
IV | 2 |
| 2026 | Fine-Grained Domain Alignment for Face Anti-Spoofing With Asymmetric Pseudo-Labels
Jing Yang 0014, Xusheng Cui, Yuehai Chen, Shaoyi Du, Badong Chen, Yuewen Liu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | SConvNSys: Accelerating Spiking Convolutional Neural Networks With a Reconfigurable Neuromorphic Architecture for Diverse ApplicationsabstractIn recent years, spiking neural networks (SNNs) have progressively closed the performance gap with convolutional neural networks (CNNs), which are renowned for their success in complex artificial intelligence tasks. However, most existing SNN accelerators suffer from limited flexibility, lacking support for diverse convolutional topologies and struggling to deploy complex SNN models. To address these challenges, this article proposes a reconfigurable neuromorphic architecture, SConvNSys, tailored for accelerating spiking CNNs (SCNNs) across a wide range of applications. A fast, sparse detection technique, coupled with a dedicated sparse response unit, is introduced to effectively exploit the spatio-temporal sparsity inherent in spike-based computation. On this basis, a reconfigurable spiking convolution dataflow is designed to optimize computation across various SCNN structures. The proposed architecture supports multiple convolution types, including standard, transposed, dilated, and residual convolutions. Implemented on a field-programmable gate array (FPGA), SConvNSys achieves competitive results: for image classification, it attains recognition accuracies of 91.48% on CIFAR-10 and 68.54% on CIFAR-100, with a power consumption of just 1.8 W and a processing rate of 73 frames/s. In image segmentation tasks, it reaches 99.00% segmentation accuracy at 153 frames/s. For object detection, the proposed detection model achieves a mean intersection over union (MIoU) of 74.20%, with 0.11 giga operations (GOP) of convolutional computation, resulting in a throughput of 27.4 giga operations per second (GOPS) and an energy efficiency of 15.23 GOPS/W. Wujian Ye, Yingzhang Liang, Yijun Liu 0005, Youfeng Cui, Yuehai Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Weak-edge sample extension for enhancing unsupervised feature learning
Yuehai Chen, Shanying Chen, Jing Yang 0014, Badong Chen, Shaoyi Du, Yuewen Liu |
Neurocomputing | 1 |
| 2025 | The architecture design and training optimization of spiking neural network with low-latency and high-performance for classification and segmentation
Wujian Ye, Shaozhen Chen, Haoxian Liu, Yijun Liu 0010, Yuehai Chen, Youfeng Cui |
Neural Networks | 5 |
| 2025 | CSCC: Cross-Scene Crowd Counting via Learning to Diversify for Domain GeneralizationabstractIt is challenging for crowd counting models to generalize to new scenes due to domain shifts in training and test data. Although domain adaptation approaches have made notable progress in bridging the domain gap, they require target domain data. In this paper, we propose a novel framework for cross-scene crowd counting, which unifies domain generalization and adaptation. For domain generalization, we train a model only using single-domain data and the model can be generalized to any scene with satisfying performance. Regarding domain adaptation, we use both source and target domain data to further improve the performance. We first design a generation network that diversifies the generated samples to cover the unseen target domains as much as possible by minimizing mutual information. This approach simulates training data in various domains, thereby enhancing the model's generalization ability. Then we develop a pixel-wise supervised contrastive loss function that pulls the human heads in the source images and generated images closer to each other and pushes them further away from the background. This loss helps extract a domain-invariant feature representation, thus improving the model's generalization ability. Moreover, if information about the target domain is available, our generalization method can be easily applied as an adaptation method by replacing the mutual information minimization loss with the mutual information maximization loss. This can further improve cross-scene crowd counting performance. The experimental results demonstrate the strong generalizability of our method across different datasets. Yuehai Chen, Qingzhong Wang, Jing Yang 0014, Badong Chen, Haoyi Xiong, Shaoyi Du |
IEEE Trans. Multim. | 1 |
| 2025 | Research on Hardware Acceleration of Traffic Sign Recognition Based on Spiking Neural Network and FPGA PlatformabstractMost of the existing methods for traffic sign recognition exploited deep learning technology such as convolutional neural networks (CNNs) to achieve a breakthrough in detection accuracy; however, due to the large number of CNN’s parameters, there are problems in practical applications such as high power consumption, large calculation, and slow speed. Compared with CNN, a spiking neural network (SNN) can effectively simulate the information processing mechanism of biological brain, with stronger parallel processing capability, better sparsity, and real-time performance. Thus, we design and realize a novel traffic sign recognition system [called SNN on FPGA-traffic sign recognition system (SFPGA-TSRS)] based on spiking CNN (SCNN) and FPGA platform. Specifically, to improve the recognition accuracy, a traffic sign recognition model spatial attention SCNN (SA-SCNN) is proposed by combining LIF/IF neurons based SCNN with SA mechanism; and to accelerate the model inference, a neuron module is implemented with high performance, and an input coding module is designed as the input layer of the recognition model. The experiments show that compared with existing systems, the proposed SFPGA-TSRS can efficiently support the deployment of SCNN models, with a higher recognition accuracy of 99.22%, a faster frame rate of 66.38 frames per second (FPS), and lower power consumption of 1.423 W on the GTSRB dataset. Huarun Chen, Wujian Ye, Jialiang Ye, Yuehai Chen, Shaozhen Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | SiBrain: A Sparse Spatio-Temporal Parallel Neuromorphic Architecture for Accelerating Spiking Convolution Neural Networks With Low LatencyabstractCurrently, the performance of spiking neural networks (SNNs) under complicated tasks is gradually close to that of convolutional neural networks (CNNs). However, the existing neuromorphic computing hardware architectures (NCHAs) for accelerating SNNs lack effective sparse detection and cannot realize effective spatio-temporal parallel computing, resulting in high inference latency. In this paper, we propose an improved neuromorphic computing hardware architecture (called SiBrain) with high accuracy, low power and low latency, consisting of a sparse spatio-temporal parallel processing element (S$^{2}$TP-PE) array for spike convolution and pooling computation and a fully-connected (FC) Core for spike FC computation. In the novel S$^{2}$TP-PE array, a S$^{2}$TP computation unit is designed by using time-step and channel parallel computation techniques to reduce the latency caused by multiple time-steps of SNN, and a sparse detection and response unit is presented based on channel-cache and block-multiplex to achieve the spike detection, response, and reuse. Combining the above kernel components, the SiBrain is built and implemented on Virtex-7 FPGA with 200MHz. The experimental results show that the SiBrain can effectively support the deployment of SCNN models with different sizes, and the deployed large-scale Spiking Visual-Geometry-Group (VGG) model can achieve the recognition accuracies of 90.25% and 66.97% on CIFAR-10 and CIFAR-100 with the power consumption of 1.5 W and the energy efficiency of 83 GSOPs/W. Compared with existing FPGA-based SNN accelerators, SiBrain has a maximum increase in inference speed of nearly 11 times and a maximum decrease in energy consumption of nearly 34 times, respectively. Yuehai Chen, Wujian Ye, Yijun Liu 0005 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | IA-LSTM: Interaction-Aware LSTM for Pedestrian Trajectory PredictionabstractPredicting the trajectory of pedestrians in crowd scenarios is indispensable in self-driving or autonomous mobile robot field because estimating the future locations of pedestrians around is beneficial for policy decision to avoid collision. It is a challenging issue because humans have different walking motions, and the interactions between humans and objects in the current environment, especially between humans themselves, are complex. Previous researchers focused on how to model human-human interactions but neglected the relative importance of interactions. To address this issue, a novel mechanism based on correntropy is introduced. The proposed mechanism not only can measure the relative importance of human-human interactions but also can build personal space for each pedestrian. An interaction module, including this data-driven mechanism, is further proposed. In the proposed module, the data-driven mechanism can effectively extract the feature representations of dynamic human-human interactions in the scene and calculate the corresponding weights to represent the importance of different interactions. To share such social messages among pedestrians, an interaction-aware architecture based on long short-term memory network for trajectory prediction is designed. Experiments are conducted on two public datasets. Experimental results demonstrate that our model can achieve better performance than several latest methods with good performance. Jing Yang 0014, Yuehai Chen, Shaoyi Du, Badong Chen, José C. Príncipe |
IEEE Trans. Cybern. | 2 |
| 2024 | Learning Discriminative Features for Crowd CountingabstractCrowd counting models in highly congested areas confront two main challenges: weak localization ability and difficulty in differentiating between foreground and background, leading to inaccurate estimations. The reason is that objects in highly congested areas are normally small and high-level features extracted by convolutional neural networks are less discriminative to represent small objects. To address these problems, we propose a learning discriminative features framework for crowd counting, which is composed of a masked feature prediction module (MPM) and a supervised pixel-level contrastive learning module (CLM). The MPM randomly masks feature vectors in the feature map and then reconstructs them, allowing the model to learn about what is present in the masked regions and improving the model's ability to localize objects in high-density regions. The CLM pulls targets close to each other and pushes them far away from background in the feature space, enabling the model to discriminate foreground objects from background. Additionally, the proposed modules can be beneficial in various computer vision tasks, such as crowd counting and object detection, where dense scenes or cluttered environments pose challenges to accurate localization. The proposed two modules are plug-and-play, incorporating the proposed modules into existing models can potentially boost their performance in these scenarios. Yuehai Chen, Qingzhong Wang, Jing Yang 0014, Badong Chen, Haoyi Xiong, Shaoyi Du |
IEEE Trans. Image Process. | 1 |
| 2023 | Video Self-Supervised Cross-Pathway Training Based on Slow and Fast PathwaysabstractIn the field of video self-supervised learning, contrastive instance learning methods suffer from a lack of semantic information, resulting in inadequate generalization in downstream tasks. Although optical flow can provide some semantic information, it requires significant computational cost prior to training. To address this, we propose a Video self-supervised Cross-pathway training model based on Slow and Fast pathways (VCSF). This model separately extracts temporal and spatial features from pure RGB video frames, and uses the complementary representations of the two pathways to conduct cross-pathway training. Additionally, we propose a motion perception module in the low-frame-rate space to enhance the network's ability to perceive rapidly changing human motion. We conducted extensive experiments in downstream missions of UCF101 and HMDB51, and obtained state-of-the-art results in models using the UCF101 data set for self-supervised pre-training, including motion recognition and nearest neighbor retrieval. Jing Yang 0014, Zhou Jiang 0001, Yuehai Chen, Shaoyi Du |
SMC | 4 |
| 2023 | Multi-style transfer and fusion of image's regions based on attention mechanism and instance segmentation
Wujian Ye, Chaojie Liu, Yuehai Chen, Yijun Liu 0010, Chenming Liu |
Signal Process. Image Commun. | 3 |
| 2023 | The Implementation and Optimization of Neuromorphic Hardware for Supporting Spiking Neural Networks With MLP and CNN TopologiesabstractSpiking neural network (SNN) has attracted extensive attention in large-scale image processing tasks. To obtain higher computing efficiency, the development of hardware architecture suitable for SNN computing has become a hot research topic. However, the existing hardware of spike neurons still has high computational complexity and they do not perform well enough on complicated datasets, and the neuromorphic system cannot support SNNs with different convolutional topologies, resulting in low efficiency of the system. To address the above problems, an optimized leaky integrated-and-fire (LIF) neuron called EPC-LIF and a neuromorphic hardware acceleration system (ELIF-NHAS) are designed and implemented based on the field-programmable gate array (Xilinx Kintex-7). First, the classical LIF neuron is designed using the optimization method of extended prediction correction (EPC), which can reduce the computation complexity and hardware resources with a maximum frequency of 439.95 MHz. The ELIF-NHAS is constructed and optimized with parallel and pipeline techniques for effectively running SNNs, working with a maximum frequency of 135.6 MHz. Then, the genetic algorithm is applied to adjust the membrane threshold of neurons for further improving the accuracy of SNNs. Furthermore, the ELIF-NHAS can support different SNNs with multilayer perceptron and convolutional neural network topologies (called SCNN), including traditional, depth-separate, and residual convolutions. The accuracy of multilayer SCNNs can achieve 99.10%, 90.29%, and 82.15% on MNIST, Fashion-MNIST, and SVHN datasets, respectively; and the speed and energy consumption achieve 1.21 ms/image and 1.19 mJ/image. Compared with existing systems, the ELIF-NHAS is more suitable for the deployment and inference of SNNs with higher speed and lower consumption. Wujian Ye, Yuehai Chen, Yijun Liu 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Counting Varying Density Crowds Through Density Guided Adaptive Selection CNN and Transformer EstimationabstractIn real-world crowd counting applications, the crowd densities in an image vary greatly. When facing density variation, humans tend to locate and count the targets in low-density regions, and reason the number in high-density regions. We observe that CNN focus on the local information correlation using a fixed-size convolution kernel and the Transformer could effectively extract the semantic crowd information by using the global self-attention mechanism. Thus, CNN could locate and estimate crowds accurately in low-density regions, while it is hard to properly perceive the densities in high-density regions. On the contrary, Transformer has a high reliability in high-density regions, but fails to locate the targets in sparse regions. Neither CNN nor Transformer can well deal with this kind of density variation. To address this problem, we propose a CNN and Transformer Adaptive Selection Network (CTASNet) which can adaptively select the appropriate counting branch for different density regions. Firstly, CTASNet generates the prediction results of CNN and Transformer. Then, considering that CNN/Transformer is appropriate for low/high-density regions, a density guided adaptive selection module is designed to automatically combine the predictions of CNN and Transformer. Moreover, to reduce the influences of annotation noise, we introduce a Correntropy based optimal transport loss. Extensive experiments on four challenging crowd counting datasets have validated the proposed method. Yuehai Chen, Jing Yang 0014, Badong Chen, Shaoyi Du |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Tolerating Annotation Displacement in Dense Object Counting via Point Annotation Probability MapabstractCounting objects in crowded scenes remains a challenge to computer vision. The current deep learning based approach often formulate it as a Gaussian density regression problem. Such a brute-force regression, though effective, may not consider the annotation displacement properly which arises from the human annotation process and may lead to different distributions. We conjecture that it would be beneficial to consider the annotation displacement in the dense object counting task. To obtain strong robustness against annotation displacement, generalized Gaussian distribution (GGD) function with a tunable bandwidth and shape parameter is exploited to form the learning target point annotation probability map, PAPM. Specifically, we first present a hand-designed PAPM method (HD-PAPM), in which we design a function based on GGD to tolerate the annotation displacement. For end-to-end training, the hand-designed PAPM may not be optimal for the particular network and dataset. An adaptively learned PAPM method (AL-PAPM) is proposed. To improve the robustness to annotation displacement, we design an effective transport cost function based on GGD. The proposed PAPM is capable of integration with other methods. We also combine PAPM with P2PNet through modifying the matching cost matrix, forming P2P-PAPM. This could also improve the robustness to annotation displacement of P2PNet. Extensive experiments show the superiority of our proposed methods. Yuehai Chen, Jing Yang 0014, Badong Chen, Shaoyi Du, Gang Hua 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Region-aware network: Model human's Top-Down visual perception mechanism for crowd counting
Yuehai Chen, Jing Yang 0014, Dong Zhang 0009, Badong Chen, Shaoyi Du |
Neural Networks | 1 |
| 2022 | FPGA-NHAP: A General FPGA-Based Neuromorphic Hardware Acceleration Platform With High Speed and Low PowerabstractSpiking neural network (SNN) can process discrete spikes and offers a high degree of real-time performance and excellent energy efficiency ratio. However, most current neuromorphic hardware platforms lack efficient driven algorithms and only support a single type of neuron model, which has slow speed and poor scalability. This paper proposes a general FPGA-based neuromorphic hardware acceleration platform (FPGA-NHAP), supporting the effective inference and acceleration of SNN network with low power, high speed and good scalability. First, a neuron computing unit is designed to simulate the both LIF and Izhikevich (IZH) neurons with the parallel spike caching and scheduling technique. Second, a novel integrated driven update algorithm is proposed to complete the spike encoding of external data, reducing the waiting time of neuron state update effectively. Third, the proposed platform is implemented using a RISC-V processor and a Xilinx FPGA, simulating 16,384 neurons and 16.8 million synapses with a power consumption of 0.535 W. Finally, two different three-layer SNN networks are deployed on the proposed platform for recognition tasks on the MNIST and Fashion-MNIST datasets, achieving the accuracy of 97.70%, 85.14% (LIF) and 97.81%, 83.16% (IZH), frame rates of 208 frame/s, 128 frame/s (LIF) and 206 frame/s, 141 frame/s (IZH), respectively. Yijun Liu 0005, Yuehai Chen, Wujian Ye, Yu Gui |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Robust Trajectory Prediction of Multiple Interacting Pedestrians via Incremental Active Learning
Yi Xi, Dongchun Ren, Mingxia Li, Yuehai Chen, Mingyu Fan, Huaxia Xia |
ICONIP (5) | 4 |