Yulong Yan

dblp:210/2307 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-9845-0899ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PicoSleepNet: An Ultra Lightweight Sleep Stage Classification by Spike Neural Network Using Single-Channel EEG Signal
abstract
This study introduces PicoSleepNet, an ultra-lightweight sleep stage classification method that utilizes a spiking neural network (SNN) with single-channel electroencephalogram (EEG) signals. Traditional methods use multi-bit Nyquist sampling and dense computing, which result in high complexity and power consumption, hindering their deployment on wearable devices. To address these limitations, we propose an innovative pipeline combining single-bit sub-Nyquist level-crossing sampling (LCS) and sparse computing based on SNN. First, LCS adaptively encodes EEG signals into event-driven spike sequences, reducing data volume by 6.98× while preserving essential signal characteristics compared to Nyquist sampling. Second, a sparse recurrent spiking neural network (RSNN) architecture, optimized by the masked backpropagation and sparse regularization (Masked-BPSR) technique, improves performance and reduces computational costs. Third, quantization-aware training (QAT) ensures that the model maintains high accuracy with low-bitwidth quantization, significantly reducing computational power consumption and enabling hardware-friendly deployment. Compared with current state-of-the-art sleep staging approaches, PicoSleepNet achieves competitive performance on three public datasets (Sleep-EDF-20, Sleep-EDF-78, and ISRUC-Sleep) with accuracies of 83.5%, 77.9%, 79.4% and macro-F1 scores of 75.2%, 68.1%, 77.2%, respectively. Meanwhile, by leveraging the computational sparsity design of RSNN and the joint optimization of Masked-BPSR and QAT, PicoSleepNet achieves an ultra-lightweight model with only 14.0-25.8 K parameters (reduced by nearly 2×) and 681.4-842.0 K operations (reduced by 27×), reducing computational power consumption by 1480×. This approach demonstrates the feasibility of deploying ultra-lightweight sleep staging systems in wearable devices and neuromorphic hardware, paving the way for broader applications in real-time health monitoring.
Shengnan Liu, Haoming Chu, Yukun Feng, Yulong Yan, Yuxiang Huan
IEEE J. Biomed. Health Informatics4
2025 NLBP: Efficient Training Spiking Neural Networks with Neuron-Level Back Propagation
abstract
Training brain-inspired spiking neural networks (SNNs) with multi-timestep backpropagation imposes substantial memory and computational overhead, hindering their scalability and deployment. To address these challenges, we propose a Neuron-Level Back Propagation (NLBP) method, which eliminates the need for temporal unfolding while maintaining competitive performance. We fit the relationship between a neuron’s average input current and its firing rate using an adaptive sigmoid function. Leveraging this mapping, we derive a neuron-level, single-step backpropagation rule that avoids explicit temporal unfolding. This design significantly reduces computational and memory costs when training deep, large-scale SNNs. Furthermore, NLBP unifies the training of integrate-and-fire (IF) and leaky integrate-and-fire (LIF) neuron models within a framework, and supports both soft and hard reset mechanisms, enhancing generality and practicality. Extensive experiments on pattern classification and object detection demonstrate that NLBP achieves competitive accuracy while reducing memory usage by up to 73.07% and training time by 66.04%.
Zikai Zhu, Yulong Yan, Longrun Xu, Jinqiao Yang, Lirong Zheng 0001, Zhuo Zou
ECAI2
2023 Practical periodic strategy for 40/100 Gbps Energy Efficient Ethernet
Wanchun Jiang, Renfu Yao, Kaiqin Liao, Yulong Yan, Jiawei Huang 0001, Weiping Wang 0003, Jianxin Wang 0001
Comput. Networks4
2023 A Domain-Specific Accelerator for Ultralow Latency Market Data Distribution System
abstract
Ultralow latency parsing of financial data is gaining significance in the high-frequency trading of the security exchange market. Hardware-aided systems exhibit superior improvement of latency over traditional software solutions, but flexibility may suffer when processing the financial protocol. This article presents a domain-specific accelerator for the market data distribution system, which integrates a financial information exchange adapted for streaming (FAST) decoder, a 10-Gbps network interface, and a high-speed PCIe host interface into a single field programmable gate array (FPGA) acceleration card. The proposed FAST decoder adopts the finite state machines-coordinated sequence mapping table to achieve run-time reconfigurability over fine-grained FPGA programming, and 16 fields can be decoded simultaneously in a pipelined manner, resulting in a decoding latency of only 33 ns. Evaluated on the Xilinx Alveo U200 acceleration card, this work improves the latency of decoding a FAST message by 26%–72% than state-of-the-art FPGA designs, and outperforms the software baseline by$>27\times$in terms of latency covering both decoding and communication.
Yuxiang Huan, Chen Ding 0010, Yulong Yan, Jianjun Cui, Jiachen Wang 0008, Chuhuang Cai, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Ind. Informatics4
2023 Consistent Low Latency Scheduler for Distributed Key-Value Stores
abstract
Nowadays, the distributed key-value stores have become the basic building block for large-scale cloud applications. In large-scale distributed key-value stores, many key-value access operations, which will be processed in parallel on different servers, are usually generated for a single end-user request. Accordingly, the completion time of an end-user request is determined by the last completed key-value access operation. Scheduling the order of serving key-value access operations can effectively reduce the completion times of end requests, thereby improving the user experience. However, existing scheduling algorithms hardly achieve consistent low latency due to the following challenges: the large overhead of cooperating clients and servers, the time-varying load and performance of servers, the traffic distribution can be either heavy-tailed or light-tailed and both the mean and the tail completion time are expected to be low. In this paper, we formalize the problem of scheduling key-value access operations and show it is NP-hard. Furthermore, we heuristically design the distributed adaptive scheduler (DAS), which distributively combines the largest remaining processing time last and the shortest remaining process time first algorithms. Theoretical analysis shows that DAS is adaptive to the time-varying traffic and server performance and can achieve consistent low mean and tail latency regardless of traffic distributions. Extensive simulations show that DAS reduces the mean request completion time by$17 \! \sim \! 50\%$with heavy-tailed traffic and$2 \! \sim 26 \! \%$with light-tailed traffic, while keeping the smallest tail completion time, compared to the default first come first served algorithm. Moreover, DAS outperforms the existing Rein-SBF algorithm under various scenarios.
Wanchun Jiang, Haoyang Li 0006, Yulong Yan, Fa Ji, Jiawei Huang 0001, Jianxin Wang 0001, Tong Zhang 0018
IEEE Trans. Parallel Distributed Syst.3
2022 A Hybrid-Mode On-Chip Router for the Large-Scale FPGA-Based Neuromorphic Platform
abstract
Large-scale neuromorphic computing requires the multi-chip network to provide high computing power. Efficient routing schemes and on-chip router design are necessary for handling various inter-chip transmission patterns. In this paper, we propose a hybrid-mode on-chip router that supports both multicast and unicast routing for the large-scale neuromorphic simulation. Two routing schemes, namely Cache-like Spike Weight Indexing and General Unicast Flow Control, are proposed to accommodate the chip-to-chip transmission of spike and non-spike data. This work is evaluated on a neuromorphic platform built with an$8\times 8$FPGA chips array. Running a simulation of 1M neurons at 200MHz, the proposed router achieves a processing latency of 25ns and a chip-to-chip latency of 287ns. Working in the unicast mode, the router can synchronize status flags of all chips within$5 ~\mu \text{s}$. Moreover, it reduces the peak spike traffic by 25.65% with the help of Load-aware Multicast Routing, compared with other multicast routing strategies.
Chen Ding 0010, Yuxiang Huan, Yulong Yan, Fanxi Yang, Lizheng Liu, Meigen Shen, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 Edge-Based Collaborative Training System for Artificial Intelligence-of-Things
abstract
The descending of intelligence from the cloud to the heterogeneous and low-power edge in the Artificial Intelligence-of-Things prevents uploading user-sensitive information to the cloud. It brings an urgent demand for deploying training tasks collaboratively in industrial scenarios to manage data locally. This article proposes an edge-based collaborative training system for the smart factory which harnesses the intelligence of edge devices by balancing the computational and communicational resources and improving system dependability. Two typical scenarios of parts recognition and defect inspection are evaluated as a case study with our system. The feasibility and dependability of the presented system are verified with a platform composed of eight high-performance (Nvidia Jetson Nano) and eight low-performance edge devices (Raspberry Pi 4B). The efficiency under tradeoff between computational resource and network condition constraints in a cluster is tested to simulate real-case performance in smart factory scenarios. Our platform reaches the peak performance of 1167 images/s training efficiency on ResNet32 under a 125 MB/s bandwidth. Experimental results demonstrate that the proposed design can collaboratively perform training tasks with optimized efficiency and provide dependable collaborations for system fault detection and cluster extension.
Yi Jin 0007, Yulong Yan, Yuxiang Huan, Jiawei Xu 0002, Shancang Li, Prosanta Gope, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Ind. Informatics3
2021 Cutting the Request Completion Time in Key-value Stores with Distributed Adaptive Scheduler
abstract
Nowadays, the distributed key-value stores have become the basic building block for large scale cloud applications. In large-scale distributed key-value stores, many key-value access operations, which will be processed in parallel on different servers, are usually generated for the data required by a single end-user request. Hence, the completion time of the end request is determined by the last completed key-value access operation. Accordingly, scheduling the order of key-value access operations of different end requests can effectively reduce their completion time, improving the user experience. However, existing algorithms are either hard to employ in distributed key-value stores due to the relatively large cooperation overhead for centralized information or unable to adapt to the time-varying load and server performance under different traffic patterns. In this paper, we first formalize the scheduling problem for small mean request completion time. As a step further, because of the NP-hardness of this problem, we heuristically design the distributed adaptive scheduler (DAS) for distributed key-value stores. DAS reduces the average request completion time by a distributed combination of the largest remaining processing time last and shortest remaining process time first algorithms. Moreover, DAS is adaptive to the time-varying server load and performance. Extensive simulations show that DAS reduces the mean request completion time by more than 15 ~ 50% compared to the default first come first served algorithm and outperforms the existing Rein-SBF algorithm under various scenarios.
Wanchun Jiang, Haoyang Li 0006, Yulong Yan, Fa Ji, Jianxin Wang 0001, Tong Zhang 0018
ICDCS3
2021 Self-aware distributed deep learning framework for heterogeneous IoT edge devices
Yi Jin 0007, Jiawei Cai, Jiawei Xu 0002, Yuxiang Huan, Yulong Yan, Yongliang Guo, Lirong Zheng 0001, Zhuo Zou
Future Gener. Comput. Syst.5
2021 An IoT-Based Anti-Counterfeiting System Using Visual Features on QR Code
abstract
This article presents an Internet-of-Things (IoT) anti-counterfeiting system that uses visual features combined with the quick response (QR) code. The visual features guarantee the authenticity of a product with the QR code for tracking and tracing. Two visual features, i.e., natural texture features and printed micro features are exploited in the proposed system. The natural texture features use the texture of fiber paper to achieve physical unclonable function (PUF), while the micro features are artificially generated for improved industrial manufacturability and reliability. Features are generated and registered in the production phase when the QR code is printed. In the anti-counterfeiting verification phase, the feature obtained through the feature extraction algorithm is compared with the record to calculate similarity, which indicates the verification result. Such an approach is fully compatible with the QR code-based logistic process without any additional manufacturing cost. A user-friendly application has been developed on a mobile platform that facilitates easy-to-use and affordable devices for verification, such as a mobile phone or a handheld code reader. The experimental results show 99.6% and 99.9% accuracy of anti-counterfeiting verification for texture features and micro features, respectively. The system with corresponding algorithms and software has been demonstrated in real-life products.
Yulong Yan, Zhuo Zou, Yu Gao 0042, Lirong Zheng 0001
IEEE Internet Things J.1
2020 PS: Periodic Strategy for the 40-100Gbps Energy Efficient Ethernet
abstract
The 40-100Gbps Energy Efficient Ethernet (EEE) standardized by IEEE 802.3bj contains not only the DeepSleep mode, which is 10% of the normal power consumption, but also the FastWake mode, which has much less state transition time and 70% of the normal power consumption. Correspondingly, the strategy of determining the state transitions is crucial to both the power consumption and the incurred latency of frames. As EEE applications always place tail latency constraint on frame transmission, existing strategies for the 40-100Gbps EEE are rarely configured with proper parameters for good power saving under the given tail latency constraint, especially in face of the variable traffic load. To address this problem, we design the Periodic Strategy (PS) for the 40-100Gbps EEE. Specifically, PS works periodically and optimizes the power saving independently in each cycle. In this way, PS decouples the latency constraint and energy saving. 1) It limits the largest incurred latency by the length of the cycle to satisfy the given latency constraint. 2) It achieves good power saving by selecting the proper low power state based on the prediction of incoming frames and entering the selected state just once in each cycle. What’s more, the good performance of PS is insensitive to the traffic pattern. Extensive simulations confirm that PS achieves better energy saving than all the existing strategies under the given latency constraint, under several kinds of different traffic patterns.
Wanchun Jiang, Kaiqin Liao, Yulong Yan, Jianxin Wang 0001
ICPP3