Jiawei Nian

dblp:315/6901 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-9693-3467ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deep reinforcement learning for energy-efficient workflow scheduling in edge computing
Mengyao Wen, Xiufeng Liu 0001, Xin Ning 0001, Cong Liu 0012, Jiawei Nian, Long Cheng 0003
Comput. Networks6
2026 Prototype Retrieval-Augmented Federated Learning System for Robust Intrusion Detection
abstract
Detecting malicious attacks is essential for protecting computer systems and ensuring device security. Federated Learning (FL)-based Intrusion Detection Systems (IDS) have emerged as promising solutions, enabling multiple clients (i.e., data owners) to collaboratively train intrusion detection models without sharing private data. However, current FL studies typically assume that each client’s training and test label distribution is identical. This assumption is overly idealistic and rarely holds in real-world scenarios, leading to suboptimal performance when label distribution shifts occur between the training and testing data. To address this challenge, we propose FedPRO, a plug-and-play framework designed to improve the test-time performance of existing FL methods, without modifying their original training pipelines or fine-tuning the trained FL models. Specifically, we develop a unique prototype generation and optimization mechanism to produce semantically meaningful class prototypes. These prototypes constitute a prototype memory bank, serving as an external knowledge repository. At test time, a prototype retrieval-augmented inference strategy is employed to query relevant prototypes and refine predictions on each client, effectively alleviating the label distribution shift issues and boosting prediction accuracy. We evaluate FedPRO by integrating it with various off-the-shelf FL methods on benchmark datasets. Extensive results consistently demonstrate its effectiveness in diverse settings. Notably, applying FedPRO to the state-of-the art method FedDBE improves its test accuracy from 79.25% to 86.66% on the CICIDS-2018 dataset, while introducing only approximately 32KB of additional communication overhead.
Hanlin Zhou, Huiru Yan, Jiawei Nian, Cong Liu 0012, Ying Wang 0001, Georgios Theodoropoulos 0001, Long Cheng 0003
IEEE Trans. Computers3
2026 MCU-MixQ: A HW/SW Co-optimized Mixed-precision Neural Network Design Framework for MCUs
abstract
Mixed-precision neural network (MPNN) that utilizes just enough data width for the neural network processing is an effective approach to meet the stringent resources constraints including memory and computing of MCUs. Nevertheless, there is still a lack of sub-byte and mixed-precision SIMD operations in MCU-class ISA and the limited computing capability of MCUs remains underutilized, which further aggravates the computing bound encountered in neural network processing. As a result, the benefits of MPNNs cannot be fully unleashed. In this work, we propose to pack multiple low-bitwidth arithmetic operations within a single instruction multiple data (SIMD) instructions in typical MCUs, and then develop an efficient convolution operator by exploring both the data parallelism and computing parallelism in convolution along with the proposed SIMD packing. Finally, we further leverage Neural Architecture Search (NAS) to build a HW/SW co-designed MPNN design framework, namely MCU-MixQ. This framework can optimize both the MPNN quantization and MPNN implementation efficiency, striking an optimized balance between neural network performance and accuracy. According to our experiment results, MCU-MixQ achieves 2.1× and 1.4× speedup over CMix-NN and MCUNet respectively under the same resource constraints. MCU-MixQ is also open sourced on GitHub. 1
Junfeng Gong, Long Cheng 0003, Jiawei Nian, Cheng Liu 0008, Huawei Li 0001
ACM Trans. Embed. Comput. Syst.3
2025 Spatio-Temporal Feature Extraction for Predicting Large-Scale Train Delay Propagation in High-Speed Railway Networks
abstract
The safe, punctual, and reliable operation of high-speed railway (HSR) networks is crucial for ensuring system efficiency and enhancing passenger experience. However, due to the complexity of HSR systems and long operation routes, failures in system components can lead to unexpected incidents, causing deviations from the scheduled operations and, subsequently, delays. These delays, particularly large delays, significantly impact the overall performance and passenger experience of the network. Thus, accurate prediction of train delays, especially in large-delay scenarios, is essential for optimizing train scheduling and restoring normal operations in a timely manner. To address this challenge, this study proposes a deep learning architecture based on spatio-temporal feature extraction for accurate train delay prediction. A graph attention network-long short-term memory block is introduced to capture the spatio-temporal evolution features of different trains. In addition, a sequence forecasting approach and a mixture of experts module are integrated to model the complex relationships between the target train’s delay and its previous states. The delay evolution features are incorporated into the loss function, allowing for more accurate predictions. Experimental results show that the proposed model outperforms the baseline models, achieving at least an improvement of 50.62% and 27.89% in root-mean-squared error and mean absolute error, respectively. When the error tolerance is set within 3 min, the prediction accuracy reaches 96.68%. Experimental results demonstrate the superior performance of the proposed model.
Xingtang Wu, Fang Fang 0007, Min Zhou 0003, Jiawei Nian, Hairong Dong 0001
IEEE Trans. Ind. Informatics5
2024 ReBEC: A replacement-based energy-efficient fault-tolerance design for associative caches
Xin Gao 0016, Naiyuan Cui, Jiawei Nian, Zongnan Liang, Jiaxuan Gao, Hongjin Liu, Mengfei Yang
Future Gener. Comput. Syst.3
2024 Improving the Security of Audio CAPTCHAs With Adversarial Examples
abstract
CAPTCHAs (completely automated public Turing tests to tell computers and humans apart) have been the main protection against malicious attacks on public systems for many years. Audio CAPTCHAs, as one of the most important CAPTCHA forms, provide an effective test for visually impaired users. However, in recent years, most of the existing audio CAPTCHAs have been successfully attacked by machine learning-based audio recognition algorithms, showing their insecurity. In this article, a generative adversarial network (GAN)-based method is proposed to generate adversarial audio CAPTCHAs. This method is implemented by using a generator to synthesize noise, a discriminator to make it similar to the target and a threshold function to limit the size of the perturbation; then, the synthetic perturbation is combined with the original audio to generate the adversarial audio CAPTCHA. The experimental results demonstrate that the addition of adversarial examples can greatly reduce the recognition accuracy of automatic models and improve the robustness of different types of audio CAPTCHAs. We also explore ensemble learning strategies to improve the transferability of the proposed adversarial audio CAPTCHA methods. To investigate the effect of adversarial CAPTCHAs on human users, a user study is also conducted.
Ping Wang 0027, Haichang Gao, Zhongni Yuan, Jiawei Nian
IEEE Trans. Dependable Secur. Comput.5
2023 An Efficient Fault-Tolerant Protection Method for L0 BTB
abstract
Branch prediction structures are increasingly used in space processors due to their crucial role in improving processor performance. Due to radiation effects such as Single Event Upset (SEU) causing system failures, it is necessary to provide protection techniques for branch prediction modules against complex spatial environments. In this paper, we proposed a simple, efficient, low-power, and fault-tolerant design scheme for L0 BTB consisting of a master-slave and a check-decision module. The strategy sets up the L0 BTB structure as a master-slave structure and adds an error check to detect single-bit errors. Experimental results show that we achieve 100% fault tolerance without a loss hit rate. The increase in resource usage is 1.1x, and the path delay increases by 8.1%, superior to other methods. The L0 BTB is a fully-associative structure and multiple entries are accessed simultaneously, which introduces significant power consumption. We added a low-power design for the master-slave module to reduce the query power consumption by 65.9%. In addition, we also applied the scheme to the RAS module. The experimental results demonstrate that our approach is an efficient, generic, fault-tolerant design scheme that can be deployed to different register files.
Jiawei Nian, Zongnan Liang, Hongjin Liu, Mengfei Yang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 C-DMR: a cache-based fault-tolerant protection method for register file
Zongnan Liang, Jiawei Nian, Hongjin Liu, Xuru Wang, Mengfei Yang
J. Supercomput.2
2022 A deep learning-based attack on text CAPTCHAs by using object detection techniques
abstract
Abstract Text‐based CAPTCHAs have been widely deployed by many popular websites, and many have been attacked. However, most previous cracks were based on classification algorithms that typically rely on a series of preprocessing operations or on many training samples, thus making such attacks complicated and costly. In this study, a simple, generic, fast and end‐to‐end attack based on advanced object detection technologies is introduced. The proposed attack combines a feature extraction module, a character location and recognition module and a coordinate matching module. The experiments show that the attack can break a wide range of real‐world text CAPTCHAs deployed by the 50 most popular websites on Alexa.com and that the method achieves a high attack accuracy with only 2000 samples at an attack speed of less than 0.10 s. The attack was also evaluated on four click‐based CAPTCHAs that cannot be attacked in the end‐to‐end manner used by previous attacks, and the results demonstrated that within one step, the proposed approach achieves high success rates on both click‐based CAPTCHAs and schemes based on large‐scale character sets, such as Chinese character sets.
Jiawei Nian, Ping Wang 0027, Haichang Gao
IET Inf. Secur.1