Jinyu Zhan

dblp:10/6219 · DBLP profile ↗
← Back
40ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0002-0214-7124ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 10 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Data poisoning-based backdoor attacks against supervised learning rules of Spiking Neural Networks
Lingxin Jin, Wei Jiang 0016, Jinyu Zhan, Meiyu Lin, Letian Chen, Boran Quan, Lin Zuo, Xingzhi Zhou 0001, Maregu Assefa, Naoufel Werghi
J. Syst. Archit.3
2025 A Computation-Quantized Training Framework to Generate Accuracy Lossless QNNs for One-Shot Deployment in Embedded Systems
abstract
Quantized Neural Networks (QNNs) have received increasing attention, since they can enrich intelligent applications deployed on embedded devices with limited resources, such as mobile devices and AIoT systems. Unfortunately, the numerical and computational discrepancies between training systems (i.e., servers) and deployment systems (e.g., embedded ends) may lead to large accuracy drop for QNNs in real deployments. We propose a Computation-Quantized Training Framework (CQTF), which simulates deployment-time fixed-point computation during training to enable one-shot, lossless deployment. The training procedure of CQTF is built upon a well-formulated quantization-specific numerical representation that quantifies both numerical and computational discrepancies between training and deployment. Leveraging this representation, forward propagation executes all computations in quantization mode to simulate deployment-time inference, while backward propagation identifies and mitigates gradient vanishing through an efficient floating-point gradient update scheme. Benchmark-based experiments demonstrate the efficiency of our approach, which can achieve no accuracy loss from training to deployment. Compared with existing five frameworks, the deployed accuracy of CQTF can be improved by up to 18.41%.
Xingzhi Zhou 0001, Wei Jiang 0016, Jinyu Zhan, Lingxin Jin, Lin Zuo
IEEE Trans. Computers3
2025 Accelerate Point Cloud Structuring for Deep Neural Networks via Fast Spatial-Searching Tree
abstract
Due to the disorder of points, point clouds need to be structured by sampling and neighbor query before feeding to Deep Neural Networks (DNNs). Structuring point clouds costs high computation overhead, which limits the deployment of DNNs on embedded devices such as autonomous vehicles and robots. To address this problem, we design a novel data structure, i.e., Fast Spatial-Searching Tree (FSSTree), to accelerate point cloud structuring for DNNs on embedded devices. The FSSTree is constructed based on density distribution of point clouds to achieve semantic segmentation, which can guarantee that points with similar spatial positions are stored in adjacent storage sets. Based on FSSTree, we propose a point-sparsity-aware sampling method and a leafwise k-nearest neighbor query method to reduce the computation overhead of structuring point clouds. Meanwhile, the point-sparsity-aware sampling method achieves fair sampling on both dense and sparse parts, which can overcome the nonuniform distribution of point clouds caused by occlusion, lighting and other factors. The leafwise k-nearest neighbor query method skips a large number of dissimilar points to quickly obtain the neighbor points, which can significantly reduce the search scope. We also present a layerwise self-pruning algorithm to automatically adjust the FSSTree after each layer’s operation to match the hierarchical architecture of DNNs. Finally, we conduct extensive experiments on KITTI, S3DIS and ModelNet40 datasets and three devices (including an RTX 3090 server, a Jetson AGX Xavier and an Apple M2). The experimental results demonstrate the efficiency of our approach, which can reduce the time overhead by up to 97.46% compared with the other five methods. The code is released athttps://github.com/EmbeddedAILab-UESTC/fsstree.
Jinyu Zhan, Shiyu Zou, Wei Jiang 0016, Youyuan Zhang, Suidi Peng, Ying Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 A Lightweight-Window-Portion-Based Multiple Imputation for Extreme Missing Gaps in IoT Systems
abstract
Intelligent techniques, including artificial intelligence and deep learning, normally perform on complete data without missing data. Multiple imputation is indispensable for addressing missing data resulting in unbiased estimates and dealing with uncertainty by providing more valid results. Most state-of-the-art techniques focus on high-missing rates (around 50%–60%) and short missing gaps, while imputation for extreme missing gaps and missing rates is an important challenge for multivariate time-series data generated through the Internet of Things (IoT). Hence, we propose an lightweight-window-portion-based multiple imputation (LWPMI) based on multivariate variables, correlation, data fusion, regression, and multiple imputations. We conduct extensive experiments by generating extreme missing gaps and high-missing rates ranging from 10% to 90% on data generated by sensors. We also investigate different sets of feature to examine how LWPMI works when features have high, weak, or a mixture of high and weak correlation. All the obtained results prove LWPMI outperforms baseline techniques in preserving pattern, structure, and trend in both 90% extreme missing gap and missing rates.
Deepak Adhikari, Wei Jiang 0016, Jinyu Zhan, Maregu Assefa, Hadi Akbarzadeh Khorshidi, Uwe Aickelin, Danda B. Rawat
IEEE Internet Things J.3
2024 Highly Evasive Targeted Bit-Trojan on Deep Neural Networks
abstract
Bit-Trojan attacks based on Bit-Flip Attacks (BFAs) have emerged as severe threats to Deep Neural Networks (DNNs) deployed in safety-critical systems since they can inject Trojans during the model deployment stage without accessing training supply chains. Existing works are mainly devoted to improving the executability of Bit-Trojan attacks, while seriously ignoring the concerns on evasiveness. In this paper, we propose a highly Evasive Targeted Bit-Trojan (ETBT) with evasiveness improvements from three aspects, i.e., reducing the number of bit-flips (improving executability), smoothing activation distribution, and reducing accuracy fluctuation. Specifically, key neuron extraction is utilized to identify essential neurons from DNNs precisely and decouple the key neurons between different classes, thus improving the evasiveness regarding accuracy fluctuation and executability. Additionally, activation-constrained trigger generation is devised to eliminate the differences between activation distributions of Trojaned and clean models, which enhances evasiveness from the perspective of activation distribution. Ultimately, the strategy of constrained target bits search is designed to reduce bit-flip numbers, directly benefits the evasiveness of ETBT. Benchmark-based experiments are conducted to evaluate the superiority of ETBT. Compared with existing works, ETBT can significantly improve evasiveness-relevant performances with much lower computation overheads, better robustness, and generalizability. Our code is released athttps://github.com/bluefier/ETBT.
Lingxin Jin, Wei Jiang 0016, Jinyu Zhan, Xiangyu Wen 0001
IEEE Trans. Computers3
2024 Improving Dependability of Distributed Real-Time Applications via Safety and Security Co-Design
abstract
With the increasing deployment in mission-critical domains, it is of foremost importance to improve dependability of distributed real-time applications for cyber–physical systems (CPSs) with safety & security-critical threats. Different from existing works addressing the security or safety design separately, this article makes efforts to achieve the safety and security co-design from system-level perspective, especially considering the interplay between fault tolerance and security harden techniques. To guarantee the safety of real-time applications, fault-tolerant techniques, e.g., task re-execution and active replica, are leveraged to tolerate faults in task executions. To improve the security of distributed applications, cryptography is deployed to resist confidentiality attacks on messages delivered over the communication media. We analyze the impact of task’s fault tolerance on secure message communication, and then formulate the design problem as a multiobjective optimization problem, i.e., to minimize the failure probability and security vulnerability of the application while subject to given fault-tolerant constraints, execution constraints and deadline constraints. Since the optimization problem is NP-hard, we then propose an improved multiobjective optimization algorithm, called decomposition-based dependability co-optimization (DeDeCo) algorithm, to search for the optimal Pareto solutions of security and reliability harden assignments for messages and tasks, respectively. Extensive experiments and an industrial case evaluate the efficiency of DeDeCo, indicating that our design and optimization algorithm are suitable for improving the dependability of real-time applications running on security & safety-critical CPSs.
Jinyu Zhan, Wei Jiang 0016, Xinke Liao, Deepak Adhikari
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Detecting Spoofed Noisy Speeches via Activation-Based Residual Blocks for Embedded Systems
abstract
Spoofed noisy speeches seriously threaten the speech-based embedded systems, such as smartphones and intelligent assistants. Consequently, we present an anti-spoofing detection model with activation-based residual blocks to identify spoofed noisy speeches with the requirements of high accuracy and low time overhead. Through theoretic analysis of noise propagation on shortcut connections of traditional residual blocks, we observe that different activation functions can help reducing the influence of noise under certain situations. Then, we propose a feature-aware activation function to weaken the influence of noise and enhance the anti-spoofing features on shortcut connections, in which a fine-grained processing is designed to remove noise and strengthen significant features. We also propose a variance-increasing-based optimization algorithm to find the optimal hyperparameters of the feature-aware activation function. Benchmark-based experiments demonstrate that the proposed method can reduce the average equal error rate of anti-spoofing detection from 21.72% to 4.51% and improve the accuracy by up to 37.06% and save up to 91.26% of time overhead on Jetson AGX Xavier compared with ten state-of-the-art methods.
Jinyu Zhan, Suidi Peng, Wei Jiang 0016
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Audio-Visual Contrastive and Consistency Learning for Semi-Supervised Action Recognition
abstract
Semi-supervised video learning is an increasingly popular approach for improving video understanding tasks by utilizing large-scale unlabeled videos along with a few labels. Recent studies have shown that multimodal contrastive learning and consistency regularization are effective techniques for generating high-quality pseudo-labels for semi-supervised action recognition. However, existing pseudo-labeling approaches are solely based on the model's class predictions and can suffer from confirmation biases due to the accumulation of false predictions. To address this issue, we propose exploiting audio-visual feature correlations to achieve high-quality pseudo-labels instead of relying on model confidence. To achieve this goal, we introduce Audio-visual Contrastive and Consistency Learning (AvCLR) for semi-supervised action recognition. AvCLR generates reliable pseudo-labels from audio-visual feature correlations using deep embedded clustering to mitigate confirmation biases. Additionally, AvCLR introduces two contrastive modules: intra-modal contrastive learning (ImCL) and cross-modal contrastive learning (XmCL) to discover complementary information from audio-visual alignments. The ImCL module learns informative representations within audio and video independently, while the XmCL module aims to leverage global high-level features of audio-visual information. Furthermore, the XmCL is constrained by introducing intra-instance negatives from one modality to the other. We jointly optimize the model with ImCL, XmCL, and consistency regularization in an end-to-end semi-supervised manner. Experimental results have demonstrated that the proposed AvCLR framework is effective in reducing confirmation biases and outperforms existing confidence-based semi-supervised action recognition methods.
Maregu Assefa, Wei Jiang 0016, Jinyu Zhan, Kumie Gedamu, Getinet Yilma, Melese Ayalew, Deepak Adhikari
IEEE Trans. Multim.3
2024 Critical Path-Based Backdoor Detection for Deep Neural Networks
abstract
Backdoor attack to deep neural networks (DNNs) is among the predominant approaches to bring great threats into artificial intelligence. The existing methods to detect backdoor attacks focus on the perspective of distributions in DNNs, however, limited by its ability of generalization across DNN models. In this article, a critical-path-based backdoor detector (CPBD) is proposed, which approaches to detect backdoor attacks via DNN's interpretability. CPBD is designed to efficiently discover the characteristics of backdoors, which distinguish the critical paths in the attacked DNNs. To deal with the intractably large number of neurons, we propose to simplify the neurons, and the preserved key nodes are integrated into a set of critical paths. Thus, a DNN model can be formulated as a combination of several critical paths. Afterward, the detection of backdoors is performed based on the analysis of critical paths corresponding to different classes. Then, combining all the above steps, the CPBD algorithm is integrated to present the results in a standard and systematic manner. In addition, CPBD is able to locate neurons associated with malicious triggers, the combination of which is named as trigger propagation path. Extensive experiments are conducted, which testify the efficiency of the proposed method on multiple DNNs and different trigger sizes.
Wei Jiang 0016, Xiangyu Wen 0001, Jinyu Zhan, Xupeng Wang 0001, Chen Bian
IEEE Trans. Neural Networks Learn. Syst.3
2023 DESCO: Decomposition-Based Co-Design to Improve Fault Tolerance of Security-Critical Tasks in Cyber Physical Systems
abstract
Confidentiality-Specific Faults (CSFs) will put cyber physical systems in threat, since they can result in corrupted information or even retrieve the cryptographic key of security-critical applications. In this paper, we will look into fault-tolerant co-design optimization for security-critical cyber physical systems with resource constraints, such that the encryption/decryption of confidential messages are protected against transient CSF faults. We consider imperfect fault detection mechanisms to identify transient CSF faults happened on confidentiality protection, and utilize duplication code to recovery from such faults. We utilize FPGA to accelerate the executions of security tasks, reducing the overheads of fault-tolerant implementations. The system-level design problem is formulated as a two-objective optimization problem, i.e., to minimize the average reliability degradation of the fault tolerant assignments and to minimize the balanced degree of the reliability degradation, subject to available FPGA budget, deadline, and application execution constraints. Since finding Pareto-optimal solutions is NP-hard, we propose an improved multi-objective optimization algorithm, called DEcomposition-based Security Co-design Optimization (DESCO), to search for Pareto-optimal solutions of fault-tolerant assignments. Experimental results demonstrate that DESCO is effective and can outperform other candidates, proving that our approach is promising in dealing with system-level optimization problem for security-critical applications on cyber physical systems.
Wei Jiang 0016, Xinke Liao, Jinyu Zhan, Deepak Adhikari
IEEE Trans. Computers3
2023 Query-Efficient Generation of Adversarial Examples for Defensive DNNs via Multiobjective Optimization
abstract
Due to the inherent vulnerability of deep neural networks (DNNs), the adversarial example (AE) attack has become a serious threat to intelligent systems, e.g., the failure cause of an image classification system. Different to existing works, in this article we are interested in the generation of AEs for DNNs with defensive mechanisms. To make the attack more practical, we exploit a query-based method to generate image AEs in a black-box attack setting. Considering that the generation of AEs is inherently a constrained optimization problem, this article first formulates three objectives regarding defensive DNNs, i.e., attack effectiveness, attack evasiveness and attack coverage. Then, this article proposes a query-efficient AE attack based on the genetic algorithm (GA) and particle swarm optimization (PSO) to address the perturbation optimization problem. To improve the efficiency of search and query, AE-specific operators including block-level and pixel-level crossovers, discrete perturbation mutation and direction-driven reproduction are designed within the GA-based search framework. In addition, predication-based adaptation of reproduction-related parameters is implemented to speed up the search convergence. PSO-based jumping process is further devised to avoid stuck in local optimum. Benchmark-based experiments evaluated the efficiency of our method, which can achieve an attack success rate of 100% with averagely 52.95% reduced queries in contrast to existing black-box attacks on nondefensive models. For defensive DNN models, our method can obtain top attack performance with the query reduction up to 70.92% comparing with the candidates.
Wei Jiang 0016, Shen You, Jinyu Zhan, Xupeng Wang 0001, Deepak Adhikari
IEEE Trans. Evol. Comput.3
2022 NIC-QF: A design of FPGA based Network Interface Card with Query Filter for big data systems
Jinyu Zhan, Wei Jiang 0016, Ying Li 0130, Junting Wu, Jianping Zhu 0003, Jinghuan Yu
Future Gener. Comput. Syst.1
2022 Interpretability-Guided Defense Against Backdoor Attacks to Deep Neural Networks
abstract
As an emerging threat to deep neural networks (DNNs), backdoor attacks have received increasing attentions due to the challenges posed by the lack of transparency inherent in DNNs. In this article, we develop an efficient algorithm from the interpretability of DNNs to defend against backdoor attacks to DNN models. To extract critical neurons, we deploy sets of control gates following neurons in layers, and the function of a DNN model can be interpreted as semantic sensitivities of neurons to input samples. A backdoor identification approach, derived from the activation frequency distribution on critical neurons, is proposed to reveal anomalies of particular neurons produced by backdoor attacks. Subsequently, a feasible and fine-grained pruning strategy is introduced to eliminate backdoors hidden in DNN models, without the need of retraining. Extensive experiments demonstrate that the proposed algorithm can identify and eliminate malicious backdoors efficiently in both single-target and multitarget scenarios with the performance of a DNN model retained to a large extent.
Wei Jiang 0016, Xiangyu Wen 0001, Jinyu Zhan, Xupeng Wang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Accelerating Queries of Big Data Systems by Storage-Side CPU-FPGA Co-Design
abstract
As a promising technology of big data systems, storage and computing separated architecture has attracted increasing attention of famous companies, such as Tencent, IBM, Facebook, and Microsoft. Under this new architecture, conventional query engines like Hive and Presto choose all the original data from storage nodes and send them to computing nodes to be filtered, causing high data transmission overhead and great I/O bandwidth fluctuation. To address this problem, we design a novel data processing framework to prefilter data on storage side, and then propose a CPU-FPGA (field-programmable gate array) co-design to accelerate the queries with the purpose of reducing the communication overheads and the workloads of computing nodes. To obtain the optimal efficiency of CPU-FPGA co-processing, a workload-aware task scheduler is designed to allocate query tasks to CPU or FPGA according to the estimation of the filtering data size and processing time of query tasks. A data projection scheme is designed to support data in RCFile format which is widely used in modern systems, such as Tencent and Facebook applications. To make full use of the high parallelism of FPGA, we formulate the SQL conditions of combined predicates into Boolean parameters, and design two filtering schemes on FPGA (i.e., parallel sequential filter for fix-length data type and parallel pipeline filter for variable-length data type). Experiments on the TPC-H benchmark and Tencent data set demonstrate the efficiency of our approach, which can save up to 72.28% and 80.16% of time overheads compared with Presto and Hive, respectively.
Jinyu Zhan, Wei Jiang 0016, Ying Li 0130, Junting Wu, Jianping Zhu 0003, Jinghuan Yu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Detecting Spoofed Speeches via Segment-Based Word CQCC and Average ZCR for Embedded Systems
abstract
Intelligent speech recognition is increasingly used in embedded systems, which is also seriously threatened by malicious speech spoofing attacks. Different from the conventional methods, this article proposes a segment-based anti-spoofing detection (SASD) method for the quick detection of spoofed speeches against embedded speech recognition, which focuses on the anti-spoofing features rather than the contexts of speeches and the voiceprints of speakers. The speeches are divided into word segments and silent segments. Based on constant$Q$cepstral coefficients (CQCCs), a word CQCC (WCQCC) extraction is first designed for the word segments of speeches. Then, based on short-term zero crossing rate (ZCR), an average ZCR (AZCR) extraction is devised for the silent segments. Combining the WCQCC of word segments and AZCR of silent segments, a biased decision strategy is proposed to quickly determine whether a speech is spoofed. Based on ASVspoof 2021 datasets, extensive experiments are conducted to evaluate the effectiveness of the proposed method. Specifically, our SASD can improve the accuracy of anti-spoofing detection by up to 33.47% and save up to 69.10% of time overhead on embedded devices compared with the existing methods.
Jinyu Zhan, Zhibei Pu, Wei Jiang 0016, Junting Wu, Yongjia Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Improving Fault Tolerance for Reliable DNN Using Boundary-Aware Activation
abstract
In this article, we approach to construct reliable deep neural networks (DNNs) for safety-critical artificial intelligent applications. We propose to modify rectified linear unit (ReLU), a commonly used activation function in DNNs, to tolerate the faults incurred by bit-flip perturbation on weights. Through theoretic analysis of the fault propagation in the layers with ReLU activation, we observe that bounding the output of ReLU activation can help to tolerate the weight faults. Then, we propose a novel ReLU design called boundary-aware ReLU (BReLU) to improve the reliability of DNNs, in which an upper bound of ReLU is determined such that the deviation between the boundary and original outputs cannot affect the final result. We propose a gradient-ascent-based algorithm to find the boundaries for BReLU activations of all DNN layers. Without retraining the network, our approach is cost effective and practical when deployed in safety-critical artificial intelligent systems. Detailed experiments and real-life application benchmarking demonstrate that our approach can improve the accuracy of DNN VGG16 from 16.7% to 82.6% on average assuming the practical weight faults, with only 13% memory and 2.78% time overhead, respectively.
Jinyu Zhan, Ruoxu Sun, Wei Jiang 0016, Yucheng Jiang, Xunzhao Yin, Cheng Zhuo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Layerwise Security Protection for Deep Neural Networks in Industrial Cyber Physical Systems
abstract
Although deep neural networks (DNNs) have been increasingly applied in industrial cyber physical systems (ICPSs), they are vulnerable to security attacks due to the tight interaction between cyber elements and physical elements. In this article, we aim to protect the core IP of DNNs, i.e., the model weights, against security attacks. Different from conventional approaches, a layerwise protection framework is proposed to ensure the confidentiality of DNN model weights during the inference procedure such that the security quality is maximized, while satisfying the latency constraint of the DNN task. Based on the layerwise execution characteristics of DNN tasks, the encrypted layer-related weights are decrypted and fed to the next layer of DNN in plaintext. CPU-field programmable gate array (FPGA) coscheduling is considered to accelerate the execution of confidentiality protection, where CPU is utilized to conduct the decryption of weights and FPGA is used to perform the layer execution of DNN. Considering to provide optimal confidential protection for each layer, the problem is transformed into a quality of security maximization problem subject to layerwise execution constraint and deadline constraint of the DNN application. Due to the problem being NP-hard, a fast approximation algorithm is proposed to obtain the near-optimal solution under given real-time and security constraints. Extensive experiments and a real-life ICPS application evaluate the efficiency of the proposed techniques.
Wei Jiang 0016, Jinyu Zhan, Di Liu 0002, Jiafu Wan
IEEE Trans. Ind. Informatics3
2021 Detecting deepfake videos by visual-audio synchronism: work-in-progress
abstract
Different to traditional works on frame-level features and temporal characteristics, we propose a deepfake video detection method based on visual-audio synchronism, which compares the audio stream and the visual stream by an improved siamese neural network. We combine the audio stream and visual stream as a 2-channel input and design a 2-branches network to achieve the visual-audio synchronism detection. Preliminary experiments demonstrate the efficiency of the proposed method, which can achieve the highest accuracy compared with other existing methods.
Zhufeng Fan, Jinyu Zhan, Wei Jiang 0016
EMSOFT2
2021 Improving fault tolerance of DNNs through weight remapping based on gaussian distribution: work-in-progress
abstract
In this paper, we approach to improve the fault tolerance of Deep Neural Networks (DNNs) for safety-critical artificial intelligent applications. We propose to remap the range of 32-bit float to weights to reduce the influence of invalid weights caused by bit-flip faults. From preliminary experiments, we observe that weakening bit-flip faults which make positive weights larger can help to improve the reliability of DNNs. Then, we propose a gaussian distribution based mapping method to prevent weights from being influenced by bit-flip faults, in which a novel function is formulated to remap the relation between 32-bit float and the values of weights. Extensive experiments demonstrate that our approach can improve the accuracy of VGG16 from 13.5% to 80.5%, which is better than the other six tolerance approaches of DNNs.
Ruoxu Sun, Jinyu Zhan, Wei Jiang 0016, Yucheng Jiang
EMSOFT2
2021 Generative strategy based backdoor attacks to 3D point clouds: work-in-progress
abstract
3D deep learning has been applied in safety-critical scenarios, e.g., autonomous driving. Several works have raised the security problems of 3D deep learnings mainly from the perspective of adversarial attacks. In this paper, we propose a novel backdoor attack method to threaten 3D deep learning without the original training data. Several neurons are selected and made sensitive to backdoor triggers. The backdoor triggers are generated by reversing neural network, and the shape of which is constrained to map the objects in the physical world. Sufficient training data can be also generated by reverse engineering. Finally, retraining with the generated 3D trigger and training data is applied to inject backdoors, which is in no need of accessing the original training process and data.
Xiangyu Wen 0001, Wei Jiang 0016, Jinyu Zhan, Chen Bian
EMSOFT3
2021 Work-in-Progress: Improving Resilience of Distributed Real-Time Applications via Security and Fault Tolerance Co-Design
abstract
In this paper, we focus on improving the resilience of distributed real-time applications for Cyber Physical Systems. To guarantee the safety related resilience, fault-tolerant techniques, e.g., task re-execution and active replica, are leveraged to tolerate faults in task executions. To provide security related resilience, cryptography is deployed to resist confidentiality attack on messages delivered over the communication media. We analyze the impact of task’s fault tolerance on secure message communication, and then formulate the design problem as a multi-objective optimization problem, i.e., to minimize the failure probability and security vulnerability of the application while subject to given fault-tolerant constraints, execution constraints and deadline constraints. We propose an improved multi-objective optimization algorithm, called Decomposition-based Security and Fault tolerance Co-Optimization (DeSFCO) algorithm, to search for the optimal Pareto solutions of security and reliability harden assignments for messages and tasks. Two preliminary experiments evaluate efficiency of our approach.
Wei Jiang 0016, Xinke Liao, Jinyu Zhan
RTSS3
2021 ArchNet: A data hiding design for distributed machine learning systems
Wei Jiang 0016, Jinyu Zhan, Zicheng Gong, Weijia Pan
J. Syst. Archit.3
2021 Field programmable gate array-based all-layer accelerator with quantization neural networks for sustainable cyber-physical systems
abstract
Summary Low‐Bit Neural Network (LBNN) is a promising technique to enrich intelligent applications running on sustainable Cyber‐Physical Systems (CPS). Although LBNN has the advantages of low memory usage, fast inference and low power consumption, Low‐bit design requires additional computation units and may cause large accuracy drop. In this paper, we approach to design Field Programmable Gate Array (FPGA)‐based LBNN accelerator to support sustainable CPS. First, we propose a method to quantize the neural networks into 2‐bit weights, 8‐bit activations and 8‐bit biases with few accuracy loss. The mapping function is presented to approximate discrete space of weights gradually and quantize the activations and biases through the improved straight‐through estimator. Second, we design the bitwise FPGA‐based accelerator to speed up the LBNN. Different from traditional accelerating techniques (mainly focused on convolution layer), the dataflows of fully connected layer, pooling layer and convolution layer are considered to accelerate all layers of neural networks. The 2×8 bitwise multiplier implemented by AND/XOR operation is devised to replace 32×32‐bit multiplication unit, which can bring faster inference and lower power consumption. We conduct extensive experiments on benchmarks of MNIST, CIFAR‐10, CIFAR‐100 and ImageNet to evaluate the efficiency of our approach. The LBNN obtained by our quantization method can save 93.75% memory with 2.26% accuracy loss on average compared with original networks. The FPGA‐based accelerator achieves a peak performance of 427.71 GOPS under 100 MHz working frequency, which outperforms previous approaches significantly.
Jinyu Zhan, Xingzhi Zhou 0001, Wei Jiang 0016
Softw. Pract. Exp.1
2021 Attack-Aware Detection and Defense to Resist Adversarial Examples
abstract
This article approaches to design an attack-aware detection and defense framework to resist adversarial attacks on the security-critical artificial intelligent systems. We first make efforts to test the performances of adversarial attacks and present classifying and grading rule (CGR) for the fine-grained grouping of adversarial example attacks. According to CGR, adversarial attacks can be divided into six groups. Then, we propose a feature squeezing and CGR-based detector to detect adversarial attacks, which can be aware of the detailed attack group and is evaluated to be effective by extensive experiments. We also test the defense performances of typical defense methods against these six groups of adversarial attacks, and finally give the defense recommendations for each type of adversarial attack.
Wei Jiang 0016, Zhiyuan He 0001, Jinyu Zhan, Weijia Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 Research Progress and Challenges on Application-Driven Adversarial Examples: A Survey
abstract
Great progress has been made in deep learning over the past few years, which drives the deployment of deep learning–based applications into cyber-physical systems. But the lack of interpretability for deep learning models has led to potential security holes. Recent research has found that deep neural networks are vulnerable to well-designed input examples, called adversarial examples . Such examples are often too small to detect, but they completely fool deep learning models. In practice, adversarial attacks pose a serious threat to the success of deep learning. With the continuous development of deep learning applications, adversarial examples for different fields have also received attention. In this article, we summarize the methods of generating adversarial examples in computer vision, speech recognition, and natural language processing and study the applications of adversarial examples. We also explore emerging research and open problems.
Wei Jiang 0016, Zhiyuan He 0001, Jinyu Zhan, Weijia Pan, Deepak Adhikari
ACM Trans. Cyber Phys. Syst.3
2020 An FPGA based Network Interface Card with Query Filter for Storage Nodes of Big Data Systems
abstract
In this paper, we are interested in improving the data processing of storage and computing separated Big Data systems. We propose an Field Programmable Gate Array (FPGA) based Network Interface Card with Query Filter (NIC-QF) to accelerate the data query efficiency of storage nodes and reduce the workloads of computing nodes and the communication overheads between them. NIC-QF designed with PCIe core, query filter and NIC communication can filter the original data on storage nodes as an implicit coprocessor and directly send the filtered data to computing nodes of Big Data systems. Filter units in query filter can perform multiple SQL tasks in parallel, and each filter unit is internally pipelined, which can further speed up the data processing. Filter units can be designed to support general SQL queries on different data formats and we implement two schemes for TextFile and RCFile separately. Based on TPC-H benchmark and Tencent data set, we conduct extensive experiments to evaluate our design, which can achieve averagely up to 46.91% faster than the traditional approach.
Ying Li 0130, Jinyu Zhan, Wei Jiang 0016, Junting Wu, Jianping Zhu 0003
ASP-DAC2
2020 An Insight into Fault Propagation in Deep Neural Networks: Work-in-Progress
abstract
Reliability is of critical importance for Deep Neural Networks (DNNs) applied in safety-critical applications. Traditional analysis of fault propagation in DNNs is not suitable for such applications. In this paper we approach to give the theory-driven analysis of fault propagation in DNN. Specifically, the perturbation on weights of layers are formulated and the propagation conditions of faults are obtained through theoretical derivation. All the analysis is based on DNNs with 32-bit float numbers. Finally, initial experiments on three typical DNNs are conducted to evaluate our theoretical results.
Ruoxu Sun, Jinyu Zhan, Wei Jiang 0016
EMSOFT2
2020 Interpretability Derived Backdoor Attacks Detection in Deep Neural Networks: Work-in-Progress
abstract
Backdoor attacks to deep neural networks (DNNs) have received increasing attentions, particularly in applications from edge computing. The detection of backdoor attacks is a challenging task, due to the lack of transparency in DNN. In this paper, we design a novel method to detect backdoor attacks in deep neural networks, which is derived from the interpretability of a DNN. A comprehensive analysis of the critical path in DNN is conducted, based on which two indicators are proposed, including the correlation coefficient and the discrete degree. Conseqently, an efficient backdoor detection algorithm is proposed, which only needs a few runtime images to identify the backdoor attacks. Initial experiments indicated the efficiency.
Xiangyu Wen 0001, Wei Jiang 0016, Jinyu Zhan, Xupeng Wang 0001, Zhiyuan He 0001
EMSOFT3
2020 Optimized co-scheduling of mixed-precision neural network accelerator for real-time multitasking applications
Wei Jiang 0016, Jinyu Zhan, Zhiyuan He 0001, Xiangyu Wen 0001
J. Syst. Archit.3
2020 Design optimization of confidentiality-critical cyber physical systems with fault detection
Wei Jiang 0016, Liang Wen, Jinyu Zhan
J. Syst. Archit.3
2020 Branch-aware data variable allocation for energy optimization of hybrid SRAM+NVM SPM☆
Jinyu Zhan, Wei Jiang 0016, Jiayu Yu, Jinghuan Yu
J. Syst. Archit.1
2019 Vehicle data management with specific wear-levelling and fault tolerance for hybrid DRAM-NVM memory
Jinyu Zhan, Junhuan Yang, Wei Jiang 0016
J. Syst. Archit.1
2018 Writing-aware data variable allocation on hybrid SRAM+NVM SPM: work-in-progress
Jinyu Zhan, Wei Jiang 0016, Ying Li 0130
CASES2
2018 Persistence improvement for distributed cache with NVM based storage system: work-in-progress
Wei Jiang 0016, Jinyu Zhan, Jinghuan Yu, Liugen Xu
CASES3
2018 Design of security-critical distributed real-time applications with fault-tolerant constraint: work-in-progress
abstract
We approach the design of security-critical distributed applications with task-level fault-tolerant techniques. We focus on the impact of fault tolerance on secure message communication, which was seriously overlooked before. Fault-tolerant techniques, e.g., task re-execution and active replica, are leveraged to tolerate faults in task executions, while cryptography is deployed to protect the confidentiality of messages delivered over the communication media. The design problem is to minimize the schedule length and security vulnerability of the application, subject to given fault-tolerant constraints. We then propose a multi-objective optimization method to find the best solutions. Initial experiments indicated the efficiency.
Wei Jiang 0016, Jinyu Zhan
EMSOFT3
2018 Energy optimization of security-sensitive mixed-criticality applications for distributed real-time systems
Jinyu Zhan, Xia Zhang 0001, Wei Jiang 0016, Yue Ma 0001
J. Parallel Distributed Comput.1
2018 Energy-aware page replacement and consistency guarantee for hybrid NVM-DRAM memory systems
Jinyu Zhan, Yiming Zhang 0007, Wei Jiang 0016, Junhuan Yang, Lin Li 0051
J. Syst. Archit.1
2017 Energy-aware page replacement for NVM based hybrid main memory system
abstract
With the advantage of low power consumption, Non-Volatile Memories (NVMs) has been widely used in hybrid memory architecture. This paper presents a page replacement method based on NVM-DRAM hybrid main memory system for low power and consistency guarantee, called EAPR The energy consumption of page access in DRAMs and NVMs can be calculated according to the memory access, and the pages are migrated according to their energy consumption, by which the pages are determined to migrate from NVM to DRAM or from DRAM to NVM. Instead of deleting the logs directly, our approach guarantees the consistency of the hybrid memory architecture by optimizing the structure of logs after the transactions of the applications are submitted. Finally, the experimental results show that EAPR can not only reduce the energy consumption at least 20% compared with other page replacement algorithms but also guarantee the consistency of transactions in the NVM-DRAM hybrid memory system.
Yiming Zhang 0007, Jinyu Zhan, Junhuan Yang, Wei Jiang 0016, Lin Li 0051
RTCSA2
2017 Design optimization of secure message communication for energy-constrained distributed real-time systems
Wei Jiang 0016, Xia Zhang 0001, Jinyu Zhan, Yue Ma 0001
J. Parallel Distributed Comput.3
2013 A Vulnerability Optimization Method for Security-Critical Real-Time Systems
abstract
In this paper, we focus on task scheduling problems in security-critical real-time systems. We consider that all of the critical tasks are equipped with RC5 algorithm to reduce the vulnerability when facing security attacks. The relationships among security level, vulnerability and execution time of each task are firstly deduced. Then, a Multiple Task Vulnerability Optimization Method (MTVOM), which is based on dynamic programming algorithm, is devised to obtain minimal vulnerability under strict timing constraints. Furthermore, we take task criticality into consideration because each task is not equally important to its system. Finally, simulation experiments demonstrate the effectiveness of this method.
Xia Zhang 0001, Jinyu Zhan, Wei Jiang 0016, Yue Ma 0001
NAS2