Wei Guo 0018

dblp:71/6601-18 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-1023-7277ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PTME: A Regular Expression Matching Engine Based on Speculation and Enumerative Computation on FPGA
abstract
Fast regular expression matching is an essential task for deep packet inspection. In previous works, the regular expression matching engine on FPGA struggled to achieve an ideal balance between resource consumption and throughput. Speculation and enumerative computation exploits the statistical properties of deterministic finite automata, allowing for more efficient pattern matching. Existing related designs mostly revolve around vector instructions and multiple processors/cores or SIMD instruction sets, with a lack of implementation on FPGA platforms. We design a parallelized two-character matching engine on FPGA for efficiently fast filtering off fields with no pattern features. We transform the state transitions with sequential dependencies to the existing problem of elements in one set, enabling the proposed design to achieve high throughput with low resource consumption and support dynamic updates. Results show that compared with the traditional DFA matching, with a maximum resource consumption of 25% for on-chip FFs (74323/1045440) and LUTs (123902/522720), there is an improvement in throughput of 8.08–229.96× speedup and 87.61–99.56% speed-up(percentage improvement) for normal traffic, and 11.73–39.59× speedup and 91.47–97.47% speed-up(percentage improvement) for traffic with high-frequency match hits. Compared with the state-of-the-art similar implementation, our circuit on a single FPGA chip is superior to existing multi-core designs.
Mingqian Sun, Guangwei Xie, Fan Zhang 0044, Wei Guo 0018, Xitian Fan, Jiayu Du
ACM Trans. Reconfigurable Technol. Syst.4
2025 FPGA-Based Large-Scale Sorting with Optimized Bandwidth Utilization
abstract
Fast sorting of large-scale data is an essential task for data centers. In previous works, the existing computational model of sorting kernel still results in lower bandwidth utilization on the external memory bus. And the execution of merge operations in merge sort circuit on FPGAs depends on control commands from the host CPU. In this case, the merge sort circuit is not fully offloaded to hardware layer for acceleration, resulting in a performance loss. We design an on-chip merge sort controller to efficiently command the merge sort process. The proposed controller has the ability to schedule multiple on-chip computing kernels simultaneously in a more efficient mode, thus ensuring that the circuit has a better bandwidth utilization. Meanwhile, fundamental factors affecting the performance of merge sort are studied and analyzed, and we propose a high-performance merge sort architecture. Results show that using the proposed controller-centered architecture, an overall improvement of 20%–30% in sorting throughput can be achieved. Compared with the state-of-the-art previous merge sorting implementation on FPGA, our circuit can achieve 1.22/1.46 \(\times\) speedup.
Mingqian Sun, Guangwei Xie, Fan Zhang 0044, Wei Guo 0018, Xitian Fan, Jiayu Du
ACM Trans. Reconfigurable Technol. Syst.4
2024 Diversity supporting robustness: Enhancing adversarial robustness via differentiated ensemble predictions
Xi Chen 0112, Wei Huang 0035, Ziwen Peng, Wei Guo 0018, Fan Zhang 0044
Comput. Secur.4
2024 Correction: Adversarial defence by learning differentiated feature representation in deep ensemble
Xi Chen 0112, Wei Huang 0035, Wei Guo 0018, Fan Zhang 0044, Jiayu Du, Zhizhong Zhou
Mach. Vis. Appl.3
2024 Adversarial defence by learning differentiated feature representation in deep ensemble
Xi Chen 0112, Huang Wei, Wei Guo 0018, Fan Zhang 0044, Jiayu Du, Zhizhong Zhou
Mach. Vis. Appl.3
2024 Boosting Multimode Ruling in DHR Architecture With Metamorphic Relations
abstract
ABSTRACT The DHR architecture provides a revolutionary security defense structure for cyberspace. The multimode ruling in DHR is expected to alleviate the oracle problem, which still suffers from the existence of common model vulnerability. In this work, we design a test segmentation method to transform multimode ruling to a metamorphic testing problem. The text test input that causes inconsistency of heterogeneous executors is converted to a condition set, and we extract subsets of conditions based on its syntax tree. The original test can exploit a specific vulnerability, the follow‐up tests are composed by different subsets of conditions within the original test. We collect the execution matrix for the follow‐up tests to analyse the impact of each subset of conditions on ruling decision. Metamorphic relations are extracted based on the localization of independent condition, that is, the subsets of conditions that can impact ruling decision independently. The executors in an inconsistent ruling should be examined with metamorphic testing methods, rather than traditional majority voting mechanism. The proposed test segmentation and improved multimode ruling methods are evaluated on two DHR‐based cases, SQL injection in cyber‐range system and deserialization attack in ‐ project. The experimental results show that our test segmentation can help to locate malicious expressions and the metamorphic testing‐based multimode ruling can generate more correct results than majority voting mechanism with an average 15.8% performance loss.
Ruosi Li, Xianglong Kong, Wei Guo 0018, Jingdong Guo, Hongfa Li, Fan Zhang 0044
Softw. Test. Verification Reliab.3
2024 ETRS: efficient turn restrictions setting method for boundary routers in chiplet-based systems
Zhipeng Cao 0001, Wei Guo 0018, Zhiquan Wan, Peijie Li, Qinrang Liu, Caining Wang, Yangxue Shao
J. Supercomput.2
2023 Unified Accelerator for Attention and Convolution in Inference Based on FPGA
abstract
Many models combining Transformers with convolutional neural networks (CNNs) for computer vision tasks have achieved state-of-the-art results. However, due to the different computation patterns between attention and convolution, using a dedicated Transformer or CNN accelerator will inevitably reduce the computing efficiency of the other. To overcome this problem, we propose a unified architecture for attention and convolution on FPGA. We reduce runtime overhead by offloading part of self-attention computations offline before inference. Furthermore, we present a unified mapping method according to the computing characteristics of attention-based and convolution-based models. This accelerator implements multi-head attention in Transformer, independent ResNet-50 and hybrid blocks of attention and con-volution in BoTNet-50 at 200MHz on Xilinx Virtex Ultrascale+ XCVU37P. Experimental results show that the solution is nearly 3.62 times more energy-efficient than the NVIDIA V100 GPU, and the computational efficiency is 11.86% and 28.29% higher than the state-of-the-art Transformer and ResNet-50 accelerators, respectively.
Fan Zhang 0044, Xitian Fan, Jianliang Shen, Wei Guo 0018, Wei Cao 0002
ISCAS5
2023 GAMBD: Generating adversarial malware against MalConv
Wei Guo 0018, Fan Zhang 0044, Jiayu Du
Comput. Secur.2
2023 SFTN: Fast object detection for aerial images
abstract
Abstract The task of remote sensing image object detection in low latency scenes is of great research significance. To address the problem that the current high‐precision object detection algorithm based on a feature pyramid network is slow due to a large number of parameters and complicated computation, a fast remote sensing image object detection method based on a Single‐scale Feature Transformation Network (SFTN) is proposed. Firstly, based on the single‐scale remote sensing image features, the new channel features are quickly generated by a linear transformation of the original features and convolution kernel clustering optimization using cosine similarity; secondly, in order to obtain multi‐scale receptive fields, a parallel residual hole convolution module is designed to cover multi‐category remote sensing object scales on the feature map; finally, angle variables are introduced and optimized using angle similarity to effectively improve the object orientation accuracy. The experimental results on different datasets show that the method in this paper improves the detection speed rapidly while ensuring the accuracy of remote sensing image object detection, which is better than many remote sensing image object detection methods. The results demonstrate the reliability and robustness of the method.
Fan Zhang 0044, Wei Guo 0018, Mingqian Sun
IET Image Process.3
2020 MDLB: a metadata dynamic load balancing mechanism based on reinforcement learning
abstract
With the growing amount of information and data, object-oriented storage systems have been widely used in many applications, including the Google File System, Amazon S3, Hadoop Distributed File System, and Ceph, in which load balancing of metadata plays an important role in improving the input/output performance of the entire system. Unbalanced load on the metadata server leads to a serious bottleneck problem for system performance. However, most existing metadata load balancing strategies, which are based on subtree segmentation or hashing, lack good dynamics and adaptability. In this study, we propose a metadata dynamic load balancing (MDLB) mechanism based on reinforcement learning (RL). We learn that the Q_learning algorithm and our RL-based strategy consist of three modules, i.e., the policy selection network, load balancing network, and parameter update network. Experimental results show that the proposed MDLB algorithm can adjust the load dynamically according to the performance of the metadata servers, and that it has good adaptability in the case of sudden change of data volume.
Zhaoqi Wu, Fan Zhang 0044, Wei Guo 0018, Guangwei Xie
Frontiers Inf. Technol. Electron. Eng.4
2020 Erratum to: MDLB: a metadata dynamic load balancing mechanism based on reinforcement learning
abstract
Unfortunately the corresponding author’s ORCID was incorrect. It should be: Fan ZHANG, https://orcid.org/0000-0001-7456-8377
Zhaoqi Wu, Fan Zhang 0044, Wei Guo 0018, Guangwei Xie
Frontiers Inf. Technol. Electron. Eng.4
2019 A Reliable Physical Unclonable Function Based on Differential Charging Capacitors
abstract
Physical Unclonable Function (PUF) is an emerging security primitive for cryptography applications. However, achieving a very high reliability against the environmental variations remains a main challenge in PUF design and a key barrier for its commercialization. This paper presents a new PUF design based on the charging of a symmetric MOS capacitor pair by constant current with cross-coupled positive feedback inverters. The proposed weak PUF features high raw response reliability against variations in power supply and temperature without power-up reset noise and other issues due to the power-down and up of an array of cells. Extensive Monte-Carlo simulations have been performed using a standard 110nm CMOS process technology. The simulated results show an almost ideal uniqueness of 50.03% and superior reliability of 97.70% over a temperature range from 0 °C to 80 °C, and 96.20% with the supply voltage varies from 1.2 V to 1.8 V. The response bit can be generated at a rate of 27.78 Mbps with an average power consumption of 20.86 μW at 1.5V, and the energy consumption is only 750 fJ/bit.
Wei Guo 0018, Chip-Hong Chang, Yuan Cao 0003, Shaojun Wei, Shouyi Yin, Chenchen Deng, Leibo Liu, Fan Zhang 0044
ISCAS2