Wei Cao 0002

dblp:54/6265-2 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0003-0339-7093ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FHPSAC: FPGA-based High-Parallelism SAC Accelerator
abstract
Reinforcement learning (RL) enables autonomous decision-making in applications such as robotics and control, and Soft Actor-Critic (SAC) is a leading model-free algorithm for continuous tasks. However, SAC’s small-batch training updates and fine-grained computation lead to heavy scheduling and kernel-launch overheads on GPU, limiting efficiency. In this work, we present FHPSAC (FPGA-based High-Parallelism SAC Accelerator), the first FPGA-accelerated architecture dedicated to SAC training. First, we propose a hardware–software co-designed on-chip memory hierarchy to statically partition and allocate SAC’s training data for conflict-free parallel access. Second, we build a high-parallelism accelerator with a tensor core for GEMM (General Matrix Multiply) and a lightweight unit for irregular elementwise/reduction kernels. Finally, we implement the full system on a Xilinx XCVU9P FPGA and demonstrate significant speedup with low power. Experimental results show that FHPSAC obtains 5.33–14.85× speedup compared with the Intel Xeon Gold 6130 CPU, while outperforming an NVIDIA A100-SXM4 GPU by 2.90–10.43× in training latency with an average power of 40.17 W. FHPSAC substantially reduces SAC training latency, providing a computational foundation for large-scale SAC deployments.
Jiabin Xu, Wang Fan, Xuegong Zhou, Wei Cao 0002, Fengzhe Zhang, Fan Zhang 0044, Xinsheng Yu 0001
FCCM4
2026 An Agile Deployment System for Password Recovery on FPGA
abstract
Hardware-based acceleration of password recovery remains a pressing challenge, as CPUs and GPUs struggle to efficiently process modern cryptographic primitives. Although FPGAs offer superior performance-per-watt, their widespread adoption is limited by long development cycles, manual optimization, and the absence of an end-to-end deployment framework that jointly accelerates password generation and verification. To address this gap, we propose the first agile, end-to-end FPGA-based password recovery system that unifies deep-learning-driven password generation and cryptographic verification within a single deployment workflow. The framework consists of: (1) a customized Neural Processing Unit (NPU) that accelerates GAN-based password generation models such as PassGAN; (2) an automated, template-based accelerator generator for verification kernels, built on reusable Chisel hardware primitives; and (3) a multi-objective Design Space Exploration (DSE) engine that co-optimizes kernel-level parameters (e.g., loop unrolling) and system-level parallelism to determine globally optimal FPGA configurations. We deploy the system on a heterogeneous platform combining a Zynq MPSoC with dual Virtex UltraScale+ FPGAs. Experimental results show that the NPU outperforms an NVIDIA Tesla V100 by 82.16% in PassGAN inference throughput. The full system achieves 1.90× higher end-to-end throughput and 2.32× better energy efficiency than GPU-based implementations, and delivers an average 32.58% speedup over state-of-the-art FPGA-only verification designs. These results demonstrate the practicality and scalability of our architecture for real-world password recovery workflows.
Liming Deng, Guowei Zhu, Xitian Fan, Guangwei Xie, Mingqian Sun, Xuegong Zhou, Wei Cao 0002, Fan Zhang 0044, Xinsheng Yu 0001
IEEE Trans. Computers7
2025 Agile Design Flow for Cryptographic Hardware Accelerators
abstract
This paper presents an agile design flow for cryptographic hardware accelerators, which supports the automated generation of RTL code for efficient hardware accelerators intended for FPGA deployment. An automated design method for hardware accelerators based on domain-specific hardware design templates (DST) is proposed. Based on the characteristics of cryptographic algorithms, we designed a parameterized DST and constructed a corresponding specific hardware operator library (SHWOL). We adopted a novel operator matching strategy based on subgraph isomorphism to realize the mapping of user input algorithms in the operator library, thereby deriving the DST parameters for code generation thus finalizing the RTL code generation with design space exploration (DSE). When implemented on an FPGA, compared with the existing high-level synthesis (HLS) tools, the code generated by our proposed design flow has an LUT efficiency (throughput/number of LUTs) of up to 483× and energy efficiency (throughput/power) of up to 676×.
Liming Deng, Guowei Zhu, Wei Cao 0002, Xitian Fan, Xuegong Zhou
ICCD3
2025 AHCA: Agile Design Framework for Hashcat Acceleration Based on FPGA
abstract
This article presents AHCA, an agile design framework for Field Programmable Gate Array (FPGA)-based Hashcat acceleration that automates the generation of optimized register transfer level (RTL) code. Our approach is centered on a proposed automated design method using a parameterized domain-specific template (DST) and a specific hardware operator library. The framework analyzes an algorithm’s graph to extract key hardware operators and their interconnection network. To support diverse user inputs, we introduce an innovative operator matching strategy using subgraph isomorphism, which maps algorithms to our operator library. This matched information, combined with design space exploration (DSE), is used to configure the DST and generate the final RTL code, avoiding redundancy for previously implemented algorithms. Compared to state-of-the-art high-level synthesis (HLS) tools, AHCA demonstrates a maximum performance enhancement of 797×, a Look-Up table (LUT) efficiency improvement of up to 105×, and an energy efficiency gain of up to 676×. When deployed on an FPGA for password cracking, the AHCA-generated hardware achieves a 63.95× enhancement in energy efficiency over CPUs and a 4.71× improvement over GPUs.
Liming Deng, Guowei Zhu, Xitian Fan, Wei Cao 0002, Xuegong Zhou, Fan Zhang 0044, Shaobo Yang
ACM Trans. Reconfigurable Technol. Syst.4
2025 DVHetero: A Framework for Designing and Validating Heterogeneous SoC with RISC-V Processor and CGRA
abstract
CGRA, as a coprocessor in SoCs, has been widely studied. However, there is limited research on how to efficiently debug and verify SoCs composed of CGRAs and processors during the design process. To address this gap, we introduce DVHetero. DVHetero incorporates a simulation and validation framework, SoCDiff, which enables comprehensive SoC simulation, debugging, and rapid error localization. Using this verification framework, we successfully implemented and validated the entire SoC. The SoC includes a Chisel-based CGRA generator and provides a pipelined CGRA architecture template. The CGRA is tightly integrated with the RISC-V processor, allowing for efficient DMA-based data transfer and MMIO support within the SoC. The pipelined CGRA architecture generated by DVHetero shows a 1.27× improvement in area efficiency and a 10.54× increase in mapping speed compared to the state-of-the-art CGRA framework, HierCGRA. Additionally, compared to state-of-the-art CGRA-SoC systems FDRA, DVHetero demonstrates a 1.67× increase in execution speed and a 4.34× improvement in area efficiency.
Guowei Zhu, Liming Deng, Kaisen Zhang, Wang Fan, Boyin Jin, Wei Cao 0002, Fengzhe Zhang, Xuegong Zhou, Fan Zhang 0044, Xinsheng Yu 0001
ACM Trans. Reconfigurable Technol. Syst.6
2023 Unified Accelerator for Attention and Convolution in Inference Based on FPGA
abstract
Many models combining Transformers with convolutional neural networks (CNNs) for computer vision tasks have achieved state-of-the-art results. However, due to the different computation patterns between attention and convolution, using a dedicated Transformer or CNN accelerator will inevitably reduce the computing efficiency of the other. To overcome this problem, we propose a unified architecture for attention and convolution on FPGA. We reduce runtime overhead by offloading part of self-attention computations offline before inference. Furthermore, we present a unified mapping method according to the computing characteristics of attention-based and convolution-based models. This accelerator implements multi-head attention in Transformer, independent ResNet-50 and hybrid blocks of attention and con-volution in BoTNet-50 at 200MHz on Xilinx Virtex Ultrascale+ XCVU37P. Experimental results show that the solution is nearly 3.62 times more energy-efficient than the NVIDIA V100 GPU, and the computational efficiency is 11.86% and 28.29% higher than the state-of-the-art Transformer and ResNet-50 accelerators, respectively.
Fan Zhang 0044, Xitian Fan, Jianliang Shen, Wei Guo 0018, Wei Cao 0002
ISCAS6
2021 A High-Precision Flexible Symmetry-Aware Architecture for Element-Wise Activation Functions
abstract
Nonlinear activation functions (NAFs) play an essential role in deep neural networks (DNNs). Since versatile DNN accelerators need to support various DNNs which contain different NAFs, the flexible hardware design supporting those NAFs has become crucial. However, there are few high-precision flexible hardware architectures, and the symmetries of different NAFs have not been fully studied. This paper proposes a high-precision symmetry-aware architecture based on piecewise linear approximation. Through the reconfigurable data path, the architecture can support various typical NAFs. The efficient non-uniform segmentation scheme is proposed to achieve high precision for each NAF. Besides, the utilization of unified symmetry for NAFs can save half the memory. To reduce the computational cost, a 25×18 DSP is shared by two INT 7×9 multipliers with two independent inputs. The architecture is implemented on Xilinx ZC706 at a frequency of 410MHz. Compared with the state-of-the-art flexible nonlinear core, our flexible architecture costs fewer hardware resources with higher precision. Applying the design to BERT-BASE, MobileNetV3, and EfficientNet-B3 on the PyTorch platform, experimental results show that the accuracy loss is either 0 for BERT-BASE, or 0.002% for EfficientNet-B3. For MobileNetV3, the accuracy is even improved by 0.01%.
Jingbo Gao, Wei Cao 0002, Lingli Wang
FPT5
2021 LETA: A lightweight exchangeable-track accelerator for efficientnet based on FPGA
abstract
Lightweight convolutional neural networks (CNNs) have become increasingly popular due to their lower computational complexity and fewer memory accesses with equivalent accuracy compared to previous CNN models. However, the newly proposed networks bring new challenges to efficient hardware design, such as, in EfficientNet, depthwise convolution, squeeze-and-excitation (SE) module, and swish/sigmoid functions. Although individual engine architecture could achieve a high computing efficiency for the standard convolution or the depth-wise convolution, it is still not efficient for EfficientNet because the workload imbalance between two types of convolutional engines causes inevitable idling. To overcome this problem, we present a lightweight reconfigurable computational kernel based on FPGA with an exchangeable-track datapath scheme. In addition, a low-accuracy-loss function replacement strategy is proposed for swish/sigmoid functions. Furthermore, the low-cost hardware architecture to implement the replaced functions is designed. The proposed accelerator (LETA) can implement EfficientNet on Xilinx XCVU37P with a 300 MHz system clock and a 600 MHz kernel clock. The linear growth of resource usage in the 4-kernel implementation in 1 super logic region (SLR) with the same clock frequencies justifies the scalability of LETA. The experimental results show that LETA can achieve 2× throughput/DSP compared to the latest FPGA-based accelerator with 1.6% (0.7%) top-1 (top-5) accuracy loss on EfficientNet-B3.
Jingbo Gao, Yihan Hu 0003, Xitian Fan, Wai-Shing Luk, Wei Cao 0002, Lingli Wang
FPT6
2021 MRI-based brain tumor segmentation using FPGA-accelerated neural network
abstract
BACKGROUND: Brain tumor segmentation is a challenging problem in medical image processing and analysis. It is a very time-consuming and error-prone task. In order to reduce the burden on physicians and improve the segmentation accuracy, the computer-aided detection (CAD) systems need to be developed. Due to the powerful feature learning ability of the deep learning technology, many deep learning-based methods have been applied to the brain tumor segmentation CAD systems and achieved satisfactory accuracy. However, deep learning neural networks have high computational complexity, and the brain tumor segmentation process consumes significant time. Therefore, in order to achieve the high segmentation accuracy of brain tumors and obtain the segmentation results efficiently, it is very demanding to speed up the segmentation process of brain tumors. RESULTS: Compared with traditional computing platforms, the proposed FPGA accelerator has greatly improved the speed and the power consumption. Based on the BraTS19 and BraTS20 dataset, our FPGA-based brain tumor segmentation accelerator is 5.21 and 44.47 times faster than the TITAN V GPU and the Xeon CPU. In addition, by comparing energy efficiency, our design can achieve 11.22 and 82.33 times energy efficiency than GPU and CPU, respectively. CONCLUSION: We quantize and retrain the neural network for brain tumor segmentation and merge batch normalization layers to reduce the parameter size and computational complexity. The FPGA-based brain tumor segmentation accelerator is designed to map the quantized neural network model. The accelerator can increase the segmentation speed and reduce the power consumption on the basis of ensuring high accuracy which provides a new direction for the automatic segmentation and remote diagnosis of brain tumors.
Siyu Xiong, Guoqing Wu 0003, Xitian Fan, Zhongcheng Huang, Wei Cao 0002, Xuegong Zhou, Shijin Ding, Jinhua Yu 0003, Lingli Wang, Zhifeng Shi
BMC Bioinform.6
2021 SWM: A High-Performance Sparse-Winograd Matrix Multiplication CNN Accelerator
abstract
Many convolutional neural network (CNN) accelerators are proposed to exploit the sparsity of the networks recently to enjoy the benefits of both computation and memory reduction. However, most accelerators cannot exploit the sparsity of both activations and weights. For those works that exploit both sparsity opportunities, they cannot achieve the stable load balance through a static scheduling (SS) strategy, which is vulnerable to the sparsity distribution. In this work, a balanced compressed sparse row format and a dynamic scheduling strategy are proposed to improve the load balance. A set-associate structure is also presented to tradeoff the load balance and hardware resource overhead. We propose SWM to accelerate the CNN inference, which supports both sparse convolution and sparse fully connected (FC) layers. SWM provides Winograd adaptability for large convolution kernels and supports both 16-bit and 8-bit quantized CNNs. Due to the activation sharing, 8-bit processing can achieve theoretically twice the performance of the 16-bit processing with the same sparsity. The architecture is evaluated with VGG16 and ResNet50, which achieves: at most 7.6 TOP/s for sparse-Winograd convolution and three TOP/s for sparse matrix multiplication with 16-bit quantization on Xilinx VCU1525 platform. SWM can process 310/725 images per second for VGG16/ResNet50 with 16-bit quantization. Compared with the state-of-the-art works, our design can achieve at least 1.53 × speedup and 1.8 × energy efficiency improvement.
Xitian Fan, Wei Cao 0002, Lingli Wang
IEEE Trans. Very Large Scale Integr. Syst.3
2018 RNA: An Accurate Residual Network Accelerator for Quantized and Reconstructed Deep Neural Networks
abstract
With the continuous refinement of Deep Neural Networks (DNNs), a series of deep and complex networks such as Residual Networks (ResNets) show impressive prediction accuracy in image classification tasks. Unfortunately, the structural complexity and computational cost of residual networks make hardware implementation difficult. In this paper, we present the quantized and reconstructed deep neural network (QR-DNN) technique, which first inserts batch normalization (BN) layers in the network during training, and later removes them to facilitate efficient hardware implementation. Moreover, an accurate and efficient residual network accelerator (RNA) is presented based on QR-DNN with batch-normalization-free structures and weights represented in a logarithmic number system. RNA employs a systolic array architecture to perform shift-and-accumulate operations instead of multiplication operations. QR-DNN is shown to achieve a 1% ~ 2% improvement in accuracy over existing techniques, and RNA over previous best fixed point accelerators. An FPGA implementation on a Xilinx Zynq XC7Z045 device achieves 804.03 GOPS, 104.15 FPS and 91.41% top-5accuracyfortheResNet-50benchmark, andstate-of-the-art results are also reported for AlexNet and VGG.
Wei Cao 0002, Philip H. W. Leong, Lingli Wang
FPL3
2018 A Novel Low-Communication Energy-Efficient Reconfigurable CNN Acceleration Architecture
abstract
Winograd algorithm is an efficient approach to alleviate the computation burden of deep CNNs. Firstly, we introduce a fast matrix algorithm to combine with Winograd algorithm to further reduce the computation complexity and adapt the Winograd algorithm to large-stride convolution with a kernel-partitioning method. Secondly, computation efficiency improvement due to the fast algorithms aggravates the off-chip communication. DRAM access of different data-flows varies significantly with different CNN patterns. Dynamic configurations of both data-flows and on-chip shared memory can reduce the DRAM access effectively. A quantitative analysis is established on the design space to guide the configurations. Finally, a reconfigurable architecture that supports three categories of data-flows is presented. For evaluation, VGGNet16, AlexNet and ResNet50 are implemented respectively which can achieve the state-of-art DSP efficiency. Overall performance of 685.6GOP/s, 1250GOP/s and 507GOP/s for AlexNet, VGGNet16 and ResNet50 respectively on ZC706 platform and better energy efficiency are achieved compared with representative prior works.
Wei Cao 0002, Lingli Wang
FPL3
2018 High Throughput CNN Accelerator Design Based on FPGA
abstract
Due to the fact that FPGA on-chip memory capacity increases significantly, the feature maps and weights of convolutional layers can be stored on chip, which can reduce the data movement between on-chip memory and off-chip memory. Hence, the bottleneck can shift from the bandwidth to the computing resources in convolutional layers, which will improve the performance dramatically. Under this circumstance, this paper quantitatively analyzes how to design the hardware architecture based on the roofline model to optimize the performance under the constraints of available on-chip computing resources and propose an efficient architecture. Our accelerator is implemented on Xilinx UltraScale+ FPGA with the performance of 9.39 TOPS and 6.86 TOPS for 8-bit data width with 100MHz main frequency and 400MHz DSP frequency on ResNet-50 and AlexNet, which outperforms the existing FPGA-based CNN accelerator.
Liang Xie 0009, Xitian Fan, Wei Cao 0002, Lingli Wang
FPT3
2018 Stream Processing Dual-Track CGRA for Object Inference
Xitian Fan, Wei Cao 0002, Wayne Luk, Lingli Wang
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Accelerating low bit-width convolutional neural networks with embedded FPGA
abstract
Convolutional Neural Networks (CNNs) can achieve high classification accuracy while they require complex computation. Binarized Neural Networks (BNNs) with binarized weights and activations can simplify computation but suffer from obvious accuracy loss. In this paper, low bit-width CNNs, BNNs and standard CNNs are compared to show that low bit-width CNNs is better suited for embedded systems. An architecture based on the two-stage arithmetic unit (TSAU) as the basic processing element is proposed to process each layer iteratively for low bit-width CNN accelerators. Then the DoReFa-Net which is trained with weights and activations represented in 1 bit and 2 bits respectively is implemented on Zynq XC7Z020 FPGA with a 410.2 GOPS performance. The accelerator can meet the real-time requirement of embedded applications with a 106 FPS throughput and a 73.1% top-5 accuracy on the ImageNet dataset. The accelerator outperforms existing FPGA-based CNN accelerators in the tradeoff among accuracy, energy and resource efficiency.
Wei Cao 0002, Xuegong Zhou, Lingli Wang
FPL3
2016 DT-CGRA: Dual-track coarse-grained reconfigurable architecture for stream applications
abstract
This paper presents a new type of coarse-grained reconfigurable architecture (CGRA) for the object inference domain in machine learning. The proposed CGRA is optimized for stream processing and a correspondent programming model called dual-track model is proposed. The CGRA is realized in Verilog HDL and implemented in SMIC 55 nm process, with the footprint of 3.79 mm2and consuming 1.79 W at 500 MHz. To evaluate the performance, eight machine-learning algorithms including HOG, CNN, k-means, PCA, SPM, linear-SVM, Softmax and Joint-Bayesian are selected as benchmarks. These algorithms cover a general machine learning flow in object inference domain: feature extraction, feature selection and inference. The experimental results show that the proposed CGRA can gain 1443× average energy efficiency comparing to the Intel i7-3770 CPU and 7.82× energy efficiency comparing to a high performance FPGA solution [19].
Xitian Fan, Huimin Li 0005, Wei Cao 0002, Lingli Wang
FPL3
2016 A high performance FPGA-based accelerator for large-scale convolutional neural networks
abstract
In recent years, convolutional neural networks (CNNs) based machine learning algorithms have been widely applied in computer vision applications. However, for large-scale CNNs, the computation-intensive, memory-intensive and resource-consuming features have brought many challenges to CNN implementations. This work proposes an end-to-end FPGA-based CNN accelerator with all the layers mapped on one chip so that different layers can work concurrently in a pipelined structure to increase the throughput. A methodology which can find the optimized parallelism strategy for each layer is proposed to achieve high throughput and high resource utilization. In addition, a batch-based computing method is implemented and applied on fully connected layers (FC layers) to increase the memory bandwidth utilization due to the memory-intensive feature. Further, by applying two different computing patterns on FC layers, the required on-chip buffers can be reduced significantly. As a case study, a state-of-the-art large-scale CNN, AlexNet, is implemented on Xilinx VC709. It can achieve a peak performance of 565.94 GOP/s and 391 FPS under 156MHz clock frequency which outperforms previous approaches.
Huimin Li 0005, Xitian Fan, Wei Cao 0002, Xuegong Zhou, Lingli Wang
FPL4
2016 High performance Deformable Part Model accelerator based on FPGA
abstract
Deformable Part Model (DPM) is one of the best algorithms for image-based object detection. However, the high computation intensity leads to relatively long detecting time. Even with the powerful CPU or GPU computing system, it is still too slow for practical applications. To solve this problem, this paper proposes a high performance DPM accelerator based on FPGA, where a dedicated JPEG decoder is integrated to process the images with the 1080p JPEG format. Pipelined architecture and data reuse strategies are developed to achieve the high throughput and energy efficiency. The proposed accelerator can process input images with 22 fps at the frequency of 156MHz on Xilinx VC709 board, which outperforms previous approaches.
Qi Zhan, Wei Cao 0002, Xuegong Zhou, Lingli Wang
FPT4
2013 Implementation of high performance hardware architecture of OpenSURF algorithm on FPGA
abstract
This paper proposes a high performance hardware architecture of Speeded Up Robust Features (SURF) algorithm based on OpenSURF. In order to achieve high processing frame rate, the hardware architecture is designed with several characteristics. Firstly, a sliding window method is proposed to extract feature points in parallel at selected scale levels. As a result, the time cost in feature extraction can be greatly reduced. Secondly, data reuse strategy is proposed in orientation generation and descriptor generation to reduce the memory access times. In this way, 3.87x and 2.25X speedup are achieved respectively. Thirdly, the integral image is segmented to buffer in different memory blocks in order to support multiple data accessing in one clock cycle, which will further reduce the whole calculating time of our implementation. The hardware architecture is implemented on an XC6VSX475T FPGA with 156 MHz and its maximal frame rate for VGA format image can reach 356 frames per second (fps), which is 6.25 times frame rate of OpenSURF running on a server with a Xeon 5650 processor, and 6 times the reported frame rate of the recent implementation on three Vritex4 FPGAs [8].
Xitian Fan, Chenlu Wu, Wei Cao 0002, Xuegong Zhou, Shengye Wang, Lingli Wang
FPT3
2013 An FPGA-cluster-accelerated match engine for content-based image retrieval
abstract
In this paper, a high-performance match engine for content-based image retrieval is proposed. Highly customized floating-point(FP) units are designed, to provide the dynamic range and precision of standard FP units, but with considerably less area than standard FP units. Match calculation arrays with various architectures and scales are designed and evaluated. An CBIR system is built on a 12-FPGA cluster. Inter-FPGA connections are based on standard 10-Gigabyte Ethernet. The whole FPGA cluster can compare a query image against 150 million library images within 10 seconds, basing on detailed local features. Compared with the Intel Xeon 5650 server based solution, our implementation is 11.35 times faster and 34.81 times more power efficient.
Chenlu Wu, Xuegong Zhou, Wei Cao 0002, Shengye Wang, Lingli Wang
FPT4
2013 A hardware implementation of Bag of Words and Simhash for image recognition
abstract
Algorithms such as Bag of Words and Simhash have been widely used in image recognition. To achieve better performance as well as energy-efficiency, a hardware implementation of these two algorithms is proposed in this paper. To the best of our knowledge, it is the first time that these algorithms have been implemented on hardware for image recognition purpose. The proposed implementation is able to generate a fingerprint of an image and find the closest match in the database accurately. It is implemented on Xilinx's Virtex-6 SX475T FPGA. Tradeoffs between high performance and low hardware overhead are obtained through proper parallelization. The experimental result shows that the proposed implementation can process 1,018 images per second, approximately 17.8x faster than software on Intel's 12-thread Xeon X5650 processor. On the other hand, the power consumption is 0.35x compared to software-based implementation. Thus, the overall advantage in energy-efficiency is as much as 46x. The proposed architecture is scalable, and is able to meet various requirements of image recognition.
Shengye Wang, Xuegong Zhou, Wei Cao 0002, Chenlu Wu, Xitian Fan, Lingli Wang
FPT4