EDBT 2026 Demo / reviewers in the wild / expert
Xitian Fan
dblp:140/2203
· DBLP profile ↗
15ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-8698-6584ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Agile Deployment System for Password Recovery on FPGAabstractHardware-based acceleration of password recovery remains a pressing challenge, as CPUs and GPUs struggle to efficiently process modern cryptographic primitives. Although FPGAs offer superior performance-per-watt, their widespread adoption is limited by long development cycles, manual optimization, and the absence of an end-to-end deployment framework that jointly accelerates password generation and verification. To address this gap, we propose the first agile, end-to-end FPGA-based password recovery system that unifies deep-learning-driven password generation and cryptographic verification within a single deployment workflow. The framework consists of: (1) a customized Neural Processing Unit (NPU) that accelerates GAN-based password generation models such as PassGAN; (2) an automated, template-based accelerator generator for verification kernels, built on reusable Chisel hardware primitives; and (3) a multi-objective Design Space Exploration (DSE) engine that co-optimizes kernel-level parameters (e.g., loop unrolling) and system-level parallelism to determine globally optimal FPGA configurations. We deploy the system on a heterogeneous platform combining a Zynq MPSoC with dual Virtex UltraScale+ FPGAs. Experimental results show that the NPU outperforms an NVIDIA Tesla V100 by 82.16% in PassGAN inference throughput. The full system achieves 1.90× higher end-to-end throughput and 2.32× better energy efficiency than GPU-based implementations, and delivers an average 32.58% speedup over state-of-the-art FPGA-only verification designs. These results demonstrate the practicality and scalability of our architecture for real-world password recovery workflows. Liming Deng, Guowei Zhu, Xitian Fan, Guangwei Xie, Mingqian Sun, Xuegong Zhou, Wei Cao 0002, Fan Zhang 0044, Xinsheng Yu 0001 |
IEEE Trans. Computers | 3 |
| 2025 | Agile Design Flow for Cryptographic Hardware AcceleratorsabstractThis paper presents an agile design flow for cryptographic hardware accelerators, which supports the automated generation of RTL code for efficient hardware accelerators intended for FPGA deployment. An automated design method for hardware accelerators based on domain-specific hardware design templates (DST) is proposed. Based on the characteristics of cryptographic algorithms, we designed a parameterized DST and constructed a corresponding specific hardware operator library (SHWOL). We adopted a novel operator matching strategy based on subgraph isomorphism to realize the mapping of user input algorithms in the operator library, thereby deriving the DST parameters for code generation thus finalizing the RTL code generation with design space exploration (DSE). When implemented on an FPGA, compared with the existing high-level synthesis (HLS) tools, the code generated by our proposed design flow has an LUT efficiency (throughput/number of LUTs) of up to 483× and energy efficiency (throughput/power) of up to 676×. Liming Deng, Guowei Zhu, Wei Cao 0002, Xitian Fan, Xuegong Zhou |
ICCD | 4 |
| 2025 | AHCA: Agile Design Framework for Hashcat Acceleration Based on FPGAabstractThis article presents AHCA, an agile design framework for Field Programmable Gate Array (FPGA)-based Hashcat acceleration that automates the generation of optimized register transfer level (RTL) code. Our approach is centered on a proposed automated design method using a parameterized domain-specific template (DST) and a specific hardware operator library. The framework analyzes an algorithm’s graph to extract key hardware operators and their interconnection network. To support diverse user inputs, we introduce an innovative operator matching strategy using subgraph isomorphism, which maps algorithms to our operator library. This matched information, combined with design space exploration (DSE), is used to configure the DST and generate the final RTL code, avoiding redundancy for previously implemented algorithms. Compared to state-of-the-art high-level synthesis (HLS) tools, AHCA demonstrates a maximum performance enhancement of 797×, a Look-Up table (LUT) efficiency improvement of up to 105×, and an energy efficiency gain of up to 676×. When deployed on an FPGA for password cracking, the AHCA-generated hardware achieves a 63.95× enhancement in energy efficiency over CPUs and a 4.71× improvement over GPUs. Liming Deng, Guowei Zhu, Xitian Fan, Wei Cao 0002, Xuegong Zhou, Fan Zhang 0044, Shaobo Yang |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | PTME: A Regular Expression Matching Engine Based on Speculation and Enumerative Computation on FPGAabstractFast regular expression matching is an essential task for deep packet inspection. In previous works, the regular expression matching engine on FPGA struggled to achieve an ideal balance between resource consumption and throughput. Speculation and enumerative computation exploits the statistical properties of deterministic finite automata, allowing for more efficient pattern matching. Existing related designs mostly revolve around vector instructions and multiple processors/cores or SIMD instruction sets, with a lack of implementation on FPGA platforms. We design a parallelized two-character matching engine on FPGA for efficiently fast filtering off fields with no pattern features. We transform the state transitions with sequential dependencies to the existing problem of elements in one set, enabling the proposed design to achieve high throughput with low resource consumption and support dynamic updates. Results show that compared with the traditional DFA matching, with a maximum resource consumption of 25% for on-chip FFs (74323/1045440) and LUTs (123902/522720), there is an improvement in throughput of 8.08–229.96× speedup and 87.61–99.56% speed-up(percentage improvement) for normal traffic, and 11.73–39.59× speedup and 91.47–97.47% speed-up(percentage improvement) for traffic with high-frequency match hits. Compared with the state-of-the-art similar implementation, our circuit on a single FPGA chip is superior to existing multi-core designs. Mingqian Sun, Guangwei Xie, Fan Zhang 0044, Wei Guo 0018, Xitian Fan, Jiayu Du |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2025 | FPGA-Based Large-Scale Sorting with Optimized Bandwidth UtilizationabstractFast sorting of large-scale data is an essential task for data centers. In previous works, the existing computational model of sorting kernel still results in lower bandwidth utilization on the external memory bus. And the execution of merge operations in merge sort circuit on FPGAs depends on control commands from the host CPU. In this case, the merge sort circuit is not fully offloaded to hardware layer for acceleration, resulting in a performance loss. We design an on-chip merge sort controller to efficiently command the merge sort process. The proposed controller has the ability to schedule multiple on-chip computing kernels simultaneously in a more efficient mode, thus ensuring that the circuit has a better bandwidth utilization. Meanwhile, fundamental factors affecting the performance of merge sort are studied and analyzed, and we propose a high-performance merge sort architecture. Results show that using the proposed controller-centered architecture, an overall improvement of 20%–30% in sorting throughput can be achieved. Compared with the state-of-the-art previous merge sorting implementation on FPGA, our circuit can achieve 1.22/1.46 \(\times\) speedup. Mingqian Sun, Guangwei Xie, Fan Zhang 0044, Wei Guo 0018, Xitian Fan, Jiayu Du |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2023 | Unified Accelerator for Attention and Convolution in Inference Based on FPGAabstractMany models combining Transformers with convolutional neural networks (CNNs) for computer vision tasks have achieved state-of-the-art results. However, due to the different computation patterns between attention and convolution, using a dedicated Transformer or CNN accelerator will inevitably reduce the computing efficiency of the other. To overcome this problem, we propose a unified architecture for attention and convolution on FPGA. We reduce runtime overhead by offloading part of self-attention computations offline before inference. Furthermore, we present a unified mapping method according to the computing characteristics of attention-based and convolution-based models. This accelerator implements multi-head attention in Transformer, independent ResNet-50 and hybrid blocks of attention and con-volution in BoTNet-50 at 200MHz on Xilinx Virtex Ultrascale+ XCVU37P. Experimental results show that the solution is nearly 3.62 times more energy-efficient than the NVIDIA V100 GPU, and the computational efficiency is 11.86% and 28.29% higher than the state-of-the-art Transformer and ResNet-50 accelerators, respectively. Fan Zhang 0044, Xitian Fan, Jianliang Shen, Wei Guo 0018, Wei Cao 0002 |
ISCAS | 3 |
| 2021 | LETA: A lightweight exchangeable-track accelerator for efficientnet based on FPGAabstractLightweight convolutional neural networks (CNNs) have become increasingly popular due to their lower computational complexity and fewer memory accesses with equivalent accuracy compared to previous CNN models. However, the newly proposed networks bring new challenges to efficient hardware design, such as, in EfficientNet, depthwise convolution, squeeze-and-excitation (SE) module, and swish/sigmoid functions. Although individual engine architecture could achieve a high computing efficiency for the standard convolution or the depth-wise convolution, it is still not efficient for EfficientNet because the workload imbalance between two types of convolutional engines causes inevitable idling. To overcome this problem, we present a lightweight reconfigurable computational kernel based on FPGA with an exchangeable-track datapath scheme. In addition, a low-accuracy-loss function replacement strategy is proposed for swish/sigmoid functions. Furthermore, the low-cost hardware architecture to implement the replaced functions is designed. The proposed accelerator (LETA) can implement EfficientNet on Xilinx XCVU37P with a 300 MHz system clock and a 600 MHz kernel clock. The linear growth of resource usage in the 4-kernel implementation in 1 super logic region (SLR) with the same clock frequencies justifies the scalability of LETA. The experimental results show that LETA can achieve 2× throughput/DSP compared to the latest FPGA-based accelerator with 1.6% (0.7%) top-1 (top-5) accuracy loss on EfficientNet-B3. Jingbo Gao, Yihan Hu 0003, Xitian Fan, Wai-Shing Luk, Wei Cao 0002, Lingli Wang |
FPT | 4 |
| 2021 | MRI-based brain tumor segmentation using FPGA-accelerated neural networkabstractBACKGROUND: Brain tumor segmentation is a challenging problem in medical image processing and analysis. It is a very time-consuming and error-prone task. In order to reduce the burden on physicians and improve the segmentation accuracy, the computer-aided detection (CAD) systems need to be developed. Due to the powerful feature learning ability of the deep learning technology, many deep learning-based methods have been applied to the brain tumor segmentation CAD systems and achieved satisfactory accuracy. However, deep learning neural networks have high computational complexity, and the brain tumor segmentation process consumes significant time. Therefore, in order to achieve the high segmentation accuracy of brain tumors and obtain the segmentation results efficiently, it is very demanding to speed up the segmentation process of brain tumors. RESULTS: Compared with traditional computing platforms, the proposed FPGA accelerator has greatly improved the speed and the power consumption. Based on the BraTS19 and BraTS20 dataset, our FPGA-based brain tumor segmentation accelerator is 5.21 and 44.47 times faster than the TITAN V GPU and the Xeon CPU. In addition, by comparing energy efficiency, our design can achieve 11.22 and 82.33 times energy efficiency than GPU and CPU, respectively. CONCLUSION: We quantize and retrain the neural network for brain tumor segmentation and merge batch normalization layers to reduce the parameter size and computational complexity. The FPGA-based brain tumor segmentation accelerator is designed to map the quantized neural network model. The accelerator can increase the segmentation speed and reduce the power consumption on the basis of ensuring high accuracy which provides a new direction for the automatic segmentation and remote diagnosis of brain tumors. Siyu Xiong, Guoqing Wu 0003, Xitian Fan, Zhongcheng Huang, Wei Cao 0002, Xuegong Zhou, Shijin Ding, Jinhua Yu 0003, Lingli Wang, Zhifeng Shi |
BMC Bioinform. | 3 |
| 2021 | SWM: A High-Performance Sparse-Winograd Matrix Multiplication CNN AcceleratorabstractMany convolutional neural network (CNN) accelerators are proposed to exploit the sparsity of the networks recently to enjoy the benefits of both computation and memory reduction. However, most accelerators cannot exploit the sparsity of both activations and weights. For those works that exploit both sparsity opportunities, they cannot achieve the stable load balance through a static scheduling (SS) strategy, which is vulnerable to the sparsity distribution. In this work, a balanced compressed sparse row format and a dynamic scheduling strategy are proposed to improve the load balance. A set-associate structure is also presented to tradeoff the load balance and hardware resource overhead. We propose SWM to accelerate the CNN inference, which supports both sparse convolution and sparse fully connected (FC) layers. SWM provides Winograd adaptability for large convolution kernels and supports both 16-bit and 8-bit quantized CNNs. Due to the activation sharing, 8-bit processing can achieve theoretically twice the performance of the 16-bit processing with the same sparsity. The architecture is evaluated with VGG16 and ResNet50, which achieves: at most 7.6 TOP/s for sparse-Winograd convolution and three TOP/s for sparse matrix multiplication with 16-bit quantization on Xilinx VCU1525 platform. SWM can process 310/725 images per second for VGG16/ResNet50 with 16-bit quantization. Compared with the state-of-the-art works, our design can achieve at least 1.53 × speedup and 1.8 × energy efficiency improvement. Xitian Fan, Wei Cao 0002, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | High Throughput CNN Accelerator Design Based on FPGAabstractDue to the fact that FPGA on-chip memory capacity increases significantly, the feature maps and weights of convolutional layers can be stored on chip, which can reduce the data movement between on-chip memory and off-chip memory. Hence, the bottleneck can shift from the bandwidth to the computing resources in convolutional layers, which will improve the performance dramatically. Under this circumstance, this paper quantitatively analyzes how to design the hardware architecture based on the roofline model to optimize the performance under the constraints of available on-chip computing resources and propose an efficient architecture. Our accelerator is implemented on Xilinx UltraScale+ FPGA with the performance of 9.39 TOPS and 6.86 TOPS for 8-bit data width with 100MHz main frequency and 400MHz DSP frequency on ResNet-50 and AlexNet, which outperforms the existing FPGA-based CNN accelerator. Liang Xie 0009, Xitian Fan, Wei Cao 0002, Lingli Wang |
FPT | 2 |
| 2018 | Stream Processing Dual-Track CGRA for Object Inference
Xitian Fan, Wei Cao 0002, Wayne Luk, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | DT-CGRA: Dual-track coarse-grained reconfigurable architecture for stream applicationsabstractThis paper presents a new type of coarse-grained reconfigurable architecture (CGRA) for the object inference domain in machine learning. The proposed CGRA is optimized for stream processing and a correspondent programming model called dual-track model is proposed. The CGRA is realized in Verilog HDL and implemented in SMIC 55 nm process, with the footprint of 3.79 mm2and consuming 1.79 W at 500 MHz. To evaluate the performance, eight machine-learning algorithms including HOG, CNN, k-means, PCA, SPM, linear-SVM, Softmax and Joint-Bayesian are selected as benchmarks. These algorithms cover a general machine learning flow in object inference domain: feature extraction, feature selection and inference. The experimental results show that the proposed CGRA can gain 1443× average energy efficiency comparing to the Intel i7-3770 CPU and 7.82× energy efficiency comparing to a high performance FPGA solution [19]. Xitian Fan, Huimin Li 0005, Wei Cao 0002, Lingli Wang |
FPL | 1 |
| 2016 | A high performance FPGA-based accelerator for large-scale convolutional neural networksabstractIn recent years, convolutional neural networks (CNNs) based machine learning algorithms have been widely applied in computer vision applications. However, for large-scale CNNs, the computation-intensive, memory-intensive and resource-consuming features have brought many challenges to CNN implementations. This work proposes an end-to-end FPGA-based CNN accelerator with all the layers mapped on one chip so that different layers can work concurrently in a pipelined structure to increase the throughput. A methodology which can find the optimized parallelism strategy for each layer is proposed to achieve high throughput and high resource utilization. In addition, a batch-based computing method is implemented and applied on fully connected layers (FC layers) to increase the memory bandwidth utilization due to the memory-intensive feature. Further, by applying two different computing patterns on FC layers, the required on-chip buffers can be reduced significantly. As a case study, a state-of-the-art large-scale CNN, AlexNet, is implemented on Xilinx VC709. It can achieve a peak performance of 565.94 GOP/s and 391 FPS under 156MHz clock frequency which outperforms previous approaches. Huimin Li 0005, Xitian Fan, Wei Cao 0002, Xuegong Zhou, Lingli Wang |
FPL | 2 |
| 2013 | Implementation of high performance hardware architecture of OpenSURF algorithm on FPGAabstractThis paper proposes a high performance hardware architecture of Speeded Up Robust Features (SURF) algorithm based on OpenSURF. In order to achieve high processing frame rate, the hardware architecture is designed with several characteristics. Firstly, a sliding window method is proposed to extract feature points in parallel at selected scale levels. As a result, the time cost in feature extraction can be greatly reduced. Secondly, data reuse strategy is proposed in orientation generation and descriptor generation to reduce the memory access times. In this way, 3.87x and 2.25X speedup are achieved respectively. Thirdly, the integral image is segmented to buffer in different memory blocks in order to support multiple data accessing in one clock cycle, which will further reduce the whole calculating time of our implementation. The hardware architecture is implemented on an XC6VSX475T FPGA with 156 MHz and its maximal frame rate for VGA format image can reach 356 frames per second (fps), which is 6.25 times frame rate of OpenSURF running on a server with a Xeon 5650 processor, and 6 times the reported frame rate of the recent implementation on three Vritex4 FPGAs [8]. Xitian Fan, Chenlu Wu, Wei Cao 0002, Xuegong Zhou, Shengye Wang, Lingli Wang |
FPT | 1 |
| 2013 | A hardware implementation of Bag of Words and Simhash for image recognitionabstractAlgorithms such as Bag of Words and Simhash have been widely used in image recognition. To achieve better performance as well as energy-efficiency, a hardware implementation of these two algorithms is proposed in this paper. To the best of our knowledge, it is the first time that these algorithms have been implemented on hardware for image recognition purpose. The proposed implementation is able to generate a fingerprint of an image and find the closest match in the database accurately. It is implemented on Xilinx's Virtex-6 SX475T FPGA. Tradeoffs between high performance and low hardware overhead are obtained through proper parallelization. The experimental result shows that the proposed implementation can process 1,018 images per second, approximately 17.8x faster than software on Intel's 12-thread Xeon X5650 processor. On the other hand, the power consumption is 0.35x compared to software-based implementation. Thus, the overall advantage in energy-efficiency is as much as 46x. The proposed architecture is scalable, and is able to meet various requirements of image recognition. Shengye Wang, Xuegong Zhou, Wei Cao 0002, Chenlu Wu, Xitian Fan, Lingli Wang |
FPT | 6 |