Mingqian Sun

dblp:127/2849 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-4304-2836ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Agile Deployment System for Password Recovery on FPGA
abstract
Hardware-based acceleration of password recovery remains a pressing challenge, as CPUs and GPUs struggle to efficiently process modern cryptographic primitives. Although FPGAs offer superior performance-per-watt, their widespread adoption is limited by long development cycles, manual optimization, and the absence of an end-to-end deployment framework that jointly accelerates password generation and verification. To address this gap, we propose the first agile, end-to-end FPGA-based password recovery system that unifies deep-learning-driven password generation and cryptographic verification within a single deployment workflow. The framework consists of: (1) a customized Neural Processing Unit (NPU) that accelerates GAN-based password generation models such as PassGAN; (2) an automated, template-based accelerator generator for verification kernels, built on reusable Chisel hardware primitives; and (3) a multi-objective Design Space Exploration (DSE) engine that co-optimizes kernel-level parameters (e.g., loop unrolling) and system-level parallelism to determine globally optimal FPGA configurations. We deploy the system on a heterogeneous platform combining a Zynq MPSoC with dual Virtex UltraScale+ FPGAs. Experimental results show that the NPU outperforms an NVIDIA Tesla V100 by 82.16% in PassGAN inference throughput. The full system achieves 1.90× higher end-to-end throughput and 2.32× better energy efficiency than GPU-based implementations, and delivers an average 32.58% speedup over state-of-the-art FPGA-only verification designs. These results demonstrate the practicality and scalability of our architecture for real-world password recovery workflows.
Liming Deng, Guowei Zhu, Xitian Fan, Guangwei Xie, Mingqian Sun, Xuegong Zhou, Wei Cao 0002, Fan Zhang 0044, Xinsheng Yu 0001
IEEE Trans. Computers5
2025 S3Det: a fast object detector for remote sensing images based on artificial to spiking neural network conversion
abstract
Artificial neural networks (ANNs) have made great strides in the field of remote sensing image object detection. However, low detection efficiency and high power consumption have always been significant bottlenecks in remote sensing. Spiking neural networks (SNNs) process information in the form of sparse spikes, creating the advantage of high energy efficiency for computer vision tasks. However, most studies have focused on simple classification tasks, and only a few researchers have applied SNNs to object detection in natural images. In this study, we consider the parsimonious nature of biological brains and propose a fast ANN-to-SNN conversion method for remote sensing image detection. We establish a fast sparse model for pulse sequence perception based on group sparse features and conduct transform-domain sparse resampling of the original images to enable fast perception of image features and encoded pulse sequences. In addition, to meet accuracy requirements in relevant remote sensing scenarios, we theoretically analyze the transformation error and propose channel self-decaying weighted normalization (CSWN) to eliminate neuron overactivation. We propose S3Det, a remote sensing image object detection model. Our experiments, based on a large publicly available remote sensing dataset, show that S3Det achieves an accuracy performance similar to that of the ANN. Meanwhile, our transformed network is only 24.32% as sparse as the benchmark and consumes only 1.46 W, which is 1/122 of the original algorithm’s power consumption.
Fan Zhang 0044, Guangwei Xie, Yanzhao Gao, Xiaofeng Qi, Mingqian Sun
Frontiers Inf. Technol. Electron. Eng.6
2025 PTME: A Regular Expression Matching Engine Based on Speculation and Enumerative Computation on FPGA
abstract
Fast regular expression matching is an essential task for deep packet inspection. In previous works, the regular expression matching engine on FPGA struggled to achieve an ideal balance between resource consumption and throughput. Speculation and enumerative computation exploits the statistical properties of deterministic finite automata, allowing for more efficient pattern matching. Existing related designs mostly revolve around vector instructions and multiple processors/cores or SIMD instruction sets, with a lack of implementation on FPGA platforms. We design a parallelized two-character matching engine on FPGA for efficiently fast filtering off fields with no pattern features. We transform the state transitions with sequential dependencies to the existing problem of elements in one set, enabling the proposed design to achieve high throughput with low resource consumption and support dynamic updates. Results show that compared with the traditional DFA matching, with a maximum resource consumption of 25% for on-chip FFs (74323/1045440) and LUTs (123902/522720), there is an improvement in throughput of 8.08–229.96× speedup and 87.61–99.56% speed-up(percentage improvement) for normal traffic, and 11.73–39.59× speedup and 91.47–97.47% speed-up(percentage improvement) for traffic with high-frequency match hits. Compared with the state-of-the-art similar implementation, our circuit on a single FPGA chip is superior to existing multi-core designs.
Mingqian Sun, Guangwei Xie, Fan Zhang 0044, Wei Guo 0018, Xitian Fan, Jiayu Du
ACM Trans. Reconfigurable Technol. Syst.1
2025 FPGA-Based Large-Scale Sorting with Optimized Bandwidth Utilization
abstract
Fast sorting of large-scale data is an essential task for data centers. In previous works, the existing computational model of sorting kernel still results in lower bandwidth utilization on the external memory bus. And the execution of merge operations in merge sort circuit on FPGAs depends on control commands from the host CPU. In this case, the merge sort circuit is not fully offloaded to hardware layer for acceleration, resulting in a performance loss. We design an on-chip merge sort controller to efficiently command the merge sort process. The proposed controller has the ability to schedule multiple on-chip computing kernels simultaneously in a more efficient mode, thus ensuring that the circuit has a better bandwidth utilization. Meanwhile, fundamental factors affecting the performance of merge sort are studied and analyzed, and we propose a high-performance merge sort architecture. Results show that using the proposed controller-centered architecture, an overall improvement of 20%–30% in sorting throughput can be achieved. Compared with the state-of-the-art previous merge sorting implementation on FPGA, our circuit can achieve 1.22/1.46 \(\times\) speedup.
Mingqian Sun, Guangwei Xie, Fan Zhang 0044, Wei Guo 0018, Xitian Fan, Jiayu Du
ACM Trans. Reconfigurable Technol. Syst.1
2023 SFTN: Fast object detection for aerial images
abstract
Abstract The task of remote sensing image object detection in low latency scenes is of great research significance. To address the problem that the current high‐precision object detection algorithm based on a feature pyramid network is slow due to a large number of parameters and complicated computation, a fast remote sensing image object detection method based on a Single‐scale Feature Transformation Network (SFTN) is proposed. Firstly, based on the single‐scale remote sensing image features, the new channel features are quickly generated by a linear transformation of the original features and convolution kernel clustering optimization using cosine similarity; secondly, in order to obtain multi‐scale receptive fields, a parallel residual hole convolution module is designed to cover multi‐category remote sensing object scales on the feature map; finally, angle variables are introduced and optimized using angle similarity to effectively improve the object orientation accuracy. The experimental results on different datasets show that the method in this paper improves the detection speed rapidly while ensuring the accuracy of remote sensing image object detection, which is better than many remote sensing image object detection methods. The results demonstrate the reliability and robustness of the method.
Fan Zhang 0044, Wei Guo 0018, Mingqian Sun
IET Image Process.5